Skip to content
Cognitecta

Training / Architecture

Advanced · 2 days

Production AI Architecture

A two-day architecture course for senior engineers and technical leaders: reference architectures, model gateways, RAG, agents, security, evaluation and the operational concerns that decide whether a system can leave the lab.

2 days · Advanced · Architecture · Architects, senior engineers and platform engineers

Talk to us

Overview

Production AI systems fail less often because the model is weak, and more often because the surrounding architecture was never designed: no gateway, no evaluation, no identity story, no cost control, and no owner for failure. This course is for people who have to draw that architecture and defend it.

The course works through reference architectures for hosted and self-hosted models, model abstraction, RAG, agents, event-driven and asynchronous workloads. Data boundaries, identity, permissions, secrets and governance are treated as structural, not as a later security review. Observability, prompt versioning, evaluation, rate limits, fallbacks and resilience are part of the same design.

The work is architectural. You complete design exercises and design reviews, including multi-model systems and the latency, scaling and deployment choices those systems force.

You leave with a production-shaped reference you can take back to a real programme, and a review method for designs that still look like a demo.

Audience

Software architects, senior engineers, platform engineers and technical leaders responsible for production shape, not only for a prototype.

Prerequisites

Experience designing or delivering production software. Familiarity with LLMs is expected. You do not need to implement models, but you should be able to read an architecture and argue about interfaces, failure and operations.

Duration

2 days

Can be combined with the enterprise platform course, or focused on a live customer architecture.

Learning outcomes

  1. 01

    Design a reference architecture with a model gateway, clear data boundaries and an application-facing interface.

  2. 02

    Compare hosted, self-hosted and multi-model inference against latency, cost, control and failure domains.

  3. 03

    Place RAG, agents and event-driven workloads in an architecture without collapsing them into a single chat service.

  4. 04

    Specify identity, permissions, secrets and governance for model access and retrieved data.

  5. 05

    Define observability, evaluation and prompt-version management as platform concerns.

  6. 06

    Design for resilience: rate limits, fallbacks, timeouts and degraded operation.

  7. 07

    Review a proposed AI system and identify the production gaps that would block release.

Outline

  1. 01

    Reference architectures

    • The components that recur: gateway, orchestration, retrieval, tools, evaluation, observability.
    • Model abstraction and why applications should not own provider details.
    • Hosted versus self-hosted inference, and when both exist in one estate.
    • Exercise: sketch a target architecture for a product team that currently calls a model directly.
  2. 02

    Workloads: RAG, agents and events

    • Synchronous assistants versus asynchronous generation and evaluation jobs.
    • RAG as a service: indexes, tenancy, and the application’s contract with retrieval.
    • Agents as software with tools, budgets and an authority boundary.
    • Event-driven AI: queues, retries, idempotency and human review queues.
    • Exercise: place three workloads on the same platform without sharing the wrong state.
  3. 03

    Security, identity and governance

    • Identity for users, services and tools; permissions on retrieval and actions.
    • Secrets, provider keys, and isolating untrusted content from privileged tools.
    • Data boundaries, tenancy, residency and logging that does not become a leak.
    • Governance: who can change prompts, models, tools and evaluation criteria.
    • Exercise: threat-model a gateway-plus-RAG design.
  4. 04

    Operations: cost, latency, scale and release

    • Observability: traces, token economics, retrieval records and user-visible errors.
    • Evaluation, prompt versions and model comparison as release machinery.
    • Cost controls, caching, rate limits and quotas per team or product.
    • Fallbacks, resilience, scalability, deployment and multi-model routing.
    • Design review: a complete production architecture against a realistic brief.

Practical work

Architecture exercises and design reviews, not coding labs. You produce architecture options, review a weak design, and complete a production architecture for a realistic system including gateway, RAG or agents, security and operations. Reviews are run as they would be in an engineering organisation.

Takeaways

  • Reference architecture diagrams and component checklists
  • A production-readiness review method
  • Decision notes for hosted versus self-hosted and multi-model routing
  • An observability, evaluation and cost-control baseline

Delivery

Cognitecta delivers private corporate training, on-site or as remote live training. Courses can be run as published, or adapted to your organisation’s stack, domain and experience level.

Instructor

Nicholas Johnson, AI architect and software engineer. He has a degree in Artificial Intelligence and around twenty years of professional technology training, including hundreds of courses for engineering teams and large organisations. About.

Related courses

  1. 2 days · Advanced

    Architecting Enterprise Generative AI

    A two-day course for architects and senior technical leaders: design enterprise-scale AI capability as a platform, not as a sequence of isolated projects.

  2. 2 days · Intermediate

    Building Enterprise AI Assistants

    A two-day course for developers and platform teams: build a secure internal assistant that can use organisational knowledge and systems, with identity, permissions, citations and evaluation.

Talk to us