AI Agent foundations, architecture and workflow design

Build a practical mental model of agents from models, context, tools, state, workflows and verification.

What is an agent?

A normal LLM call maps context to an answer. An agent places the model in a controlled loop: decide, call a tool, read the result, update state and stop only when a delivery condition is met.

  • Conversation: one response for a question.
  • RAG: retrieve context and generate an evidence-based answer.
  • Workflow: code controls known nodes and branches.
  • Single agent: dynamically selects restricted tools.
  • Multi-agent: several specialized roles collaborate behind explicit boundaries.

Reference architecture

User/API -> session and auth -> orchestration and state -> model routing -> RAG/Memory/Tools/MCP -> business data -> trace, evaluation, cost and guardrails -> answer or approved action

Explain the architecture through the request lifecycle: identity, state, model decision, restricted tool execution, trace, output validation and business acceptance.

Execution loop and state machine

agent-looppseudo
while not finished and steps < max_steps:
    decision = model(messages, tools, state)
    if decision requests a tool:
        validate(tool, arguments)
        result = execute(tool, arguments)
        state = reduce(state, result)
    else:
        return validate_and_deliver(decision.content)
raise StepLimitExceeded()

Set explicit limits for steps, time, permissions and delivery. Persist task ID, current node, input summary, tool records, pending approvals and failure reason when a task must resume later.

ReAct, planning and workflows

  • ReAct: flexible step-by-step action, but potentially expensive.
  • Plan and execute: easier to inspect for long tasks, but plans can become stale.
  • Supervisor: a central agent routes work to specialists, adding context complexity.
  • Fixed workflow: stable and auditable when the process or risk boundary is known.

Do not add multi-agent complexity just to appear intelligent. Establish a measurable single-agent or workflow baseline first.

Context and memory

  • Short-term context: messages, tool results and current constraints.
  • Working memory: plan, intermediate variables and pending tasks.
  • Long-term memory: stable preferences and facts with sources and timestamps.
  • External knowledge: documents and business records entered through retrieval and permissions.

Do not concatenate all history blindly. Keep required source text, summarize older context and re-check relevance and access when reading durable facts.

Architecture selection rules

  • Fixed, approval-heavy process: deterministic workflow with a few LLM nodes.
  • Enterprise knowledge: RAG with structured answers and citations.
  • Dynamic but low-risk tools: one agent with restricted capabilities.
  • Write, payment or deletion: deterministic code plus human approval.

Minimum exercise

Build a project-status assistant with two read-only tools and explicit uncertainty handling.

  • Validate tool arguments with JSON Schema.
  • Set a four-step limit, a three-second tool timeout and a ten-second task timeout.
  • Persist model requests, tool inputs, results and the final answer trace.
  • Return structured errors instead of pretending a failed tool succeeded.
  • Test success, missing permission, missing record, timeout and conflicting data.

Continue with RAG knowledge-base engineering.