AI Agent foundations, architecture and workflow design
Build a practical mental model of agents from models, context, tools, state, workflows and verification.
What is an agent?
A normal LLM call maps context to an answer. An agent places the model in a controlled loop: decide, call a tool, read the result, update state and stop only when a delivery condition is met.
- Conversation: one response for a question.
- RAG: retrieve context and generate an evidence-based answer.
- Workflow: code controls known nodes and branches.
- Single agent: dynamically selects restricted tools.
- Multi-agent: several specialized roles collaborate behind explicit boundaries.
Reference architecture
User/API -> session and auth -> orchestration and state -> model routing -> RAG/Memory/Tools/MCP -> business data -> trace, evaluation, cost and guardrails -> answer or approved actionExplain the architecture through the request lifecycle: identity, state, model decision, restricted tool execution, trace, output validation and business acceptance.
Execution loop and state machine
while not finished and steps < max_steps:
decision = model(messages, tools, state)
if decision requests a tool:
validate(tool, arguments)
result = execute(tool, arguments)
state = reduce(state, result)
else:
return validate_and_deliver(decision.content)
raise StepLimitExceeded()Set explicit limits for steps, time, permissions and delivery. Persist task ID, current node, input summary, tool records, pending approvals and failure reason when a task must resume later.
ReAct, planning and workflows
- ReAct: flexible step-by-step action, but potentially expensive.
- Plan and execute: easier to inspect for long tasks, but plans can become stale.
- Supervisor: a central agent routes work to specialists, adding context complexity.
- Fixed workflow: stable and auditable when the process or risk boundary is known.
Do not add multi-agent complexity just to appear intelligent. Establish a measurable single-agent or workflow baseline first.
Context and memory
- Short-term context: messages, tool results and current constraints.
- Working memory: plan, intermediate variables and pending tasks.
- Long-term memory: stable preferences and facts with sources and timestamps.
- External knowledge: documents and business records entered through retrieval and permissions.
Do not concatenate all history blindly. Keep required source text, summarize older context and re-check relevance and access when reading durable facts.
Architecture selection rules
- Fixed, approval-heavy process: deterministic workflow with a few LLM nodes.
- Enterprise knowledge: RAG with structured answers and citations.
- Dynamic but low-risk tools: one agent with restricted capabilities.
- Write, payment or deletion: deterministic code plus human approval.
Minimum exercise
Build a project-status assistant with two read-only tools and explicit uncertainty handling.
- Validate tool arguments with JSON Schema.
- Set a four-step limit, a three-second tool timeout and a ten-second task timeout.
- Persist model requests, tool inputs, results and the final answer trace.
- Return structured errors instead of pretending a failed tool succeeded.
- Test success, missing permission, missing record, timeout and conflicting data.
Continue with RAG knowledge-base engineering.