Capstone — Building a Real-World AI Agent
Every module built or hardened one piece. Here's how they compose into a real agent — plus three production case studies that show these same pieces solving actual business problems.
Learning objectives
- Beginner: Point to which module contributed each piece of the finished agent (chat, memory, RAG, tools, observability).
- Intermediate: Explain how a real-world case study maps this site's patterns onto an actual business problem.
- Advanced: Extend the capstone agent with a genuinely new capability (a new tool, a new provider) without breaking its existing reliability guarantees.
The reference agent for this module: an MCP server exposing domain tools (e.g. order lookup, inventory check, refund initiation), and an MCP client application (built with ChatClient) that connects to it, holds a conversation with a user, and autonomously decides which tools to call to satisfy each request.
The full stack: memory and RAG advisors ground and contextualize the request, the agent loop decides which MCP tools to call (with elicitation gating anything destructive), and an evaluator checks the final answer before it reaches the user — with observability (not pictured) instrumenting every hop.
| Capability | Modules |
|---|---|
| Explain what an LLM actually does mechanically — tokens, embeddings, attention, statelessness | 01 |
| Wire up any of six+ model providers behind one portable API | 02, 18 |
| Build correct prompts with proper message roles, templates, and structured typed output | 03, 04, 05 |
| Give a stateless model both short-term conversational memory and long-term grounded knowledge | 07, 09 |
| Choose and operate a production vector store from among 11 supported integrations | 08, 11 |
| Let a model safely take real action via tools, with proper error handling | 12 |
| Expose and consume tools as a portable, standardized MCP server/client, including sampling and elicitation | 13, 14 |
| Compose multiple LLM calls into workflows and agents, choosing the least autonomy that solves the problem | 15 |
| Catch hallucinations and off-topic answers before they reach a user | 16 |
| Monitor token cost, latency, and evaluation pass rate as first-class production metrics | 17 |
| Keep every credential out of source control, resolved from a real secrets manager | 18 |
Three production-shaped systems, each combining modules from this site differently:
| System | Combines |
|---|---|
| Multi-service career platform — separate Candidate, Job, Hiring, and AI Career-Advisor services, where the advisor service calls the other services as HTTP clients and synthesizes AI-generated resume/job-comparison output | Modules 03–06 (prompts/advisors), 12 (tool-style HTTP calls), 15 (Chain workflow for multi-step generation), 18 (secrets management across services) |
| Hotel booking customer-support assistant — a chatbot handling room availability, booking, and modification requests against a real backend | Modules 07 (memory), 09 (RAG over policy docs), 12 (tool calling against the booking backend), 06 (SafeGuardAdvisor) |
| Call-center transcription & triage — inbound call audio transcribed, summarized, and routed | Module 19 (transcription), 15 §4 (Routing workflow), 16 (evaluating summary accuracy before it reaches a human agent) |
None of these needed a fundamentally different toolkit — each is a different composition of the same modules, which is the whole point of building on a small set of well-understood primitives (Advisors, Tools, Evaluators, Workflows) rather than a bespoke architecture per project.
◆ Beyond what's covered here
Suggested extensions past what's covered here: Multi-agent orchestration — have one agent's tool calls trigger a second, independently running agent (e.g. via its own MCP server), rather than a single agent with a flat tool list. Fine-tuning — train a smaller, task-specific model on your own labeled data instead of relying entirely on prompting a general-purpose model. Guardrails as a dedicated service — extract the SafeGuardAdvisor/Evaluator logic into a shared, independently-deployed policy service multiple AI applications in your organization can call, instead of duplicating safety logic per application. A full eval harness — a regression-test suite of golden question/answer pairs run against every model/prompt change, so prompt engineering has the same "did this break anything" safety net as regular code changes.
✓ Final check
If you can look at a real product requirement — "let users ask natural-language questions about their order history, safely, with citations" — and immediately decompose it into Advisors, a VectorStore, Tools, an Evaluator, and a workflow shape, rather than reaching for one giant prompt and hoping, this site has done its job. Go build something.
Want a visual for this concept?
Generate a diagram tailored to “Capstone — Building a Real-World AI Agent” — the AI picks whichever visual (flowchart, comparison, sequence, etc.) best fits.
Sign in to generate a visual →