intermediate~4h

Agentic Workflow Patterns

"Agent" gets used to mean almost anything. This module gives you five concrete, composable workflow patterns — and the vocabulary to say precisely which one you're building.

Learning objectives

  • Beginner: A 2-step Chain workflow for a task with an obvious, fixed order of operations (draft, then polish).
  • Intermediate: A Routing workflow for a support-ticket triage system, where BILLING/TECHNICAL/GENERAL genuinely need different prompts and possibly different tools.
  • Advanced: A full agent (Module 20's capstone) with several MCP-exposed tools, bounded by a max-turns limit and an evaluator (Module 16) checking the final answer before it's returned — combining nearly every module on this site into one system.

Everything through Module 14 was still fundamentally "one request, one response" (possibly with tool calls inside it). An agentic workflow composes multiple LLM calls (and possibly tool calls) into a structured multi-step process — sometimes with a fixed structure your code defines (a "workflow"), sometimes with the model itself deciding the next step at runtime (an "agent"). Both are useful; knowing which one you're building changes how predictable and debuggable the result is.

The simplest pattern: decompose a task into a fixed sequence of LLM calls, each step's output feeding the next step's input. Good when a task naturally breaks into ordered sub-steps that don't need to branch.

Chain: a fixed sequence of steps, each call's output feeding the next.

String outline = chatClient.prompt().user("Outline a blog post about: " + topic).call().content(); String fullPost = chatClient.prompt() .user("Expand this outline into a full blog post:\n" + outline) .call().content();

💻 Code example

String outline = chatClient.prompt().user("Outline a blog post about: " + topic).call().content(); String fullPost = chatClient.prompt() .user("Expand this outline into a full blog post:\n" + outline) .call().content();

Run multiple independent LLM calls concurrently instead of sequentially, then combine their results — used when sub-tasks genuinely don't depend on each other's output, purely to reduce total latency.

CompletableFuture<String> grammar = CompletableFuture.supplyAsync(() -> chatClient.prompt().user("Check grammar issues in: " + text).call().content()); CompletableFuture<String> tone = CompletableFuture.supplyAsync(() -> chatClient.prompt().user("Assess the tone of: " + text).call().content()); CompletableFuture.allOf(grammar, tone).join(); String combined = "Grammar: " + grammar.get() + "\nTone: " + tone.get();

▲ Pitfall

Parallel calls multiply your concurrent API usage against the same provider — running headfirst into rate limits under load is a common first surprise when moving from a sequential chain to a parallel workflow. Plan for provider rate limits, not just wall-clock latency, when parallelizing.

💻 Code example

CompletableFuture<String> grammar = CompletableFuture.supplyAsync(() -> chatClient.prompt().user("Check grammar issues in: " + text).call().content()); CompletableFuture<String> tone = CompletableFuture.supplyAsync(() -> chatClient.prompt().user("Assess the tone of: " + text).call().content()); CompletableFuture.allOf(grammar, tone).join(); String combined = "Grammar: " + grammar.get() + "\nTone: " + tone.get();

One LLM call classifies the input, then routes it to a different, specialized prompt (or even a different model entirely) based on that classification — useful when different input categories genuinely need different handling, rather than one generic prompt trying to handle all of them adequately.

public enum Category { BILLING, TECHNICAL, GENERAL } Category category = chatClient.prompt() .user("Classify this support ticket as BILLING, TECHNICAL, or GENERAL: " + ticket) .call().entity(Category.class); String response = switch (category) { case BILLING -> billingChatClient.prompt().user(ticket).call().content(); case TECHNICAL -> technicalChatClient.prompt().user(ticket).call().content(); case GENERAL -> generalChatClient.prompt().user(ticket).call().content(); };

💻 Code example

public enum Category { BILLING, TECHNICAL, GENERAL } Category category = chatClient.prompt() .user("Classify this support ticket as BILLING, TECHNICAL, or GENERAL: " + ticket) .call().entity(Category.class); String response = switch (category) { case BILLING -> billingChatClient.prompt().user(ticket).call().content(); case TECHNICAL -> technicalChatClient.prompt().user(ticket).call().content(); case GENERAL -> generalChatClient.prompt().user(ticket).call().content(); };

◆ The problem

Routing (§15.4) assumes you know the categories up front and each input goes to exactly one path. Some tasks need to be dynamically broken into an unknown number of sub-tasks, each handled independently, then synthesized back together.

An orchestrator LLM call decides how to decompose the task into sub-tasks at runtime; independent worker calls handle each sub-task; a final synthesis step combines their results — more flexible than a fixed chain or a fixed set of routes, at the cost of being harder to predict and test.

Orchestrator-workers: an orchestrator call decides the sub-tasks at runtime; workers execute them independently; results are synthesized back together.

One LLM call generates a candidate result; a second call (the evaluator) critiques it against explicit criteria; if it fails, the critique feeds back into another generation attempt — looping until the evaluator is satisfied or a max-attempts limit is hit. This directly previews Module 16's Evaluators, applied here as a generation-quality loop rather than a pure safety check.

String draft = chatClient.prompt().user("Write a product description for: " + product).call().content(); for (int attempt = 0; attempt < 3; attempt++) { String critique = chatClient.prompt() .user("Critique this product description for clarity and persuasiveness. Reply PASS or list issues:\n" + draft) .call().content(); if (critique.startsWith("PASS")) break; draft = chatClient.prompt() .user("Revise this description to address: " + critique + "\n\nOriginal:\n" + draft) .call().content(); }

▲ Pitfall

Always cap the loop with a maximum attempt count. An evaluator that never returns PASS (a too-strict rubric, or a task the generator genuinely can't satisfy) turns this pattern into an unbounded, silently expensive loop with no natural stopping condition otherwise.

💻 Code example

String draft = chatClient.prompt().user("Write a product description for: " + product).call().content(); for (int attempt = 0; attempt < 3; attempt++) { String critique = chatClient.prompt() .user("Critique this product description for clarity and persuasiveness. Reply PASS or list issues:\n" + draft) .call().content(); if (critique.startsWith("PASS")) break; draft = chatClient.prompt() .user("Revise this description to address: " + critique + "\n\nOriginal:\n" + draft) .call().content(); }

Every pattern above has a structure your code defines. An agent inverts this: the model itself decides, at each step, what tool to call next and when it has enough information to stop — your code just runs the loop (call model → if it requested a tool, run it and feed back the result → repeat) until the model produces a final answer instead of another tool request. This is exactly Module 12 §12.3's mechanism, simply allowed to repeat multiple times instead of stopping after one tool call.

◆ Under the hood — agents vs. workflows, precisely

The distinction that actually matters in practice: a workflow's control flow is fixed by your code (which is why Chain/Parallel/Routing/Orchestrator are all predictable and easy to test), while an agent's control flow is decided by the model at runtime (which is more flexible for open-ended tasks, but harder to fully test or bound in advance — the same model non-determinism from Module 01 now applies to which steps run, not just what text comes back). Choose the least agentic pattern that actually solves your problem; reach for full agent autonomy only when the task genuinely can't be decomposed into a fixed or orchestrated structure ahead of time.

✓ Quick recap

What's the core difference between a workflow and an agent? A workflow's control flow (which steps run, in what order) is fixed by your code; an agent's control flow is decided by the model at runtime. When would Orchestrator-workers be a better fit than a fixed Routing workflow? When the task needs to be decomposed into an unknown number of sub-tasks at runtime, rather than classified into one of a few known fixed categories. Why must an Evaluator-optimizer loop have a hard attempt cap? An evaluator that never passes turns the loop unbounded and silently expensive with no natural exit condition.

Want a visual for this concept?

Generate a diagram tailored to “Agentic Workflow Patterns” — the AI picks whichever visual (flowchart, comparison, sequence, etc.) best fits.

Sign in to generate a visual →

Practice quiz

Next Step

Continue to Evaluators — Testing & Safety← Back to all Spring AI chapters