Chat Memory & Conversation Management
Module 01 established that LLMs are stateless. This module builds the illusion of memory on top — correctly, per-user, and without silently blowing your context window.
Learning objectives
- Beginner: In-memory MessageWindowChatMemory with a small maxMessages, fine for a single-session demo that doesn't need to survive a restart.
- Intermediate: JDBC-backed memory, conversation ID derived from the authenticated user, so real users get durable, isolated conversation history.
- Advanced: Combine bounded chat memory (recent context) with RAG (Module 09) for durable facts, so the system has both short-term conversational fluency and long-term recall without an ever-growing token bill.
ChatMemory is a repository abstraction — add() a message, get() the stored messages for a conversation. It has nothing to do with the model itself; it's plain storage that a MessageChatMemoryAdvisor reads from and writes to on every call, directly applying the advisor mechanism from Module 06.
public interface ChatMemory { void add(String conversationId, List<Message> messages); List<Message> get(String conversationId); void clear(String conversationId); }
💻 Code example
public interface ChatMemory { void add(String conversationId, List<Message> messages); List<Message> get(String conversationId); void clear(String conversationId); }
@Bean public ChatMemory chatMemory() { return MessageWindowChatMemory.builder().maxMessages(20).build(); // in-memory, §07.5 } @Bean public ChatClient chatClient(ChatClient.Builder builder, ChatMemory chatMemory) { return builder .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build()) .build(); }
With this wired up, every call through chatClient automatically has the conversation's prior messages prepended before the request is sent, and the new exchange appended to storage afterward — none of which your controller method has to know about.
💻 Code example
@Bean public ChatMemory chatMemory() { return MessageWindowChatMemory.builder().maxMessages(20).build(); // in-memory, §07.5 } @Bean public ChatClient chatClient(ChatClient.Builder builder, ChatMemory chatMemory) { return builder .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build()) .build(); }
◆ The problem
Without any way to distinguish conversations, every user of your application would share one giant, ever-growing history — user B's questions would leak into user A's context, and vice versa.
Every call must specify a CONVERSATION_ID so the advisor knows which stored history to read and append to — typically derived from the authenticated user's session or ID, not a value the client is trusted to set arbitrarily.
String answer = chatClient.prompt() .user(question) .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, currentUser.getId())) .call() .content();
▲ Pitfall
If a conversation ID is ever sourced from unauthenticated client input rather than the server-verified session, a malicious client could pass another user's conversation ID and read (or pollute) their stored chat history. Always derive the conversation ID server-side from an authenticated identity.
💻 Code example
String answer = chatClient.prompt() .user(question) .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, currentUser.getId())) .call() .content();
The in-memory ChatMemory from §07.2 is lost on every application restart — fine for a demo, unacceptable for a real product. JdbcChatMemoryRepository persists conversation history to a real database instead.
@Bean public ChatMemory chatMemory(JdbcChatMemoryRepository repository) { return MessageWindowChatMemory.builder() .chatMemoryRepository(repository) .maxMessages(20) .build(); }
spring.datasource.url=jdbc:postgresql://localhost:5432/chatdb
Spring AI auto-creates the schema it needs on top of your existing datasource — the same auto-configuration philosophy as everything else in this framework: point it at a DataSource and it handles the plumbing.
💻 Code example
@Bean public ChatMemory chatMemory(JdbcChatMemoryRepository repository) { return MessageWindowChatMemory.builder() .chatMemoryRepository(repository) .maxMessages(20) .build(); }
◆ The problem
A conversation that runs long enough will eventually accumulate more tokens of stored history than the model's context window can hold — every call resends the entire retained history (Module 01 §8), so an unbounded memory store is a ticking time bomb, not just a cost concern.
maxMessages on MessageWindowChatMemory caps how many of the most recent messages are retained and resent — older messages beyond that window are dropped, a sliding-window strategy that trades long-term recall for a bounded, predictable token cost per call.
◆ Under the hood — maxMessages is a message count, not a token budget
Setting maxMessages(20) caps message count, not token count — 20 short messages and 20 very long messages consume wildly different token budgets. For genuinely token-aware capacity planning, you need to reason about maxMessages together with your model's actual context window size and typical message length, not treat the message count alone as a safety guarantee.
▲ Pitfall
A sliding window silently drops the oldest messages first, including anything important the user established early in the conversation (e.g. "my account number is X," stated once at the start). If your application needs durable facts to survive indefinitely regardless of conversation length, that's a signal you need retrieval (RAG, Module 09) or explicit structured state, not sliding-window chat memory alone.
✓ Quick recap
What does ChatMemory actually do to the underlying model? Nothing directly — it's storage that an advisor reads/writes, prepending history to each new stateless call. Why must CONVERSATION_ID come from server-verified identity, not client input? An untrusted client-supplied ID could let one user read or pollute another user's stored conversation. Does maxMessages cap token usage directly? No — it caps message count; token usage still depends on how long each retained message is.
Want a visual for this concept?
Generate a diagram tailored to “Chat Memory & Conversation Management” — the AI picks whichever visual (flowchart, comparison, sequence, etc.) best fits.
Sign in to generate a visual →