Evaluators — Testing & Safety
Module 01 established LLMs are non-deterministic and can hallucinate confidently. This module is the engineering discipline that catches it — before your users do.
Learning objectives
- Beginner: Log evaluator failures without acting on them, to establish a baseline of how often your system produces off-topic or ungrounded answers before building automated handling.
- Intermediate: Return a safe fallback message on evaluation failure instead of the raw (possibly wrong) answer, without automatic retry.
- Advanced: Bounded automatic retry via Spring Retry for high-stakes endpoints specifically, combined with observability (Module 17) tracking evaluation pass/fail rates as a first-class production metric, not just an ad-hoc log line.
◆ The problem
A normal unit test calls a method and asserts an exact expected output. Call the same LLM prompt twice and you can get two different (both individually plausible) responses — assertEquals(expected, actual) simply doesn't work as a testing strategy for LLM output the way it does for deterministic code.
Non-determinism means testing an AI application requires a different tool: instead of asserting exact output, you assert properties of the output — is it relevant to the question, is it grounded in the supplied context, does it avoid a list of disallowed topics. That's exactly what an Evaluator does.
RelevancyEvaluator checks whether a response actually addresses the user's query — typically by using a (possibly separate, cheaper) LLM call to judge the relationship between question and answer, rather than a keyword-overlap heuristic.
RelevancyEvaluator evaluator = new RelevancyEvaluator(chatClientBuilder); EvaluationRequest request = new EvaluationRequest(userQuestion, retrievedDocs, aiResponse); EvaluationResponse result = evaluator.evaluate(request); if (!result.isPass()) { log.warn("off-topic response detected: {}", result.getFeedback()); }
💻 Code example
RelevancyEvaluator evaluator = new RelevancyEvaluator(chatClientBuilder); EvaluationRequest request = new EvaluationRequest(userQuestion, retrievedDocs, aiResponse); EvaluationResponse result = evaluator.evaluate(request); if (!result.isPass()) { log.warn("off-topic response detected: {}", result.getFeedback()); }
FactCheckingEvaluator checks whether a response's claims are actually supported by a given context (typically the documents retrieved by RAG) — directly targeting Module 01's core hallucination risk by verifying grounding, not just plausibility.
FactCheckingEvaluator factChecker = new FactCheckingEvaluator(chatClientBuilder); EvaluationResponse result = factChecker.evaluate( new EvaluationRequest(userQuestion, retrievedDocs, aiResponse)); if (!result.isPass()) { // the response contains claims not supported by the retrieved context — treat as a hallucination }
◆ Under the hood
Both evaluators are themselves LLM calls judging another LLM call's output — "use an LLM to check an LLM" is a real, widely-used pattern (sometimes called "LLM-as-judge"), not a contradiction. It works because judging whether a given answer is supported by a given, fixed piece of text is a narrower, easier task than generating a correct answer from scratch — the same reason grading a multiple-choice-justification is easier than writing the essay.
💻 Code example
FactCheckingEvaluator factChecker = new FactCheckingEvaluator(chatClientBuilder); EvaluationResponse result = factChecker.evaluate( new EvaluationRequest(userQuestion, retrievedDocs, aiResponse)); if (!result.isPass()) { // the response contains claims not supported by the retrieved context — treat as a hallucination }
Chaining both evaluators after a RAG call gives you two independent checks: is the answer relevant to what was asked, and is it actually grounded in what was retrieved. A response can fail either independently — a relevant-but-unsupported answer, or an on-topic answer that drifted from the retrieved facts partway through.
◆ The problem
Catching a bad response is only half the job — what should actually happen when an evaluator fails in production, on a real user's request? Silently returning the bad answer defeats the point; hard-failing the request for something a retry might fix is often unnecessarily harsh.
Combining an evaluator with Spring Retry lets you automatically regenerate a response that fails evaluation, bounded by a max attempt count — the same generate-check-retry shape as Module 15's Evaluator-Optimizer pattern, applied here as a production safety net rather than a quality-improvement loop.
@Retryable(retryFor = UngroundedResponseException.class, maxAttempts = 3, backoff = @Backoff(delay = 500)) public String getGroundedAnswer(String question, List<Document> context) { String answer = ragChatClient.prompt().user(question).call().content(); EvaluationResponse check = factChecker.evaluate(new EvaluationRequest(question, context, answer)); if (!check.isPass()) { throw new UngroundedResponseException(check.getFeedback()); // triggers @Retryable } return answer; } @Recover public String fallback(UngroundedResponseException e, String question, List<Document> context) { return "I couldn't produce a confidently grounded answer to that. Please rephrase or contact support."; }
▲ Pitfall
Every retry re-runs the full RAG + generation + evaluation pipeline — evaluation is itself an LLM call, so a retry loop multiplies both latency and cost by the number of attempts. Reserve automatic retry-on-failed-evaluation for genuinely high-stakes answers (e.g. medical, legal, financial guidance) rather than applying it blanket across every endpoint.
✓ Quick recap
Why doesn't assertEquals work as a primary testing strategy for LLM output? LLM output is non-deterministic — the same prompt can produce different, individually valid responses, so you assert properties of the output rather than exact equality. What's the difference between RelevancyEvaluator and FactCheckingEvaluator? Relevancy checks whether the answer addresses the question; fact-checking verifies the answer's claims are actually supported by the given context. What's the main cost trade-off of automatic retry-on-evaluation-failure? Each retry re-runs the full generation and evaluation pipeline, multiplying both latency and cost — reserve it for genuinely high-stakes responses.
💻 Code example
@Retryable(retryFor = UngroundedResponseException.class, maxAttempts = 3, backoff = @Backoff(delay = 500)) public String getGroundedAnswer(String question, List<Document> context) { String answer = ragChatClient.prompt().user(question).call().content(); EvaluationResponse check = factChecker.evaluate(new EvaluationRequest(question, context, answer)); if (!check.isPass()) { throw new UngroundedResponseException(check.getFeedback()); // triggers @Retryable } return answer; } @Recover public String fallback(UngroundedResponseException e, String question, List<Document> context) { return "I couldn't produce a confidently grounded answer to that. Please rephrase or contact support."; }
Want a visual for this concept?
Generate a diagram tailored to “Evaluators — Testing & Safety” — the AI picks whichever visual (flowchart, comparison, sequence, etc.) best fits.
Sign in to generate a visual →