advanced~2.5h

Prompt Guarding, Jailbreak Defense & Moderation

An LLM-backed endpoint is reachable by anyone who can send it text, and some of that text will be an attempt to override its instructions or extract content it shouldn't produce. This module covers defending against prompt injection and jailbreaking at runtime, moderating input and output, and where that defense sits relative to the testing-focused Evaluators module.

Learning objectives

  • Beginner: Define prompt injection and jailbreaking, and explain how they differ from each other.
  • Beginner: Distinguish input moderation from output moderation and state why both are needed.
  • Intermediate: Harden a system prompt against common injection patterns without making it unusable for legitimate requests.
  • Intermediate: Implement a runtime guard-rail that checks a request before it reaches the model and a response before it reaches the user.
  • Advanced: Explain the difference between this module's runtime-defense angle and the Evaluators module's pre-deployment testing angle, and how the two combine.
  • Advanced: Design a layered defense (input filter, hardened system prompt, output filter) and justify why no single layer is sufficient alone.

This is a Pro chapter

Sign in, then upgrade to Pro or Power to unlock this and the full Spring Ecosystem Mastery library.

Prompt Guarding, Jailbreak Defense & Moderation

Next Step

Practice interview questions on this topic →← Back to all Spring AI chapters