intermediate~2h

Error Handling and Resilience for Spring Boot on AWS

Building a Spring Boot application that gracefully absorbs transient AWS SDK failures, returns consistent API error responses, and gives on-call engineers a fast path from a CloudWatch alert to the root cause.

Want a visual for this topic?

Generate a diagram tailored to Error Handling and Resilience for Spring Boot on AWS — the AI picks whichever visual (architecture, flowchart, ER diagram, etc.) best fits this specific AWS concept.

Sign in to generate a visual →
0
Subtopics

🎓 Learning objectives

  • •Distinguish transient AWS SDK failures (throttling, timeouts) from permanent client errors
  • •Configure SDK v2 retry policies and backoff behavior for resilient AWS calls
  • •Centralize API error responses with @ControllerAdvice and @ExceptionHandler
  • •Correlate application errors with CloudWatch Logs and X-Ray traces during incident response
  • •Apply circuit-breaker thinking to downstream AWS service calls

What is it?

This is the practice of layering three things on a Spring Boot application running on AWS: (1) retry/backoff handling for AWS SDK calls that fail for transient reasons, (2) a single, consistent place — @ControllerAdvice — where any exception that reaches the web layer gets translated into a well-formed API error response, and (3) structured logging and tracing that lets a production error be traced back to its root cause quickly, using CloudWatch Logs and X-Ray rather than guesswork.

Why it exists

Distributed systems fail partially and intermittently by nature — a downstream AWS service being briefly overloaded is normal, expected behavior, not an exceptional one. Error handling and resilience patterns exist because treating every failure as fatal produces a brittle application that returns errors to users for problems that would have resolved themselves in milliseconds, while treating every failure as retryable produces an application that can make outages worse by hammering an already-struggling dependency.

Problem it solves

It solves the gap between 'the AWS SDK threw an exception' and 'the user saw a sensible, actionable response while an engineer could find the root cause in minutes' — turning the normal, expected noise of a distributed system into something both resilient for users and diagnosable for operators.

Intuition

The guiding question for every failure path is: 'if this exact call is retried in one second, will it likely succeed?' Throttling, connection timeouts, and 500-level AWS responses are usually yes — the service is temporarily overloaded, not fundamentally rejecting the request. Validation errors, permission errors, and 'resource doesn't exist' errors are always no — retrying just repeats the same failure. Resilience code exists to apply exactly that distinction automatically instead of leaving a developer to reason about it ad hoc during an incident.

Analogy

A transient AWS throttling error is like a busy signal on a phone line — calling back a moment later usually gets through. A permanent error, like an access-denied response, is like finding out the number was disconnected: retrying the exact same call forever accomplishes nothing and just wastes everyone's time. Good resilience code tells these two situations apart instead of treating every failure the same way.

Technical explanation

AWS SDK v2 ships with an adaptive retry strategy by default, which combines exponential backoff with jitter and a token-bucket-style rate limiter that backs off further if retries keep failing, specifically to avoid amplifying load on a struggling service; this is configured via ClientOverrideConfiguration.builder().retryStrategy(...) on the client builder. Each SdkException subtype exposes whether it's considered retryable, and the SDK's internal retry condition checks this before attempting a retry, so a ValidationException is never retried while a ThrottlingException or 5xx service error is. On the Spring side, @ControllerAdvice works by registering @ExceptionHandler methods globally via Spring MVC's HandlerExceptionResolver chain, intercepting any exception that propagates out of a @Controller method before it reaches the servlet container's default error page, allowing a single class to own the mapping from internal exception types to the external HTTP/JSON contract. X-Ray trace propagation works by injecting a trace header (X-Amzn-Trace-Id) that's passed through HTTP calls and SDK calls alike, letting X-Ray assemble a full service map of segments and subsegments for a single logical request even as it crosses the ALB, the Spring Boot app, and any AWS service calls it makes.

Architecture

Requests enter through API Gateway or an ALB, hit the Spring Boot application, which calls out to AWS services (DynamoDB, S3, SQS, etc.) via the SDK. Each SDK client is configured with a ClientOverrideConfiguration carrying retry policy and timeout settings. Uncaught exceptions from the service layer propagate up to a global @ControllerAdvice, which maps them to HTTP responses. In parallel, the application emits structured logs (with trace/correlation IDs) to CloudWatch Logs via the CloudWatch Logs agent or the container's stdout (captured automatically by ECS/Fargate), and emits trace segments to X-Ray, letting an engineer pivot from a CloudWatch Logs Insights query to the matching X-Ray trace and back.

Workflow

  1. Classify AWS SDK exceptions: software.amazon.awssdk.core.exception.SdkException subtypes carry an isThrottlingException() / retryable() signal the SDK itself uses internally. 2) Configure the SDK client's RetryPolicy (SDK v2 uses a default adaptive retry strategy, but it can be tuned via ClientOverrideConfiguration) so transient errors (throttling, 5xx responses, socket timeouts) are retried with exponential backoff and jitter, while 4xx client errors are not retried. 3) Define a @ControllerAdvice class with @ExceptionHandler methods for specific exception types (e.g. ResourceNotFoundException → 404, a custom validation exception → 400, any unretried SdkException that escaped → 503) returning a consistent JSON error body (error code, message, correlation ID). 4) Use SLF4J with a structured logging encoder (e.g. logback JSON encoder) so logs shipped to CloudWatch Logs are queryable by field, not just grep-able text. 5) Enable AWS X-Ray tracing (the X-Ray SDK or OpenTelemetry with the AWS X-Ray exporter) so each request gets a trace ID, and include that trace ID in every log line and in the error response body so a user-reported error can be searched directly in X-Ray. 6) Set CloudWatch Alarms on error-rate and latency metrics (from Actuator/Micrometer or custom metrics) so a spike triggers an alert before a user has to report it.

Example

A Spring Boot service calling DynamoDB occasionally receives a ProvisionedThroughputExceededException during traffic spikes. Rather than letting that exception bubble up as a 500 error to the end user, the AWS SDK v2's built-in retry strategy (exponential backoff with jitter) automatically retries the call a few times transparently; if it's still failing after those retries, a @ControllerAdvice class catches the resulting SDK exception and returns a clean 503 Service Unavailable with a Retry-After header, while a structured log line tagged with the request's X-Ray trace ID lets the on-call engineer jump straight from the CloudWatch alarm to the exact request in X-Ray's service map.

Real-world usage

Production Spring Boot services on AWS almost universally rely on the SDK v2's built-in adaptive retry strategy for routine throttling, combined with a project-wide @ControllerAdvice and a logging/tracing setup (often via Micrometer + CloudWatch + X-Ray, or an equivalent observability stack) as a baseline resilience setup. Teams typically tune retry limits and circuit-breaker-style fallbacks only for specific downstream dependencies that have proven flaky in production, rather than hand-tuning every single call from day one.

Trade-offs

Centralizing error handling and retries makes behavior consistent and removes a huge class of duplicated try/catch logic, but it adds a layer of indirection that can obscure a specific failure if the @ControllerAdvice mapping is too generic (e.g. mapping every exception to a blanket 500). Aggressive retrying improves apparent reliability for the caller but can amplify load on an already-struggling downstream service if backoff and jitter aren't configured carefully — retries without jitter can create synchronized retry storms across many instances.

Visual explanation

Picture a triage nurse at the front of a hospital. A patient who just needs to wait a moment and try again (a throttling error) gets asked to sit back down briefly and come back. A patient with a condition no amount of waiting will fix (a permission or validation error) gets routed immediately to a specialist instead of being told to wait and retry. Every patient's visit is tagged with the same case number (the trace ID) no matter which path they took, so staff reviewing the day's incidents can follow any single visit start to finish.

Advantages

  • —

    Transient failures become invisible to end users instead of surfacing as application errors

  • —

    A single, consistent error response format makes the API predictable for client developers

  • —

    Correlation IDs and trace IDs turn 'something failed for some user somewhere' into a specific, searchable incident in minutes rather than hours

  • —

    Properly classified retries reduce unnecessary alerting and on-call noise for self-resolving issues

Disadvantages

  • —

    Misconfigured retries (too many attempts, no jitter) can turn a minor downstream blip into a self-inflicted overload

  • —

    Overly generic exception handlers can mask the real cause of an error behind a vague 500 response

  • —

    Structured logging and tracing add a small amount of per-request overhead and require disciplined instrumentation to stay useful

  • —

    Retrying non-idempotent operations (e.g. a payment call without an idempotency key) can cause duplicate side effects if not designed carefully

Common mistakes

  • —

    Retrying every exception type indiscriminately, including validation and permission errors that will never succeed on retry

  • —

    Catching Exception broadly in a @ExceptionHandler and returning a generic 500 for everything, erasing the distinction between client errors and server errors

  • —

    Logging an error without any correlation or trace ID, making it nearly impossible to connect a user's bug report to the specific log lines and trace

  • —

    Configuring retries with fixed delay instead of exponential backoff with jitter, which can synchronize retries across instances into a thundering herd

  • —

    Not setting a request timeout on SDK calls at all, letting a hung downstream dependency block a thread indefinitely

In the AWS Console

  1. 1

    X-Ray → Traces (to view after instrumentation is deployed)

    Enable X-Ray tracing for the application (via the X-Ray daemon as an ECS sidecar, or AWS Distro for OpenTelemetry).

  2. 2

    CloudWatch → Logs → Logs Insights

    Create a CloudWatch Logs Insights saved query filtering by the application's log group and a `traceId` field to quickly pull every log line for one request.

  3. 3

    CloudWatch → Alarms → Create alarm

    Create a CloudWatch Alarm on the application's 5xx/error-rate metric (from Container Insights or custom Micrometer metrics) with a threshold and SNS notification for on-call.

    Pairing this alarm's SNS topic with the incident runbook link in its description saves the on-call engineer a search during a live incident.

🎤 Interview questions

How do you decide whether an AWS SDK exception should be retried automatically versus surfaced immediately as an error? (Listen for: transient errors — throttling, timeouts, 5xx responses — are retryable because the same request is likely to succeed shortly after; client errors like validation or permission failures are not, since retrying repeats an identical failure)

What's the risk of configuring retries with a fixed delay instead of exponential backoff with jitter? (Listen for: fixed delay across many instances retrying simultaneously can synchronize into a thundering herd that worsens an already-struggling downstream service; jitter spreads retries out over time to avoid that)

Why centralize exception-to-HTTP-response mapping in a @ControllerAdvice rather than handling errors in each controller method? (Listen for: consistency of the API error contract across endpoints, avoiding duplicated try/catch logic, and a single place to update the response format or add new exception types)

A user reports an intermittent 500 error but gives no other detail. How does trace/correlation ID logging change how you'd investigate that compared to plain text logs? (Listen for: a correlation or X-Ray trace ID lets you pull every log line and the full service-call trace for that exact request instead of searching through unrelated log noise)

Why can retrying be dangerous for a non-idempotent operation like charging a payment, and how do you make a retry safe in that case? (Listen for: a naive retry can duplicate the side effect if the first attempt actually succeeded but the response was lost; an idempotency key lets the downstream service recognize and deduplicate a repeated request)

How would you use CloudWatch Logs Insights and X-Ray together during an incident instead of using either alone? (Listen for: CloudWatch Logs Insights is good for aggregate querying/filtering across many requests by field, while X-Ray shows the full call graph and timing of one specific request — pivoting between the two via a shared trace ID narrows from 'what's happening broadly' to 'what exactly happened in this one case')

What's a circuit-breaker pattern, and when would you add one on top of SDK-level retries for a call to a downstream AWS service? (Listen for: a circuit breaker stops attempting calls to a dependency entirely for a cooldown period after repeated failures, rather than retrying every single request individually, which prevents piling up latency and load against a service that's clearly down rather than just occasionally throttling)

💬 Deep Dive with AI

Related concepts

spring-boot-on-ec2cloudwatch-logs-and-metricsaws-sdk-v2-essentialssqs-dead-letter-queues-deep-dive