A failed listener invocation needs a bounded error policy, because an unacknowledged record alone does not decide retry or recovery.
Spring Kafka listener failures: separate transient retries from poison records
Classify the rejection
The local listener throws before acknowledging when the key disagrees with the payload, and it throws after a rolled-back transaction when stock is insufficient. Those are different failures. A mismatched key is unlikely to improve with time; a database outage may. Spring Kafka's container error handler can apply bounded backoff and a recovery destination, but the source kit deliberately does not register one or run a broker. The checked assertion is only that the method left the record unacknowledged.
Define what a recovered record means operationally: topic, original event ID, reason category, attempt count, and a way to inspect or replay it. A dead-letter publication can fail too. If the recovery handler returns success before a durable recovery record exists, the consumer can lose work. A broker integration test must check offset movement, recovery publication and restart behavior.
Do not retry a business rejection forever
The checked stock update requires available units. A reservation for 20 units against 19 fails and rolls back its event marker. Retrying that exact record indefinitely monopolizes a partition unless a later inventory change is intentionally part of the business contract. Decide whether to reject it, park it, or keep it pending with a deadline. The failure policy is a product rule with an operator path, not just a backoff number.
Code boundary
int updated = jdbc.update(
"update stock_balance set available = available - ? where sku = ? and available >= ?",
reservation.units(), reservation.sku(), reservation.units());
if (updated != 1) throw new IllegalStateException("stock item absent or insufficient");Verification boundary
The Java 21 / Spring Boot 4 source kit checks insufficientStockRollsBackMarkerAndDoesNotAcknowledge and mismatchedKeyIsRejectedBeforeDatabaseWork; broker retry and dead-letter behavior are not locally verified.
Cost and limits
A long retry interval delays records behind the failing record in its partition; very short retries can amplify load on an unhealthy database. The correct bound depends on service latency, retention and operator recovery time. No broker policy or dead-letter topic is tested here.
Common Mistakes
- Do not treat every exception as transient.
- Do not assume lack of acknowledgement alone defines a finite retry policy.
- Do not acknowledge a poison record before durable recovery if loss is unacceptable.
Read next
Spring Kafka manual acknowledgement: commit work before advancing the offset, Spring Kafka testing layers: know what a direct listener test cannot prove, Spring consumer idempotency: reserve an event ID with the stock mutation, Spring outbox retries: due time, delay cap, and attempt accounting.
Related boundary
Spring Cloud Stream Kafka DLQ: stop poison messages without losing the trail
