Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Spring AI ChatClient request budgets: bound latency, tokens and work per caller

Last updated: 1 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

A ChatClient call can wait on a remote model and consume variable tokens; the application must put limits around caller work.

Download Spring source kit

Admit work before the model call

Set a request deadline, input byte cap, per-tenant concurrency cap and output-token ceiling. A streaming response needs cancellation propagation when the browser disconnects; otherwise the model may keep producing unused output. Measure accepted, rejected, cancelled and timed-out calls separately. Local admission control addresses the same capacity problem for worker threads, though a model provider adds remote latency and quota.

Keep business state outside the response text

Model output is not a transaction record. If a receipt workflow requires a write, call a service that validates the request and commits an idempotent operation; do not parse a sentence as proof that stock changed. Spring AI provides synchronous and streaming ChatClient paths, but this lesson is a boundary design, not a running provider integration in the source kit.

Boundary sketch

Java
if (requestText.length() > 4_096 || inFlightForTenant >= 3) {
    throw new RequestNotAdmittedException();
}
// Only after admission: invoke the configured ChatClient with a deadline.

Cost and verification

Token and latency costs vary by provider and prompt length. Admission limits spend less on rejected work but can return 429 during bursts; a queue without a cap only moves the overload. This sketch is not executed by the current Spring source kit; verify it against the chosen dependencies and deployment.

Common Mistakes

  • Do not let abandoned streams consume provider quota indefinitely.
  • Do not infer a durable business write from generated text.
  • Do not publish these limits as measured provider behavior; they are application policy examples.

Read next

Spring task executors: reject work when every slot is occupied, Spring AI tools: authorize each requested action after model selection, Spring AI conversation memory: partition by verified tenant and caller, Spring Boot liveness versus readiness: do not restart on every dependency outage.

spring
spring-boot
spring-ai
ai-chat-budget
Storage details