A ChatClient call can wait on a remote model and consume variable tokens; the application must put limits around caller work.
Spring AI ChatClient request budgets: bound latency, tokens and work per caller
Admit work before the model call
Set a request deadline, input byte cap, per-tenant concurrency cap and output-token ceiling. A streaming response needs cancellation propagation when the browser disconnects; otherwise the model may keep producing unused output. Measure accepted, rejected, cancelled and timed-out calls separately. Local admission control addresses the same capacity problem for worker threads, though a model provider adds remote latency and quota.
Keep business state outside the response text
Model output is not a transaction record. If a receipt workflow requires a write, call a service that validates the request and commits an idempotent operation; do not parse a sentence as proof that stock changed. Spring AI provides synchronous and streaming ChatClient paths, but this lesson is a boundary design, not a running provider integration in the source kit.
Boundary sketch
if (requestText.length() > 4_096 || inFlightForTenant >= 3) {
throw new RequestNotAdmittedException();
}
// Only after admission: invoke the configured ChatClient with a deadline.Cost and verification
Token and latency costs vary by provider and prompt length. Admission limits spend less on rejected work but can return 429 during bursts; a queue without a cap only moves the overload. This sketch is not executed by the current Spring source kit; verify it against the chosen dependencies and deployment.
Common Mistakes
- Do not let abandoned streams consume provider quota indefinitely.
- Do not infer a durable business write from generated text.
- Do not publish these limits as measured provider behavior; they are application policy examples.
Read next
Spring task executors: reject work when every slot is occupied, Spring AI tools: authorize each requested action after model selection, Spring AI conversation memory: partition by verified tenant and caller, Spring Boot liveness versus readiness: do not restart on every dependency outage.
