An AI chatbot app is a web product that enables users to converse with an LLM in real time. After a user submits a message, the app forwards it to the model and progressively renders the reply as tokens arrive, so the response appears to be typed out live. ChatGPT, Claude, and Gemini are common examples.
We are designing backend and streaming infrastructure for an ephemeral AI chatbot. A person signs in, sends a text prompt, and receives model output streamed token-by-token over Server-Sent Events. The server never persists messages: the browser holds the entire conversation, and a refresh clears it. The central challenge is providing a fast, progressive token stream through stateless infrastructure while handling stream failures, unpredictable upstream model behavior, and long-lived connections protected by authentication.
High-level architecture of the stateless AI chatbot streaming backend.
Excluded from this design
500 ms at p95. After the stream begins, token-to-render latency should be imperceptible.99.9% uptime, with graceful fallbacks when the external LLM is slow or unavailable.50K concurrent streams at peak, including bursts when many users send prompts at once.[!TIP] During an interview, an early clarification should be whether messages are persisted or conversations are ephemeral. That decision affects backend complexity, the data model, and how failures are handled.
| Metric | Value | What it affects |
|---|---|---|
| Peak concurrent streams | 50K | Determines SSE fleet size and load balancer configuration |
| Average stream duration | 15-30s | Sets how long connections are held and how JWT expiry is handled |
| Tokens per response (avg) | ~500 | Influences context window limits and cost estimates |
| Token throughput at peak | ~1.5M tokens/min | Determines required upstream AI capacity and rate limits |
| Prompts per user per day | ~20 | Guides per-user prompt throttling and usage tracking |
| Context window limit | ~128K tokens (model-dependent) | Limits conversation length and truncation approach |
The number of simultaneous streams shows that one server cannot keep every SSE connection open, so the system must run as a horizontally scaled set of instances behind a load balancer. Since streams average 15-30 seconds, a connection can outlive a short-lived JWT, but the duration is still easier to handle than persistent WebSocket sessions. Token volume is the most important factor for cost: moving roughly 1.5M tokens per minute through an external API makes per-user budgets and rate limits essential.