Saoirse Mulligan 8 min readA language model may appear to answer one user at a time. The server beneath it is doing something more intricate: coordinating many requests whose prompts and responses have different lengths, while trying to keep expensive accelerators occupied.
Traditional batching is poorly suited to this workload. If a server groups several requests and waits for every sequence to finish, short answers remain trapped behind long ones. Continuous batching replaces that rigid batch with a changing set of active sequences. Requests enter, advance, finish, and leave while generation continues.
This sounds like a scheduling refinement. In practice, it is one of the mechanisms that determines whether an AI product feels immediate, becomes economical at scale, or collapses under uneven traffic.
The mismatch between static batches and generated text
Conventional neural-network inference often uses static batches. A system collects inputs, combines them into a tensor, runs the model, and returns the results. This works well when inputs require roughly the same amount of computation.
Autoregressive language generation behaves differently. After processing a prompt, the model produces one token, appends it to the sequence, and repeats. Each request may stop after a different number of iterations because it reaches an end token, a length limit, or an application-specific stopping rule.
Imagine a static batch containing three requests:
- Request A: a short prompt followed by a five-token classification answer.
- Request B: a medium prompt followed by a paragraph.
- Request C: a long prompt followed by a detailed report.
A finishes first, but its position in the batch cannot immediately be reused if the runtime treats the batch as fixed. Computation may continue with padding or masked positions until C finishes. The accelerator remains active, yet some of its work no longer advances a user request.
Continuous batching changes the unit of scheduling. Instead of admitting work only between complete batches, the server can reconsider the batch between generation iterations. When A finishes, another waiting request can take its place.
Two phases, two very different workloads
Every standard autoregressive request has a prefill phase and a decode phase. Treating them as equivalent obscures the central scheduling problem.
| Phase | What happens | Typical computational shape | User-facing metric |
|---|---|---|---|
| Prefill | The model processes all prompt tokens and creates attention state | Many tokens processed together; substantial matrix computation | Time to first token |
| Decode | The model generates subsequent tokens using cached state | Usually one new token per active sequence per iteration | Inter-token latency |
During prefill, attention must account for the prompt and construct key-value state for its layers. During decode, the server reuses that state, commonly called the KV cache, so it does not recompute the entire sequence for every new token.
A long prefill can monopolize computation and delay decode iterations for requests already streaming. Conversely, prioritizing decode indefinitely can leave new requests waiting without a first token. A production scheduler must therefore decide not merely which request runs next, but how prefill and decode work share each scheduling window.
How the moving batch operates
A simplified continuous-batching loop has four steps:
- Inspect running requests, waiting requests, and available KV-cache capacity.
- Select token work that fits the iteration’s compute and memory budget.
- Run a model forward pass for the selected sequences.
- Sample tokens, retire completed sequences, update cache allocation, and admit more work.
The active batch can change at each iteration. It is therefore better understood as a scheduling decision than as a persistent tensor.
Suppose A, B, and C are decoding. A emits its final token during the current iteration. The scheduler marks A complete, releases its KV-cache blocks, and examines the queue. If D’s prompt and anticipated state fit the available budget, D can begin prefill without waiting for B and C to finish.
Runtimes still need compatible tensor operations underneath. They may concatenate token positions, maintain request-to-cache mappings, and use specialized attention kernels that understand noncontiguous memory. Continuous batching does not eliminate batching; it makes membership dynamic.
The KV cache is the hidden admission constraint
Compute is only half the problem. Each active sequence accumulates keys and values for attention layers as its context grows. This KV cache can consume substantial accelerator memory, particularly with long prompts, large batches, or long generated outputs.
A scheduler cannot safely admit requests based only on current token counts. It must reserve or progressively allocate enough cache space for sequences to continue. If memory is exhausted mid-generation, the server may need to pause, evict, recompute, or reject work.
Modern serving systems often divide KV-cache memory into blocks or pages rather than reserving one contiguous region for each request. A sequence receives additional blocks as it grows. This reduces waste from over-reservation and avoids requiring physically contiguous memory for an expanding context.
Block-based allocation introduces its own trade-offs. Smaller blocks reduce internal waste but increase bookkeeping and mapping overhead. Larger blocks simplify management but may strand unused capacity at sequence boundaries. Shared prompt prefixes can sometimes reuse cache blocks, although safe reuse depends on exact token identity, model configuration, and cache lifecycle.
The practical consequence is important: an apparently idle accelerator may still be unable to admit a request because the relevant scarce resource is cache memory, not arithmetic capacity.
Chunked prefill prevents long prompts from taking the road
Consider one user submitting a very long document while dozens of other users are receiving streamed answers. Processing the entire new prompt in one uninterrupted prefill can delay every active decode sequence. Their text appears to freeze, even if aggregate throughput remains respectable.
Chunked prefill divides the prompt into smaller token segments. The scheduler can process one segment, return to decoding existing requests, then process another segment later. This converts a long blocking operation into interleavable work.
The choice of chunk size creates a direct trade-off:
- Larger chunks can use accelerator compute efficiently and finish prompt processing sooner, but may increase pauses between streamed tokens.
- Smaller chunks improve responsiveness for active decodes, but add scheduling overhead and may use kernels less efficiently.
- Adaptive chunks can respond to queue pressure, though they make behavior harder to predict and tune.
For example, an interactive assistant may favor decode continuity because visible pauses damage the experience. An offline summarization service may permit larger prefills because total completion time matters more than smooth streaming. The same model and hardware can therefore require different scheduling policies.
Throughput, first-token latency, and fairness pull apart
There is no universally optimal continuous-batching policy because serving objectives conflict. Maximizing the number of tokens produced per unit of time generally favors fuller batches. Waiting briefly to gather work may improve hardware utilization, but it also delays the earliest request.
A decode-first policy protects inter-token latency for active users but can starve queued prefills. A prefill-heavy policy admits newcomers quickly but may make existing streams stutter. Shortest-job-first scheduling can improve average completion time while repeatedly postponing large requests whose cost is difficult to estimate in advance.
Fairness also operates at several levels. Should the scheduler be fair to requests, users, organizations, or service tiers? A single user can submit many small requests and occupy more scheduling slots than another user with one long task. Request-level fairness does not necessarily produce tenant-level fairness.
Useful controls include per-tenant concurrency limits, weighted queues, maximum waiting thresholds, and aging policies that gradually raise the priority of delayed work. Each control changes both efficiency and product behavior. Scheduling is not merely infrastructure; it encodes who gets responsiveness when resources tighten.
Where continuous batching reaches its limits
Continuous batching cannot remove the model’s underlying computation. If demand exceeds sustainable capacity, scheduling can only choose how congestion appears. Queues still grow, latency still rises, and admission control may still be necessary.
It also works best when requests can share the same model execution path. Different models, numerical formats, adapter configurations, attention implementations, or decoding constraints may fragment traffic into separate compatibility groups. Dynamic adapter loading can preserve flexibility, but it adds memory pressure and switching complexity.
Sampling can create divergence too. Some requests use greedy decoding; others require multiple candidates, beam search, or structured constraints. Their token expansion patterns and stopping behavior differ, complicating batch construction.
Failures introduce another boundary. If a worker loses its KV cache, generation state is not automatically recoverable from the emitted text alone with identical performance. The server may need to reconstruct state by replaying the context, move the request using transferred cache data, or restart it. Continuous scheduling does not itself provide durability.
The open design questions
The next advances are likely to come from schedulers that reason across more than one accelerator and more than one moment. They may predict output lengths, account for service-level deadlines, coordinate disaggregated prefill and decode workers, or move KV-cache state across devices.
Prediction creates risk. Output length is unknown until generation stops, and prompt length alone is an unreliable proxy. A scheduler that trusts inaccurate forecasts can reserve too much memory or repeatedly disadvantage requests misclassified as expensive.
Disaggregating prefill and decode could let each phase use hardware and batching policies suited to its shape. Yet transferring KV-cache state introduces bandwidth, placement, and failure-recovery questions. The boundary is attractive precisely because it moves complexity rather than erasing it.
The deeper revelation is that language-model serving is not a sequence of isolated answers. It is a live market for token work, cache memory, and waiting time. Continuous batching is its matching mechanism. The quality of that mechanism becomes visible not in benchmark peaks, but in whether first tokens arrive promptly, streams remain steady, long jobs eventually finish, and scarce capacity is assigned according to the product’s actual promises.
This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.
From our own rounds
Measured on The Curator, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 171
- Questions per round
- 1.7
Rate this article
Discussion
Comments are moderated. Read our editorial policy.