The Curator

Inside Speculative Decoding: How a Small Model Makes a Large Model Faster Without Choosing the Words

Last updated: 10/4/2026

Back to blog
Camila Reyes avatarCamila Reyes 7 min read
Cover image for Inside Speculative Decoding: How a Small Model Makes a Large Model Faster Without Choosing the Words
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these →

Large language models generate text through an awkwardly sequential process. They predict one token, append it to the context, run another forward pass, and repeat. The arithmetic inside each pass is highly parallel, but the dependency between tokens is not: token seven cannot be sampled until tokens one through six exist.

Speculative decoding changes the rhythm. A fast model proposes a short continuation; a stronger model evaluates that entire proposal in one pass. Accepted tokens advance generation several positions at once. Crucially, the smaller model does not replace the stronger one. It acts as a speculative author whose work is admitted only through a verification procedure designed to preserve the stronger model’s distribution.

The bottleneck is not simply model size

During autoregressive generation, the prompt is usually processed first in a parallelizable prefill phase. The system then enters decode, producing tokens one at a time while reading previously cached attention keys and values.

Decode can leave accelerators underused, particularly for individual requests or small batches. Each step must move model weights and cache data, yet it emits only one token per sequence. Adding hardware does not remove the dependency chain. A conventional decoder still needs another sequential step for every new token.

Speculative decoding attacks the number of sequential target-model steps rather than merely making each step cheaper. If a verifier accepts several proposed tokens, one expensive pass performs the work that would otherwise require several passes. The gain comes from converting serial generation into parallel verification.

The draft-and-verify loop

Let the target model be the system whose behavior the product intends to serve. A smaller draft model receives the same current context and generates a block of candidate tokens autoregressively. Suppose it proposes: “The permit expires after thirty days.”

The target model then processes those proposed positions together. Because the full candidate block is available, a causal attention mask lets the target calculate its next-token distribution at every position in parallel. The verifier compares the draft probabilities with its own probabilities and determines how much of the candidate prefix can be retained.

  1. The draft model proposes a fixed or adaptive number of tokens.
  2. The target model scores all proposed positions in one forward pass.
  3. An acceptance rule scans candidates from left to right.
  4. At the first rejection, the system discards that token and every draft token after it.
  5. A replacement is sampled according to the correction rule, and drafting begins again from the new prefix.

The prefix rule matters. Once a candidate is rejected, later candidates were conditioned on a history that no longer exists. They cannot safely survive independently.

Why verification is more than agreement

A tempting implementation is to keep a draft token whenever it equals the target model’s most likely token. That can accelerate greedy decoding, but ordinary sampling requires a more careful rule. Naive agreement checks can distort which outputs occur and how often.

For an exact speculative sampling method, consider candidate token x, with probability q(x) under the draft and p(x) under the target. The candidate is accepted with probability equal to the smaller of one and p(x)/q(x). A candidate favored at least as strongly by the target as by the draft is therefore accepted. If the draft overproduces it, only the appropriate fraction is retained.

When rejection occurs, the replacement is sampled from a corrected distribution based on the positive difference between target and draft probabilities. Informally, this restores probability mass that the draft failed to represent correctly. The combination of probabilistic acceptance and residual sampling reproduces the target distribution rather than merely approximating its favorite sequence.

Speed comes from letting the draft predict the target’s work; correctness comes from refusing to let the draft redefine that work.

This distinction separates exact speculative decoding from heuristic acceleration. Some production systems deliberately use approximate verification because small output changes are acceptable. That is an engineering choice, not an inherent property of speculation.

A worked example

Imagine the context ends with “Water freezes at”. The draft proposes four tokens: “ zero degrees Celsius .” The target scores all four candidate positions at once.

PositionDraft candidateVerification resultConsequence
1zeroAcceptedAdvance one position
2degreesAcceptedAdvance again
3CelsiusRejectedDiscard this and later candidates
4.Not usableIts conditioning prefix was invalidated

The verifier then samples a corrected token at position three—perhaps “Fahrenheit” in an unusual context, or another valid continuation—and returns to drafting. Two draft tokens were accepted, so the expensive target pass advanced generation by more than one position. Yet the draft’s period could not be retained because it followed a rejected history.

If all four candidates are accepted, implementations can often obtain an additional token from the target model’s distribution at the position after the block. This further improves the amount of progress made per verification pass.

What determines whether it is faster

Acceptance rate is central but insufficient. A useful draft must be both aligned with the target and inexpensive to run. A tiny but inaccurate model proposes quickly and wastes verification capacity. A large, accurate draft earns acceptance but may consume most of the latency it was meant to save.

  • Task predictability: Formulaic prose, code patterns, and repeated structures can be easier to draft than surprising creative continuations.
  • Draft length: Longer blocks offer more potential progress but increase wasted computation after an early rejection.
  • Batching: Highly optimized, heavily batched serving may already use the accelerator effectively, reducing speculation’s advantage.
  • Memory pressure: Hosting another model and its cache can reduce available batch capacity or force less favorable placement.
  • Sampling settings: High-entropy decoding makes draft and target trajectories harder to align.
  • Kernel efficiency: Verification only helps when the serving stack can process the candidate block efficiently.

The correct metric is end-to-end latency or throughput under the intended workload, not tokens accepted in isolation. A strong acceptance ratio can still lose if draft generation, synchronization, cache management, or cross-device communication is expensive.

Variants that remove the separate draft model

A second standalone model is not mandatory. Some systems attach lightweight prediction heads to intermediate layers of the target model. These heads propose future tokens using representations already being computed. Other methods build candidate trees, allowing the verifier to inspect several possible continuations in one structured pass.

Prompt-derived approaches search the existing context for likely continuations. This can work well when text contains repetition, as code and structured documents often do. The draft mechanism becomes a lookup process rather than a neural model.

Each design moves cost differently. Auxiliary heads require training and architecture changes. Candidate trees complicate attention masks and cache handling. Prompt lookup is cheap but limited to continuations recoverable from prior context. The shared principle is unchanged: produce candidate future tokens cheaply, then amortize target-model verification across them.

Where speculation fails

Speculative decoding cannot eliminate autoregression. Rejections reintroduce serial steps, and accepted blocks still form a sequence of verification rounds. It also does not reduce the target model’s memory footprint; the full verifier remains present. When a separate drafter is used, total memory requirements may rise.

Operational complexity is another boundary. Serving systems must coordinate two probability streams, preserve exact sampling semantics where required, manage compatible tokenization, and update caches without duplicating excessive work. Dynamic batches create further tension because different requests accept different numbers of tokens, causing sequences to advance unevenly.

There is also no universally best drafter. Agreement varies by language, domain, prompt style, temperature, and position within a response. A drafter selected on general text may perform poorly on a specialized coding workload. Evaluation must therefore reflect actual traffic rather than a single aggregate benchmark.

The open design question: speculation as a policy

The most promising frontier is not merely a better small model, but a controller deciding how to speculate. It could choose draft length from recent acceptance history, switch drafters by domain, disable speculation under memory pressure, or use shallow target layers when confidence is high and a separate model when it is not.

This turns speculative decoding into an online resource-allocation problem. The system must estimate whether another candidate will save more target work than it costs to produce and verify. That estimate changes with batching, hardware load, output entropy, and the emerging text itself.

The deeper idea extends beyond language models. Whenever an expensive computation can verify several cheap hypotheses in parallel, prediction can become a systems primitive. The small model’s role is not to be right enough to replace authority. It is to be right often enough that authority can work in larger, more efficient steps.

This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.

speculative decodingLLM inferencelatencydraft modelsAI infrastructure

From our own rounds

Measured on The Curator, from real sessions people played on this site — not a third-party dataset.

Rounds played here
163
Questions per round
1.7
Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.