NVIDIA opens up its 'speculative decoding' playbook: the AI race is now being fought over cost per token

🕒 Published on Zendoric: September 3, 2026 · 10:20
✨ AI-generated · how it's made
NVIDIA has published five engineering rules for accelerating LLM inference with 'speculative decoding', along with its own benchmark, SPEED-Bench, to measure it. Behind the technical jargon lies a deeper thesis: the battle is no longer just about training the biggest model, but about serving it more cheaply.
By Zendoric · September 3, 2026.
NVIDIA has published the third installment of its technical series on AI model "co-design," this time focused on speculative decoding, a technique for accelerating the slowest phase of inference in large language models (LLMs). The team, led by engineer Tiyasa Mitra, proposes five practical rules for choosing how many tokens to "guess" per iteration and with which mechanism, along with its own benchmark, SPEED-Bench, to measure effectiveness on tasks such as coding or summarization. The training code, according to NVIDIA itself, is already available in the NVIDIA/Model-Optimizer repository, with examples on its Nemotron 3.5 Lightning model.
The underlying idea is simple to explain even if the execution is not: a small "draft" model proposes several tokens (the minimal units of text an LLM processes) in advance, and the large "target" model verifies them all in one go, in parallel, instead of generating them one by one. The final result is identical to what would be obtained by generating token by token —only the tokens the target accepts are kept— but in fewer steps and therefore faster. The key technical contribution of this post is how many tokens it is worth proposing at once (the "draft length," D) depending on batch size, the model's attention architecture and the point on the performance curve at which it operates, and how to choose among half a dozen drafting mechanisms —from standalone external models to integrated techniques such as EAGLE-3, MTP, DFlash or DSpark— with training costs ranging from a handful of billions of tokens to more than a trillion.
This kind of optimization matters more than its engineering-manual appearance suggests. As the article itself explains, mixture-of-experts models (MoE, architectures that activate only a fraction of their parameters per query) are increasingly sparse and the contexts they handle increasingly long, which reduces effective concurrency per expert. Any technique that squeezes more performance out of the same GPU without touching the quality of the answers translates directly into a lower cost of serving those models to millions of users. It is no coincidence that NVIDIA is the one publishing this: besides selling the chips, the company is consolidating its software layer (TensorRT-LLM, Model-Optimizer) as the de facto standard for squeezing them, which reinforces its central position across the entire AI value chain against rival hardware alternatives or frameworks.
Our reading is that pieces like this are the least flashy but most decisive barometer of where the industry is heading. While the public debate focuses on which model is "smarter," the real fight over margins is waged over how many cents each generated token costs, and that marginal cost is exactly the mechanism by which AI gets cheaper and becomes accessible to more people and more use cases: every improvement of this kind pushes a little further toward the scenario of computational abundance we defend as a long-term horizon. In the short term, however, it is worth keeping in mind that these efficiency gains are captured first by those who already have the engineering muscle to implement them —the large labs and the hyperscale clouds— which widens, at least temporarily, the distance from those who cannot afford that level of custom optimization.
🔗 Related on Zendoric
- AI's new bottleneck isn't in the model: it's in what we're able to imagine for it · 2026-06-24
- The China–U.S. gap in AI narrows: why the real contest is no longer technical but about the business model · 2026-06-25
- Europe does not need to win the chatbot race, but the autonomy one: Domyn's bet · 2026-06-27


