Batching eight agents leaves each one two speculative tokens on RDNA4
Discuss on XThe last two posts sold speculative decoding as nearly free. A token tree rides along under the weight stream a decode step was going to pay anyway, and one verification pass on an RX 9070 XT has slack for about two dozen candidate tokens before it costs more than a single token. That number, two dozen, was the whole argument. It is also a batch-of-one number, and a local agent swarm is never a batch of one.
Run eight agents through the same card and the picture changes before speculation even starts. The eight agents already share one decode step, because batching amortizes the weight stream: the weights stream across the bus once and every sequence in the batch reads them. That is exactly why batching works. But it is also why the free speculation budget is already half spent by the time any agent drafts a token. The batch and the draft are drawing on the same account.
The account is compute slack, and it is small. A decode step is bound by streaming weights and paying launch overhead, not by arithmetic, so most of the card’s matrix throughput sits idle during a token. Speculation spends that idle arithmetic. So does batching. They are not additive, and treating them as if they were is how you tune a draft length that looks great on a single sequence and quietly does nothing for the swarm.
The budget is fixed, and the batch is inside it
Start from the one measured floor. A decode step for Qwen3.5-9B on one RX 9070 XT lands about 39.6 tokens per second, roughly 25 milliseconds, almost all of it weight streaming and launch cost. Compute on the same card runs at the 962 tokens per second prefill rate, about 1.04 milliseconds per token of work. Divide the decode floor by the per-token compute and you get the budget: about 24 tokens of arithmetic hide under one weight stream before the compute grows past it. Below 24 the extra work is free; above it each token is priced at the prefill rate. That is the roofline the token-tree post ended on.
A batch of B agents lives inside that same 24-token budget. Each agent contributes one committed token per step, so the batch spends B of the 24 before a single speculative candidate is added. Whatever is left, 24 minus B, is the room for speculation, shared across the whole batch.
Read the bars top to bottom. At one agent, 23 of the 24 tokens are free to draft, which is the number the last post spent on a token tree. At two agents each gets 11. At four, five each. At eight agents the committed tokens take a third of the budget and leave 16 to share, about two speculative tokens per agent. At sixteen the batch fills the budget on its own, and there is nothing left before the pass costs more than a plain decode step. The free ride collapses roughly as one over the batch size.
This is not a ZINC quirk. It is the batch-size dependence that Su and colleagues characterized directly in the synergy of speculative decoding and batching: the optimal speculation length depends on the batch size, and it falls as the batch grows, because larger batches already use the hardware that speculation was exploiting. Their conclusion was to make the speculation length adaptive rather than fixed, and the roofline above is why.
The batch also goes ragged
Shrinking budget is the first cost. The second is that the tokens you do draft do not come back cleanly. Speculation is only useful because the model verifies several candidates in one pass and accepts a prefix of them, but different sequences accept different amounts. One agent copying a familiar line from its context might confirm all six drafted tokens; another mid-reasoning confirms one and rejects the rest. After a single pass the sequences in the batch have advanced by different distances.
That is the ragged tensor problem, and a February 2026 study of batched speculative decoding found it is not a minor inefficiency. Varying accept counts desynchronize position IDs, attention masks, and KV-cache lengths across the batch, and the paper shows that every existing batch speculative implementation it tested violated output equivalence, producing repetition or gibberish when the misalignment was handled wrong. Their corrected algorithm restores equivalence but measures the alignment overhead at up to 40 percent of the pass, growing superlinearly with batch size, before a scheduler that regroups same-length sequences claws some of it back.
For a swarm this is the expensive corner. The whole point of batching agents is that they share the weight stream, but speculation pushes them out of lockstep exactly when you most want them aligned. You can pad every sequence to the longest accepted run, which wastes the fast agents’ slots, or you can pay to realign the tensors. Neither is free, and both scale the wrong way with the number of agents. Speculative decoding must stay lossless, identical in output to plain decoding, so there is no shortcut of just letting the misalignment slide.
What the budget says to build
The design consequence is that the per-agent draft length is not a constant to tune once. It is a function of the live batch.
| Live batch | Committed tokens | Free to speculate (total) | Sensible per-agent draft |
|---|---|---|---|
| 1 agent active | 1 | 23 | a small tree, ~8 to 16 deep |
| 4 agents active | 4 | 20 | ~4 to 5 tokens each |
| 8 agents active | 8 | 16 | ~2 tokens each |
| 16 agents active | 16 | 8 | draft off, or shortest branch only |
The table reads as one rule: keep the total candidate count across the batch under the 24-token budget, and let the batch size decide how that total is split. A local swarm makes this easy to get wrong, because its batch size never sits still. Agents block on a file read, a grep, or a test run and drop out of the batch, then resume and rejoin, so the live count swings between one and a dozen within a single task. A draft length fixed for a full batch wastes almost all of the win when only one agent is awake, and a draft length tuned for a single sequence blows the budget and goes ragged the moment the swarm fills up. Production servers face the same swing and default to conservative speculation for this reason; the vLLM speculative decoding docs note that the gains are largest at low request rates, when the batch is small, and taper as concurrency climbs.
So ZINC’s scheduler already knows the one number that sets the draft length: how many agents are currently in the decode batch. The speculation stage should read that count each step and scale the per-agent draft so the batch stays under budget, wide when the swarm is quiet and off when it is full. The token-tree index from yesterday still supplies the candidates; the batch size just decides how many of them are worth proposing.
The framing I want to keep is that batching and speculation are two ways to spend one pool of idle arithmetic, not two independent wins to stack. On a bandwidth-bound card the pool is fixed at about two dozen tokens per step, and the batch is the first claim on it. Speculation gets what is left. At eight agents that is two tokens each, which is still worth taking, but it is a tenth of what a single agent gets, and pretending otherwise is how a swarm ends up paying ragged-batch overhead for a speedup that was never in the budget.