Agent / Code / Inference
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Authors: Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon
arXiv ID: 2610.01108
Problem: Coding agents repeatedly emit text that already exists somewhere - code from opened files, logs, and the agent's own earlier attempts - which makes them an unusually good fit for retrieval-based speculative decoding, where a cheap drafter copies token continuations from existing text and the target model verifies them in parallel. Existing retrieval drafters fall short inside agent pipelines for two reasons: much of the reusable text is missing from their corpora or stored in a format that differs from what the agent emits, and their draft lengths assume an acceptance profile that actually varies across agents and drifts over turns.
Key Methodology:
- Retrieval from three corpora: the ongoing session trajectory, the workspace, and a global corpus, with opened files indexed in the agent's own emission format so drafted text matches what the model tends to write.
- An offline-profiled draft-length cap per agent, adapted online from verification feedback, instead of a fixed draft length.
- The framework leaves the target model, drafter weights, and decoding rule untouched and works on top of existing retrieval-based speculative decoding engines.
Key Results:
- On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings.
- Generation throughput rises up to 4.37x over autoregressive decoding at batch size 1 and up to 4.76x at batch size 16.
- Gains persist on benchmarks without a repository or a multi-agent pipeline, indicating the effect generalizes across coding-agent shapes rather than only repo-level tasks.
Applied Context: Agent inference is dominated by text the agent itself has already produced or read, and that repetition is a serving asset: retrieval-based speculative decoding turns it into throughput without changing outputs. The second lesson is that acceptance behavior is per-agent and non-stationary, so draft-length policy has to be profiled and adapted rather than set once - a placement decision that fits the compute-placement pattern of spending on the policy around the model, not the model.
Paper: arXiv:2610.01108