Skip to main content
Watch: I Asked Claude to Build Me a Business

Agent / Code / Inference

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon

arXiv:2610.011084 upvotes

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Authors: Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon

arXiv ID: 2610.01108

Problem: Coding agents repeatedly emit text that already exists somewhere - code from opened files, logs, and the agent's own earlier attempts - which makes them an unusually good fit for retrieval-based speculative decoding, where a cheap drafter copies token continuations from existing text and the target model verifies them in parallel. Existing retrieval drafters fall short inside agent pipelines for two reasons: much of the reusable text is missing from their corpora or stored in a format that differs from what the agent emits, and their draft lengths assume an acceptance profile that actually varies across agents and drifts over turns.

Key Methodology:

  • Retrieval from three corpora: the ongoing session trajectory, the workspace, and a global corpus, with opened files indexed in the agent's own emission format so drafted text matches what the model tends to write.
  • An offline-profiled draft-length cap per agent, adapted online from verification feedback, instead of a fixed draft length.
  • The framework leaves the target model, drafter weights, and decoding rule untouched and works on top of existing retrieval-based speculative decoding engines.

Key Results:

  • On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings.
  • Generation throughput rises up to 4.37x over autoregressive decoding at batch size 1 and up to 4.76x at batch size 16.
  • Gains persist on benchmarks without a repository or a multi-agent pipeline, indicating the effect generalizes across coding-agent shapes rather than only repo-level tasks.

Applied Context: Agent inference is dominated by text the agent itself has already produced or read, and that repetition is a serving asset: retrieval-based speculative decoding turns it into throughput without changing outputs. The second lesson is that acceptance behavior is per-agent and non-stationary, so draft-length policy has to be profiled and adapted rather than set once - a placement decision that fits the compute-placement pattern of spending on the policy around the model, not the model.

Paper: arXiv:2610.01108