Skip to main content
Watch: I Asked Claude to Build Me a Business

Reinforcement Learning / Training / Agents

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li · Meta Superintelligence Labs; University of Wisconsin-Madison; Stanford University; NYU

arXiv:2610.0150974 upvotes

Sharpening Tax in Post-Training

Authors: Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li

arXiv ID: 2610.01509

Problem: The "RL only sharpens the base model" hypothesis holds that post-training improves single-shot accuracy at the cost of solution coverage. That trade-off has been observed in math and coding tasks, but it was not known whether it extends to agentic tasks, where multi-turn tool use and interaction may require capabilities acquired during post-training. If it does, then pass@1 leadership is not the same as being the better policy for a harness that samples, retries, or explores.

Key Methodology:

  • Fourteen base/post-trained checkpoint pairs from four model families, evaluated on three agentic benchmarks that require multi-turn tool calling: BFCL v4 multi-turn base split, WebShop, and ACEBench (42 model-benchmark cases).
  • Harness-equipped base models compared against their post-trained counterparts on pass@1 and pass@K (test-time scalability), with uncertainty estimated across rollouts.
  • The Sharpening Tax metric summarizes the post-training effect on test-time scalability in a single number, decomposed into ceiling and saturation components and correlated against other evaluation metrics.
  • Posterior-tempered group sampling (PTGS), a Bayesian sampler that adapts per-prompt temperature to estimated difficulty, applied during RL training in two agentic environments.

Key Results:

  • Pre-trained LLMs with a light inference harness can serve as capable agents: despite far lower pass@1, they often surpass post-trained counterparts in pass@K given a sufficient test-time budget.
  • Post-training pushes tasks toward two extremes, always solved or never solved, improving sampling efficiency and consistency at the cost of solution coverage; the tax is substantial and pervasive, predictable from a few rollouts, and closely tied to other evaluation metrics.
  • PTGS pays a smaller tax than fixed-temperature sampling, solves more tasks under repeated sampling, and also improves single-shot accuracy.

Applied Context: Model choice for an agent fleet becomes a two-metric decision once the harness samples or retries: report pass@1 and pass@K together, and keep the base checkpoint available as a coverage option, which the paper supports with a tax-guided routing analysis. The mechanism also reframes post-training evaluation: a checkpoint that loses at pass@1 can be the better policy for a sampling loop, and a per-prompt temperature schedule is the cheap training-time intervention the paper offers.

Paper: arXiv:2610.01509