Skip to main content
Watch: I Asked Claude to Build Me a Business

Agents / Evaluation / Benchmark

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia · Microsoft

arXiv:2609.25804111 upvotes

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Authors: Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia

arXiv ID: 2609.25804

Problem: Long-horizon agents make many decisions along a run - which hypothesis to test, which implementation to build on - and those decisions determine the outcome. Existing benchmarks grade end-to-end success only; none measures the quality of the decisions made along the way, which the authors call "taste."

Key Methodology:

  • Decision forks are mined automatically from agent trajectories (parallel attempts at the same task, and detours inside a single trajectory), with no human annotation
  • Each fork asks the evaluated model to choose among directions without seeing what happens after the fork; the better direction is known from the trajectory outcome
  • Frontier models evaluated on Taste-Bench; forks labeled by how late the deciding evidence appears in the trajectory
  • Taste training: distill the judgment of a teacher that saw the outcome into a student model, then measure student decisions on unseen tasks and end-to-end SWE-bench Pro success

Key Results:

  • Best frontier model answers 59.7% of taste questions correctly
  • Forks whose deciding evidence appears later in the trajectory are much harder for every model, and a larger reasoning budget does not improve accuracy
  • Outcome-distilled students make better decisions on unseen tasks and improve end-to-end success on held-out SWE-bench Pro tasks

What it means for developers: End-to-end pass rates and decision quality along the path are separable axes, and the under-measured one is now instrumentable. If you rely on an agent for long-running engineering or research work, fork-level judgment - not just final success - decides outcomes, and more reasoning budget at a fork does not buy better decisions. The trainer's recipe (distill decisions from a teacher that saw the outcome) is directly applicable to improving a student agent's mid-run judgment without changing the runtime.

Paper: arXiv:2609.25804 (code: github.com/wbopan/tastebench)