Agents / Skills / Evaluation
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
Yongqi Tong, Pan Wang, Hang Wang, Jianshe Li, Xin Zhang, Jiang-Ming Yang, Wei Wu · Ant International
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
Authors: Yongqi Tong, Pan Wang, Hang Wang, Jianshe Li, Xin Zhang, Jiang-Ming Yang, Wei Wu
arXiv ID: 2609.05571
Problem: Reusable skills give agents transferable procedural knowledge, but scalable acquisition is the bottleneck. Trajectory-based synthesis requires interacting with specific environments first, and document-derived skills lack executable evidence and verification. Both leave agents cold-starting with no reusable knowledge.
Key Methodology:
- Code2Skill: a fully automated pipeline that selects code units and converts them into implementation-anchored records - atomic operations, composite workflows, and recurring patterns
- Each record is verified through source-body-blind reconstruction (rebuild the behavior from the record alone) and source-aware comparison, so accepted records carry executable evidence
- Applied to 19,769 popular, actively maintained GitHub repositories, producing CodeSkillBank: 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata
- Evaluated via 72 protocol-matched evaluations covering nine model settings and eight benchmarks, plus a unified downstream interface comparing against trajectory-derived skill banks
Key Results:
- Retrieved CodeSkillBank skills improve models by 11.7% on average over matched baselines, winning in 57 of 72 evaluation settings
- Under a unified downstream interface, Code2Skill outperforms trajectory-derived skill banks on all seven shared benchmarks
- Skills synthesized from tested AI-generated code achieve a 93.50% pass rate vs 93.00% for human-written code, so the pipeline scales with the growing volume of AI-generated software
What it means for developers: Source code is the missing acquisition path for skill libraries: it needs no agent experience, yet provides executable evidence for grounding abstractions. Verified code-derived skills beat trajectory-derived banks before agents have accumulated any interaction history, and the verification gate (blind reconstruction) is a reusable pattern for any skill admission pipeline. The near-parity between human- and AI-written source means the skill mine replenishes itself as codegen volume grows.
Paper: arXiv:2609.05571