Skip to main content
Watch: I Asked Claude to Build Me a Business

Agents / Skills / Evaluation

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

Yongqi Tong, Pan Wang, Hang Wang, Jianshe Li, Xin Zhang, Jiang-Ming Yang, Wei Wu · Ant International

arXiv:2609.0557184 upvotes

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

Authors: Yongqi Tong, Pan Wang, Hang Wang, Jianshe Li, Xin Zhang, Jiang-Ming Yang, Wei Wu

arXiv ID: 2609.05571

Problem: Reusable skills give agents transferable procedural knowledge, but scalable acquisition is the bottleneck. Trajectory-based synthesis requires interacting with specific environments first, and document-derived skills lack executable evidence and verification. Both leave agents cold-starting with no reusable knowledge.

Key Methodology:

  • Code2Skill: a fully automated pipeline that selects code units and converts them into implementation-anchored records - atomic operations, composite workflows, and recurring patterns
  • Each record is verified through source-body-blind reconstruction (rebuild the behavior from the record alone) and source-aware comparison, so accepted records carry executable evidence
  • Applied to 19,769 popular, actively maintained GitHub repositories, producing CodeSkillBank: 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata
  • Evaluated via 72 protocol-matched evaluations covering nine model settings and eight benchmarks, plus a unified downstream interface comparing against trajectory-derived skill banks

Key Results:

  • Retrieved CodeSkillBank skills improve models by 11.7% on average over matched baselines, winning in 57 of 72 evaluation settings
  • Under a unified downstream interface, Code2Skill outperforms trajectory-derived skill banks on all seven shared benchmarks
  • Skills synthesized from tested AI-generated code achieve a 93.50% pass rate vs 93.00% for human-written code, so the pipeline scales with the growing volume of AI-generated software

What it means for developers: Source code is the missing acquisition path for skill libraries: it needs no agent experience, yet provides executable evidence for grounding abstractions. Verified code-derived skills beat trajectory-derived banks before agents have accumulated any interaction history, and the verification gate (blind reconstruction) is a reusable pattern for any skill admission pipeline. The near-parity between human- and AI-written source means the skill mine replenishes itself as codegen volume grows.

Paper: arXiv:2609.05571