
SWE-NFI: The Benchmark That Catches What Coding Agents Miss
A new 188-task benchmark for non-functional improvements finds coding agents hit 70% on functional correctness but lag humans on refactors and structural changes - the quality gap that becomes tech debt.






