Clean Code Makes AI Agents 34% More Efficient - New Research

TL;DR
A controlled study of 660 Claude Code trials shows clean codebases reduce token usage by 7-8% and file revisitations by 34%, while pass rates stay the same. Traditional maintainability principles still matter in the age of AI coding.
Last updated: July 6, 2026
A new research paper from SonarSource examines whether the structural and stylistic quality of code affects how AI coding agents perform. The answer is nuanced: clean code does not change whether agents succeed at tasks, but it dramatically changes how efficiently they work.
The Study#
The researchers constructed 33 tasks across six repository pairs, testing Claude Code through hidden application-level tests. The key innovation was using "minimal pairs" - repositories identical in architecture but differing in code cleanliness. This isolates code quality as the variable being measured.
Across 660 trials:
- Pass rate: No significant difference between clean and messy codebases
- Token usage: 7-8% reduction on cleaner code
- File revisitations: 34% fewer on cleaner code
The methodology involved using static analyzer rule violations (50-100+ per repository) as the measure of "messiness." To create clean versions, they had agents systematically remove these violations while preserving functionality.
What HN is Saying#
The discussion on Hacker News has over 78 comments with significant debate about the methodology and implications.
On practical experience: Many developers report that code quality has a noticeable impact on agent performance in their own work. One commenter noted: "In my experience, the delta in agent performance is substantial if the codebase is littered with dead code, redundant code, unreachable fallbacks, leaking abstractions and half-baked design patterns."
On methodology concerns: Several commenters questioned the approach of using AI to "clean" messy codebases and then measuring AI performance on those cleaned versions. One skeptic wrote: "I simply am not going to trust any conclusion that requires assuming these AI 'cleaned' repos are in any way representative of actually-good codebases."
The first author responded directly to concerns, clarifying that their notion of "clean" was not asking agents to write better code, but giving them lists of static analyzer rule violations and asking them to remove those specific issues.
On the real implications: The most upvoted practical insight was around linting and deterministic guardrails. Multiple commenters shared that setting up strict linters, pre-commit hooks, and automated code quality checks has been the most effective way to improve agent performance in their workflows.
A recurring theme: if agents work more efficiently on clean code, you can use agents to clean the code first. Prompts like "Refactor the Python code to make it more Pythonic" or "Refactor the Rust codebase to fit code organization standards expected of popular open-source Rust code" appear to both improve code quality and agent performance on subsequent tasks.
On the control group issue: The study explicitly does not check whether agents break unrelated tests already present in the repository. Critics argued this is a significant gap - any conclusions about efficiency are less meaningful if the quality of final output is not controlled for.
Why This Matters for Developers#
The finding that pass rates stay constant but efficiency improves has practical implications for how you structure AI-assisted development workflows.
Cost optimization#
If you are paying per token (API pricing) or have limited context windows (Claude Code quotas), cleaner code directly reduces your costs. A 7-8% token reduction across a full development session adds up.
Iteration speed#
The 34% reduction in file revisitations means agents are finding what they need faster. In agentic coding workflows where each file read is a round trip (queue time, prefill, decode, output, parsing, tool call, tool response), this compounds into meaningful time savings.
Legacy code strategy#
The study suggests a two-phase approach for messy codebases:
- Use agents to systematically clean up violations flagged by static analyzers
- Then use agents for feature work on the improved codebase
This is not unlike how you would prepare a codebase for a new team member - except the "team member" is an AI agent that will measurably benefit from the cleanup.
Limitations to Keep in Mind#
The study has several acknowledged limitations:
- Single model tested: Only Claude Code was evaluated. Other agents may respond differently to code quality.
- Synthetic cleanup: Half the repository pairs were created by AI-based cleanup, not by experienced human developers making architectural decisions.
- No test regression checking: A solution that passes hidden tests but breaks existing tests would still count as passing.
The researchers note that models change frequently, so these results are "an historical report on artifact that will be unavailable soon." The specific numbers may not hold for future model versions.
Practical Takeaways#
-
Set up linting aggressively. Pre-commit hooks that enforce code quality standards help both humans and agents.
-
Consider cleanup sprints before feature work. If your codebase has significant technical debt, investing time in cleanup may pay dividends in faster agent-assisted development afterward. A codebase knowledge graph can also help agents keep a durable map of how files, docs, and decisions connect, on top of whatever cleanup you do.
-
File organization matters. The reduction in file revisitations suggests that clear naming conventions and logical file structures help agents navigate codebases more efficiently. Loose, permissive rules tend to erode over long sessions too - see our piece on constraint decay in AI coding agents for why explicit guardrails hold up better than soft conventions.
-
Do not expect miracles. Pass rates did not improve on cleaner code - just efficiency. If your agent is failing at tasks, code cleanliness is probably not the bottleneck.
The paper reinforces something developers have long intuited: code quality is not just about human readability. Well-organized, well-named, well-structured code is easier for any reader to work with - including AI agents that are increasingly part of the development workflow. For more on how the Hacker News community has converged on similar conclusions across other agentic-coding debates, see what Hacker News gets right about AI coding agents in 2026.
Continue Reading#
Sources#
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on Claude Code
Does Code Cleanliness Affect AI Coding Agents?
A new SonarSource study finds clean code doesn't boost agent pass rates - but it cuts token usage by 8% and file revisitations by 34%. Here's what that means for your codebase.
5 min readClaude Code Sends 33k Tokens Before Your Prompt - OpenCode Sends 7k
New research shows Claude Code's system prompt and tool scaffolding consume 4.7x more tokens than OpenCode before processing user input. The HN thread debates whether that overhead buys better outcomes.
7 min readWrite Code Like a Human Will Maintain It - The AI Era Debate
A new essay argues that letting AI generate sloppy code creates a downward spiral where future AI absorbs those bad patterns. HN's 250+ comment thread is split between believers and pure vibe-coders.
6 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








