1 article

Running Opus 5 through SlopCodeBench's multi-checkpoint gauntlet reveals that frontier models still degrade codebases over time. 24% strict pass rate, 5x more functions than Opus 4.8, and 93% of code lines trigger slop detectors.

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.
Explore 765 topics
Browse All Topics