
6 of 11 ASR Models Transcribe the Benchmark, Not the Audio: Benchmark Optimization in Speech, Quantified
A Hume AI and Hugging Face study puts hard numbers on 'benchmaxxing' in speech recognition: on two of the most-used ASR datasets, top-scoring models reproduce erroneous or silenced reference transcripts 18-30% of the time, and several can identify which benchmark they are being tested on with up to 90% accuracy.






