AI Transcribe

← All guides

What is Word Error Rate (WER)? Accuracy claims, decoded

Last updated:

Every transcription product claims a number — "99% accurate," "95%+ accuracy." Word Error Rate is the metric behind those claims, and knowing how it works makes you immune to most of the marketing.

Short answer: WER measures the percentage of words a transcription engine gets wrong — counting substitutions (wrong word), deletions (missed word) and insertions (invented word) against a human-verified reference. A 5% WER means "95% accurate." The catch: WER is always measured on some particular audio, and vendors quote their best-case audio. Your noisy meeting will score worse than their benchmark — on every engine.

The formula, without the math degree

WER = (substitutions + deletions + insertions) ÷ words actually spoken

Say a 100-word recording comes back with 3 wrong words, 1 missed and 1 invented: WER = 5%, i.e. "95% accurate." Simple — and easy to game by choosing easy audio.

How to read accuracy claims

  • "99% accurate" on clean, close-mic, native-speaker English is achievable by modern engines — including on your recordings if they match those conditions.
  • The same engine on a windy street interview might hit 85% — that's your audio, not a worse engine. What moves WER most: microphone distance, noise, crosstalk.
  • Guaranteed accuracy (like Rev's 99%+ human service at $1.99/min) is a different product: a human fixes the tail. AI numbers are typical-case, not guaranteed.

What WER doesn't capture

  • Punctuation and formatting — a transcript can have low WER and still read badly.
  • Speaker attributiondiarization errors aren't counted in WER.
  • Which words are wrong — a 2% WER that mangles every drug name is worse for a doctor than a 4% WER that misses filler words.

The practical takeaway

Between modern tools, published WER differences are small and benchmark-dependent — choose tools by workflow, and spend your effort on recording conditions, where the real accuracy is won. Two minutes fixing names in a transcript from AI Transcribe beats an hour comparing vendor benchmark pages.

Frequently asked questions

Is there an official WER benchmark for consumer apps? No standard public leaderboard covers consumer iPhone apps on realistic audio; academic benchmarks use datasets that don't resemble your voice memos.

What's a "good" WER in 2026? Low single digits on clean audio is state of the art; 10–20% on genuinely hard audio (noise, accents, crosstalk) is normal for every engine.

AI Transcribe — speech to text on iPhone.

AI Transcribe is developed by Engcraft, LLC, the team behind several AI productivity apps on the App Store. Contact: useaitranscribe@gmail.com