Preliminary report · last updated 18 July 2026
Real human recordings replayed through the app's exact pipeline, scored word-by-word against human-made transcripts of the same audio:
| Source | Style | Word accuracy |
|---|---|---|
| Tech commentary (fast, casual US) | Rapid conversational | 97.1% |
| Explainer video (natural UK) | Natural narrative | 95.3% |
| Audiobook narration × 3 narrators | Long-form prose | 99–100% |
| Synthetic business documents × 3 voices | Email & contract dictation | 97–100% |
Most "misses" against captions are actually filler words the captions dropped but the engine faithfully heard. About 30 minutes of unique audio and ~7,900 words scored to date.
The deterministic formatting rules are verified by an automated suite of 235 checks that runs before every release — phone numbers, currency, thousands separators, ordinals, clock times, percentages, fractions, decimals, IP addresses, version numbers, web and email addresses, letters, and lists. AI formatting is additionally held to a word-for-word guarantee: it may rearrange, never delete or invent.