Accuracy

What Fabled Flow actually is, why it gets your words right, and what it will never do to them. This page is the story. The numbers live on the Benchmarks page.

What Fabled Flow actually is

Every dictation product on the market starts from a speech engine. Most stop there: they wrap an engine — theirs or a cloud vendor's — and paste its raw guess at your cursor. That's a wrapper. And raw engine output has never been a finished product, from any vendor, ever.

We're open about our base ears, because there's nothing to hide. English recognition runs on NVIDIA's Parakeet model family — the engine lineage at the top of the industry's public Open ASR Leaderboard — and Chinese, Japanese and Korean run on SenseVoice. Both execute entirely on your machine, on the Neural Engine. Anyone can download those engines. That part is not the product.

Fabled Flow is the voice-to-text engine plus the Fabled Flow Editor. What does an editor do? Improves your text without changing what you meant. That's the whole job: it edits like a professional, without ever changing a word you said. Hundreds of hours of our own time spent developing the Editor, built from real dictations that went wrong on real days, pinned by an automated suite of more than 700 checks. No wrapper ships that layer.

And it exists because of one founding premise: we wanted to build a product with the best accuracy and speed, together with real-world usable text. Not just words that don't go missing and words that don't get invented — rules, so that SKU numbers, year numbers, technical terms, everything comes out clear and accurate, as usable text, the way a normal person would write it. We never buy one with the other.

Our aim is blunt: the most accurate dictation product you can put on a desk. We back that the only honest way we know — the strongest open base engines, machinery no wrapper has, complete-set benchmark numbers published where competitors publish none, and samples below you can try yourself, today, against anything else.

The two promises

Everything on this page serves two rules that outrank every feature we ship:

What the base engine cannot do alone

Base engines are superb at continuous, clean prose. That's what they're trained on and scored on. Real dictation isn't that. Real dictation is pauses, stumbles, codes, addresses, names, numbers, and a paste that has to land in a real app. The Fabled Flow Editor is the shipped machinery that covers the distance:

Every one of these ships with guards against firing where it shouldn't, and each is pinned by the regression suite. A rule that misbehaves once gets a permanent test.

We also incorporated multiple dictionaries and term lexicons — abbreviations, brand names, technical terms, your own personal dictionary — things you would think are included in the base voice-to-text engine. Apparently not. Localized language packs and multilingual dictation too. All of it going toward one thing: making the absolute best super app that exists.

The Editor

The Editor that never leaves your machine. Cloud dictation products get clean transcripts by sending your voice and words to their servers, where an AI model rewrites what you said. Fabled Flow gets the same cleanliness a different way — the Second Listen re-hears your audio on a second local engine, evidence gates catch words that were never spoken, and the Formatting Engine handles the rest with deterministic rules. All of it on your machine. Your voice never goes anywhere.

And one difference we hold sacred: we never rewrite your words. What you said is what appears — cleaned, formatted, accurate, and always yours.

Every speech engine invents

When the audio evidence is weak — a stumble, a breath, a noisy pause — every recognition engine on earth fills the gap with its best statistical guess. Including the engine we use. Including the engines the cloud products use. Most dictation software ships those guesses straight to your screen.

Fabled Flow gives every dictation a second hearing: an independent engine re-listens, and where the evidence is weak, invented words are challenged instead of pasted. Cross-examination for your transcript — entirely on your machine.

Same engine. No waiting.

The speech engine everyone builds on transcribes a second of speech in a fraction of a second. The real question is when it does the work. Most apps wait until you stop talking — then start. Dictate for ten minutes, and you wait while the whole thing processes.

Fabled Flow transcribes while you speak. By the time you release the key, the work is already done — your text lands in about a second whether you spoke for ten seconds or an hour. We call it Rolling Decode: the same engine everyone uses, run the way nobody else runs it. Short dictations are quick everywhere; the difference is long-form — no waiting, no matter how long you talked.

Measured: an hour and eighteen minutes of continuous speech in one dictation, delivered 4.2 seconds after release (2026-08-11). In the same head-to-head series, a 74-second dictation: 867 ms for Fabled Flow against 3.8–5.0 s for a competitor product on the same engine family — a gap that grows with every second you keep talking, because their delivery time scales with length and ours stays flat. Numbers and method on the Benchmarks page.

The honest numbers

The industry's standard accuracy score is Word Error Rate (WER) — lower is better. What that score does and doesn't measure is spelled out on the Benchmarks page, with examples you can rerun yourself. Here are ours, measured on the complete public test sets through the exact pipeline the app ships, beside the published numbers of the open-model leaders:

Product / modelLibriSpeech test-clean (read speech)LibriSpeech test-other (noisy)Benchmark named?
Fabled Flow (full shipping pipeline)3.13% WER · 96.87% accuracy (COMPLETE set — all 2,620 utterances, 52,576 words)4.84% WER · 95.16% accuracy (COMPLETE set — all 2,939 utterances, 52,343 words)Yes — public datasets, full sets, reproducible
2026 leaderboard leaders (published, full test sets; read 2026-08-12)0.92%1.88%Yes — Open ASR Leaderboard
Wispr Flowno public benchmark publishedNo
Willow Voice"98%+" — self-scored on a private synthetic test set; not public, not reproducibleNo
Spokenlyno public benchmark publishedNo

Measured 2026-08-08, build 10.0 (b734), every utterance, zero excluded. Full tables, methodology links and speed results on the Benchmarks page.

Now the honest explanation. LibriSpeech is 19th-century audiobook prose, read aloud, scored after stripping formatting. The things our Editor adds — names done right, numbers as numerals, codes assembled, punctuation — either don't count there or actively count against us. Write "182 teeth" where the 1850s reference spells "one hundred and eighty two teeth" and the scorer charges five errors for a correct rendering. We audited every single scored error in both runs. Roughly 40–48% of our charged "errors" are benchmark noise — US/UK spelling variants, archaic proper nouns no living dictation user will ever say, period spellings like "to day", and our own designed number formatting being penalized for being right.

Real dictation is codes, addresses, brand names and numbers. That test contains none of them. Our torture tests do. We publish the standard score anyway, on the complete set, because honest numbers beside the leaders beat perfect-sounding numbers beside nothing.

Try these yourself

Dictate these into Fabled Flow, then into any product that pastes the engine's raw guess. Count the corrections. Every output below is the actual shipped behavior, pinned by the automated suite:

You sayFabled Flow types
"A forty-seven dash one one zero two forward slash B"A47-1102/B
"forty six dot seventeen dot one seventy five dot one twenty"46.17.175.120
"Login at yomied dot com"login@yomied.com
"Soi thirty-two" (Thailand pack on)Soi 32
"one hundred HKD"100 HKD
"159 plus 86 minus 22"159 + 86 - 22

And dictate any of them twice: you'll get the same thing twice. In our head-to-head testing, the leading cloud competitor rendered identical spoken codes five different ways — and once silently dropped a dictated element. ("Soy sauce", for the record, stays soy sauce — every rule carries guards like that.)

What we're building now

In development — not in the shipping build, and never blurred with what is: the full evidence standard, where every word — not just evidence-flagged spans — must survive two independent listeners before it reaches your screen; personal accent and inflection calibration — the self-learning voice profile; formal per-language dictation standards; and industry and technical term matching. Each of these will appear in the changelog — and its measured results on the Benchmarks page — when it ships, not before.

Every claim on this page corresponds to shipped, tested behavior, and every number to a dated, recorded measurement. When something is preliminary or in development, this page says so explicitly.