A few years ago, good speech recognition meant one of two things: a big model, or an API call. I wanted to know how much of that is still true, so I ran a controlled benchmark of ten speech-to-text systems on one ordinary machine and measured everything I could: accuracy, speed, latency, memory, cost and where the audio ends up.

The machine is a mid-range Windows laptop with a 2019-era 4 GB GPU. The lineup runs from a 61M-parameter edge model and a 178 MB ternary model up to Whisper large-v3 and three commercial cloud APIs. Everything below is English only, measured on 2026-10-05.

What I tested

Seven open-weights configurations ran locally. Three proprietary models were called through their APIs.

  • Parakeet Redux (Moondream): NVIDIA's Parakeet TDT 0.6B v3 compressed to ternary weights, where every weight is −1, 0 or +1. 178 MB on disk.

  • Parakeet TDT 0.6B v3, the model Redux was made from, at 16-bit (1,256 MB) and at standard 8-bit (740 MB).

  • Parakeet Unified EN 0.6B, an English-only model that can also stream, at 8-bit (731 MB).

  • Whisper large-v3 (1.55B parameters, 1,669 MB at 8-bit) and Whisper large-v3-turbo (0.8B, 886 MB).

  • Moonshine base, a 61M-parameter model built for edge devices (132 MB).

  • Cloud APIs: OpenAI gpt-transcribe, ElevenLabs scribe_v2, and Google's gemini-3.5-transcribe through the Vercel AI Gateway.

The three Parakeet TDT rows are the important ones for compression: the same architecture at 16-bit, 8-bit and ternary, so any difference between them comes from the weights alone. Every open model except Redux ran in the same engine, transcribe-cpp, using the GGUF conversions published by handy-computer. Redux only exists for Moondream's Photon runtime.

The hardware: an AMD Ryzen 5 5600H (6 cores, 12 threads, AVX2 but no AVX-512 or AVX-VNNI), 15.3 GB of RAM and an NVIDIA GTX 1650 with 4 GB, used through Vulkan, on Windows 11.

How I tested it

I wrote the method down before collecting a single result, and logged every decision and deviation after that with a timestamp. About 36 minutes of English audio, in five sets, each with a human-written reference:

  • Real-world (6 clips, 20.7 min): YouTube videos with human-made captions only, no auto-captions. A TED talk on CRISPR, a podcast interview, an MIT lecture, PBS NewsHour, a TEDx talk by an Indian-English speaker, and street interviews.

  • Benchmark sets (50 clips): 10 utterances each from LibriSpeech test-clean and test-other, Earnings-22, AMI meetings and VoxPopuli, 4 to 25 seconds long, drawn with a fixed seed from the Open ASR Leaderboard test sets.

  • Accents (12 clips): 12 speakers with 12 different first languages reading the same passage, including Hindi, Bengali and Tamil, plus a US English control. From the Speech Accent Archive (Weinberger, George Mason University, CC BY-NC-SA 2.0).

  • Noise (10 clips): the 10 LibriSpeech test-clean clips again, with pink noise mixed in at 5 dB SNR.

  • Dictation (4 paragraphs, 225 words): me reading an email, a technical note, a paragraph full of numbers and places, and a casual message into the laptop microphone.

Fairness controls mattered more than I expected:

  • Every system received byte-identical audio. Clips over 30 seconds were cut once, at pauses, into 7 to 28 second chunks that every system used.

  • Every CPU run was pinned to 6 threads. Each configuration ran in a fresh process, in randomized order, with a 60-second cool-down in between, on AC power with background apps closed.

  • Speed is the median of 3 timed passes after a warm-up pass.

Accuracy is word error rate (WER): substituted, deleted and inserted words divided by the words in the reference. Lower is better; 4% means about one word in 25 is wrong. Both texts pass through the Whisper English normalizer first, the convention the Open ASR Leaderboard uses, so formatting like "3:30" vs "three thirty" isn't counted as an error. Each number comes with a 95% confidence interval from 10,000 bootstrap resamples over clips, and the compression comparisons use a paired bootstrap on identical clips.

Accuracy: five systems tied at the top

Bar chart of word error rate with 95% confidence intervals: ElevenLabs Scribe v2 4.03%, OpenAI gpt-transcribe 4.09%, Parakeet Unified EN 0.6B 4.16%, Whisper large-v3-turbo 4.34%, Gemini 3.5 Transcribe 4.38%, Parakeet TDT 8-bit 4.52%, Parakeet TDT 16-bit 4.54%, Whisper large-v3 5.04%, Parakeet Redux 5.33%, Moonshine base 7.02%
78 public clips, 5,531 reference words. The top five overlap almost completely.

On the 78 public clips (5,531 words), ElevenLabs scored 4.03%, OpenAI 4.09%, Parakeet Unified 4.16%, Whisper turbo 4.34% and Gemini 4.38%. That's five systems within 0.35 points of each other, with confidence intervals that overlap almost entirely. On this test they're statistically tied, and two of the five are open models under 1 GB running on a laptop.

Below them: the Parakeet TDT original at 4.52% (8-bit) and 4.54% (16-bit), Whisper large-v3 at 5.04%, Parakeet Redux at 5.33% and Moonshine at 7.02%.

Two things stood out. First, the biggest model here, Whisper large-v3 at 1.55B parameters, was not more accurate than its own 0.8B turbo variant. Second, the confidence intervals are wide: about 1 to 2.4 points either side of each number. With 78 clips, a gap under about one point overall isn't something I'd bet on.

Punctuation is a different story. When I kept punctuation as separate tokens (scored only on the sets with punctuated references), OpenAI produced the cleanest text at 12.2%, followed by Whisper turbo at 12.3%, while ElevenLabs and Moonshine were around 16%. Those numbers are high for everyone because human captions punctuate inconsistently, so they're only useful for comparing systems with each other.

Size no longer predicts accuracy

Scatter plot of word error rate against weights on disk on a log scale. Parakeet Unified (731 MB, 4.16%) and Whisper turbo (886 MB, 4.34%) sit inside the band of the three cloud APIs (4.03 to 4.38%). Whisper large-v3 (1,669 MB) scores 5.04%, Parakeet Redux (178 MB) 5.33% and Moonshine (132 MB) 7.02%
Weights on disk against WER. The grey band is the range of the three cloud APIs, whose sizes aren't published.

Plotted against size, the relationship is weak. The two best open models are 731 and 886 MB. The largest file, Whisper large-v3 at 1.67 GB, scores worse than three models half its size. At the small end, Redux at 178 MB lands between Whisper large-v3 and Moonshine, which is barely smaller at 132 MB but makes about a third more errors.

For English transcription, architecture and training data clearly matter more than parameter count past about 0.6B.

No system wins everywhere

Heat table of word error rate for each system on each audio source: six YouTube clips, five benchmark sets, accents, noise and dictation. The lowest error in each column is outlined, and the best system changes from column to column
WER per source. A green outline marks the lowest error in each column.

The overall tie hides a lot of movement. Broken down by source, the best system changes almost every column. Each source has only 147 to 715 words, so a single column is a hint rather than a ranking, but some patterns are clear enough to mention:

  • Meetings were the hardest benchmark. On AMI, ElevenLabs had the lowest error at 6.8% while Gemini reached 14.3% and OpenAI 12.9%. Cross-talk and casual speech still separate systems that tie elsewhere.

  • The lecture favoured Whisper turbo (4.0%), while both Parakeet TDT variants were around 10%. On earnings calls it was the other way round: Parakeet TDT and Redux at 5.5%, Parakeet Unified at 9.8%.

  • Whisper large-v3 lost points on the podcast (9.4%) for a reason that isn't really a mistake. It writes clean text and drops repeated fillers like "you know" and "so", while the podcast's captions are verbatim. I checked chunk by chunk: nothing was skipped. Part of Whisper's real-world WER is style, not errors.

  • Clean read speech is essentially solved. On LibriSpeech test-clean, OpenAI and Gemini made zero errors, and every system except Moonshine stayed at or under 1.1%.

  • The street interviews were hard for everyone, from 6.8% (Whisper turbo) to 12.4% (Moonshine).

Accents, noise and dictation

Accents. On the 12 speakers reading the same passage, most strong systems stayed between 0.85% (Whisper large-v3) and 2.42% (OpenAI). Parakeet Redux, at 3.62%, was the clear outlier among them: 1.45 points worse than its own original (2.17%). With 12 speakers this is directional, but the dictation results point the same way.

Noise. With pink noise at 5 dB SNR on clean speech, every system except Moonshine stayed under 2.2%. The Parakeet TDT models made exactly the same number of errors with and without the noise. Stationary noise at this level no longer separates good systems. Real-world noise does, as the street clip shows.

Dictation. I didn't read my script perfectly, so I scored each system twice: against the script, and against a verbatim reference of what I actually said. A change counted as spoken only where at least 5 of the 6 strong local models agreed on the same everyday word; names kept their script spelling, since recognising them is part of the test. Against what was said:

  • Whisper large-v3: 1.3%

  • ElevenLabs: 2.2%

  • Parakeet Unified, Gemini and Parakeet TDT 8-bit: 2.7%

  • OpenAI, Parakeet TDT 16-bit and Whisper turbo: 3.1%

  • Parakeet Redux: 4.9%

  • Moonshine: 5.8%

At 225 words, one error is 0.44 points, so anything under about 1.5 points apart is a tie. Redux's extra mistakes were phonetic: "flight lines" for "lands", "revenue grow" for "grew". And every one of the ten systems wrote the name "Ananya" as "Aranya". Rare names are still a weakness across the board, open or cloud.

What squeezing a model into 178 MB costs

Parakeet Redux case study. Weights: 1,256 MB at 16-bit, 740 MB at 8-bit, 178 MB ternary. Extra word error rate against the 16-bit original: +0.20 on clean benchmark clips, +0.78 on real-world YouTube, +1.09 with 5 dB noise, +1.45 on accents, +1.78 on dictation. CPU speed: Redux 15.1 times real time, the 8-bit original 10.7 in the same container and 9.3 on Windows, the 16-bit original 7.6
Parakeet Redux against the model it was made from, on identical audio.

This was the question that started the whole experiment. Parakeet Redux is Parakeet TDT 0.6B v3 with ternary weights, and its model card claims accuracy close to the original at 178 MB. Here is what I measured:

  • Size: 178 MB against 740 MB (8-bit) and 1,256 MB (16-bit). Seven times smaller than the original.

  • Accuracy: 5.33% against 4.54%, so +0.80 points overall, or about 17% more errors. In a paired bootstrap over the same 78 clips, Redux was worse in 96% of 10,000 resamples, with a 95% interval of roughly −0.1 to +1.5 points.

  • Where it loses: barely on clean benchmark audio (+0.2 points), more on accented speech (+1.45) and dictation (+1.8 against the 16-bit original, +2.2 against 8-bit).

  • Speed: 15.1× real time on 6 CPU threads, against 10.7× for the 8-bit original in the same Linux container and 7.6× for the 16-bit original on Windows.

  • Memory: about 980 MB of RAM at runtime, against 1,232 MB for the 8-bit original. Small weights did not mean small memory: Photon's PyTorch-based runtime takes up most of the footprint.

Before running anything, I had defined "no meaningful loss" as an upper bound below +0.5 points. The bound reached +1.5, so that claim isn't supported. The loss also isn't significant at the strict 95% level on 78 clips. The honest reading is that the cost is likely real but modest.

That's consistent with the model card, which reports a 6.55% average WER across seven Open ASR Leaderboard sets and has no accent breakdown. On clean, leaderboard-style audio the gap here was only 0.2 points. The cost shows up on harder inputs, which averages hide.

The two compression steps behave very differently. Going from 16-bit to 8-bit removed 41% of the size and changed WER by −0.02 points, which is nothing: the two versions are statistically identical. Going from 8-bit to ternary removed another 76% and bought roughly 1.4 to 2× CPU speed, at a measurable cost.

Speed: an hour of audio in a minute and a half

Real-time factor on GPU and CPU on a log scale. Moonshine 65.7 times real time on GPU and 20.3 on CPU; Parakeet TDT 16-bit 38.7 and 7.6; Parakeet Unified 38.5 and 9.3; Parakeet TDT 8-bit 36.7 and 9.3; Parakeet Redux CPU only, 15.1; Whisper turbo 6.2 and 0.8; Whisper large-v3 4.8 and 0.6. One hour of audio takes 55 seconds to 1.6 minutes on the GPU for Moonshine and Parakeet, and 93 minutes on the CPU for Whisper large-v3
Real-time factor: seconds of audio per second of compute. Median of 3 timed passes, spread under 1%.

On the GTX 1650, the Parakeet models ran at 37 to 39× real time, which is about 1.6 minutes for an hour of audio. Moonshine was faster still at 66× (55 seconds an hour), but with clearly worse accuracy. On the CPU alone, Parakeet ran at 7.6 to 9.3× and Redux at 15.1×, so four minutes for an hour of audio without any GPU.

Whisper is the outlier. Large-v3 managed 4.8× on the GPU (12.5 minutes an hour) and 0.64× on the CPU, which is slower than real time: 93 minutes for an hour of audio. Turbo was faster but still below real time on the CPU at 0.81×.

Accuracy didn't depend on the device. I reran 10 clips on the CPU and every local model produced word-for-word the same text it had produced on the GPU.

Latency: how long one sentence takes

Latency per clip, median and 95th percentile, with cold start and memory. Moonshine GPU 0.10 seconds, Parakeet Unified GPU 0.20, Parakeet TDT GPU 0.21, Moonshine CPU 0.30, Parakeet Redux CPU 0.48, Parakeet Unified CPU 0.81, Parakeet TDT CPU 1.00, ElevenLabs 1.03, OpenAI 1.05, Whisper turbo GPU 1.45, Whisper large-v3 GPU 1.81, Gemini 3.03, Whisper large-v3 CPU 13.75
One 4 to 25 second clip at a time with the model loaded. Cloud times include the upload from my connection.

For dictation, what matters is the time per sentence. With the model already loaded, Parakeet on the GPU took a median 0.20 seconds per clip. On the CPU it was 0.81 seconds for Parakeet Unified and 0.48 for Redux. The cloud APIs took 1.03 (ElevenLabs), 1.05 (OpenAI) and 3.03 seconds (Gemini), upload included, from my connection.

A few observations from this part:

  • Whisper's latency doesn't shrink with the clip. It always processes a 30-second window, so a 5-second sentence costs as much as a 30-second one. On this CPU that was 11 to 14 seconds per sentence, which rules it out for dictation there regardless of accuracy.

  • Memory is modest on the GPU. Parakeet Unified used about 772 MB of GPU memory and Moonshine 181 MB. Whisper large-v3 needed almost 2 GB, half of this card.

  • Cold starts vary. Parakeet loaded in 2 to 4 seconds; Redux took 8 to 9 seconds. The very first GPU run after installing added 4 to 8 seconds once, for shader compilation.

  • More threads were not faster. Redux ran at 14.6× on 4 threads (its default), 15.1× on 6 and 13.3× on 12. Hyper-threads hurt.

Cloud throughput isn't comparable here: I sent requests one at a time, and an API can be called in parallel. The fair cloud comparison is latency per request.

Cost and privacy

Price per audio hour: local open models $0 plus own hardware and electricity, ElevenLabs Scribe v2 $0.22, OpenAI gpt-transcribe $0.27, Gemini 3.5 Transcribe about $0.37. For 1,000 hours of audio, $220 to $370 in the cloud, or about 26 GPU-hours with Parakeet on a GTX 1650
List prices under default account terms, as of 2026-10-05.
  • OpenAI gpt-transcribe: $0.27 per audio hour ($0.0045 a minute). Per OpenAI's data controls, audio sent to the transcription endpoint isn't stored for abuse monitoring or used for training by default; zero retention is available to approved customers.

  • ElevenLabs Scribe v2: $0.22 per hour. Audio is retained under the default privacy terms with history on; zero-retention mode is for enterprise plans.

  • Gemini 3.5 Transcribe via Vercel: about $0.37 per hour. It's billed per token ($2 per million input, $12 per million output), so this is an estimate assuming 32 audio tokens a second and about 12,000 output tokens per hour of speech. The gateway keeps nothing; what happens upstream depends on routing, and zero retention is opt-in on paid plans.

  • Local models: no per-hour price, and the audio never leaves the device. That isn't the same as free: I didn't measure the hardware or electricity cost.

At these prices, 1,000 hours of audio costs $220 to $370 in the cloud. Locally, the same volume is about 26 GPU-hours on a GTX 1650 with Parakeet, or about 66 CPU-hours with Redux. The whole experiment cost me under $0.50 in API calls.

Things that went wrong along the way

  • Redux wouldn't run on Windows. Photon's Windows build reported no AVX2 support on a CPU that has it, and forcing the AVX2 path failed with "no 'avx2' matrix-multiply path on this machine". In a Linux container on the same CPU it detected AVX2 and ran fine, so this is a packaging gap, not a hardware limit. Its GPU path needs an NVIDIA Ampere card or newer, so the GTX 1650 was out too.

  • The container was faster than Windows. To check the container wasn't flattering Redux, I ran the 8-bit Parakeet in the same container. It ran 16% faster there (10.7× against 9.3× natively), so Redux was not given an unfair advantage; the fair comparison is 15.1× against 10.7×.

  • The YouTube audio was clipping. Some clips had samples above full scale (up to 1.86), from YouTube's lossy encoding. One runtime rejected them, and my cloud uploader clipped them when writing 16-bit WAV, which meant the systems weren't hearing identical audio. I normalized the peaks and reran all ten systems on that set.

What the results mean

Architecture and training now matter more than parameter count. The most accurate open models here have 0.6 to 0.8B parameters. Scaling past that, at least for English transcription, didn't buy accuracy on its own.

For English, the accuracy gap between open models and the cloud has closed. On public data and on live dictation, the best open models were statistically tied with three commercial APIs. They tie; they don't beat them. The cloud still differs in output style (OpenAI's punctuation was the cleanest) and in features I didn't test: many languages, speaker diarization, word timestamps and custom vocabulary.

Compression has two regimes. 8-bit quantization is effectively free and a sensible default for local use. Ternary weights are the new frontier: a 0.6B model in 178 MB, close to Moonshine's footprint, with far better accuracy than Moonshine (5.33% against 7.02%). But the cost concentrates on harder inputs like accents, so anyone shipping a ternary model should test it on their own speakers rather than on leaderboard averages.

Software is now the bottleneck. On this machine, the practical ranking was decided as much by runtimes as by weights. Redux had no Windows kernel for this CPU and no support for this GPU. Whisper's fixed 30-second window makes short sentences slow on any hardware. Redux's runtime erased much of its memory advantage. Better kernels and packaging are the next unlock for small models, more than better models.

Modest hardware is enough. A 4 GB GPU from 2019 ran cloud-level English recognition at 38× real time with 0.2 seconds per sentence. For English speech, cost and privacy are no longer good reasons on their own to call an API.

Where each kind of model fits

  • Real-time dictation on a PC with a GPU: a 0.6B Parakeet-class model at 8-bit. Cloud-level accuracy, about 0.2 seconds per sentence, offline, with roughly 0.8 GB of GPU memory.

  • Dictation on a CPU-only laptop: 8-bit Parakeet, at 0.5 to 0.8 seconds per sentence. Redux where its runtime is supported (Linux, macOS or a supported CPU), keeping its weaker accent results in mind.

  • Edge and low-memory devices: Moonshine (132 MB) when footprint matters most, at around 7% WER. Redux (178 MB of weights) when accuracy matters more, but it currently needs about 1 GB of RAM.

  • Bulk transcription of archives: Parakeet on a GPU, about 1.6 minutes per hour of audio with no per-hour cost.

  • Sensitive audio (health, legal, internal meetings): any local model; the audio never leaves the machine. Cloud zero-retention options are enterprise-only or approval-gated.

  • Many languages, diarization, a managed service: the cloud APIs, or multilingual open models like Whisper and Parakeet TDT v3. I tested none of that here.

  • Heavily accented speakers: the systems with the best accent and dictation results here (Whisper, ElevenLabs, Parakeet Unified). Avoid heavily compressed models without testing on your own speakers.

Limitations

  • Sample size. 78 public clips (about 36 minutes) and 225 dictated words. Differences under about 1 point overall, or 1.5 points on dictation, aren't reliable.

  • English only, even though several of these systems are multilingual.

  • One machine. CPUs with AVX-VNNI or AVX-512, Apple Silicon and Ampere-class GPUs would change the speed picture, especially for Redux, whose fast paths weren't available here.

  • One runtime per model. NeMo, faster-whisper or CUDA backends could change speed and, slightly, output.

  • Reference style. Captions are partly verbatim and partly edited, which penalises Whisper's clean output.

  • Cloud models change without notice. These results are for the three named models as of 2026-10-05, under default account terms, from one network.

What I'd test next

  • The speed runs on an AVX-VNNI or AVX-512 CPU, Apple Silicon, an Ampere-or-newer GPU and an ARM board, where Redux's fast kernels apply.

  • Streaming latency: Parakeet Unified can stream in chunks down to 160 ms, so time to first word is the real dictation number.

  • Multilingual accuracy on FLEURS or Common Voice.

  • Accents at scale, 100+ speakers, to see whether the ternary penalty holds and where it concentrates.

  • Energy per audio hour, local against cloud.

  • 4-bit GGUFs, which sit between 8-bit and ternary in size.

  • Custom vocabulary and prompting for proper nouns, the one weakness every system shared.

  • Redux again on Windows once Photon ships an AVX2 kernel there.

Thanks to NVIDIA for Parakeet, Moondream for Parakeet Redux, OpenAI for Whisper, Useful Sensors for Moonshine, the Open ASR Leaderboard maintainers for the test sets, and Steven Weinberger and George Mason University for the Speech Accent Archive.