Parakeet Redux: 178 MB of speech recognition, one proprietary kernel
On 22 September 2026 Moondream, the team behind the small vision-language models of the same name, published two derivatives of NVIDIA's Parakeet TDT 0.6B v3. Redux forces every encoder weight down to three possible values, minus 1, 0 or plus 1, and lands at 177.8 MB instead of roughly 1.2 GB. Ultra keeps full precision and got further training. Both transcribe 25 languages, and the weights of both are CC-BY-4.0.
The Redux model card reports 6.55 per cent word error rate across seven English test sets against the original's 6.26, and it improves on the 25-language FLEURS evaluation, from 11.62 to 10.56. On paper this is a model that shrinks to about a seventh of its size and barely loses accuracy. Three parts of that reading do not hold in our environment, and one more question needs answering before any benchmark.
What the quantisation actually changed
Ternary means no weight in the encoder may take more than three states. The card puts it at 1.58 bits. The repository holds a model.safetensors of 177.8 MB, a config.json, a ternary.json carrying the masks, a tokenizer.json and a folder of evaluation results. There is no GGUF or ONNX branch in it. Tokenizer, language handling, punctuation, casing and number formatting are the NVIDIA original's.
All accuracy figures below come from the vendor's own measurement series, scored with the Open ASR Leaderboard pipeline as of September 2026, Redux running in Photon on an NVIDIA GPU and the original in NeMo in bf16. Word error rate in per cent, lower is better, with our calculated deltas:
| Evaluation | Original | Redux | relative |
|---|---|---|---|
| Seven English test sets | 6.26 | 6.55 | +4.6 % |
| FLEURS, 25 languages | 11.62 | 10.56 | −9.1 % |
| Business speech | 6.15 | 6.96 | +13.2 % |
| Noise, nine MUSAN conditions | 6.72 | 9.04 | +34.5 % |
| Long-form talks, TED-LIUM | 2.71 | 2.51 | −7.4 % |
The pattern is consistent. Redux improves on multilingual speech and on long recordings, stays close on English, and pays a clear price in noise. For transcription inside a working plant, that last column is the one that matters.
The 113x figure assumes AVX-512
Every speed number on the card belongs to Photon's own kernels, and the x86 variant needs AVX-512 VNNI. The machine behind the headline figure is an AMD EPYC 9575F, eight cores of one chiplet, Zen 5. On 4 October 2026 we looked at the CPU flags on all five of our AI hosts, read-only through /proc/cpuinfo.
Four hosts came back without a single AVX-512 flag: .180.1, .180.2, .180.202 and .180.211. They are EPYC 7F52 parts plus one EPYC 7502, CPU family 23 model 49, which is Zen 2. That generation has no AVX-512 instructions at all. On these machines, including our 64-core box .180.202, 113x is not an available result. It belongs to a CPU generation we never bought. The 38x on a MacBook Air M2 is the other number, and nobody has measured that on our fleet either.
The exception is .180.3. Its EPYC 8024P reports avx512_vnni and avx512_bf16, so the fast path is available there in principle. Its four RTX 5090 cards are fully occupied by the production chat lane, and .180.202 is the fallback for every chat lane plus embeddings. A speech model placed there takes away production's way back.
We did not install Redux on any host. What throughput a Zen 2 machine really delivers remains unmeasured on our side.
Open weights, proprietary kernel
The CC-BY-4.0 on the weights is correct and matches the NVIDIA original. What is open is the question of which software you may run them with. We read the PyPI metadata behind pip install moondream.
kestrel-kernels version 0.7.5 carries the full licence text in its PyPI license field, 5 096 characters, plus the classifier Other/Proprietary License. It declares the package and every artifact it installs proprietary and confidential, licensed rather than sold, and available only under the terms of a separate written agreement. The sentence that decides it for us: If you have not entered into such an Agreement, you have no license to use this software. On top of that come prohibitions on reverse engineering, on working around the .kstlc container format, and on any redistribution. The prohibition clause names automated systems explicitly, AI assistants included.
The same vendor's pricing page advertises Photon with one word: Free. We read both on 4 October 2026. There is no legal contradiction in that, because free describes the price list and the licence text describes permission. Operationally it lands in the same place for us. Whether we may run these kernels inside the company is written in a document that has to be signed first, and neither the model card nor the licence badge on the repository mentions it. That belongs in procurement before the trial, and it is a bigger hurdle than the accuracy table.
German is the worst case in these tables
In FLEURS, German word error rate rises from 4.13 to 5.42 per cent, a 31.2 per cent relative increase. English in the same table goes 4.25 to 4.90, which is 15.3 per cent. The German regression is roughly double the English one. Add MUSAN noise at 0 dB and German reads 14.45 against 19.18, again 32.7 per cent. The card gives its own explanation: the ternary encoder has a thinner acoustic margin, and at low signal levels it substitutes similar-sounding words more often. Omissions and inventions do not become more frequent.
The second look is the unpleasant one. NVIDIA's own model card publishes a model-index value of 5.04 per cent for FLEURS German with the original. Redux at 5.42 sits 7.5 per cent above the number NVIDIA published for German. Anyone transcribing German telephone calls does not end up at 4.13. They end up behind the published state of the original.
Absolute values from two test benches are not comparable
NVIDIA's model-index lists AMI 11.31, GigaSpeech 9.59, Earnings-22 11.42, VoxPopuli 6.14 and SPGISpeech 3.97 for the original. Moondream measures that same original with the leaderboard pipeline and reports 10.86, 8.05, 10.75, 5.88 and 3.63, lower in all five cases. Deltas inside one card are usable. Absolute values across two cards are not. Putting Redux next to a Whisper figure lifted from another context compares two test benches and two normalisers, and it usually flatters whichever model the author likes.
The support matrix does not cover our cards either
Photon's production support grid lists 102 configurations. Every entry in the tested-hardware column is an NVIDIA accelerator: A10, A40, A100, B200, GH200, H100, H200, three Jetson Orin variants, L4, L40S, RTX 30 series, RTX 3090, RTX 4090 and the RTX PRO 6000 Blackwell. The precision column reads BF16 throughout. No CPU target appears in the grid, no Apple silicon target, and the Parakeet row names TDT 0.6B v3, the NVIDIA original, rather than the ternary Redux. Our RTX 5090 is not on the list at all.
None of that disproves the model. It does mean the one number the vendor validates as production-ready is a different model on a different accelerator than the one we own, while the headline CPU result sits outside the grid entirely.
What we do with this
Gateway .180.2 carried eleven model lanes on 4 October 2026, from thinking and nothink through embed and rerank to mocr. A speech lane is not among them, and no ASR container runs on .180.211 alongside the embedding, rerank and document services. Redux would replace nothing here. It would add a capability to an estate whose GPU capacity is already committed.
We keep it on the watch list, with three conditions for a second look. First, a licence that covers self-hosted internal use without being tied to a subscription. Second, an open runtime path for the ternary checkpoint as GGUF or ONNX, so that one closed kernel is not the only program able to read weights we are formally free to redistribute. Third, our own measurement on German audio, because a 31.2 per cent relative increase is not a figure to push into a telephone system without testing it.
The arithmetic, the per-host CPU flags and the whole licence chain are recorded with their commands in our measurement file MESSUNG-2026-10-04-parakeet-redux.md. The open question is whether M87 Labs issues a kernel licence for pure self-hosting at all, or whether the $350 Team plan is the practical price of entry.
Further reading
- Moondream Parakeet Redux model card, with all benchmark tables and measurement conditions: Moondream Parakeet Redux model card, with all benchmark tables and measurement conditions
- Parakeet Ultra model card, the full-precision GPU variant: Parakeet Ultra model card, the full-precision GPU variant
- NVIDIA parakeet-tdt-0.6b-v3, the base model and its model-index with NVIDIA's own published values: NVIDIA parakeet-tdt-0.6b-v3, the base model and its model-index with NVIDIA's own published values
- Moondream's release post of 22 September 2026: Moondream's release post of 22 September 2026
- Photon production support matrix, 102 configurations on NVIDIA accelerators: Photon production support matrix, 102 configurations on NVIDIA accelerators
- Moondream pricing page, listing Photon as free and the Team plan: Moondream pricing page, listing Photon as free and the Team plan
- PyPI entry for kestrel-kernels with its licence text: PyPI entry for kestrel-kernels with its licence text
- Better Stack's overview video on the same question: Better Stack's overview video on the same question
Would Parakeet Redux run on our existing AI infrastructure?+
Not on the fast path on four of five hosts. .180.1, .180.2, .180.202 and .180.211 report no AVX-512 flag in /proc/cpuinfo at all; they are EPYC family 23 model 49, which is Zen 2. The card's 113x real-time figure comes from an EPYC 9575F with AVX-512 VNNI. Host .180.3 does have avx512_vnni, but its four RTX 5090 cards are full with the production chat lane. We did not install Redux, so the number for Zen 2 is still open.
Are we allowed to use Parakeet Redux commercially?+
The weights are CC-BY-4.0, the same licence as the NVIDIA original. The engine kernels are not. The PyPI package kestrel-kernels 0.7.5 ships the full text of a proprietary M87 Labs licence which states that without a separate written agreement you have no licence to use the software. The same vendor's pricing page advertises Photon as free. That question has to be settled before the benchmark.
Is Redux good enough for German?+
For clean studio audio probably yes, for everyday telephone audio not yet. In the vendor's FLEURS run, German word error rate rises from 4.13 to 5.42 per cent, a 31.2 per cent relative increase, against 15.3 per cent for English. With MUSAN noise at 0 dB it is 14.45 against 19.18. NVIDIA's own published German figure for the original is 5.04, so Redux lands above the number the original was published with.
senn-tech