Moondream Releases Parakeet Redux — a 1.58-bit Speech Recognition Model with 149M Parameters
On HuggingFace, the moondream/parakeet-redux model for automatic-speech-recognition was published on September 18, 2026. It is a 1.58-bit version of parakeet-tdt-0.6b-v3: the same architecture and tokenizer, but every encoder weight takes the value -1, 0 or +1. The weights take up 178 MB, the model runs 113 times faster than real time on eight x86 CPU cores and keeps WER within 0.3 of the original on English, while on the 25-language FLEURS set and on long audio it beats it.
What it means
The model contains 149 million parameters, the largest weight file is model.safetensors at 0.2 GB, safetensors format. License — cc-by-4.0: you can use and modify it with attribution. It runs through Photon, which reads the packed weights directly: AVX-512 VNNI on x86, NEON on ARM, Metal on Apple GPU. All the speed measurements on the model page were made in Photon.
For comparison: the full version of parakeet-tdt-0.6b-v3 weighs 1.2 GB, Redux — 178 MB. The author notes that on background noise from nine MUSAN conditions Redux is behind the original (9.04 versus 6.72 WER), and on business speech too (6.96 versus 6.15). But on long audio TED-LIUM Redux is ahead: 2.51 versus 2.71 WER.
The model is aimed at CPU inference and is noticeably more compact than the full-precision counterpart, which makes it an option for speech recognition on local hardware without a discrete GPU. You can follow the news about tools for running on your own hardware in the CPU3D news feed.How the method works. The diagram is drawn based on this news note.