Whistle: Speech Recognition Model for CPU Released on HuggingFace
On September 30, 2026, the model Cactus-Compute/whistle was published on HuggingFace with the automatic-speech-recognition task. The author describes it as a speech-to-text model for mobile devices, wearables, robots, smart homes, cars and microcontrollers.
What it means
The model is distributed under the apache-2.0 license. The largest weights file — checkpoints/whistle.safetensors — takes up 0.2 GB, while the author claims the entire model is packed into a single 16.9 MB file. The weights are published in safetensors format. It runs via the cactus-needle engine, which, according to the author, works on CPU without a GPU and without external dependencies.
Whistle performs three tasks on-device: transcription of 16 kHz mono audio up to 30 seconds long in seven languages — Russian is not mentioned among them — English, German, French, Spanish, Italian, Dutch and Polish; output of timestamps for each word with probabilities; extraction of a speech embedding — the encoder output, one row per every 80 ms frame. The model reuses the .cact container, Cactus Quants quantization, SIMD kernels and the KV cache from Needle, so on a device where Needle already runs, Whistle launches in the same binary without a second runtime. The decoder depth can be selected at load time with the --audio-depth flag.
For those looking for speech recognition tools that run on your own hardware without a GPU, Whistle is one of the candidates with an open license and claimed CPU support. You can follow the arrival of similar models in the CPU3D news feed.How the method works. The diagram was drawn based on this news note.