Voxtral Mini 4B Realtime Arabic Released for Streaming Arabic Speech Recognition
Mistral AI published the Voxtral-Mini-4B-Realtime-Arabic model on HuggingFace — a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic. It is fine-tuned from Voxtral-Mini-4B-Realtime-2602 and tailored to live audio with code-switching, when a speaker alternates languages within a single conversation. At a transcription latency of 480 ms, the model shows an average Character Error Rate of 8.82% across seven Arabic benchmarks.
What it means
The model contains 4.4 billion parameters, with weights released in BF16. The largest file is consolidated.safetensors at 8.3 GB, and the entire repository takes up 16.5 GB. To run it, the author suggests vLLM with Voxtral Realtime support and the mistral-common audio dependencies: audio is sent over WebSocket to the /v1/realtime endpoint. The developer did not specify a license on the page.
Running it locally will require a GPU capable of holding the 16.5 GB of BF16 weights, plus headroom for activations and streaming output. There are no quantized builds in the repository — only full-precision safetensors. This will primarily suit those building live subtitles or voice interfaces in Arabic who already work with vLLM. More news about tools that run on your own hardware is in the CPU3D news feed.How the method works. The diagram was drawn based on this news note.