CPU3DAI tools for 3D, video and audio

Voxtral Mini 4B Realtime Arabic Released for Streaming Arabic Speech Recognition

Mistral AI published the Voxtral-Mini-4B-Realtime-Arabic model on HuggingFace — a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic. It is fine-tuned from Voxtral-Mini-4B-Realtime-2602 and tailored to live audio with code-switching, when a speaker alternates languages within a single conversation. At a transcription latency of 480 ms, the model shows an average Character Error Rate of 8.82% across seven Arabic benchmarks.

What it means

The model contains 4.4 billion parameters, with weights released in BF16. The largest file is consolidated.safetensors at 8.3 GB, and the entire repository takes up 16.5 GB. To run it, the author suggests vLLM with Voxtral Realtime support and the mistral-common audio dependencies: audio is sent over WebSocket to the /v1/realtime endpoint. The developer did not specify a license on the page. Running it locally will require a GPU capable of holding the 16.5 GB of BF16 weights, plus headroom for activations and streaming output. There are no quantized builds in the repository — only full-precision safetensors. This will primarily suit those building live subtitles or voice interfaces in Arabic who already work with vLLM. More news about tools that run on your own hardware is in the CPU3D news feed.
Voxtral Mini 4B Realtime Arabic streaming Arabic speech recognition Live audio Arabic dialects code-switching WebSocket endpoint /v1/realtime streaming V Voxtral Mini 4B 4.4B parameters BF16 weights fine-tuned from Realtime-2602 T text Key specs Latency 480 ms CER 8,82% File size 8.3 GB Repository 16.5 GB Use: live captions and Arabic voice interfaces Run: vLLM + GPU with headroom for activations
How the method works. The diagram was drawn based on this news note.

See also