Cloudflare releases clef-omni multimodal model with text-to-audio task
The model Cloudflare/clef-omni has been published on HuggingFace. It is a mixture-of-experts model with 35.3 billion parameters that accepts state as text, JSON, images, audio or video and returns probabilities for answer options to typed questions. The model has no free-form text generation — only scoring of options. The weights are published in safetensors format, the largest file is 4.7 GB, the entire repository is 65.9 GB. License — apache-2.0, runs via transformers.
What it means
The text-to-audio task on the model page does not mean speech or sound generation in the usual sense. Clef-Omni does not create audio files: it reads state, including audio, and outputs probabilities for predefined answer options. It is a decision-making tool, not a synthesizer.
Our section of audio generators collects models with a different mechanic. ChatTTS is a text-to-audio speech generator that creates audio clips and requires from 4 GB of VRAM for 30 seconds of sound. XTTS and Fish Speech solve the text-to-speech task: they synthesize speech from text, support voice cloning and run on CUDA. All three output audio, whereas clef-omni returns only logits for answer options.
For those looking for a sound or speech generation tool, clef-omni does not replace the existing models from the reference. It belongs to a different class of systems — multimodal classifiers — and its practical application lies in decision automation, not audio synthesis.How the method works. The diagram is drawn based on this news note.