CPU3DAI tools for 3D, video and audio

Cloudflare releases clef-omni multimodal model with text-to-audio task

The model Cloudflare/clef-omni has been published on HuggingFace. It is a mixture-of-experts model with 35.3 billion parameters that accepts state as text, JSON, images, audio or video and returns probabilities for answer options to typed questions. The model has no free-form text generation — only scoring of options. The weights are published in safetensors format, the largest file is 4.7 GB, the entire repository is 65.9 GB. License — apache-2.0, runs via transformers.

What it means

The text-to-audio task on the model page does not mean speech or sound generation in the usual sense. Clef-Omni does not create audio files: it reads state, including audio, and outputs probabilities for predefined answer options. It is a decision-making tool, not a synthesizer. Our section of audio generators collects models with a different mechanic. ChatTTS is a text-to-audio speech generator that creates audio clips and requires from 4 GB of VRAM for 30 seconds of sound. XTTS and Fish Speech solve the text-to-speech task: they synthesize speech from text, support voice cloning and run on CUDA. All three output audio, whereas clef-omni returns only logits for answer options. For those looking for a sound or speech generation tool, clef-omni does not replace the existing models from the reference. It belongs to a different class of systems — multimodal classifiers — and its practical application lies in decision automation, not audio synthesis.
Clef-Omni: multimodal classifier 35.3B parameters, mixture-of-experts Input: state text, JSON, image, audio or video Clef-Omni mixture-of-experts 35.3B parameters Output: logits probabilities for answer options Comparison with audio generators ChatTTS text-to-audio generator creates audio clips from 4 GB for 30 seconds output: audio file XTTS text-to-speech text-to-speech voice cloning output: audio file Fish Speech text-to-speech text-to-speech voice cloning output: audio file Clef-Omni is a classifier, not a synthesizer: logits only, no audio output
How the method works. The diagram is drawn based on this news note.

See also