CPU3DAI tools for 3D, video and audio

Cloudflare Releases clef Model for Decision-Making on Question Schemas

On HuggingFace, on September 30, 2026, the model Cloudflare/clef appeared with the image-text-to-text task. The model accepts state as text, JSON, images, or video and returns probabilities for each valid option of each question in the schema — without free-form text generation and without parsing the output. In the first days, the model gained 5,416 downloads and 1,528 likes.

What it means

The model contains 27.4 billion parameters, the weights in the repository occupy 51.2 GB in safetensors format, the largest file is 4.6 GB. The license is apache-2.0. It runs via transformers. The platform has 17 quantized builds: GGUF, NVFP4, FP8, MLX, EXL, AWQ, GPTQ. The smallest working build is IQ4_XS from bartowski, weighing 14.2 GB. The author notes that clef is post-trained from Qwen/Qwen3.8-27B with its vision encoder, and on top of the backbone a small transformer head is added that routes features from the state to each question and scores all options jointly. In the same days, a reduced version, clef-flash, was also released. Running the full model locally will require a substantial amount of VRAM, but quantized builds lower the entry threshold. You can follow news about tools for 3D and video in the CPU3D news feed.
Cloudflare clef — question-scheme decision making Input text, JSON, images, video Backbone Qwen3.8-27B 27.4B parameters Transformer head routing features to questions Output probabilities by option Key specs 51.2 GB safetensors weights in repo 17 quantized builds GGUF, NVFP4, FP8, MLX 14.2 GB IQ4_XS smallest build apache-2.0 license Popularity on HuggingFace 5416 downloads 1528 likes image-text-to-text
How the method works. The diagram was drawn based on this news note.

See also