Cloudflare Releases clef Model for Decision-Making on Question Schemas
On HuggingFace, on September 30, 2026, the model Cloudflare/clef appeared with the image-text-to-text task. The model accepts state as text, JSON, images, or video and returns probabilities for each valid option of each question in the schema — without free-form text generation and without parsing the output. In the first days, the model gained 5,416 downloads and 1,528 likes.
What it means
The model contains 27.4 billion parameters, the weights in the repository occupy 51.2 GB in safetensors format, the largest file is 4.6 GB. The license is apache-2.0. It runs via transformers. The platform has 17 quantized builds: GGUF, NVFP4, FP8, MLX, EXL, AWQ, GPTQ. The smallest working build is IQ4_XS from bartowski, weighing 14.2 GB.
The author notes that clef is post-trained from Qwen/Qwen3.8-27B with its vision encoder, and on top of the backbone a small transformer head is added that routes features from the state to each question and scores all options jointly. In the same days, a reduced version, clef-flash, was also released.
Running the full model locally will require a substantial amount of VRAM, but quantized builds lower the entry threshold. You can follow news about tools for 3D and video in the CPU3D news feed.How the method works. The diagram was drawn based on this news note.