CPU3DAI tools for 3D, video and audio

Aleph Alpha releases open-source 78B-parameter MoE model Kolibri-1-BF16

On HuggingFace, on October 2, 2026, the model Aleph-Alpha/Kolibri-1-BF16 was published with the text-generation task. It is a mixture-of-experts model for reasoning with a focus on German and English, with an explicit reasoning mode and tool calling. The context is 1,048,576 tokens; for complex tasks the author recommends up to 262,144 tokens.

What it means

The model weighs 145.5 GB in safetensors, the largest file is 4.7 GB. Total parameters are 78.1 billion, active per token — 3.46 billion. License is Apache 2.0. It runs via vllm. For BF16 weights, the author specifies a minimum of 4x A100 80 GB, 4x H100 SXM5, 2x H200, 1x B200 or 1x B300. On the site there are 5 quantized builds in FP8 and GGUF, the smallest working one is Q4_K_XL at 44.2 GB by InsidiousFiddler. For local running on a single consumer graphics card, the model in BF16 is not suitable: the footprint is about 156 GB. The quantized Q4_K_XL at 44.2 GB is closer to reality, but the author does not specify the memory requirements for it. You can follow the appearance of new builds and details in the CPU3D news feed.
Kolibri-1-BF16: 78.1B-parameter MoE model Input German text or English up to 1,048,576 tokens Kolibri-1-BF16 mixture-of-experts reasoning mode tool calling 78.1B total / 3.46B active Output text-generation reasoned response tool result Requirements to run BF16 weights 145.5 GB safetensors footprint ~156 GB 4x A100 80 GB 4x H100 SXM5 Quantized builds 5 FP8 and GGUF builds Q4_K_XL — 44.2 GB by InsidiousFiddler for a single GPU Hardware 2x H200 1x B200 1x B300 running via vllm Apache 2.0 license · HuggingFace · text-generation task
How the method works. The diagram is drawn based on this news note.

See also