CPU3DAI tools for 3D, video and audio

SGLang v0.5.21: New Models, DeepSeek-V4.1 Acceleration and Decisions API

SGLang v0.5.21, a tool for launching models, has been released. The release includes 779 pull requests from 227 contributors. Among the new models are DeepSeek-V4.1, GigaChat 3.5, IQuest-Q1, MiMo-V2.6 and MiMo-V2.6-Pro, Ling-3.0-flash-VL, as well as the diffusion models DiffusionGemma, Qwen-Image 2.1, Anima Base v1.0, Ming-Image 0.1 and FLUX 3 Action. Details are in the release description on GitHub.

What it means

For those running models on their own hardware, the release brings several practical changes. PD instances can now switch between prefill and decode on the fly, without a restart. The prefix cache now runs on a Rust core by default. For DeepSeek-V4.1, a 22% speedup of the first token on long prompts is claimed, and for Kimi K3, a 20.6% increase in prefill throughput in PD serving. The new Decisions API (/v1/decisions) turns an LLM or VLM into a low-latency classifier and scorer, while the Score API (/v1/score) scores all candidates in a single request. Running MiniMax-H3 with SGLang Diffusion inside ComfyUI is also claimed to be possible. To update, the command uv pip install --prerelease=allow sglang==0.5.21 is suggested. Docker images are available for NVIDIA (CUDA 13), AMD MI35x, AMD MI30x, Intel GPU and Intel CPU. You can follow news about tools for running neural networks locally in the CPU3D news feed.

See also