CPU3DAI tools for 3D, video and audio

llama.cpp v0.6.0: extended batch API, new models, and acceleration on Metal and Vulkan

llama.cpp v0.6.0 is out. The release adds an extended batch API llama_batch_ext with the llama_process() function for mixed token/embedding batches and per-token state embeddings for MTP and deepstack models, support for the hybrid GLM-5.3-Flash (GLM5-Next) model with 320B parameters and text plus vision, Clef decision models, MTP speculative decoding for Qwen4Exp, a new server API /v1/systemone for five decision models, and an update of ggml to v0.26.0.

What it means

For those running models locally, the release brings several practical changes. The new batch API allows mixing tokens and embeddings in a single batch — this matters for MTP and deepstack models that need to pass state between generation steps. GLM-5.3-Flash support opens access to a 320-billion hybrid model with text and vision, although the developer does not specify the VRAM requirements for running it. On Apple devices, the new flash attention kernel for F16 KV and few-row MMA mat-mul kernels speed up matrix multiplications by up to three times in speculative and batched decoding scenarios. On Vulkan, sparse flash attention for quantized K/V has been added. MTP speculative decoding for Qwen4Exp gives roughly a one-and-a-half-fold decoding speedup on DGX Spark. The server API /v1/systemone supports five decision models: laya, julia-1, lev, openjev (with vision support) and kev. The licenses of the new models are not listed in the release description. The update touches several areas at once — from developer APIs to inference acceleration on specific hardware. You can follow news about running neural networks locally in the CPU3D news feed.
llama.cpp v0.6.0 — key changes Batch API llama_batch_ext mixed batches tokens + embeddings GLM-5.3-Flash hybrid model 320B parameters text + vision MTP decoding Qwen4Exp speculative decoding x1.5 speedup Metal flash attention F16 KV up to x3 speedup Vulkan sparse flash attention quantized K/V /v1/systemone 5 decision models laya, julia-1, lev openjev, kev Clef decision model API support ggml update to v0.26.0 The update affects the API, models and inference speedup on specific hardware Licenses of the new models are not listed in the release notes
How the method works. The diagram is drawn based on this news note.

See also