llama.cpp v0.6.0: extended batch API, new models, and acceleration on Metal and Vulkan
llama.cpp v0.6.0 is out. The release adds an extended batch API llama_batch_ext with the llama_process() function for mixed token/embedding batches and per-token state embeddings for MTP and deepstack models, support for the hybrid GLM-5.3-Flash (GLM5-Next) model with 320B parameters and text plus vision, Clef decision models, MTP speculative decoding for Qwen4Exp, a new server API /v1/systemone for five decision models, and an update of ggml to v0.26.0.
What it means
For those running models locally, the release brings several practical changes. The new batch API allows mixing tokens and embeddings in a single batch — this matters for MTP and deepstack models that need to pass state between generation steps. GLM-5.3-Flash support opens access to a 320-billion hybrid model with text and vision, although the developer does not specify the VRAM requirements for running it.
On Apple devices, the new flash attention kernel for F16 KV and few-row MMA mat-mul kernels speed up matrix multiplications by up to three times in speculative and batched decoding scenarios. On Vulkan, sparse flash attention for quantized K/V has been added. MTP speculative decoding for Qwen4Exp gives roughly a one-and-a-half-fold decoding speedup on DGX Spark.
The server API /v1/systemone supports five decision models: laya, julia-1, lev, openjev (with vision support) and kev. The licenses of the new models are not listed in the release description.
The update touches several areas at once — from developer APIs to inference acceleration on specific hardware. You can follow news about running neural networks locally in the CPU3D news feed.How the method works. The diagram is drawn based on this news note.