LocalAI v4.11.0: audio scenes, failover chains, and decision models
The LocalAI inference tool has released v4.11.0. The release dated October 2, 2026 adds audio scenes combining transcription, diarization, sound event detection and speaker name memory, ordered failover chains between local and remote targets, decision-making models via /v1/systemone, signed OCI model galleries and Kimodo text-to-animation. Details are in the release description on GitHub.
What it means
For those running models on their own hardware, the release brings several practical changes. Audio scenes let you combine speech recognition, diarization, sound event detection and speaker identification in a single scenario: Studio can diarize a recording, show clean intervals and register selected speakers by name. Failover chains make it possible to keep a single model available through multiple targets — local and remote — with retries until a response is committed, health monitoring of targets and switches, and pinning a target via API, MCP tools or the React UI.
Decision-making models now advertise their use case and answer structured questions of choice, evaluation and "noul" via /v1/systemone with validation and separate models in the gallery. Signed OCI galleries allow distributing complete model galleries as OCI artifacts and verifying them via Sigstore. The release includes 243 merged pull requests, 353 commits and 79 new gallery entries. The developer does not specify VRAM requirements or parameters of individual models.
You can follow news about tools for running neural networks on your own hardware in the CPU3D news feed.How the method works. The diagram is drawn based on this news note.