AI News: 3D, Video, and Language Models for Your Own Hardware
Short news notes about what has changed in the tools: a new version has been released, the license has changed, support for another platform has been added. No press release retelling: if a claim cannot be tied to a repository, model page, or official page, there will be no note.
Articles
134 articles
llama.cpp v0.6.0: extended batch API, new models, and acceleration on Metal and Vulkanllama.cpp v0.6.0 has been released. The release adds the extended batch API llama_batch_ext with the llama_process() function for mixed token/embeddin…
LiquidAI releases d1-3B — a decision model that answers in a single pass without generating textOn October 5, 2026, the LiquidAI/d1-3B model with the image-text-to-text task was published on HuggingFace. It is a decision model with 3.1 billion pa…
n8n 2.42.3: an agent platform focused on self-hosted stabilityRelease 2.42.3 of the n8n platform — a visual builder for AI agents and workflows — is out. Six releases have accumulated since the last review: 2.41…
Blockway releases preview of Agens-Volundr-32B model for reasoning, code and agentic workThe Blockway/Agens-Volundr-32B-Preview model with the text-generation task has been published on HuggingFace. This is a preview release from Blockway…
FlowHMR Turns Video Motion Capture into a Generation Task with Physical ControlA paper on FlowHMR has been published on arXiv — a method for recovering global 3D human motion from monocular video. The authors formulate the task a…
ManifoldSplat edits 3D head shape with text in 90 secondsResearchers present ManifoldSplat, a method for language-driven shape editing of animatable 3D heads reconstructed from monocular video. Instead of di…
Aleph Alpha releases open-source 78B-parameter MoE model Kolibri-1-BF16On October 2, 2026, the model Aleph-Alpha/Kolibri-1-BF16 was published on HuggingFace with the text-generation task. It is a mixture-of-experts reason…
Apeireth: An Agentic OS in Rust with Topological Memory and Its Own KernelIn late September 2026, the Apeireth repository appeared on GitHub — an agentic environment that the authors describe as an "AGI Operating System & Co…
SGLang v0.5.21: New Models, DeepSeek-V4.1 Acceleration and Decisions APISGLang v0.5.21, the tool for running models, has been released. The release includes 779 pull requests from 227 contributors. Among the new models are…
Generative Cinematographer Teaches a Video Model to Control Camera and Objects in 3DOn October 1, 2026, arXiv published Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene and lets you jo…
vLLM v0.31.0rc3 adds randomized dummy inputs to Model Runner V2The vLLM v0.31.0rc3 pre-release is out. It includes the Model Runner V2 change: Support randomized dummy inputs, which adds support for randomized dum…
A Survey of Post-Training and Alignment Methods for Video Generators on arXivA survey titled Video Generation Models: A Survey of Post-Training and Alignment has been published on arXiv. The authors systematize approaches to po…
Cloudflare Releases clef Model for Decision-Making on Question SchemasOn September 30, 2026, the Cloudflare/clef model appeared on HuggingFace with the image-text-to-text task. The model accepts state as text, JSON, imag…
Whistle: Speech Recognition Model for CPU Released on HuggingFaceOn September 30, 2026, the Cactus-Compute/whistle model was published on HuggingFace with the automatic-speech-recognition task. The author describes…- Rig agent platform: release 0.43.0 with local YOLOv8 pose estimationOn September 30, 2026, release 0.43.0 of the Rig platform — a Rust library for building modular and scalable LLM-based applications — was published. W…
Wind Comic v12.460.0 — First Release After 700 Updates in mainThe wind-comic tool has released v12.460.0 — the first GitHub Release since April 2026. Since then, all changes went straight to main, and more than 7…
FracGen teaches video generation to destroy objects using physical signalsThe authors of FracGen presented a video generation model that, from a single image of an intact object, creates plausible destruction dynamics driven…
EndoPrior-GS improves dynamic 3D reconstruction for endoscopyResearchers presented EndoPrior-GS, a dynamic endoscopic reconstruction method based on 3D Gaussian Splatting. The work is published on arXiv, and the…
MLX v0.32.3: Fixes for Running Models on Apple SiliconVersion v0.32.3 of the MLX runtime has been released. The update includes fixes for scan and sort operations on a zero-size axis, adds a correction pa…
OOOSplat 0.5.0: Auto-Optimization, Bridge Frame Interpolation and Re-ShootingOOOSplat 0.5.0 is out — a local tool for turning video and photos into 3D Gaussian Splatting. This release adds automatic parameter optimization, brid…
Pydantic AI 2.51.0: GPT-Live support and six releases in two weeksOn September 25, 2026, version 2.51.0 of Pydantic AI was released — a Python framework for building AI agents with a typed agent loop. Since the last…
Ollama v0.40.0 switches models to MLX on Apple Silicon by defaultOllama v0.40.0 has been released. Main change: on Apple Silicon devices, models whose architectures are supported by the MLX runtime now run through M…
YuE2 Music skill 1.2.0: instrumental generation and covers from a description, ABC, or a recordingThe YuE tool has released YuE2 Music skill 1.2.0, which adds generation of instrumental music and instrumental covers from a text description, an ABC…
Reliability-Regulated Trajectory Optimization for COLMAP-free 3DGS Released on arXivA paper on a camera trajectory optimization method for progressive COLMAP-free 3D Gaussian Splatting has been published on arXiv. The authors propose…
ViRDM Removes Teacher and Critic from Few-Step Video DistillationThe authors of ViRDM propose a post-training method for a video generator without a teacher or critic, reducing VRAM usage. The work was published on…
DeltaWAM Speeds Up Action Generation for Bimanual Robots and Releases CodeResearchers from AIGeeksGroup introduced DeltaWAM, a world-action model for controlling a robot's two hands. Instead of predicting dense future frames…
Qwen-Image-2.1 Released on HuggingFace: Image Generation and Editing with a 7B ModelThe model unsloth/Qwen-Image-2.1 has been published on HuggingFace — a unified model for text-to-image generation and editing. Publication date: Septe…
LeWAM Drops Video Diffusion for JEPA Embeddings in World Action ModelsA paper on Latent evolving World Action Model has been published on arXiv, in which the authors investigate how visual representations affect action g…
LichtFeld-Studio releases MoGe-3 ViT-L weights in native lfw formatLichtFeld-Studio has released a build with MoGe-3 ViT-L base weights in the native .lfw format. This is the default depth model for the SLAM pipeline…
Xiaomi Releases MiMo-V2.6-Flash-RL on HuggingFaceOn September 21, 2026, the XiaomiMiMo/MiMo-V2.6-Flash-RL model was published on HuggingFace with the text-generation task. The author describes it as…