AI News: 3D, Video, and Language Models for Your Own Hardware
Short news notes about what has changed in the tools: a new version has been released, the license has changed, support for another platform has been added. No press release retelling: if a claim cannot be tied to a repository, model page, or official page, there will be no note.
Articles
134 articles
Audio8-ASR-Infinite: Streaming Speech Recognition with Constant Memory UsageOn September 21, 2026, the Edge0/Audio8-ASR-Infinite model was published on HuggingFace with the text-generation task. The author describes it as nati…
WorldCrafter-Base by TencentARC: a 14.3B-parameter image-to-video modelOn September 21, 2026, the TencentARC/WorldCrafter-Base model with the image-to-video task was published on HuggingFace. This is the base part of the…
SupraLabs Releases Compact Supra2-IMG Text-to-Image ModelThe SupraLabs/Supra2-IMG model has been published on HuggingFace — a diffusion transformer for generating images from text descriptions. The author ca…
Qwen Releases a Prompt-Rewriting Model Qwen-Image-2.1-PE-T2IOn September 20, 2026, the model Qwen/Qwen-Image-2.1-PE-T2I was published on HuggingFace. It is not an image generator but an auxiliary model for prep…
FrontierAgent: a new agent environment with TUI, Agent Team mode and quality evaluationFrontierAgent is an open-source agent environment for long-running research and file tasks; the repository was opened on August 22, 2026 and gained ab…
OpenBot: agentic environment with its own browser and full action controlOpenBot is a new platform for building AI agents; its repository was open-sourced on August 17, 2026 and gained 5,112 stars on GitHub in 32 days. The…
Moondream Releases Parakeet Redux — a 1.58-bit Speech Recognition Model with 149M ParametersOn September 18, 2026, the moondream/parakeet-redux model for automatic-speech-recognition was published on HuggingFace. It is a 1.58-bit version of p…
Cube-Splat Introduces Open-Source Panoramic Gaussian Splatting SLAMA paper on Cube-Splat has been published on arXiv — a method for simultaneous localization and mapping (SLAM) based on 3D Gaussian Splatting for panor…- Pipecat 1.11.0: Resampler with Manual Reset and Function Call Evaluation in ScenariosOn September 18, 2026, Pipecat 1.11.0 was released — an open source Python framework for real-time voice and multimodal agents. Since the last review…
LocalAI v4.10.0: cluster dashboard, credentials file and benchmarkThe LocalAI running tool has released v4.10.0. The authors describe three areas of work: cluster management, loading models from private sources and b…
LichtFeld-Studio Releases RoMa v1 Weights for Image MatchingThe LichtFeld-Studio tool has a new release — RoMa v1 dense matcher weights. This is a permanent model asset in the native .lfw format, used in the de…
Xing4.0-29B-A4B: China Telecom's new 31B-parameter MoE modelOn September 16, the model XingChen-AGI/Xing4.0-29B-A4B was published on HuggingFace — a text-generation model from the Xing series (formerly TeleChat…- PhysStream Introduces Streaming Video Generation with Physics-Based Motion ControlA group of researchers has published the PhysStream paper on arXiv — an autoregressive image-to-video model that lets you control object motion during…
MoE-JEPA boosts V-JEPA with a mixture of experts for synthetic image detectionResearchers presented MoE-JEPA, an architecture for detecting generated and manipulated images based on V-JEPA 2 with an added Residual Mixture-of-Exp…
NVIDIA Releases c-foundationstereo-s Depth Estimation Model on HuggingFaceNVIDIA published the c-foundationstereo-s model on HuggingFace for the depth-estimation task. This is FoundationStereo — a model for depth estimation…
llama.cpp v0.4.1: support for Maple 20B-A1B, Tencent Hy 4 and Spark2.5The llama.cpp v0.4.1 release came out on September 14, 2026. The update adds support for three new architectures: Maple 20B-A1B — a ternary MoE for CP…
VC-Attention Speeds Up Low-Bit Attention for Video Generators on GPUThe authors of VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention propose a training-free low-bit attention method for Diffusion…
Cheap Chinese Coding Models: What to Choose in September 2026A developer of a web project asked a simple question: which inexpensive model to hook up for coding with pay-per-token billing that also performs well…
BubbleID: Image Segmentation Model for Boiling Analysis Released on HuggingFaceOn September 14, 2026, the UARK-NED3/BubbleID model with the image-segmentation task was published on HuggingFace. It is a research package for analyz…
OOOSplat v0.4.1 improves error handling and preview stabilityOOOSplat v0.4.1 is out — an update to the local 3D Gaussian Splatting generator focused on stability and failure handling. The GitHub release describe…
Agnes-AI releases open preview model Agnes-3.0-Flash with 33 billion parametersThe model Agnes-AI/Agnes-3.0-Flash with the text-generation task has been published on HuggingFace. This is an open preview checkpoint of Agnes 3.0 Fl…
whisper.cpp v1.9.4: Windows on ARM support and ggml update to 0.23.0Release v1.9.4 of the whisper.cpp runner is out. Changes published in the GitHub repository: Windows on ARM support added to the build pipeline, the g…
Tri-DehazeGS Separates Scene and Haze in 3D Reconstruction from PhotosA group of researchers presented Tri-DehazeGS, a method for recovering clean 3D scenes from smoky or hazy images. The work was published on September…
ComfyUI v0.35.0: WAN3-Prime, Omni 1.1 support and new 3D nodesComfyUI v0.35.0 was released on 09.09.2026. The release brings support for new partner models, 3D nodes for working with meshes, and several fixes rel…
YuE2 released with open source and weights under CC BY-NC 4.0The YuE tool has a new release, YuE2 · frontier music generation with symbolic planning. Release date — 09.09.2026. The GitHub release describes three…
YuE2-3B: An Open Model for Music Generation with Editable ScoreThe m-a-p/YuE2-3B model with the text-to-audio task has been published on HuggingFace. The author is m-a-p, whose tools are already covered in our ref…
Programmable World Model — a new framework for controllable video worlds with persistent stateA paper on Programmable World Model has been published on arXiv, in which the authors propose a framework for creating interactive video worlds with a…
Test-time guidance improves the geometry of image-to-3D models from partial observationsResearchers presented a training-free method that allows partial geometric observations of an object to be taken into account at the generation stage…
Streaming talking head: authors decoupled audio and rendering to speed up generationA paper titled Decoupled Self-Forcing Distillation for Streaming Talking Head Generation has appeared on arXiv. For streaming talking head generation…
Edge0 Releases 35B MoE Model That Runs in 3 GB of MemoryEdge0-35B-A3B-preview has been published on HuggingFace — a 35B-class text model with a sparse MoE architecture. The authors claim it decodes at 15 to…