AI News for 3D and Video Generators
Short news notes about what has changed in the tools: a new version released, a license changed, support for another platform added. No press-release retelling: if a claim cannot be tied to a repository, model page, or official page, there will be no note.
Articles
64 articles
DSAQuant proposes stage-aware quantization for video diffusionThe authors of an arXiv paper showed that conventional quantization-aware training degrades precisely the details and sharpness of video models, while…
OctWorld presented as a generator of long consistent videos with octree-based 3D memoryOn arXiv, September 3, 2026, a paper on OctWorld was released — a framework for video generation based on diffusion with persistent 3D memory. From a…
VCAR: Training-Free 3DGS Segmentation via View Completeness and Axis-Aware Boundary RefinementThe authors of the paper VCAR: Training-Free 3DGS Segmentation via View Completeness and Axis-Aware Boundary Refinement propose a method for semantic…
H3-World: Interactive Image-to-Video Model Based on MiniMax-H3The model DANNY621/H3-World with the image-to-video task has been published on HuggingFace. The author describes it as the first interactive world mod…
CapFrame converts a text description of a frame into a camera pose for 3D Gaussian SplattingResearchers introduced CapFrame, a method that uses a text instruction to find a 6-DoF camera pose in a 3D Gaussian Splatting scene so that the render…
SpatialCrafter turns a single image into an explorable 3D scene via a generative proxyResearchers introduced SpatialCrafter, a two-stage framework for generating explorable 3D scenes from a single image. The work was published on arXiv…
VibeVoice-1.5B-hf: Microsoft's new open-source text-to-audio model on HuggingFaceThe vibevoice/VibeVoice-1.5B-hf model has been published on HuggingFace for the text-to-audio task. It is designed to generate long conversational aud…
4DStreamCtrl combines camera, object, and depth control in streaming modeA paper on arXiv introduces 4DStreamCtrl, a method for interactive video generation that unifies camera motion, object trajectories, and depth into a…
LeFlow speeds up planning in world models by an order of magnitude with a generative latent priorResearchers introduced LeFlow, a method that replaces iterative trajectory optimization in latent world models with generative planning. Instead of ru…
Luce introduces 3D asset generation with PBR materials and relighting from a single imageResearchers published on arXiv the paper Luce: Relightable Gaussians for 3D Asset Generation, presenting a method for generating 3D models from a sing…
SceneReGen builds 3D scenes from a single image, generating objects in a shared coordinate systemOn arXiv, August 25, 2026, the paper SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image was published. The authors propose a metho…
GlanceWAM separates imagination and control in a single video DiT for robotsResearchers introduced GlanceWAM, an approach to world-action models that moves video generation off the critical control path. An asynchronous propos…
FixAnything improves 3D scene renders via video generative modelsResearchers introduced FixAnything — a unified model for fixing rendering artifacts across different 3D representations, including Gaussian Splatting…
MiniMax-H3-slim released on HuggingFace for text-to-video generationOn August 24, 2026, the model fkyyy/MiniMax-H3-slim was published on HuggingFace with the text-to-video task. At the time of publication, it had 508 d…
New H3_HD_2K_Detailer model for image-to-video on HuggingFaceOn HuggingFace, on August 24, 2026, the model zuanfilm/H3_HD_2K_Detailer was published with the image-to-video task. At the time of publication, the m…
Loopy turns the DiT temporal axis into a circle for seamless looping videosThe authors of Loopy published a paper on arXiv about generating seamless looping videos. They show that positional embedding layers in DiT control te…
H3 Cinematic Multishot Coverage Released on HuggingFace for Image-to-VideoOn August 24, 2026, the ethanfel/H3_Cinematic_Multishot_Coverage model was published on HuggingFace for the image-to-video task. The code is released…
LTX-Video, CogVideoX and Pyramid Flow fall short of simulators on eight criteriaThe authors of a systematic review on arXiv analyzed 200 papers on generative world models from 2018 to June 2026 and compared them with traditional s…
LTX-Video and CogVideoX Still Don't Solve Multi-Object Generation in a Single VideoA comparative study on arXiv examines three approaches to generating videos with multiple objects without fine-tuning: direct, parallel, and sequentia…
InfinityEdit: Adapter for Infinite Streaming Video Editing UnveiledOn August 21, 2026, the paper InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter was published on arXiv. The authors abando…
ArtiMo: agent-driven framework for animating articulated 3D meshes from textOn August 21, 2026, the paper ArtiMo: Agent-Driven Articulated Mesh Animation was published on arXiv. The authors propose a framework that animates ar…
DiffVC-ONE: One-Step Video Compression with Video Diffusion TransformerarXiv published DiffVC-ONE, a generative video compression framework based on a one-step Video Diffusion Transformer. The authors describe three compo…
MultiCube introduces 3D generation with part-level controlResearchers published a paper on arXiv, MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control, describing a method for g…
VGI-Bench evaluates visual reasoning of video generators on 810 tasksarXiv published VGI-Bench — a benchmark of 27 tasks and 810 examples for testing visual reasoning in video generation models. The authors built a two…
Stream4D adds 4D consistency to streaming autoregressive video modelsThe authors of the paper Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models (arXiv, 20.08.2026) propose a training method fo…
VA-Judger: Reward Model for Joint Video and Audio GenerationResearchers introduced VA-Judger, a reward model for post-training of joint video and audio generation systems, along with the VAPref-10K dataset of 1…
SparsePR accelerates video generators with training-free sparse attentionResearchers introduced SparsePR, a training-free block-sparse attention method for video generators and world models. The paper was published on arXiv…
HunyuanVideo gets a generation acceleration method with near-lossless quality at 5–7x speedupResearchers introduced LinCa, a method for accelerating diffusion models via learnable feature caching with component-wise decomposition. Instead of a…
New h3-vbvr model on HuggingFace tackles image-to-videoOn August 18, 2026, the Patarapoom/h3-vbvr model was published on HuggingFace with the image-to-video task. At the time of publication, it had 2,434 d…
SQuad accelerates video generation in Wan 2.2, reducing attention by 67xThe authors of SQuad introduced a sub-quadratic attention distillation method for Video Diffusion Transformers. Instead of training a model from scrat…