AI News: 3D, Video, and Language Models for Your Own Hardware
Short news notes about what has changed in the tools: a new version has been released, the license has changed, support for another platform has been added. No press release retelling: if a claim cannot be tied to a repository, model page, or official page, there will be no note.
Articles
134 articles
HelloWorld — a model for social interaction with characters in video worldsResearchers have introduced HelloWorld, a video world model in which the user can interact with characters on screen. With the press of a button, the…
GVCCTurbo proposes bitrate planning for generative compression of video and images withoutarXiv Researchers introduced GVCCTurbo, a planner that separates costly updates of the generative prior from transmitting corrections through a codebo…
SPADE accelerates attention in video diffusion transformers without fine-tuningResearchers have introduced SPADE, a sparse attention engine that accelerates inference of video diffusion transformers without model fine-tuning. The…
SUV: Scene Future as Video Generation for End-to-End DrivingOn August 4, 2026, a paper on arXiv introduced SUV — an end-to-end driving framework that frames future scene understanding as a video generation task…
Token Radius Attention Speeds Up Video Generation in Diffusion Transformers Without Fine-TuningResearchers from Peking University have introduced Token Radius Attention (TRA), a method that reduces the computational load in Video Diffusion Trans…
EchoCache accelerates sound-driven video generation using audio signal energyarXiv reports on EchoCache, a caching method for audio-driven video generation (A2V). The authors identified two types of mismatch in existing approac…
MAGI-2-preview by Sand AI: new open-source image-to-video model on HuggingFaceThe sand-ai/MAGI-2-preview model has been published on HuggingFace — a preview version of the image-to-video generator by Sand AI. The model handles t…
UniMoCa unifies human motion and camera control in a single visual representationResearchers have introduced UniMoCa, a method that translates both human motion and camera trajectory into a shared visual representation called Motio…
MoRoute proposes dynamic layer routing for multimodal video generationResearchers from MoRoute published a paper proposing to combine a frozen vision-language model (VLM) and a pretrained video diffusion transformer (DiT…
ROAD cuts the cost of 3D generation training by transferring knowledge from discriminative modelsThe authors of ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation propose a method that uses discriminative 3D m…
Wan2.2 taught to render video from an animated mesh using the DAR methodResearchers introduced the DAR method, which turns the pretrained video model Wan2.2 into a renderer driven by an animated mesh, camera trajectory, an…
FreqForcing extends autoregressive video generation to two minutes without fine-tuningOn July 29, 2026, the paper FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring appeared on arXiv. The authors investigate t…
VoxelModel-v1: a new open-source text-to-3d model on HuggingFaceThe bench-labs/VoxelModel-v1 model, which handles the text-to-3d task, has been published on HuggingFace. The model appeared on July 26, 2026, and the…
Lightricks releases LTX-2.5 with image-to-video taskLightricks published the LTX-2.5 model with an image-to-video task on HuggingFace. The publication is dated 07/23/2026, the model has 39 downloads and…