DeltaWAM Speeds Up Action Generation for Bimanual Robots and Releases Code
Researchers from AIGeeksGroup presented DeltaWAM — a world-action model for controlling a robot's two hands. Instead of predicting dense future frames, it predicts visual deltas and actions, while Streaming Delta Memory updates the cached context with compact changes. On the RoboTwin benchmark, average success rose from 81.3% to 85.4% in a clean environment and from 75.8% to 83.9% under visual randomization. Single-step inference latency dropped by 36.57%, FLOPs — by 31.55%. The code is open source, details are in the paper on arXiv.
What it means
DeltaWAM is not a video generator for the user, but a robot control model that uses video generators as a source of visual and motor priors. In our section on AI tools for video it has no direct equivalents: Maestro, LTX-Video and CogVideoX handle image-text-to-video, image-to-video and text-to-video tasks, not action generation.
Still, there is an overlap at the architecture level. All three tools from the reference are built on diffusion video models, and DeltaWAM's idea — not to generate full frames but to predict only changes — is conceptually close to what lightweight video generators already do. For example, LTX-Video claims a need for just 1 GB of VRAM, while Maestro runs on cards from 12 GB. DeltaWAM goes further: it does not process full observations with a heavy video expert at every step at all, but updates only the deltas.
The authors of DeltaWAM do not specify the license of the code and model weights, VRAM requirements or the launch platform. Judging by the paper, the model was trained and tested on robotics benchmarks, not on user GPUs, so there is nothing to compare its system requirements with the video generators from the reference yet.
For those who follow local video models, DeltaWAM is interesting as an example of where inference optimization is heading: fewer computations per step, context caching and dropping the reprocessing of unchanged content. If these techniques make their way into user video generators, they could lower VRAM requirements and speed up generation on local hardware.How the method works. The diagram was drawn for this news note.