CPU3DAI tools for 3D, video and audio

FlowHMR Turns Video Motion Capture into a Generation Task with Physical Control

A paper on FlowHMR has been published on arXiv — a method for recovering global 3D human motion from monocular video. The authors formulate the task as video-conditioned motion generation and first pretrain a flow matching model. The model is then fine-tuned using Group Relative Policy Optimization (GRPO) with two rewards: a fidelity reward responsible for consistency with the source video, and a tracking reward for how successfully a physical controller can track the generated motion. According to the authors, this approach avoids the averaged solutions typical of direct regression and makes the result physically plausible. For evaluation, they present the Wild-4K dataset of roughly 4 thousand internet videos. More details in the paper on arXiv.

What it means

FlowHMR solves a motion capture task, not a video generation one, so it does not directly overlap with the tools in our video generators section. Still, it is worth looking at how it relates to what is already in the reference. Maestro is a local studio for generating video, images and music based on the WanGP pipeline. Its task is image-text-to-video, not recovering human motion from a finished clip. The code is open source, but the WanGP Non-Commercial Evaluation License 1.1 permits only non-commercial use of the program itself; the MiniMax H3 weights are available for commercial use up to 20 million dollars in annual revenue, but the license does not apply in the EU, the UK, Korea and the US. VRAM required is 12—24 GB. There is no physical controller or motion capture module in Maestro's fact sheet. LTX-Video from Lightricks is an open source image-to-video generator under Apache-2.0. The weights are distributed under the LTXV Open Weights License 0.X: commercial use is permitted, but companies with revenue of 10 million dollars or more need a separate license. The authors claim 1 GB of VRAM is enough for the model. This is also a generative model, not a motion capture tool, and physical plausibility of the result is not claimed in its fact sheet. CogVideoX from Zhipu AI solves a text-to-video task. The code is open source under Apache-2.0, the weights under The CogVideoX License: academic use is free, commercial use requires registration, free up to 1 million visits per month. The largest weights file is 9.2 GB. As in the previous cases, this is a video generator, not a motion recovery system, and there is no overlap with FlowHMR in terms of the task. FlowHMR is interesting because it shifts the emphasis from regression accuracy to the physical trackability of the result. For users looking for tools to work with human motion, this is more a signal about the direction of research than a ready-made tool: our reference currently has no equivalents for this task.
FlowHMR: from video to physically trackable 3D motion Video monocular clip with a person Flow Matching pretraining motion generation GRPO fine-tuning with two rewards 3D pose motion Two GRPO rewards Fidelity reward consistency with the source video Tracking reward tracking success by a physical controller Difference from direct regression Direct regression: averaged solutions without physical control FlowHMR: physically plausible trackable motion Wild-4K about 4K
How the method works. The diagram was drawn for this news note.

See also