FracGen teaches video generation to destroy objects using physical signals
The authors of FracGen presented a video generation model that, from a single image of an entire object, creates plausible fracture dynamics driven by physical signals. For training, they built the FracSim simulator based on the material point method with a damage model, which outputs paired fracture videos and dense physical fields. The model is trained to predict these fields alongside RGB video and is additionally supervised by physics-informed losses. The authors did not release the code. More details in the arXiv paper.
What it means
FracGen is a research work, not a ready-made tool from our reference. In the video generators section there are no models tailored for physically accurate fracture: Maestro solves the image-text-to-video task, LTX-Video — image-to-video, CogVideoX — text-to-video. None of them predicts physical fields or is trained with physics-informed losses.
The fact sheets show that the closest in task is LTX-Video: it also works from a single image, its code is open source under Apache-2.0, and the weights require only 1 GB of VRAM. But its weights license limits commercial use by a revenue threshold, and its description says nothing about physical fracture simulation. Maestro requires 12—24 GB of VRAM and runs only on CUDA, while CogVideoX targets NVIDIA H100 and above.
As long as FracGen remains a paper without code or weights, there is nothing to compare it against local tools by hardware requirements or license: the author releases neither a repository, nor a model, nor VRAM data. If the code appears, it will become clear whether such generation can be run on your own GPU and under which license.How the method works. The diagram is drawn based on this news note.