Minimax-H3-fl2va-ref2va-hybrid-models: new text-to-video model on HuggingFace
On August 11, 2026, the model smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models was published on HuggingFace with the text-to-video task. At the time of publication, the model had zero downloads and 27 likes. The developer does not specify the architecture, hardware requirements, or output file format.
What it means
The model belongs to video generation from text descriptions, i.e., the text-to-video direction. In our reference of AI tools for 3D and video, there is currently no model with the text-to-video task — all featured generators work from images: OpenLRM, Hunyuan3D 2.1, and TRELLIS solve the image-to-3d task. There is no direct overlap in task with the new model.
Indirect comparison is also difficult: all three of our tools are open source with published weights, whereas for Minimax-H3-fl2va-ref2va-hybrid-models the developer specifies neither the license, nor code openness, nor file sizes. Until the model gains downloads and documentation appears, it is impossible to assess its place among the tools in the reference.
If the model turns out to be functional and gets a description of requirements, it could be considered as a candidate for addition to the text-to-video section — this direction is currently not covered in the reference.How the method works. The diagram is drawn based on this news note.