August 3, 2026
Lights, camera, comment war
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
People are already trash-talking rival AI video tools as creators gasp at the quality
TLDR: MiniMax H3 launched with open access and instant ComfyUI support, letting people generate short AI videos with sound from text, images, and more. The community response swung from “delete the old tools” hype to “how long does this actually take?” realism, with plenty of Hollywood-is-doomed drama in between.
MiniMax H3 landed today, and the biggest plot twist isn’t just that it makes short videos with built-in stereo sound and can use text, pictures, video, or even audio as input — it’s that ComfyUI had day-zero support ready to go. In plain English: a powerful new AI video generator showed up, and hobbyists were trying it locally almost immediately. That alone had the comment section acting like a season finale.
The loudest reaction? Pure scorched-earth hype. One commenter said they saw sample clips and “immediately deleted” competing model folders because they were now “completely worthless.” Subtle! Another went full blockbuster panic mode, declaring “Hollywood… on red alert” and, naturally, “This is AGI” — the internet’s favorite way to say, “Okay, this feels a little too real.” If you wanted calm, measured takes, this was not the thread.
But not everyone was just screaming into the void. A more grounded mini-debate broke out around speed versus wow-factor. Yes, H3 can supposedly run on consumer gear, but one user asked the question lurking behind every flashy demo: how long are we actually waiting for a full 15-second clip? Another commenter answered with the kind of reality check only a person staring at a progress bar can give: on their card, 10 minutes for a 10-second 480p video — followed by the all-important verdict, “the results are spectacular.” That’s the mood in a nutshell: equal parts awe, impatience, and cinematic doom-posting.
Key Points
- •MiniMax released its H3 video model with open weights, and ComfyUI added native day-zero support.
- •H3 accepts text, images, video, and audio inputs and can generate video with native stereo audio, up to 2K resolution and 15 seconds per clip.
- •The model supports text-to-video, image-to-video, first-and-last-frame control, and reference-driven video generation.
- •The article highlights cross-modal generation, including motion transfer from reference video and in-place editing for iterative shot creation.
- •ComfyUI says memory optimizations reduced the model footprint by 66%, from 123.6 GB to 42.5 GB for the smallest variants, enabling local inference on hardware such as an RTX 3060.