All videos
0:00 / 0:00
research

MiniMax Just Dropped a "Seedance Killer" with a Twist

Theoretically Media31 July 2026Watch on YouTube

Description

Today's video is sponsored by ElevenLabs! Try ElevenAgents and start chatting with your characters — plus grab 10,000 free credits courtesy of ElevenLabs right here: https://try.elevenlabs.io/w5183uo8pydc This is a surprise: MiniMax is BACK. After nine months of silence since Hailuo 2.3, they’ve launched Hailuo 03 (aka MiniMax H3) — a fully multimodal AI video model with native audio, up to 15-second generations at 2K, and an Omni mode that takes up to 12 image references plus video and audio inputs. Is it the Seedance killer people are claiming? I ran it through the full gauntlet to find out. In this video, we test Hailuo H3 across cinematic text-to-video (including dialogue, acting range, and multilingual support), image-to-video against the old 2.3 model, first frame/last frame, video extensions, and multi-reference character team-ups — plus what I’m seeing on pricing, which looks like it comes in around 50% under Seedance 2.0. Key takeaways: dialogue and lip sync are shockingly good, fight choreography trades speed for stability (no morphing, no teleporting), prompts cap out at a generous 7,000 characters, and the Omni model takes your storyboards very seriously — maybe more seriously than Seedance does. Note: Hailuo H3 is currently in early access and rolling out soon — consider this an opening impressions, not a full review. CHAPTERS 00:00 Intro: MiniMax Is Back 00:49 What's New in Hailuo 03 01:44 Dialogue & Acting Test (Twin Peaks Diner) 02:21 Cinematic Text-to-Video Tests 03:02 The Mob Movie Test 03:47 Multilingual Dialogue (Samurai Test) 04:39 Fight Scenes + the 7,000-Character Prompt Limit 06:10 Image-to-Video: Nine Months of Progress 07:03 Flamethrower Girl Returns 07:50 First Frame / Last Frame 08:30 The Omni Model: Multi-Reference 09:37 ElevenLabs Sponsor Segment 13:34 Sniper Girl One-Shot 14:13 Video References & Extensions 15:18 Sniper + Flamethrower Team-Up 16:02 "Reference Maxing" & Storyboards 17:25 Pricing Ballpark 18:23 Final Thoughts

What you'll learn

  • Hailuo 03 (MiniMax H3) is a multimodal AI video model with native audio, up to 15-second generations, and support for up to 12 image references simultaneously.
  • The model delivers surprisingly strong dialogue and lip-sync quality, and handles fight choreography stably without morphing or teleportation artifacts.
  • Hailuo H3 pricing comes in around 50% under Seedance 2.0, making it more accessible for budget-conscious creators.
  • The Omni mode takes storyboards and references seriously, and can combine multiple character references within the same video.
  • The model supports up to 7,000 characters per prompt and offers robust multilingual dialogue capabilities across different languages.

Frequently asked questions

What are the main improvements of Hailuo 03 over Hailuo 2.3?
Hailuo 03 is fully multimodal with native audio, supports up to 15-second generations, and can process up to 12 image references simultaneously. It shows significant improvements in dialogue, lip-sync quality, and overall video output after nine months of development.
How good is the dialogue and lip-sync quality in Hailuo H3?
Dialogue and lip-sync are described as surprisingly strong in the review. The model supports multilingual dialogue and performs well in cinematic applications, setting it apart from competitors.
How does Hailuo 03 pricing compare to Seedance 2.0?
Hailuo H3 costs around 50% less than Seedance 2.0, making it a more economical option for video generation.
Can Hailuo H3 use multiple character references in a single video?
Yes, the Omni mode can process up to 12 image references simultaneously and can combine multiple character references within the same video, enabling team-ups and complex scenes.

Topics