Hailuo AI Video Generation — Character Consistency Field Test of MiniMax Hailuo

MiniMax Hailuo 02's character consistency inside a 6-second clip depends on whether you use Start-End Frame mode: pure text-to-video multi-shot output has a hig

MiniMax Hailuo 02's character consistency inside a 6-second clip depends on whether you use Start-End Frame mode: pure text-to-video multi-shot output has a high face-drift rate, while locking the start frame with the same reference image makes character recognizability across three consecutive generations noticeably better. This is the single most important operational conclusion from testing the Hailuo line — consistency is not a built-in model capability, it is engineered on the input side. What Hailuo Is: MiniMax's Video Generation Product Line Hailuo is a video generation model from Chinese AI company MiniMax, first released as video-01 in September 2024, with Hailuo 02 following in June 2025. MiniMax was founded in 2021 and its investors include Alibaba and Tencent. "MiniMax released Hailuo 02 in June 2025, supporting native 1080p resolution and 10-second duration" (source: MiniMax official announcement) . Architecturally, Hailuo 02 uses a training architecture MiniMax calls NCR (Noise-aware Compute Redistribution). The company claims this architecture delivers "a 2.5x improvement in training and inference efficiency over the previous generation, with 3x the parameter count and 4x the training data" (source: MiniMax official announcement) . In practice this shows up in the plausibility of physics simulation — for high-dynamic scenes like gymnastic flips or splashing water, Hailuo 02's limb-breakdown rate is lower than most contemporaneous open-source models. The product line currently covers several paths: text-to-video (T2V), image-to-video (I2V), start-end frame mode, and the S2V-01 Subject Reference mode released in the second half of 2025. That last one is the heart of character consistency. Why Character Consistency Is Hard: Structural Limits of Diffusion Models Character drift in video generation models comes from the fact that every generation is an independent denoising process. The model does not "remember" what the character looked like in the previo

FAQ

What Hailuo Is: MiniMax's Video Generation Product Line

Hailuo is a video generation model from Chinese AI company MiniMax, first released as video-01 in September 2024, with Hailuo 02 following in June 2025. MiniMax was founded in 2021 and its investors include Alibaba and Tencent. "MiniMax released Hailuo 02 in June 2025, supporting native 1080p resolution and 10-second duration" (source: MiniMax official announcement) . Architecturally, Hailuo 02 uses a training architecture MiniMax calls NCR (Noise-aware Compute Redistribution). The company claim

Why Character Consistency Is Hard: Structural Limits of Diffusion Models

Character drift in video generation models comes from the fact that every generation is an independent denoising process. The model does not "remember" what the character looked like in the previous clip unless you supply an anchor on the input side. This is completely unlike an LLM's context memory — a video model's "memory" exists only inside the latent space of a single generation and resets to zero across generations. Three common drift types have to be handled separately: Intra-frame drift

Where Hailuo Sits in the Market

Hailuo's strengths are dynamic physics and cost, not character consistency. Compared with Google Veo 3, OpenAI Sora 2, and Runway Gen-4, Hailuo 02 ranks in the front tier for high-dynamic scenes, but for character stability across longer segments, Runway Gen-4's References feature and Veo 3's multi-shot consistency are more fully realized. On user scale, "MiniMax filed for a Hong Kong IPO in 2025 and disclosed global monthly active user figures for its products" (source: Reuters) . That scale gi

Related Guidebooks

Reviewed and verified by FeiYueh · Last verified 2026-09-16. Independently maintained — not AI-generated boilerplate.

← Back to Blog