Minimax Hailuo 02.3 | LLM Dance Off: How Good Are AI Video Models at Dancing?
Testing Hailuo AI with deliberately simple prompts reveals whether the model was trained on real dance footage, while also exposing serious face distortion in its fast mode.
The fastest way to tell whether an AI video model actually understands dance is to strip the prompt down to almost nothing and see what comes out. If a bare-bones prompt produces convincing choreography, the model was trained on real dance footage. If it doesn't, no amount of clever prompting will fix that underlying gap. This hands-on test puts Hailuo AI (Minimax's Hailuo 02.3, referred to throughout as "Halo") through exactly that test.
Setting up the comparison
The test starts with an uploaded image of a National Guardsman and a simple goal: get the model to dance. Hailuo's interface offers a choice between its own model and Google's model, along with a fast versus standard option; the fast option costs 15 tokens versus 25 for standard. The general pattern with fast models is that they're cheaper and quicker but lower quality, and since upscaling afterward with a tool like Topaz is standard practice anyway, starting with the fast, cheaper option and upscaling later is usually the better value than paying more for a model's built-in upscaling, which doesn't come close to what a dedicated tool like Topaz provides.
Testing with deliberately simple prompts
Rather than writing an elaborate, paragraph-long prompt, the test deliberately keeps prompts minimal, just enough to specify a hip-hop dance and change the subject count from three girls to three men. The logic is direct: if a bare-bones prompt still produces high-quality dance movement, that's strong evidence the model was actually trained on dance footage rather than just following literal instructions without any real movement fluency underneath.
What the results showed
The initial results were encouraging. The dance quality looked better than most models typically produce, suggesting Hailuo has been trained on at least some real dance data. A more specific request for a "crip walk" produced a dance that wasn't actually a crip walk, more of an Irish step-dance pattern, but the movement itself still looked clean and well-formed. The takeaway was that the model appears to have been trained on several kinds of dance without covering every style; it doesn't recognize a crip walk specifically, but its general dance training is solid enough that even an unmatched style still comes out looking like real, fluid movement rather than a garbled mess.
The preset saving feature
One interface detail stood out as genuinely useful: an "add to presets" button that saves a prompt after generation so it can be reused later, named and stored for future runs. That matters specifically for dance work, where landing on a short choreography sequence you like is valuable enough to want to reuse without rewriting the prompt from scratch. It also changes how prompts get written in the first place: knowing a good prompt might get saved as a reusable preset pushes toward writing prompts in a more generic way (using "person one, person two, person three" rather than baking in specific details like "National Guardsman") so the saved preset can be dropped onto any reference image later.
Cost and resolution tradeoffs
Pricing tiers for Hailuo run from roughly $7 to $30 a month, with $7 unlocking basic access and higher tiers required to reach longer clip lengths, including a jump to 10-second videos. Resolution has a real cost impact too: switching from 768p to 1080p more than tripled the token cost for the same generation. The video holds off on upgrading to the higher tier until testing the cheaper option more fully, treating that decision explicitly as one to make only after seeing enough results to judge whether the higher tiers are worth it.
The face distortion problem
The clearest weakness surfaced late in the test: reviewing the fast-model output closely showed the middle guardsman's face badly warped and distorted, a problem visibly worse than in the standard model. This doesn't appear to be simply a byproduct of lower-resolution upscaling, it looks like the fast model renders faces differently and worse. The practical conclusion is that fast mode may be fine for simpler, more stylized content, like a line drawing or cartoon, but it's not reliable for anything involving a recognizable human or humanlike face.
Key takeaways
- Testing with deliberately simple prompts is a fast way to check whether a video model was actually trained on real dance footage.
- Hailuo's dance output looked strong even with minimal prompts, suggesting real training data behind the movement.
- The model handled some dance styles well but didn't recognize a specific move like the crip walk, producing a different but still fluid dance instead.
- The "add to presets" feature lets a good prompt be saved and reused, which also encourages writing more generic, reusable prompts.
- Switching from 768p to 1080p resolution more than tripled the token cost of a generation.
- The fast model produced significant face distortion on human subjects, a tradeoff not worth it for anything beyond simple, stylized content.
Who this is for
This hands-on evaluation from Humanitarians AI is aimed at educators, content creators, and anyone exploring AI video generation tools who wants a realistic sense of a model's strengths and weaknesses, dance training quality, cost tiers, and rendering tradeoffs, before spending real budget on generation credits.
Full transcript(auto-generated, with timestamps)
[0:01]Okay, Professor Bear here again. I'm giving this Halo, halo, however you say it, AI another try. Uh, I create an image. So, I I want the National Guardsman to dance. What I'm trying to figure out here is it the model. So, here's their interface. You can select different models, including uh Google's model. So, you can use their model or Google's model. I right now don't know you know the difference between these models. Typically a fast model is just faster and often cheaper. Uh we'll check that. So this costs 25 tokens. If I go to fast it goes to 15 tokens. My experience with fast models
[0:48]Is I upscale anyway. So since I use a tool like Topaz to upscale the fast is usually better. So, we'll try it because it's just cheaper. I can get more for my tokens. Um, usually the non-fast models will do a little bit of upscaling for you, but not nearly what a tool like Topass does. So, you're better off just using an upscaler and, you know, saving your money by uh paying for that upscaler by just doing this. So, I'm going to try it both ways. So, they have uh a little tool here. They'll rewrite my prompt for me. [clears throat and cough] I'm leave that on for now. You have a very
[1:30]Complicated prompt that you've used other tools to help you with like Jet or Gemini, you probably don't want this, but they have their own little AI to basically optimize the prop. You can turn it off or turn it on. They have camera controls. I'll I'll talk about that in a little um a different one. They have a library something or other. Uh but that's not really what I want here. See if I can find it again. I want this. So they have prompts. This is encouraging to me. Uh because if it's prompts doing dance and the dance is good, it means their model is trained for dance, which is what I'm
[2:11]Hoping for. What I'm hoping is their model is trained to do dance. I can prompt. So I can always come up as sophisticated a prompt as possible, but I'm always fighting a model that can't dance. And so if their model is trained on dance, it makes the whole experience when you get a better quality thing. So let's see. I'm going to try hip-hop here. I want to use my So they replaced it, which is irritating. Uh they they replaced my image just automatically. But the prompt is very simple. I'm obviously going to have to change the prompt. So, I'm just going to do this. Obviously, I want to change it from
[2:53]Three girls to three uh three men. I'm going to keep it super simple. Normally, I would the prompt might be a paragraph long, but what I'm trying to test here is with a just a super simple prompt like this, what is the quality of the dance? Uh, if the quality is very high with a super simple prompt, it means their model has been trained to dance. Um, which is what I'm hoping for. So, let's see. I'm also going to do it with the uh regular one just to see the difference. 10 tokens more. If I'm using VO3, I'm just using VO3. What? So, I'm not
[3:47]Really testing that here. Halo 2, I just assume, is not as fancy as Halo 3. Halo 1, I don't really know what that is. Um, but what I'm going to do now is just go to a tool like Chat GPT. So, here here's a more sophisticated prompt. What I'd eventually do if this is a good tool is run a chat GPT plugin gel ni plugins which would allow incorporate a lot of information about dance into those GPTs and create a bunch of commands for styles of dance. So I can just use you know a a command like crypark and it'll create you know a lot of details of the
[4:42]Way I want to do it. But I'm not going to do that now. We're testing Halo. And I'm going to test this one on both models as well. Looks like we got some of our at least our first result. It's okay. [clears throat] I have to I'm going to upscale it. I'm going to put the upscaled ones at the end of this video because they're a little hard to see here. So, I'm going to upscale it and just add it add it. It shouldn't have added the watermark because I Oh, I guess I did a $7 a month plan. So, that should remove the watermark.
[5:33]That'll be kind of pissed if it doesn't, but we'll see. um to what I'm going to do after this video is I'm just going to add sort of all the generations. And what's nice here [clears throat] is they actually estimate which this is the first time I've really seen that. They're sort of giving you an estimation of you know when this is going to finish. Normally all the ones just do this. They go around and then at 95% they stay at 95% for 5 minutes. um after getting to 95% in a minute. Um we'll see if that's accurate or not. [clears throat] But so far it looks
[6:17]Intra. So I'm not going to force you to watch the generation of video number two and the more interesting one. These are from the more complicated prompts and I'm asking specifically for a crypop. Um, so we'll see if it can do that, if it understands what that means. So in order for a model to do a cripw walk, it would have to be trained with the tag walk with some examples, many examples of dancers doing walks. And if it has that, that'd be great. So, I'm liking what I'm seeing so far because the dance looks pretty good. It looks better than I think most models do.
[7:06]Dance. Again, I will upscale this and look at it in more detail. They are short. Can I do the length longer? 10 seconds. So, I have to upgrade in order to get to from the $7 a month one to get to 10 seconds. If these are good, I I would probably do that. So, in order to get a 10second dance video, it looks like I spent $7 a month on this thing. Uh, I might have to upgrade to Pro. I don't know what pro costs. Uh, but before I would do that, I I would look at these things in detail. Let me just see what that would cost me.
[7:56]It's bro pro big job. So, we're talking about $8 to $30 a month to be able to do a 10-second video. So, I'm going to hold off on that until I've uh looked at these. And what I'm really trying to evaluate here is is this model Halo 238 trained on dance? So I asked for a walk. This is not a one. This is some sort of Irish dance step thing. I like it, but it's not a one. Uh but you know, it's the dance looks nice. if you're not that that particular about and a cripwalk, I admit is not sort of the most common dance.
[8:45]And so, you know, it it doesn't know cripwalk. Um, but it looks good. So, what I'm going to do is I'm going to download these. I'm going to upscale them and put them at the end. I get a debate on whether to go to the, you know, $30 a month thing after sort of doing this a bunch with the um $7 a month thing. Yeah, this is some sort of step dance. It's not what I asked for, but the dance looks good. So, my my sense is it's been trained on a few kinds of dance, not every kind of dance. It doesn't know what a cripwalk is, but the dancing is
[9:30]Looking good. But again, what I'm going to do is I'm gonna download these. I'm gonna upscale them so I can see them in a lot more detail. And then from that, but what I'm liking here is the dance is looking good if you're not super super particular about it really following your instructions to the letter. Um, and it looks like I'm looks like it's been trained on Nance, which is what I'm hoping for. So, I'll end this now and then add those and then continue my experience with uh experiments with Halo AI. Okay, bear here. I think I figured out where these things come from. The button
[10:10]Here called add to presets. So, basically, I create a prompt. This is a fairly complex prompt, but if it turns out I really like the prompt after it generates, then I believe I just hit add to presets. I give it a name like uh gripwalk dance. I would get this would be guard kryp dance because I think national guard are in there. So I would say guard crep plot and save it. And there you go. So uh this is kind of nice. I like that the fact that if you have find a prompt that you really like this also will affect the way I write my prompts because I really
[11:01]Think because there's an image here I don't really think I need to mention their guard. So I would make it more generic for just sort of a cryon dance that I can just pop in any image for. Uh, but because the image exists, I don't think I really need to tell it that much about what's in the image. More I could just say man one, man two, man three, or even person one, person two, person three. Uh, and because it has the reference image, it should just keep the reference image. So, this will affect the little way I write for this tool. Um because I like the fact that I
[11:41]Can basically if I have a prompt that I really like and this is particularly true with dance. You'll have you'll come up with a little you know 3se second choreography thing that you really like and then you want to save it. And so saving this this is a smart thing. It's sort of an obvious thing to have but this is the first time I've seen it in any of these platforms having just you know saving a a prompt In fact, I will chase this prompt. Uh, transform the three people. I don't think there's any other specific stuff. Then I'll just save this one here, I guess. Do I have to use it to
[12:31]Save? But I think I do. I think I have to generate it again to to then get a preset. So my belief is if I just render it again and I decided to go with fast and I'm going with 1080p. So this is more expensive now because I'm remember here we have our choice of revolution. I'm bringing the resolution up and so that more than tripled my cost but I think if I did not use the fast one it would be even more expensive. But yeah, so defaults. This is a little irritating that it just sort of switches everything. Yeah. So this is 80 versus
[13:08]What would this be? 50 80 versus that's a big difference. I I don't think the I think if you have an upscale air, the the fast stuff is fine, but the 1080p versus the the 768 or whatever it is will make a difference. So, it's likely all of mine are going to be it would I need to upgrade to get to 10 seconds and that would depend on whether it can actually generate 10 seconds of good video. It's very often in these tools that they can generate even though they allow to generate 10 seconds after about 5 seconds it just gets wild and weird. So, I'll have but I
[13:48]Can't try it. So, this is a little irritating that I can't, you know, try it 10 times at 10 seconds to see if I want to go from $7 to $27 a month. Um, but for for the moment I'm just going to play with the tool for a bit. I do like a lot of things about it. It does seem that the model has been trained on dance, which is the number one thing for me. Uh, the other thing is they have a lot of nice little interface things, but I think basically these little templates that they share, which is when people publicly share their
[14:19]Phone. So like a lot of tools, you can keep it private or public. I like mid journey which forces everything to be public. It's nice having the your choice to be public or private. And so I'm going to run this and see if I can save the prompt and add it to the presets. Hopefully that'll pop up again. Add to presets after it finishes. But I like I like the fact that um I can say that but it also affects the way I write it. So, if I know that I I might save it if it turns out great, I'm going to write it in a more generic way, not
[15:00]Exactly time to this. So, that's it. Uh, I'll continue playing with this tool probably all day. Um, because I do think there's a good chance that, you know, it'll do dance really well. Okay. Thank you. Uh, this is Professor Veran. This guy's face here. Let's go back to the take a look at the middle guard's face. uh focus on his face. It's badly distorted in the fast model. So, going forward, I'm probably not going to use the fast model. It looks like it's not just upscaling, but sort of how it's rendering it is different. I mean, let's go back again to his face here with a fast model.
[15:45]Focus on his face. Completely warped and distorted. So, I'm probably not with a Halo model going to use the fast anymore. That's it's not it doesn't seem to be just upscale. It's how it's rendering it is different. So, maybe if you have a simple enough, you know, almost like a line drawn or kids cartoon or something, the fast is fine. But if you're looking at stuff like this where it's a human being's face, or not a human being, but an AI being's face, then look at his face. Completely distorted. So, going forward with the Halo thing, I like the Halo thing so far, but I I'm not going to use
[16:25]The finess moment going forward. Okay, that's it. Take care. >> I was just a shy with nothing to [singing] say. Watch her walk by now. She's far away. Her smile like the sweetest ice cream. A taste of rare things. >> >> Back in Boston and I'm thinking about time. >> Wondering if still hope praying is mine. Loveliness extreme. Sweet as [singing] ice cream. >> She's now lost but I still dream. >> Praying that she'll come back in my [singing] dreams. Oh, love is sweet as ice cream. >> She's now, >> but I still dream. [singing] >> Loiness extreme, sweetest ice cream.
[17:28]Praying that she'll come [singing] back in my dreams. I left. I lay her face in my mind. Wondering if maybe I left my chance behind. [music and singing] Oh, lovely extreme. >> Sweet as I [singing] dream. >> She's now, >> but I still dream. >> Oh, praying that she'll come back in [singing] my dreams. Her eyes, her grace, >> beauty. [singing] Supreme. Miss the moment. I hold the thing. >>
More videos
2:08Bridging the Pixel Gap in Browser Automation.
2:23How One Narrow Safety Rule Can Make an AI Less Safe Everywhere Else.
2:04Why splitting a chunk from its document makes it retrieve for the wrong question
4:20Three You Can Take Back. One You Can't.
2:21Why a 50-turn agent pays for the same screenshot 35 times unless it caches the pixels
1:53