Oz Songs 11.10.25: Planning Character Dance Sequences & AI Video Compositing Techniques
In the first real episode of the Oz Songs journal, Professor Bear studies how AI filmmakers hide the 5-second clip limit with editing, then plans dance sequences for Glinda, the Scarecrow, and more.
Adapting all fourteen of L. Frank Baum's Oz books into musical videos over the course of a year is an enormous undertaking, and this first real episode of the Oz Songs video journal starts by confronting the single biggest technical constraint standing in the way: most AI video tools can only reliably produce about five seconds of usable footage per clip.
Learning from a professional's workaround
The episode opens by studying a polished demo reel from a professional video editor and cinematographer, work good enough to pass as a commercial. What stands out isn't the individual clips, since even professional AI video tools mostly cap out around five to ten seconds of usable footage per shot. What stands out is how the cuts are handled. The camera angles and edits are built to forward the story rather than just hide a technical limitation, so the viewer doesn't consciously notice the cuts at all, and the cuts arguably push the narrative forward rather than interrupting it.
That observation led to some direct research into the actual rules professional cinematographers use to make cuts feel invisible, like alternating between a behind shot and a front view. Those rules got fed into an editing tool built to suggest which cuts might work well for a given sequence. That tool won't be used for the specific project described in this episode, but studying it clarified something important: professionals dealing with AI video's clip-length limits lean on camera work and editing as visual storytelling tools in their own right, not just as a technical patch.
Why dance is a hard case for AI video
The next challenge is more specific to this project: most AI video models are notably weak at rendering dance convincingly. Getting good results from any AI model generally requires training data specific to the task, and dance motion is complex enough that general-purpose video models tend to fall apart on it. Some newer models have been trained specifically on large amounts of dance footage, and the results from those specialized models look noticeably better, though it's worth noting that showcased demo clips are typically the best examples a company has, not a representative average.
A deliberate change in working style
Rather than running many creative projects in parallel, the usual pattern across a typical day, this series is meant to focus on finishing one thing before starting the next. That's a deliberate shift for this particular project, chosen specifically to keep the daily Oz Songs journal moving forward without getting split across too many simultaneous efforts.
The compositing plan for dance sequences
The creative plan sketched out here includes several specific dance sequences: Glinda belly dancing inside her castle, the Scarecrow, designed as a woman in this adaptation, dancing in a cornfield, the Wicked Witch of the West practicing ballet, a K-pop-style sequence featuring the Cowardly Lion, the Tin Man, and Dorothy and Toto, and a duet between the Tin Man and Dorothy.
Making any of these work well requires generating each character and background element separately rather than trying to prompt an entire complex scene at once. A single prompt asking for a specific Scarecrow in a specific cornfield in a specific version of Oz reliably produces mediocre, inconsistent results. Instead, the plan is to generate each element, a character in a specific pose, a background, separately, then composite them together into a single coherent image that becomes the starting frame for animation. Getting that starting image right is treated as the single most important factor in getting a good final video, more important than the animation prompt itself.
The tool: Weavy
The tool chosen for this compositing work is Weavy, which provides access to a large number of underlying AI models, reportedly around 30, through one interface, charging a small markup above the raw token cost of whichever model is used. That structure means trying different models for different needs doesn't require separately signing up for and paying each one individually. Weavy is also apparently in the process of joining Figma, which would put a wide range of generative model access inside a design tool many creators already use.
Key takeaways
- Most AI video tools reliably produce only about five seconds of usable footage per clip, so professional AI filmmakers use cinematographer-style cuts to hide that limit and keep the story moving.
- Dance is a particularly hard case for general AI video models, and models specifically trained on dance footage produce noticeably better results.
- The working plan for this project is to finish one thing before starting the next, rather than running many projects in parallel.
- Complex scenes work better when characters and backgrounds are generated separately and composited together, rather than prompted as one image.
- Weavy provides access to roughly 30 underlying AI models through a single interface with pay-per-use pricing.
Who this is for
Oz Songs is part of Humanitarians AI's Lyrical Literacy initiative, which combines AI-generated music and storytelling to support cognitive development and literacy skills by engaging multiple brain regions at once through musical storytelling. This episode is aimed at anyone curious about the practical, unglamorous planning work behind an AI-assisted animated music series, not just the finished results.
Full transcript(auto-generated, with timestamps)
[0:00]Hey, it's Professor Bear here. Um, this is the first real episode of the Ozong series and the Osong series is basically any sort of video journaling every day things that I work on and sharing you. So, I was going to do this. Um, let me tell you what I was going to do. So these people uh Marco Arminetti uh it looks like he is a professional like video editor, cinematographer, something like that. Good work. But what's interesting to me, let's take a quick look at this here. What's interesting to me about this video is the way AI works, AI video currently works is you can maybe make 5 seconds of
[0:43]Good video per clip. I mean, there's things that run up to 10 seconds, but usually you get maybe 5 seconds if usable video per clip. And this is a really nice little um video here. It looks definitely professional. This could be a commercial. He admits, well, not admits, he says his process. And in the process, he talks about all the tools, Hicksfield and Freep and Topaz and Clling and D. Um, but what's interesting to me about this this video here is the cuts. So, the big issue, sort of the fundamental issue of of a lot of AI video is you can only make about 5 seconds of good video per clip.
[1:25]And what he's done, which I suspect most like professional people do, is the camera angles and the cuts are forwarding the story. That is, you don't really notice the cuts. the Let me go back and play it again. You don't really notice the cuts and the cuts in fact probably pushed the story forward a bit. And so that got me intrigued. So I wrote a little uh I did some research. So I had Google Deep Mind um or Deep Research do a bunch of research on well what are the rules that cinematographers use to do this? There's a bunch. So I found them and got a bunch
[2:06]Of rules and I interpreted them some into some AI which then will help me suggest edits like we see it's from the behind and then it's front view. And so what my little tool does is it sort of suggests what edits might be good to for the story. However, I'm not we will do this and we will do this with Dorothy and Toto sort of skating through Ozam the yellow big road. Uh, but we won't do this for what I'm about to do next. But I did want to point this out because, you know, when when you make AI video, what these people are doing who are
[2:44]Clearly professionals is they're using the camera to sort of deal with the limitations of the AI and put forward the story. So, it's sort of visual storytelling by making the camera work. Let me explain what I'm going to do. And this also sort of explains the process. So I used to dance quite a bit. I still dance a bit. And a lot of these AI models, for an AI model to be good, it needs to be trained on something specific. So most AI models that do video are terrible at dance. Halo AI, it looks like they did something smart, which is pretty much what everybody is
[3:21]Doing, is they trained their models on a lot of dance sequences. And when you do that with any kind of AI, you make it better at uh and these look pretty good. Of course, the usually the ones they show, so they're they're probably picking the best of what they do. But this will talk about my process and this will talk about what we do. So, what I do not want to do, which is what I typically do during the day, is I'll do 10 different things in parallel. I'll do the skateboarding project. I'll do a dance project. I'll teach two classes. So I'll do some voice stuff, do some
[3:58]Songs, write some songs, write some poems, bunch of stuff in parallel, but I don't really want to do that in this series. What I want to do is I want to do something, finish it, then do something else, finish it. So of the things that sort of popped in my newsfeed, this is the one I think we're going to go with in the the next one. So, what I'm going to do the first step of doing any of this stuff here for me, I because this is fair sample, I could probably just take their sample and just redo it. But I don't want to do that. What I want to do is
[4:30]For example, this one belly dancing. I think I'm going to have Glenda belly dance in her Glenda castle. And to do that, I need to create a shot. So, I need to create I'm going to start by creating Glenda in a shot like this, which means I need to design Glenda. I need to design her outfit and what she's looking at. I need to design her castle. So, here I'm going to do another dance sequence. They're going to be the four, you know, the probably I'm going to have uh uh Cowardly Lion up front and Tin Man and then Dorothy and Toto in the back
[5:04]And do a little K-pop sequence. I think I'm gonna have to do do the scarecrow. My scarecrow is a woman, by the way. Um, she's going to probably do a dance like that in front of a cornfield. I think I can have the wicked witch of the west, you know, practice her balet. Uh, another K-pop thing. I think this will probably be with maybe all the witches, Flenda and the Wicked Witch, and, you know, any other witch, Princess Asthma, that kind of thing. Then I think I can have a do a duet here a duet with a character I'm still working on which is the Tin Man and Dorothy. So in order to
[5:43]Do this I can right now I could just This is a free tool by the way and well it's free for now. Basically they allow you just to sign up and you know have a limited amount for free. So you could just use it for free is what I'll do. I don't think I can even sign up. I didn't see a pricing thing here. I think they're one of the [snorts] one of the those where they get a bunch of funding and sort of get traction by making it free for a couple months. But it looks like what they've done is train a bunch of models for very specific
[6:13]Purposes, which is a good idea in AI. Uh, as you'll see throughout this series, I use a lot of different tools. And speaking of those different tools, what what is sort of a bad idea is try to get a prompt which can recreate my version of the scatcrow in my version of uh Hornfield and Oz. What I need to do is I need to generate the se separate elements. For example, I need to generate Glenda in the pose her her castle behind her and then composite those images. I need to start with an image. There's no way. And I'm really good at this stuff. There's no way I can
[6:54]Prompt to get my scarecrow, my Dorothy, my cowardly line all together looking at that. Instead, what I have to do is I have to create each one in that pose and compositing. Same thing with all of this stuff. The your starting image is critical for your video to get it right. And if all you do is sort of prompt, you're going to you're going to have stuff which you just don't like. Maybe maybe it gets the scarecrow right, but it doesn't get the cornfield right. And you can spend a lot of time trying to do that and probably with very mediocre results. Instead, we are going to use a tool called Weevy.
[7:34]Weeave is a great tool. And basically what it is, it gives you access to many many models. You can see sort of all the models. It gives you access to what basically you do. you pay and then they charge a little bit above the token cost and then you can use you know pretty much any model which has access and I think they're like 30 models if we look at the pricing for weevi I think they talk about it where here's all the models you have access to and this is nice because allows you to try different things and sort of the nature of AI just like what
[8:09]We're doing with the dance thing so hopefully the dance thing the Halo I don't think Halo was one of theirs here. No, it is actually Halo's here. So, maybe I can actually use the the dance thing directly through Weevi rather than going to their website, but maybe I have access to the dance model. Maybe I don't. Um, I haven't tried it with the dance model yet. What's great about Weevi is the interface is fantastic and you can try a lot of different models and just sort of pay for one thing. You don't have to sign up for all of these things. And and more importantly, it
[8:46]Gives you a tremendous amount of control. And and what I mean by that is what we're going to start doing, not in this video, but in the next video, is start compositing. So I can take an image of sort of the the colors of my house and use that as input. I can take my my image of a my Dorothy, use that as input. Then I can take my image of my Toto and use that as input. I can take another image of the background, use that input and care it together. Something that looks great. So, that's what we're going to do in the next one.
[9:19]I'm going to end this here. Um, but basically what we're going to do is we're going to create the starting images with Weebi um and other tools to allow us to generate just those images. And how it's going to work is, you know, for example, I'll have the scarecrow, my scarecrow, and then I'll create a background for my version of a cornfield in Os. And then once I have that image, and I can have it pretty close to this starting image, then uh I'll see if this tool works. It looks great, but the way you sort of figure out whether anything actually works is to try it. My assumption is the
[10:01]Reason why they put these up here is they're sort of the best of what they got. Why else would they they're not going to put like the things that didn't work on their website. Nevertheless, it looks great and it makes sense. And what makes sense about it is when a model is trained for a very specific purpose, and this looks like it was trained specifically for dance, it's going to do a much better job on that specific purpose. And that's why tools like Weebi are sort of going. I think Figma actually bought Weebi. I think it says something here about it's joining Figma there, which
[10:36]Means it may just be part of Figma. I don't use Figma, but I would probably start using Figma that had Wei in it. Um, is tools like these where you basically have a tool which is a great interface which allows you to access a bajillion other tools. That's definitely the way things are going. And then the way Halo will make their money is by charging for access to uh their model. Um and then what's good about that is it'll start getting people much more diversified models. What the what they're doing here makes a ton of sense. that is they they thought dance was a big market and it's probably a
[11:17]Smart move because a lot of people make bajillion sorry Tik Tok videos that it was worth the expensive training model specifically to do dance and it's very likely I haven't tried this yet but it's very likely it's going to be good at dance because it was trained to do dance if I try to use a dance model to do something else I probably won't get great results If I use a dance model to do dance, these look good. These look good. And so we'll see. So next up is I I need to create the images. And we're going to use a tool called Weeavy to do them. And
[11:56]The reason why we use Weevi to do that is creating an image is really sort of a composite of a bunch of images. If I tried to create this sort of straight, that's not that complicated image. I might be able to do that straight, but an image like this where I want Glint's castle to be the way I envision Glenn Glenn's castle. I want Glenda to look exactly the way I want Glenn to work. I want the scarecrow to look exactly the way the scarecrow should work. I want the cornfield to look like an Oz cornfield. I I don't see another way of doing that.
[12:29]I I don't think you can do that just from a simple prompt. I think is a combination of a bunch of prompts and a bunch of input which then you composite together to to get something great. So next on I don't want to keep these two too long. So this is our initial focus. Our initial focus is sort of recreating uh these videos here these six videos for us. Okay. Um I'm excited about this project. Take care.
More videos
2:08Bridging the Pixel Gap in Browser Automation.
2:23How One Narrow Safety Rule Can Make an AI Less Safe Everywhere Else.
2:04Why splitting a chunk from its document makes it retrieve for the wrong question
4:20Three You Can Take Back. One You Can't.
2:21Why a 50-turn agent pays for the same screenshot 35 times unless it caches the pixels
1:53