Aengus | Building a Personalized AI Tool for Music Video | Text-to-Video Prompt Engineering
Professor Nik Bear Brown builds a custom GPT trained on his own handwritten prompts and art, then tests the text-to-video prompts it generates for a Lyrical Literacy music video across multiple AI video platforms.
Most people who use AI tools for creative work take whatever the tool suggests and rewrite it into their own voice. This project asks a more interesting question: what happens if you build a tool that already writes in your voice, because it was trained on nothing but your own work? Professor Nik Bear Brown walks through building exactly that kind of custom GPT, and uses it, unedited, to generate a full music video.
Testing the tool with no manual rewriting
Normally, even when using a tool that suggests prompts, the habit is to try a prompt, evaluate it, and hand-rewrite it a few times before settling on a final version. For this particular experiment, that step was deliberately skipped. The prompt came straight out of the custom tool with no editing, and that unedited prompt was what generated the finished music video shown at the start, a holiday-themed piece built around the lyrics "over the river and through the wood to grandmother's house we go." Showing the result first, before explaining how the tool was built, was a deliberate choice to let the output speak for the process.
Why personalization changes what a custom GPT can do
The tool itself is built as a custom GPT plugin, and its structure reveals the actual mechanism behind why it works. Building a tool like this starts with uploading a large amount of your own material, in this case, a file full of actual handwritten prompts written over time. That step matters more than it sounds. Without genuine writing samples, a custom GPT tends to produce generic output, the same way a tool asked to write professional emails "in your style" without any real examples of your writing will default to a bland, average style rather than anything distinctive. Uploading real samples, songs actually written, prompts actually used, general reference material on relevant topics like dance, is what lets the model align its output with a specific creative voice rather than a generic one.
The same principle extends to visual style. AI-assisted art in this workflow starts from an actual hand-drawn piece, which then gets refined through prompting until it settles into a consistent visual style, tracked using what are described as P-codes to help maintain that consistency across generated images. The underlying theme across both text and image generation is personalization: the value is not in using a generic tool like ChatGPT on its own, but in combining a specific creative style with the tool's capabilities to produce something neither could produce alone.
From lyrics to video prompts
The practical workflow for this specific tool is to paste in song lyrics, paste in reference images, and have the tool suggest text-to-video prompts that connect the images and lyrics while matching an established creative style. The song used in this test was written for an artist and is part of the Lyrical Literacy project, an initiative from Humanitarians AI focused on getting people singing as a way to exercise the brain, based on the idea that biological systems generally follow a use-it-or-lose-it pattern, so activities that only exercise logic, like puzzles, leave other brain functions underdeveloped, while music engages far more of the brain at once. The tool also includes a small script that batches uploaded images into groups of ten, since many AI video and image tools cap how many reference images they can accept in a single batch.
Comparing across AI video platforms
Once prompts were generated, the same prompt and image set were tested across multiple AI video generation platforms, including Seedance and Higgs Field, with plans to extend testing to other available tools as well. This comparison surfaced real inconsistencies: the same prompt produced noticeably different results depending on the platform, and in at least one case, a platform introduced an unrelated figure into the generated video that had no basis in the prompt, prompting further investigation into whether that behavior was tied to the specific prompt wording or to that particular platform. Testing the same prompts and images across many models is treated as a separate, ongoing project in its own right, aimed at understanding which platforms handle which kinds of creative direction well.
Why building your own tool matters more than using this one
The broader point made throughout this walkthrough is that the specific tool being demonstrated is less important than the underlying approach. A tool built from someone else's writing, art, and prompt history will always produce output closer to their creative voice than to yours. The recommendation is to build a personalized version using your own material rather than relying on someone else's tool, since general-purpose AI systems do not have access to hundreds of your own prompts, samples, or stylistic choices unless you deliberately provide them.
Key takeaways
- A custom GPT trained on real handwritten prompts and creative samples produces output much closer to a specific creative voice than a generic AI tool.
- The demonstrated music video used an AI-generated prompt with zero manual rewriting, straight out of the custom tool.
- Personalization applies to both text prompts and visual art, with P-codes used to help maintain visual consistency across generated images.
- The same prompt and images were tested across multiple AI video platforms, including Seedance and Higgs Field, revealing inconsistent results between them.
- The underlying song is part of the Lyrical Literacy project, which uses singing as a whole-brain exercise rather than just entertainment.
Try it yourself
Anyone doing regular creative work with AI tools can apply the same approach: build a custom GPT or similar tool using your own writing samples, prompts, and reference material rather than relying on a generic one. This project supports the Lyrical Literacy initiative from Humanitarians AI, which uses music and singing to engage learners who struggle with more traditional teaching methods.
Full transcript(auto-generated, with timestamps)
[0:02]Okay, Professor Bear here. Uh, this is a video I generated with a custom GPT, which I'll show you how to write right after I show you the video. I decided to show you the video first and then we'll go into how to make the little tool that does this. Normally, even if I use a tool that's helping me suggest prompts, I'll try the proy. If I like it, then hand rewrite it and maybe do that a few times. For this particular experiment, I just took the prompt straight. No changing the prompt. Basically just took the the prompt straight out of the tool and created the video you're about to
[0:36]See uh in a couple seconds that after that video you'll see my discussion of how I made the tool that made the video. Over the river and through the wood to grandmother's house we go. The sleigh is packed with gifts and cheer. Christmas lights are a glow over the river and through the wood. The carols and songs we hear, the meloies ring as the joy they bring fills hearts with Christmas cheer.
[1:50]Over the river and past the trees, the starry sky shines bright. The warmth inside and the you tie. Make this a holy night. Over the river and through the snow. Harley's on the door. We gather around where the joy abounds with Christmas love in store. Over the river and to the fire, where stockings hang with care,
[3:02]With stories to tell and the midnight bell, the Christmas spirits there over the river And through the snow, the Christmas tree stands tall. Its twinkling lights in the frosty night bring joy to one and all. Over the river, the sleigh bells ring. Their music fills the air with laughter and cheer. We draw ever near to Christmas memories fair. Over the river the church bells chime, proclaiming peace. Tonight we lift up our song as we ride along
[4:14]To greet the holy light. Okay, Professor Bay here. I just wrote a little tool using uh CHPT plugin. I'll probably make Gemini and other things. Let me show you the structure of it a little bit and why these things are important. So, whenever I write a tool, I upload a lot of my own information. And so the purpose of this tool is to paste in some song lyrics and paste in some images and had it suggest uh text video prompts with the images and input that relate to the song and also fit my style. So this prompts file here has a lot of my actual
[5:14]Handwritten prompts. Why this is important for a lot of the language models is if you don't do that, if you don't give examples like for example, you could easily write a little custom GBT which wrote professional emails in your style rather than generically just by uploading a lot of samples of your writing. So these are songs that I've written and this dance thing is just generic information about dance. What is you know cripwalk? What is this? What is that? But the other files are specifically things that I wrote also true for the art. So whenever I make sort of art which AI has assisted this starts with my own sort of
[5:59]Handdrawing that's uploaded then the prompting further refineses it until it creates a style like this one. Then once this is using something called the pode. So um what it can do is the pick out along with the prompting can create the stuff and this is a very big theme in AI that big theme in AI is personalization. It's not just to have it written like chachi pt but you with chachi pt to create something better than either of you could create on your own. So what I'm going to use this for so I haven't tested this yet. I'm about to test it. Uh, and the way it works is this. I I
[6:46]Put in some song lyrics. So, this is for an artist that I write for called Mayfield King. And it's particularly for this song here. So, this song is part of uh a a project called Lyrical Liy, which is humanitarian AI. That's about us to get people singing. When you get them with singing, you can upload your own voice and have these models sing like you. It'll typically improve your voice. You can think of it like uploading a photo of you and making you 10 years younger. Well, this is what these miles can also do for your voice. But still you, it's you 10 years younger.
[7:25]And so what we do is there's a lot of evidence that music is important for training your brain. So pretty much the way all biological systems work is use it or lose it. So if all you did was bicep curls and worked on your upper arms did nothing else everything else in your body would be weak but your arms would be strong. Same for your brain doing sudo sudoku and logic puzzles is great but all of you do is logic you're not using many parts of your brain. Music turns out to stimulate many many areas of your brain. So you could basically think of this brain exercise.
[8:01]It's a good thing to do whether you're a kid or whether you're adult. And so we we have a project called lyrical literacy which encourages people just to do simple singing and tongue twisters and things like that because it's simply good. It's also a way to get people engaged in learning. We've worked with a lot of kids who sort of struggle with traditional ways of learning reading and other things and find they can get great results by doing sort of semi things with song. So this is this was a a a song contributed to that project. I'll play it at the end. So but I need to
[8:38]Make videos for it. And so this is what this tool does here. So we we have I have a command code music. If I made this for other people, I would probably give it a you know a more sensible name. Uh what I did is I have a little script which is bends my images into groups of 10 because 10 is the most that tools like chash boutique and day they can take images but not 11 images. And then I just paste in the lyrics and hit go. And then I can use this in any any uh model that can generate clicks. In fact, another thing I'm going to do
[9:22]Is uh so what I can do is just take the image uh this is called um hexfield but through hexfield I'm using a tool called cance. So what we can do is we can just upload the image, put in the prompt and hit go. I'm also going to do this in pretty much every other tool I have access to for Grock or whatever. So let me just go to imagine upload the font. it will let it it will automatically start making some assumptions but if I want to give it something I just paste it in hit go. So another thing I'm going to do is just
[10:14]Compare the different models. See if this doesn't work on a particular model and sort of how the models work. You know maybe 20 30 major models to create the same video with the same prompts and the same images. Look at the difference. That's a separate project. So this project is just uh the main thing about this project is not so much this tool. this tool. You can go to the web page here and there's a link to the tool there in the web page if you want to use it. But it'd probably be better if you wrote a tool with your own writing in your own
[10:48]Stuff because then it's going to make it more like your stuff than my stuff. And so uh this again you you just I have a bunch of videos on how to do these things. But what makes sort of these GPTs and Gemini gems and other things special is I can upload my data which uh chat or claude or whatever don't tend to have. um they're they're sort of tracking me more and more so they know more about me every day but they they don't have you know 500 of my prompts. So by doing that I train the models to sort of act more like I would. Same
[11:33]Thing with the art, same thing with the music, same thing with everything else. So that's that's the idea. What I'm going to do is at the end of this video I am going to just put uh one of these. This may have run. So for example, this is one that just I got today. I generated don't think I hit generate. Now it's generating. This is the way one of them looks. This is with C dance using Higsville. So that did a little weird thing there. Same thing. Let's see what this one did. So what I'm going to do is just pick the best of them.
[12:08]This is irritating. So this is why I'm also evaluating the models. it like completely change the style by putting some random dude in there. We actually have a third. Let's see if it does that on every one of them. Yeah. So, there's something about the prompt. Let's see if if Grock does the same thing. This is a different prompt. So, um it didn't do that with this prompt. I guess we could see it with when this finishes. Uh, but there's something about this prompt in particular that I put a random dude in and that maybe just for a se dance. So, one thing about these these GPTs is
[12:54]You I wrote this in a half, but then as you see how it actually works, you're going to refine it and refine it and refine it. And I'm already thinking I have to refine it a little bit for the model that I'm going to use. See, but this is all seed damp stuff through Hexfield. So that didn't change the art style. So I'll go back to that prompt and see what is it about the prompt which forced it to completely change the our style and put a random dude in there. And is that a hagfield/c down only thing or is that you know is that a generic
[13:32]Thing? So this changed the style a little bit after about 3 seconds. And this may I haven't used C dance before sort of why I'm trying this. Uh let's see how consistent it keeps it. So that looks pretty consistent. It didn't change it. It changed this one at the end. It literally changed the art style at the end. And here just added some random dude in there for some reason. But that might be that particular pro. But anyway, after I compile all these things, I will one evaluate the tool. Is the tool useful? The other thing is it's also basis to compare these different models and there
[14:09]Are a bazillion of them and to see uh you know which ones are good for what purposes. Um, it there's a chance that that random dude was just that particular prompt. But here's this one. See what it did. Yeah, maybe something in that prompt which just thought it needed a random node in it. I'll look at the other ones and see if the other ones interpreted in the same way. But if that's the case, I'll go back and further refine the prompting and the uh instructions. But I don't really know that until I try. Okay, that's it.
More videos
2:08Bridging the Pixel Gap in Browser Automation.
2:23How One Narrow Safety Rule Can Make an AI Less Safe Everywhere Else.
2:04Why splitting a chunk from its document makes it retrieve for the wrong question
4:20Three You Can Take Back. One You Can't.
2:21Why a 50-turn agent pays for the same screenshot 35 times unless it caches the pixels
1:53