Recreating Simon Meyer's Miniature VW Bus Commercial | With Commentary | GPT-5 & Veo 3

This is a direct comparison of GPT-5, Topaz, and Veo 3 recreating a viral tilt-shift miniature commercial, showing where each tool's consistency and physics hold up or fail.

2:08 video3 min readWatch on YouTube

Do you remember when making a commercial like this cost hundreds of thousands of dollars? Director Simon Meyer built an ultra-realistic miniature VW Bus commercial using tilt-shift photography and forced perspective, and shared the prompts publicly on Instagram. This is a direct test of what happens when you try to recreate that same commercial using only three tools: GPT-5, Topaz, and Veo 3.

The setup: recreating a viral tilt-shift commercial

The premise was simple to state and hard to execute: take Simon Meyer's published prompts for an oversized hand moving a 1960 Volkswagen Kombi Bus through a series of dramatic miniature scenes, an African savanna, a stormy coastal cliffside, an underwater coral canyon, a snow-covered alpine pass, a glowing nighttime cityscape, and see how close modern AI tools could get using only GPT-5 for images, Topaz for upscaling, and Veo 3 for video generation.

What GPT-5 gets right that Midjourney doesn't

The standout finding on the image side is consistency. GPT-5 is slow compared to alternatives, but it remembers that all of the images in a session are related, which keeps details consistent across shots, including something as specific as the correct size relationship between the oversized hand and the VW Bus itself. Midjourney doesn't hold that kind of cross-shot consistency on its own. There's a workaround: using GPT-5 images as reference inputs for Midjourney improves Midjourney's consistency noticeably. That combination, GPT-5 for consistent storyboarding, Midjourney for refinement using those images as references, became a takeaway worth carrying forward into future projects even when the GPT-5 images themselves aren't the final output.

Where Veo 3 falls short of a real production tool

Veo 3 came with real limitations. Despite an expensive Ultra subscription, generation was capped at five videos, a restriction that doesn't fit how a real production workflow needs to operate. Every output also carries a watermark, whether or not the creator wants one there. The overall read is that Google isn't designing Veo 3 for serious video production; it's being treated more like a toy for short-form content than a tool built for filmmakers.

What Veo 3 actually does well

That said, Veo 3 isn't without real strengths. Its physics simulation is genuinely strong, other tools in the comparison moved the VW Bus in ways that looked unnatural, something Veo 3 avoided. Its sound generation is also a strength, though it's likely not the best in class; Kling is probably better specifically at sound, even if Veo 3 holds the edge on physical realism.

The verdict on the current AI filmmaking toolchain

The overall conclusion is candid about tradeoffs rather than declaring a single winner. GPT-5 requires a Topaz upscaling pass to reach usable resolution, but it wins on consistency across a shot sequence. Veo 3 remains capped and watermarked in ways that limit its usefulness for serious production until Google changes how the tool is positioned. The plan going forward is to rebuild this same commercial spot using a wider set of tools, GPT-5, Midjourney, Higgsfield, Veo 3, and Kling, specifically to compare their individual strengths, weaknesses, and the points where human hands are still required to get a usable result.

Human judgment still matters

Even with all these tools changing what's technically possible, the honest note here is that any rough edges in the final edit are still the creator's own responsibility, not the tools'. As AI production tools keep evolving, the argument made is that strong creative direction becomes more valuable, not less, because the tools remove technical bottlenecks but don't remove the need for someone deciding what the shot should actually look like.

Key takeaways

  • GPT-5 maintains cross-shot consistency, including proportions like hand-to-object size, better than Midjourney does on its own.
  • Using GPT-5 images as reference inputs for Midjourney improves Midjourney's own consistency.
  • Veo 3, despite an expensive Ultra subscription, limited generation to five videos and watermarks every output regardless of preference.
  • Veo 3's strength is physical realism; other tools tested moved the VW Bus in unnatural ways that Veo 3 avoided.
  • GPT-5 requires a Topaz upscaling pass but wins on overall shot-to-shot consistency.
  • As AI tools remove technical bottlenecks, strong creative direction becomes more valuable, not less.

Who this is for

This comparison is useful for anyone experimenting with AI-driven commercial or narrative video production and trying to decide which combination of tools to invest time in. It's part of the ongoing hands-on tool testing and commentary that comes out of the broader AI creative work connected to Humanitarians AI.

Full transcript(auto-generated, with timestamps)

[0:00]Do you remember when making a commercial like this cost hundreds of thousands of dollars? Director Simon Meyer created an ultra realistic miniature car commercial and shared the prompts on his Instagram. This is a recreation of that video using only GPT5, Topaz, and VO3. GPT5 is slow, but it remembers that all of your images are related, which keeps consistency across shots. for example, maintaining the correct proportion between the hand size and the VW bus. Midjourney fails to do that. However, one can use GPT5 images as references for midjourney and the consistency improves. I will start using GPT5 for storyboarding even if those images are only used as references. V3 limited me

[0:52]To five videos despite my very expensive Ultra subscription. Google does not design their tools for real video production. They treat VO3 like a toy, even watermarking every image, whether you want it or not. Its strength is adding sound, but Cling is probably better at sound. VO3 is better at the physics. The other tools move the VW bus in unnatural ways. I'll be rebuilding this spot with GPT5, Midjourney, Higsfield, VO3, and Clling, comparing what each can do, their strengths, their weaknesses, and the edges where they still need human hands. In conclusion, GPT5 requires Topaz to upscale, but it's the consistency winner. Until Google takes AI film seriously and allow AI

[1:44]Filmmakers to really use the tool, it will remain a Tik Tok toy. Also, I'll take responsibility for the bad edits. That part's still on me. But even as production evolves, this means that great direction is more valuable than ever. This video is about testing what's now possible when the tools change everything.

More videos

Humanitarians AI Lyrical Literacy Project