ASMR Volcanic Sand Terrarium | Earth's Elements Exploration | VEO 3

A macro-scale ASMR terrarium scene generated entirely by Google's VEO 3, built from a single detailed text prompt describing sand, minerals, and sound.

0:08 video2 min readWatch on YouTube

Most ASMR content is filmed. This one was written. The entire eight-second scene of black volcanic sand, copper-flecked minerals, and quartz crystals being carved by a small copper tool exists because someone typed out a detailed prompt and handed it to Google's VEO 3 video generator.

A prompt built for sound, not just image

What makes this clip worth looking at closely is the prompt itself, which is included in full alongside the video. It does not just describe what the camera should see, a terrarium with layered sand and minerals on a wooden table in dappled forest light. It also describes what the scene should sound like: the "shush" of shifting sand, the gentle tinkle of colliding crystals, the earthy drag of metal through mineral. Asking a text-to-video model to reason about both the visual texture of stratified sand layers and the acoustic texture of an ASMR soundscape at the same time is a genuinely demanding request, and it is the reason this particular generation is interesting rather than just decorative.

Macro perspective as a technical challenge

The prompt specifically calls for an extreme macro perspective, with the camera moving between bird's-eye views of the whole miniature ecosystem and close-up shots of individual sand grains. That shift in scale, from an overview shot to something closer to individual grains, is a harder ask for a video model than a fixed wide shot, since it has to keep the material's texture and lighting consistent as the framing changes.

Key takeaways

  • The clip was generated by Google's VEO 3 from a single, fully written-out prompt rather than filmed.
  • The prompt specifies both visual detail, layered volcanic sand, copper minerals, quartz crystals, and audio detail, the specific ASMR sounds the scene should produce.
  • It uses an extreme macro perspective that shifts between wide ecosystem views and close-up grain-level detail.
  • The exercise doubles as a test of how well a generative video model can follow a prompt that asks for a coordinated sound-and-image experience.

Try it yourself

If you are experimenting with AI video generation, this is a good template to study: write a prompt that specifies texture, camera movement, and sound together, then see how closely the output matches what you asked for. Humanitarians AI shares experiments like this one as part of its broader work exploring generative AI tools.

More videos

Humanitarians AI Lyrical Literacy Project