← Read

AI · · 7 min read

Directing a video model: the shot language Veo actually understands

Shot composition, camera moves, lens effects, lighting and sound are all promptable. Google's own quickstart lists the vocabulary, and the limits around it.

The fastest way to get better video out of a model is to stop writing prompts and start writing shots. Google's own Veo quickstart says so in as many words: it lists the film vocabulary the model has been built to understand, and the list is a shot list.

This is a guide to that vocabulary, plus the four things Google documents you can control directly (lighting, camera, audio and dialogue), plus the numbers that will quietly redesign your storyboard: 4, 6 or 8 seconds, two aspect ratios, two resolutions.

Everything here comes from Google's Veo quickstart in the google-gemini cookbook, read on 20 September 2026. The linked videos are Google's own published sample outputs, not results produced for this article. Veo is a paid feature with no free tier, which the notebook states up front.

The vocabulary Google says it understands

The quickstart groups it into four families. Treat this as the set of words worth spending your prompt on.

FamilyTerms the quickstart names
Shot composition"single shot", "two shot", "over-the-shoulder shot"
Camera position and movement"eye level", "high angle", "worms eye", "dolly shot", "zoom shot", "pan shot", "tracking shot"
Focus and lens"shallow focus", "deep focus", "soft focus", "macro lens", "wide-angle lens"
Style and subject"sci-fi", "romantic comedy", "action movie", "animation", plus the subject and background you want

Two things follow from that list. First, these are craft terms, not adjectives: "over-the-shoulder" describes a camera position, "cinematic" describes nothing. Second, the genre words at the bottom are doing a lot of work, because "romantic comedy" carries a lighting scheme, a lens choice and a colour grade with it. If you are not getting the look you want, naming the genre is usually a shorter fix than describing the grade.

Write the shot list into the prompt

Google's camera control example is the most instructive thing in the notebook, because it is not a description of a scene. It is four beats in order.

a realistic video of a futuristic red sportscar speeding down a winding coastal
highway at dusk. Begin with a high-angle drone shot that slowly descends,
transitioning into a close-up, low-angle tracking shot that perfectly follows the
car as it rounds a curve, emphasizing its speed and the gleam of its paint under
the fading light. Then, execute a smooth, rapid dolly zoom, making the background
compress as the car remains the same size, conveying a sense of intense focus and
speed. Finally, end with a perfectly stable, slow-motion shot from a fixed roadside
perspective as the car blurs past, its taillights streaking across the frame.
Include the immersive sound of the engine roaring, the tires gripping the asphalt,
and the distant crash of waves.

Read the structure rather than the words: begin, transitioning into, then, finally. Four camera positions, one continuous subject, and the audio named at the end. That is a storyboard flattened into a paragraph, and it is a far better template than any list of style keywords.

storage.googleapis.com ↗Google's published camera control sample videoThe sample output Google publishes for the camera control prompt above: drone descent into a tracking shot, then a dolly zoom. Video: Google

Lighting is a brief, not a keyword

The lighting example follows the same pattern and is worth copying wholesale for anything that has to look art directed.

a solitary, ancient oak tree silhouetted against a dramatic sunset. Emphasize the
exquisite control over lighting: capture the deep, warm hues of the setting sun
backlighting the tree, with subtle rays of light piercing through the branches,
highlighting the texture of the bark and leaves with a golden glow. The sky should
transition from fiery orange at the horizon to soft purples and blues overhead,
with a single, faint star appearing as dusk deepens.

Three separate instructions are hiding in there: the direction of the key light (backlighting), what it does to the subject's texture (bark and leaves), and a gradient across the sky with a specific transition. A designer briefing a photographer would say roughly those three things. A prompt that says "beautiful golden hour lighting" says one of them badly.

storage.googleapis.com ↗Google's published lighting control sample videoGoogle's sample output for the lighting prompt above. Video: Google

The audio arrives whether you asked for it or not

This is the part that catches people who come from image models. The quickstart states that Veo 3 "generates videos with audio automatically, with no additional effort from the developer".

Automatically. Not optionally. So sound design becomes part of your brief by default, and the two examples show both halves of that.

For ambience, Google simply names the sounds in the same sentence as the picture: fireworks over a city skyline "with many different fireworks colors and sounds", plus crowd noise "surrounding the camera POV". Naming the listener's position is the useful trick there.

For dialogue, you put the lines in quotation marks inside the prompt and let the model perform them. And then you do the thing that is easy to miss: Google's dialogue example sets negative_prompt to "texts, captions, subtitles".

If you do not ask it not to, you can get burned-in subtitles across your film. That single parameter value is probably the most practically valuable line in the whole notebook.

storage.googleapis.com ↗Google's published audio control sample videoGoogle's sample output for the fireworks and crowd audio prompt. Video: Google

The numbers that redesign your storyboard

ParameterWhat the quickstart documents
duration_seconds4, 6 or 8 seconds with Veo 3.1. Always 8 with Veo 3, and 7 when extending an existing clip
aspect_ratio16:9 or 9:16. That is the whole list
resolution720p or 1080p, with the notebook's portrait example running at 720p
negative_promptWhat you do not want to see. Google's own image-to-video examples use "ugly, low quality, static, weird physics"
person_generationWhether adults may be generated. Children are always blocked

The aspect ratio row is the one to plan around. There is no 4:5 and no square, which are the two formats social feeds actually favour, so a 4:5 cutdown means shooting 9:16 and cropping, and composing with that crop in mind from the first prompt. The same discipline applies to still work, where the format ceilings are laid out in the sizes AI image models can actually give you.

The duration row is the other one. Eight seconds is not a film, it is a shot. Plan a sequence as a series of generated shots you cut together, which is also how it would be shot on a real day.

Start from a frame you art directed

Text to video is the demo. Image to video is the workflow, and the quickstart covers three routes into it: upload your own image as the first frame, generate a base image with a Gemini image model and animate that, or mix several reference images into one video.

The second route is the one to take seriously if you care what the thing looks like. You have far more control over a still: you can iterate on composition, palette, wardrobe and product detail for pennies, approve a frame, and only then spend on motion. Google's own example does exactly this, generating a photorealistic still of a cat in a red convertible before animating it.

It also matches how you would work anyway. Nobody storyboards by rolling camera. The techniques for locking a character or a product across those stills are in keeping a character consistent across a set of images.

What the quickstart will not tell you

It documents capability, not quality. Nothing in it tells you whether a face holds together across eight seconds, whether fabric behaves, whether a logo on a product survives a camera move, or how a generated grade behaves once it is in a timeline next to real footage. Those are the questions a motion designer actually has, and they can only be answered on your own brief.

The model identifiers in the notebook are also preview builds at the time of writing, and previews move. Check the identifier before you quote a rate card.

FAQ

Can I get a square or 4:5 video?

Not from the aspect ratio parameter, which the quickstart limits to 16:9 and 9:16. Generate 9:16 and crop, and compose knowing you will lose the top and bottom.

How long can a single generation be?

4, 6 or 8 seconds on Veo 3.1 per the quickstart, always 8 on Veo 3, and 7 seconds when extending. Longer pieces are edits of multiple generations, not single generations.

Can I turn the audio off?

The quickstart documents audio as automatic on Veo 3 rather than as a toggle. What it does show is steering audio with the prompt and suppressing unwanted on-screen text with negative_prompt. If you need a silent clip, plan to strip the audio track in your edit.

Is this better than the other video models?

The quickstart cannot answer that, and neither can any article written from documentation alone. What it can tell you is what Veo is documented to respond to, which is a surprisingly precise film vocabulary, and that is a useful thing to know before you spend money comparing anything.

Sources

More to read