AI
Veo 2: Cinematography Becomes Promptable
A note on why Veo 2's real shift was camera behavior, physics, and film grammar.
Question
Veo 2 made cinematographic language a more explicit model interface.
In May 2024, Veo already emphasized 1080p, durations exceeding a minute, natural-language understanding, and cinematic terms. With Veo 2, Google more explicitly discussed improved realism, physics, human movement, expression, cinematography, lenses, effects, 4K, and minute-scale duration. Together, these claims shift attention from image quality toward photographic behavior.
Technical Reading
A cinematography prompt differs from a general visual prompt:
visual prompt: mountain valley, warrior, mist, sunset
camera prompt: slow push-in, low-angle tracking, 35mm lens, shallow depth of field
The first describes what to depict; the second describes how to look. Understanding camera language allows a video model to generate viewing relationships: which subject is emphasized, how space unfolds, and how distance and movement change emotion.
Veo 2's emphasis on physics and human movement also addresses an important concern after Sora: impressive scale is insufficient if action is unconvincing. Expression, limbs, contact, gravity, inertia, and camera motion all affect whether a clip works as a shot.
Capability Boundary
An elaborate prompt still cannot settle narrative decisions. Cinematographic language can control an individual shot, while cross-shot continuity still requires references, character assets, storyboards, and human selection. Understanding lenses does not imply understanding a plot.
I see Veo 2 as advancing directing language for a single shot rather than completing multi-shot production.
Lab Judgment
After Veo 2, I would structure prompts in four layers:
subject: who is in the shot
space: the subject's spatial relationships
camera: distance, focal length, position, and movement
event: how the action starts, how it ends, and what state it leaves
This is more useful than accumulating terms such as "cinematic," "high quality," and "dramatic." The important change is making what the camera does an actionable variable.