Twenty-four shot types, grouped by what they actually control. This vocabulary predates AI by a century — which is exactly why it still works when a model release breaks everything else.
A generative model has seen millions of films. It knows what "low angle" means far better than it knows what "make it look powerful" means.
Naming the shot is not jargon. It is the difference between describing a feeling and specifying a camera position.
Angle — where the camera sits
The single most reliable lever, because it maps to how humans already read physical height and dominance.
Overhead and high angle are not the same note. High angle diminishes a person; overhead removes them from the situation entirely — they become one element in a pattern.
Distance — how much is in frame
Close-up carries a range rather than one meaning — intimacy when the subject is open, intensity when they are not. The distance sets the access; the performance decides what you do with it.
Movement — what the camera does during the take
Push-in and long take both produce tension by different means. One approaches; the other refuses to release. Knowing which you want is the difference between a shot that builds and one that merely lasts.
Handheld and shaky are also not synonyms. Handheld implies a person holding the camera — urgent, present, human. Shaky implies loss of control. Prompting for one and getting the other is a common and fixable failure.
Focus — what is sharp, and when
These three control when the audience learns something, which is why they are the hardest to specify and the most valuable when they land.
Composition — where the subject sits
Centre frame and locked shot both read as "control" — and they are not the same thing.
Locked is camera behaviour: nothing moves. Centre frame is composition: the subject owns the middle. Combine them and the control reads as absolute. Use one against the other — a locked camera on an off-centre subject — and you get something considerably more unsettled.
Using this in a prompt
- Name the shot, not the feeling. "Low angle" outperforms "make him look powerful" because it specifies a camera position rather than an outcome.
- One angle, one distance, one movement. Stacking five terms produces an average of five shots.
- Movement needs duration. "Slow push-in" is meaningless in a two-second generation.
- Composition survives model changes better than movement does. If a generation keeps failing, drop to composition and angle — those are the most reliably understood.
The full prompt structure is on anatomy of an AI video prompt, and current model capability is on the models page.