Shot Sizes, And Why "Close-Up" Isn't Specific Enough
Wide, medium, close-up, extreme close-up. How much of the subject is in frame is the single most reliable control you have over an AI shot.
Shot size answers one question: how much of the subject fills the frame?
It is the most reliable camera control in a prompt because it is unambiguous. Angle can be interpreted, mood can be ignored, but "waist up" is a measurement. Models get it right far more often than anything else you can ask for.
The ladder
Film crews use a standard set of sizes, and they work as prompt terms because the same words appear in the captions the models were trained on.
Extreme wide shot (EWS) — the subject is small in a large environment. Use it when the location is the subject. A person is a silhouette at best.
extreme wide shot, lone figure small in the frame, vast desert landscape
Wide shot (WS) — full body, head to feet, with room around them. The establishing default. You can see where they are and what they're doing.
wide shot, full body visible, subject centred in the room
Medium shot (MS) — roughly waist up. The conversation shot. Enough face to read expression, enough body to read gesture. If you don't know what to ask for, this is the safe one.
medium shot, subject framed from the waist up
Medium close-up (MCU) — chest up. The interview shot. Most talking-head video lives here.
medium close-up, subject framed from the chest up
Close-up (CU) — the face fills the frame, roughly chin to forehead. Emotion. No context, no environment, nowhere to look but at them.
close-up, face fills the frame
Extreme close-up (ECU) — one feature. An eye, hands on a keyboard, a watch face. Abstract and specific at once.
extreme close-up on the hands, fingers on the dial
"Close-up" is the word people misuse
When most people write "close-up" they mean a medium close-up — head and shoulders, the way a person looks on a video call. An actual close-up is much tighter than that and can feel confrontational when you didn't intend it.
If you write close-up of a woman in a cafe and get a face filling the entire
frame with no cafe visible, the model did what you asked. You wanted a medium
shot with a cafe behind her.
The fix is to name the crop line rather than the label:
framed from the shoulders up, cafe interior visible behind her, softly out of focus
Crop lines are more precise than size names, and they compose with everything else. "Waist up", "chest up", "knees up", "head and shoulders" — all of these are harder to misread than "medium-ish".
Avoid cutting at joints
A rule from photography that carries straight over. Framing that cuts a person at a joint — the ankle, the knee, the wrist, the elbow — looks amputated. Frame between joints instead: mid-thigh, mid-calf, mid-forearm.
framed from mid-thigh up
reads better than
framed from the knees up
Models will not do this for you. They cut wherever the composition lands.
Shot size sets the emotional distance
This is the reason to care beyond framing. Distance from a subject is how a viewer's relationship to them is set.
- Wide — observing. You are outside the scene looking in.
- Medium — conversing. You are standing with them.
- Close — intimate or confrontational. You are inside their space.
A product ad that opens wide and cuts progressively closer builds involvement. One that stays wide throughout feels like surveillance footage. This is a real editing grammar and it works the same whether a human or a model rendered the frames.
In video, shot size is per shot, not per prompt
A common mistake in video prompts is describing a shot size and then also describing action that could not possibly happen inside that frame:
Extreme close-up on her eye as she walks across the room and opens the door
Nothing about that is renderable. The eye fills the frame; there is no room, no door, no walking. The model will pick one half of the instruction and drop the other, and which half is a coin flip.
If the framing needs to change, that is either a camera move or a cut, and it needs to be written as one:
Close-up on her eye. Cut to a wide shot as she crosses the room and opens the door.
Combining size with the other three
Shot size composes with angle, lens and movement. Each answers a different question, and specifying all four gives you a shot you can reproduce:
| Control | Question it answers |
|---|---|
| Shot size | How much of the subject is in frame? |
| Angle | Where is the camera relative to the subject? |
| Lens | How compressed is the space? |
| Movement | What does the camera do during the shot? |
A worked example:
Medium close-up, low angle looking up, 35mm lens, camera slowly pushing in. Chef in a stainless steel kitchen, steam rising behind him.
Every one of those five clauses is doing a job. Compare it to "a cool dramatic shot of a chef" — same subject, and one of them you can render twice and get the same thing.
The short version
Ask for a crop line, not a mood. "Waist up" beats "medium-ish". "Shoulders up with the room visible behind" beats "close-up". And never cut at a joint — the model won't catch it, and once you've seen it you can't unsee it.
Keep reading
Camera Angles That Change What A Shot Means
Eye level, low angle, high angle, Dutch. Four words that change the meaning of a shot more than any adjective you could add.
Camera Movement In Video Prompts: One Move Per Shot
Push in, pan, tilt, orbit, handheld. Naming the move is easy — the hard part is resisting the urge to ask for three of them at once.
Lighting Is The Half Of The Prompt Everyone Skips
Direction, quality, colour. Three decisions that do more for how an image reads than any adjective about mood.