Shot Sizes for AI Image and Video Prompts

Shot Sizes, And Why "Close-Up" Isn't Specific Enough

Wide, medium, close-up, extreme close-up. How much of the subject is in frame is the single most reliable control you have over an AI shot.

11 Aug 2026 6 min read

Shot size answers one question: how much of the subject fills the frame?

It is the most reliable camera control in a prompt because it is unambiguous. Angle can be interpreted, mood can be ignored, but "waist up" is a measurement. Models get it right far more often than anything else you can ask for.

The ladder

Film crews use a standard set of sizes, and they work as prompt terms because the same words appear in the captions the models were trained on.

Extreme wide shot (EWS) — the subject is small in a large environment. Use it when the location is the subject. A person is a silhouette at best.

extreme wide shot, lone figure small in the frame, vast desert landscape

Wide shot (WS) — full body, head to feet, with room around them. The establishing default. You can see where they are and what they're doing.

wide shot, full body visible, subject centred in the room

Medium shot (MS) — roughly waist up. The conversation shot. Enough face to read expression, enough body to read gesture. If you don't know what to ask for, this is the safe one.

medium shot, subject framed from the waist up

Medium close-up (MCU) — chest up. The interview shot. Most talking-head video lives here.

medium close-up, subject framed from the chest up

Close-up (CU) — the face fills the frame, roughly chin to forehead. Emotion. No context, no environment, nowhere to look but at them.

close-up, face fills the frame

Extreme close-up (ECU) — one feature. An eye, hands on a keyboard, a watch face. Abstract and specific at once.

extreme close-up on the hands, fingers on the dial

"Close-up" is the word people misuse

When most people write "close-up" they mean a medium close-up — head and shoulders, the way a person looks on a video call. An actual close-up is much tighter than that and can feel confrontational when you didn't intend it.

If you write close-up of a woman in a cafe and get a face filling the entire frame with no cafe visible, the model did what you asked. You wanted a medium shot with a cafe behind her.

The fix is to name the crop line rather than the label:

framed from the shoulders up, cafe interior visible behind her, softly out of focus

Crop lines are more precise than size names, and they compose with everything else. "Waist up", "chest up", "knees up", "head and shoulders" — all of these are harder to misread than "medium-ish".

Avoid cutting at joints

A rule from photography that carries straight over. Framing that cuts a person at a joint — the ankle, the knee, the wrist, the elbow — looks amputated. Frame between joints instead: mid-thigh, mid-calf, mid-forearm.

framed from mid-thigh up

reads better than

framed from the knees up

Models will not do this for you. They cut wherever the composition lands.

Shot size sets the emotional distance

This is the reason to care beyond framing. Distance from a subject is how a viewer's relationship to them is set.

  • Wide — observing. You are outside the scene looking in.
  • Medium — conversing. You are standing with them.
  • Close — intimate or confrontational. You are inside their space.

A product ad that opens wide and cuts progressively closer builds involvement. One that stays wide throughout feels like surveillance footage. This is a real editing grammar and it works the same whether a human or a model rendered the frames.

In video, shot size is per shot, not per prompt

A common mistake in video prompts is describing a shot size and then also describing action that could not possibly happen inside that frame:

Extreme close-up on her eye as she walks across the room and opens the door

Nothing about that is renderable. The eye fills the frame; there is no room, no door, no walking. The model will pick one half of the instruction and drop the other, and which half is a coin flip.

If the framing needs to change, that is either a camera move or a cut, and it needs to be written as one:

Close-up on her eye. Cut to a wide shot as she crosses the room and opens the door.

Combining size with the other three

Shot size composes with angle, lens and movement. Each answers a different question, and specifying all four gives you a shot you can reproduce:

Control Question it answers
Shot size How much of the subject is in frame?
Angle Where is the camera relative to the subject?
Lens How compressed is the space?
Movement What does the camera do during the shot?

A worked example:

Medium close-up, low angle looking up, 35mm lens, camera slowly pushing in. Chef in a stainless steel kitchen, steam rising behind him.

Every one of those five clauses is doing a job. Compare it to "a cool dramatic shot of a chef" — same subject, and one of them you can render twice and get the same thing.

The short version

Ask for a crop line, not a mood. "Waist up" beats "medium-ish". "Shoulders up with the room visible behind" beats "close-up". And never cut at a joint — the model won't catch it, and once you've seen it you can't unsee it.

MarketDragon

MarketDragon

We typically reply in a few minutes

Enter to send • Shift+Enter for new line