postika

Explainers · 8 min read

Face-Aware Reframing Explained: How AI Converts 16:9 Video to 9:16

Face-aware reframing turns landscape footage into vertical video by following the speaker instead of cropping the centre. How it works, where it fails, and what to check before you export.

By Postika Team Published
Short-form field guide

Face-aware reframing is the step that turns a landscape recording into a vertical clip without cutting the person in half. It uses face detection to decide which part of the 16:9 frame to keep for 9:16, and on multi-person footage it follows whoever is speaking.

This guide explains what actually happens when a wide shot is converted to vertical, why a centre crop is the wrong default, and what to review before publishing a reframed clip.

What happens when 16:9 becomes 9:16

A 16:9 frame is roughly 1.78 times wider than it is tall. A 9:16 frame is the same shape rotated. Fitting one inside the other without letterboxing means keeping a vertical slice of the original and discarding the rest — about three-quarters of the horizontal image.

So converting to vertical is not a resize. It is a decision about where the slice goes, made once per frame. Every reframing tool, automatic or manual, is answering that one question thousands of times per clip.

There are three ways to answer it:

  1. Centre crop — the slice sits in the middle of the frame, always
  2. Manual keyframes — an editor places the slice by hand and animates it between positions
  3. Face-aware reframing — software detects faces and positions the slice around them, switching when the speaker changes

The first is fast and frequently wrong. The second is right and slow. The third is what AI video editors are trying to make both fast and right.

Why centre crop fails on real footage

Centre crop works on a single presenter who stays in the middle of a locked-off shot. Most speech-led footage is not that.

  • Two-person podcasts are shot wide with a host on each side. The centre of the frame is the gap between them — a centre crop shows two half-faces and a wall.
  • Interviews often frame the subject off-centre, following the rule of thirds. The crop lands on their shoulder.
  • Webinars and talks put the speaker to one side of a slide or a stage. The crop shows the slide.
  • Handheld or roaming footage moves. A fixed crop cannot follow.

Centre crop is not an editing choice; it is the absence of one. That is why manual reframing became a standard, tedious step in short-form workflows — and why it is worth automating.

How face-aware reframing works

Most implementations follow the same pipeline:

1. Face detection per frame

The system finds faces in each frame (or each few frames) and records their position and size. This is well-solved technology; it works on most clear, front-facing footage in reasonable light.

2. Speaker attribution

On multi-person footage, knowing where the faces are is not enough — the frame needs to sit on the one who is talking. The most reliable signal is the transcript: speaker turns in the transcript, aligned to timestamps, say who is speaking when. Some systems also use lip movement or audio direction.

3. Frame placement and smoothing

The 9:16 window is positioned around the active speaker. Smoothing stops it jittering with every small head movement, and a switch threshold stops it flipping back and forth during quick interjections. A cut between speakers is usually preferable to a pan, which looks like a security camera.

4. Review

A good tool shows the framing decision per clip and lets you override it. Automatic reframing is a proposal; the cases below are where it needs a human.

Where it goes wrong

Face-aware reframing is reliable on the footage it was designed for and brittle outside it. Plan for these:

  • Cross-talk. When two people speak at once, speaker attribution is a coin flip. Review any clip that contains overlap.
  • Faces turned away. A guest looking at the host in profile, or a presenter turned to a screen, may drop out of detection and the frame may sit on the wrong person or the last known position.
  • Screen recordings and slides. There is no face to follow. The frame will either lock to a small webcam inset or wander. Slide-heavy material needs a manual layout — a stacked composition with the slide above the speaker, for example.
  • Wide group shots. Four people at a panel table means four small faces. The slice can only hold one or two; expect to choose.
  • Low light and heavy backlighting. Detection degrades and the frame drifts.

None of these are reasons not to use automatic reframing. They are reasons to review the output rather than export blind.

What to check before exporting a reframed clip

A short checklist that catches most problems:

  1. Is the right person in frame on every line? Scrub the speaker changes specifically — they are where errors cluster.
  2. Is there headroom? A tight crop that clips the top of the head reads as amateur. Most tools let you adjust the vertical position.
  3. Do the captions overlap the face? Reframing and caption placement interact. Move the captions before you move the frame.
  4. Are the platform safe areas clear? The bottom and right edges of a Reel, Short or TikTok are covered by interface elements. Keep faces and text out of them.
  5. Does the switch timing feel right? A cut that lands a beat after the speaker changes is jarring. If the tool lets you nudge the switch, do it.

Where Postika Reelsmith fits

Reelsmith reframes landscape footage to 9:16 using face detection and the transcript to follow the active speaker. It was built around the two-shot podcast case, where a centre crop is unusable, and every framing decision is shown per clip for review before anything renders.

It does not solve slides-only recordings — those need a manual layout — and it is least reliable on overlapping speech, which is why the review step is part of the workflow rather than an afterthought. The vertical video reframer page covers the workflow in more detail, and the long video to Shorts tool is the same reframing applied to batch clip export.

FAQ: Face-aware reframing

Is face-aware reframing the same as auto reframe?

Mostly. “Auto reframe” is the generic term for software choosing the crop; “face-aware” specifies that it uses face detection to do so. Some auto-reframe features use motion or saliency instead of faces, which works better on non-speech footage and worse on conversations.

Does it work on a single presenter?

Yes, and it is the easiest case — the frame simply tracks one face. The value is highest on multi-person footage, but a single presenter who moves around still benefits.

Can I convert 16:9 to 9:16 without cropping?

Only by letterboxing — placing the full landscape frame in the middle of a vertical canvas with empty space above and below. It preserves everything and fills about a third of the screen. It is rarely the right choice for short-form, where the clip competes with native vertical video.

Do I need to re-record in vertical?

No. Reframing exists so landscape footage you already have can be reused. Shooting a second vertical version of every recording doubles the production schedule for marginal gain.

How is the speaker identified on a two-person shot?

Most reliably from the transcript: speaker turns aligned to timestamps say who is talking when, and the frame follows. Lip movement and audio direction are secondary signals. Cross-talk defeats all of them, so review those moments.

Key takeaway

Converting landscape video to vertical is a decision about where a narrow slice of the frame goes, made once per frame. Centre crop makes that decision blindly; face-aware reframing makes it based on who is on screen and who is speaking. It is reliable on speech-led footage and needs review on cross-talk, profiles and slides — so treat the automatic result as a draft and check the speaker changes before you export.

Turn the next recording into your next week of content.

Try Reelsmith →