DualCam

Does the picture-in-picture window follow your face in dual camera video?

No. The picture-in-picture window sits at whichever corner, size and shape you picked, and it stays there — it does not pan, zoom or nudge itself to keep a face centred as you move. What appears inside it is simply whatever that camera’s lens is pointed at, scaled and cropped to fit the window’s shape, the same way it would be with no face in frame at all.

It is a reasonable thing to expect otherwise. Apple ships Center Stage on some devices and in apps like FaceTime, where the front camera’s ultra-wide field of view is digitally panned and zoomed to keep a detected face in frame. Dual camera recording looks similar at a glance — a small window with a face in it — so it is a fair question whether the same kind of tracking is happening underneath. It is not, and the reason comes down to what a dual-camera recording actually is.

What is actually happening, frame by frame

A dual-camera recording is a real-time composite: two camera streams get drawn into one canvas, on the GPU, for every single frame, and that canvas is the file being written. The picture-in-picture window is a fixed rectangle at a fixed position on that canvas — one of four corners, one of three sizes, one of three shapes. Every frame, whatever the inset camera currently sees gets scaled and cropped into that same rectangle. Nothing in that pipeline asks “where is the face” and adjusts the rectangle in response — the rectangle is settled before the frame is drawn, not decided by what is in it.

That is also why the inset can look perfectly fine one moment and clip someone out the next: it is not failing to track you, because it was never trying to. It is a window with a fixed shape and position, showing whatever is in the lens’s field of view at that spot.

Why there is no auto-follow here

A face-tracking window like Center Stage is its own pipeline — it needs a wide field of view to crop into, on-device face detection running continuously, and a capture format built around leaving that digital headroom. A multi-camera session already has a narrower set of formats to choose from than a single camera does: every input has to run a format that explicitly supports multi-cam, and the combined load of both streams has to fit inside one hardware budget. Adding a continuous face-detection pass on top of a real-time two-camera composite is a second demanding job stacked on a workload that is already the heaviest thing the capture hardware does. It is not that the idea is unreasonable — it is that the budget a multi-cam session has to live inside was already spent on getting two cameras compositing in real time at all.

The upside of the simpler approach is the one dual-camera recording is actually built around: the frame you see in the viewfinder is the exact frame written to the file, with nothing decided after the fact. A tracking window would mean the recorded crop keeps changing based on a detector’s best guess, which cuts against the whole point — what you frame is what you get.

Keeping yourself in the window without the app doing it for you

Give yourself margin with size and shape

The three inset sizes exist for exactly this. A small inset looks tidy but leaves almost no room to move before an edge clips you; the largest of the three gives noticeably more margin for the same amount of shift. A portrait-shaped window also gives more vertical room than a circle or square at the same size, which matters if you tend to lean rather than drift side to side.

Mount it for anything longer than a quick shot

A fixed window depends on the phone staying aimed the way you set it up. Hand-held, small shifts in how you are holding the phone move where you land inside that window just as much as your own movement does. A mount or a stand against something stable removes one half of that problem before you even start talking.

Recenter the phone, not the window, when you drift

If you notice yourself sliding toward an edge, the fix is adjusting the phone’s angle, not looking for a setting that will do it for you — there is not one. The corner, size and shape can all be changed mid-recording, which is useful for reframing a scene between takes, but none of it repositions the crop based on where a face currently is.

Reach for split screen if you plan to move a lot

A talking-head shot where you gesture, lean into the frame, or turn to point at something behind you is a better fit for split screen than for a small inset. Each half of a split gets far more area than any picture-in-picture size does, so the same amount of movement is much less likely to put you outside it. See the guide on picture-in-picture versus split screen for when each layout actually reads better.

Common follow-up questions

Can I move the picture-in-picture window to a different corner while recording?

Yes — corner, size and shape can all be changed mid-recording, and the file follows the change immediately. None of that is automatic tracking, though; it only moves when you tell it to.

Does the full-frame camera behind the inset track a face either?

No. The only per-frame adjustment available is tap-to-focus, which sets where the lens focuses and its exposure, not what part of the frame gets shown. Framing is entirely up to how you point the phone.

Why not just add Center-Stage-style tracking to dual camera recording?

A multi-camera session already runs on a tighter hardware budget than a single camera — every input needs a multi-cam-capable format, and the combined load of both streams has to fit inside one budget. Continuous face detection on top of a live two-camera composite is a second heavy job stacked on the heaviest thing the capture hardware already does.

Want to just do this?

DualCam records the iPhone front and back cameras at the same time and writes one finished MP4 while you shoot. Free, no account, no ads, nothing leaves the phone.

Download on the App Store