How to film a sign language interpretation video with both iPhone cameras
A sign-language interpretation only works if the hands and the source stay locked to the same instant, frame for frame, for the whole video. That is a harder version of the same problem a two-camera reaction video already has, and the fix is the same one: record both feeds live into a single composited frame instead of filming two separate clips and lining them up afterwards, where drift creeps in and tends to get worse the longer the take runs.
Beyond sync, an interpretation video has requirements a reaction video does not — the hands have to stay inside the frame, stay large enough to read, and read the right way round. That last one depends on which physical camera the interpreter stands in front of, and it is the detail most setups get wrong the first time.
Why live capture matters more here than for most dual-camera content
Two people filmed on two separate cameras and synced afterwards can end up close enough for a reaction video and still be wrong for an interpretation, because a viewer relying on the signing has no source audio to fall back on if the two tracks drift even a little. A dual-camera app avoids that problem structurally rather than by being careful: the two camera images are composited into one frame on the GPU as they arrive, and that same frame is what gets written to the file, so the interpreter and the source were never two things that needed aligning in the first place.
This is the same property that makes a reaction video or a two-person interview work on one phone — timing comes for free because there is only ever one timeline. Interpretation content just has less room for error in it than either of those.
Picture-in-picture or split screen
The general rule for choosing a dual-camera layout is to use picture-in-picture when one camera is the subject and the other is secondary, and split screen when both sides carry equal weight for the whole video. An interpretation leans toward the second case: the interpreter is not commentary on top of the real content, they are carrying an equal share of the video for anyone relying on them, so a fixed half of the frame — top and bottom, or left and right — tends to serve better than a small window competing for space with whatever the source camera is doing.
Picture-in-picture is still the right call when the source material genuinely needs the full frame to be readable on its own — text on a screen, a whiteboard, small detail in a product demo — and the interpreter can work as a dedicated inset instead. If you go that way, treat size and shape as functional decisions rather than cosmetic ones; the next two sections cover why.
One caution specific to split screen: a left/right split crops roughly a quarter off each side of the source frame to fit a vertical canvas. Keep the interpreter's hands away from the outer edges of their half, since that margin is exactly what gets cut.
Put the interpreter on the back camera, not the front
The front camera in a dual-camera recording is mirrored, the same way any front-facing preview is, because that is the only way to aim a shot at your own face using the screen as a mirror. The back camera never is. For most content this difference is invisible — a face does not have a correct reading direction — but a directional or handed sign does, and mirroring flips left and right exactly the way it would in a bathroom mirror. Raise a right hand into a front-camera inset and it reads on the left to anyone watching.
That makes camera choice a real decision here, not a formality. Whoever is filming should put the interpreter in front of the back camera — the unmirrored one — and let whoever else is on camera, a presenter or narrator whose face does not carry directional meaning, take the front. In the standard two-person setup, that means standing the interpreter on the side of the phone the back lens faces, and the presenter on the selfie side, rather than the more intuitive arrangement of putting whoever is "explaining" in the main frame.
This is not something that can be fixed after the fact. The composited frame is the recording — there is no separate un-mirror step waiting at the end — so the camera assignment has to be right before you press record, not corrected in an editor afterward.
Sizing and shape, if you use picture-in-picture
An inset read from a small window loses exactly the shapes that carry meaning in a sign, the same way a reaction shot loses the point of a face read too small — so size it up rather than defaulting to the smallest of the three steps. A bigger window also needs to scale the source down less to fit, which keeps more of whatever detail the capture resolution actually gave it.
Shape matters just as much as size here. The tall rounded rectangle keeps the interpreter's hands the most room to move before running into an edge, because it crops the least from the source and shares the canvas's own proportions. A square or a circle are both tighter crops taken from the centre, and a circle in particular clips anything that drifts toward the edge of the square underneath it — exactly what a hand raised mid-sign tends to do. Save the circle for a face-only reaction shot; it is the wrong shape for hands.
If the phone can sustain it, capture at 1080p rather than 720p. The inset is a downscaled copy of whatever the source camera captured, and fingerspelling and handshapes are fine detail — a bigger window cannot show detail the capture resolution never gave it.
Corner, background and light
If the finished video is going anywhere with its own interface — captions, a username, a row of buttons — those tend to collect along the bottom edge and down the right side once the video is posted. Put the interpreter's inset toward the top of the frame rather than the bottom, and check that it is not sitting over the one part of the main shot someone actually needs to see, like a whiteboard or a demonstration.
A plain, non-busy background behind the interpreter and clothing that contrasts with skin tone is standard advice for filming any sign-language content, independent of which camera app is doing the recording, and it matters more here than in most dual-camera footage because the hands themselves are the content. Front and back cameras expose independently of each other, so if one side is noticeably darker, tap that frame to set exposure for it specifically rather than accepting whatever the main camera decided.
What tends to go wrong on a first attempt
- The inset shaped as a circle or sized too small, so a raised hand exits the frame mid-sign.
- The interpreter recorded on the front camera, so a directional or handed sign reads reversed to anyone watching.
- The inset corner sitting directly over the one part of the source frame someone needs to read.
- Recording at 720p and expecting a bigger inset to add detail the capture resolution never captured.
- Assuming the app generates captions from the spoken audio — it does not; that is a separate step after recording.
Common follow-up questions
Will the interpreter's signs come out mirrored?
Only if they are recorded on the front camera, which is mirrored the same way any front-facing preview is. The back camera is never mirrored. For a directional or handed sign, film the interpreter on the back camera to keep left and right reading correctly.
Can I switch between picture-in-picture and split screen partway through?
Yes. Layout, along with the inset's corner, size, shape and filter, can all be changed while a DualCam recording is running, and the file follows the change immediately since the preview and the recording are the same frame.
Does the interpreter get a separate audio track from the source content?
No. A dual-camera recording writes one audio track, not one per camera, so there is no way to isolate or balance an interpreter's narration separately from whatever else the microphone picks up.
Does dual camera recording add captions automatically?
No. It captures video and audio only; turning speech into text is a separate step done after recording, the same as with any other camera app.
Want to just do this?
DualCam records the iPhone front and back cameras at the same time and writes one finished MP4 while you shoot. Free, no account, no ads, nothing leaves the phone.