The core of editing a video podcast is simple: show whoever is speaking. The hard part is doing that hundreds of times in an hour-long recording. This guide covers Premiere's multi-camera tool first, then the automatic version of the same job.
Last updated:
Import every camera and every microphone recording. Select them all in the Project panel, right-click and choose Create Multi-Camera Source Sequence; set the sync point to Audio. Premiere lines the clips up by their waveforms.
Keeping each speaker's microphone on its own audio track makes both the mix and the next steps easier.
This method means watching the recording from start to finish at least once; an hour-long episode takes more than an hour.
Sylba makes the same decision from microphone levels: it cuts to the camera of whoever is speaking and goes to the wide shot when two people talk together.
If the podcast was shot as a single wide shot, the manual route is: duplicate the clip for each person, frame that person with Scale and Position on each copy, then cut between the copies following the conversation.
Sylba's Single camera mode does this by itself: it finds the faces in the wide shot, builds a virtual camera per person, picks the speaker from lip movement and makes the framing follow when someone moves. Output can be 9:16, 1:1, 4:5 or 16:9.
Cutting a ten-minute chat by hand with Multi-Camera is quick and keeps you in control. On long episodes, letting the first pass run automatically and fixing only the cuts you dislike is less tiring. With one camera you can build a framing per person by hand, but it takes longer whenever someone moves.
Cutting on every change of speaker is tiring to watch; staying on the speaker during short acknowledgements (“yeah”, “mm-hm”) feels calmer. In a long monologue, going to the wide shot or to the listener's reaction now and then keeps the pace alive.
Captions and Graphics are free. macOS and Windows.