Skip to content

Blog

Why podswitch decides every cut from the microphones

By Updated 2 min read

podswitch decides every cut in a multicam podcast from the speakers’ microphones and never analyses the video. In a conversation, the question an editor keeps asking is who is talking, and a close microphone per person answers that directly. Working from audio also means your video never has to leave your computer.

What does an editor actually decide in a podcast edit?

Most of the work in cutting a conversational multicam is repetitive: show the person talking, hold the shot long enough that it does not feel frantic, cut to a listener now and then during long answers, and go wide when everyone talks at once. Those decisions follow the conversation, and the conversation is in the audio.

podswitch makes exactly those decisions and leaves the rest to you. The result is an ordinary multicam edit in Final Cut Pro, so you can change any shot in the Angle Viewer.

Why use the microphones instead of the picture?

When each speaker has their own microphone, the loudest mic relative to its own background level is a direct measure of who is speaking. Nothing has to be inferred from faces or mouth movement.

The same signal gives you something video cannot: a cleaner dialogue mix. In a shared room every mic picks up every voice. Because podswitch knows who is speaking at each moment, it can open each mic only while its speaker talks and turn it down otherwise, which removes the bleed.

What does working from audio mean for privacy?

Video files are large, and they show people’s faces and surroundings. Audio is all podswitch needs, so it is all we take. Your browser extracts the audio from each video file on your own computer and uploads only that. The video stays on your Mac, and Final Cut Pro links to it there when you import the edit.

Uploaded audio is deleted automatically when your plan’s retention period ends, and recordings are never used to train models.

Why does the edit come out the same every time?

The decisions follow fixed rules. There is no randomness, even in where reaction shots are placed, so the same files with the same settings always produce the same cut list. When you change a setting, such as the pacing preset, you can see exactly what that change did.

It also means podswitch only re-runs the parts of the work that a change affects. Renaming a speaker needs no processing at all; changing the pacing takes seconds.

What are the limits of this approach?

It depends on having a microphone per speaker. If two people share one mic, podswitch cannot tell them apart. It also cannot know that someone did something visually interesting while silent. Those moments are where you, the editor, still make the call, and the multicam project is set up so that doing so takes a click.

If you want to try it on your own recording, the recording guide covers what to capture.

Common questions

Does podswitch use face or lip detection?

No. It never analyses video frames. Every decision comes from comparing the speakers’ microphones.

Will the same recording always produce the same edit?

Yes, for the same files and settings. Decisions follow fixed rules, including where reaction shots go, so re-running a project with unchanged settings gives the same cut list.