Editing interviews · learn extra credit

Turn a rambling interview into a video people finish.

Whether you're sitting a founder down or walking up to a stranger on the street, the value is in what you do after: finding the story hiding in the footage, making the sound clean enough that nobody clicks away, and cutting the clips that actually get seen. This is the whole craft, free and in the open — how to film so the edit is easy, the run-and-gun street interview as its own discipline, how to win on audio, how to cut, and how to turn one outing into a week of content. Sourced technique, live tools, and a 30-day plan that saves your progress in this page.

◈ Start here · your first real upgrade

Don't try to learn all of it at once.

This page is deep on purpose. But if you do one thing after reading, do this, because it's the single biggest jump from amateur to pro:

  1. On your next interview, record two audio sources — a lav clipped 6–8 inches below the chin, plus one backup (camera mic or a second recorder).
  2. Grab 30 seconds of room tone (everyone silent) before you leave.
  3. In the edit, run the cleanup chain in order and normalize to the right loudness for your platform.

Clean, consistent, loud-enough sound already puts you ahead of most creators. Everything else on this page builds from there.

How to use this

The chapters run in the order you actually work: shoot → sound → cut → clip. The loop that runs through all of it: get clean source → cut for story first → finish the audio → publish, then mine the clips. Skim it once, then work the 30-day plan — one technique a day. The practice studio has the tools you'll actually reach for while editing.

1 → 15+
clips a single interview reliably yields
~3s
to hook a viewer before they scroll
audio
the one thing viewers won't forgive — spend here first
CHAPTER 01

Shoot it right — the edit is won on set

Every clean cut and every emotional close-up is decided before you press record. The footage you hand the edit is either a gift or a life sentence. Shoot for the cut you don't yet know you'll need.

1 · Two cameras change everything

One camera is viable, but it punishes you: every time you tighten an answer or cut an "um," you get a jump cut, and your only fixes are b-roll, a dissolve, or living with it. That's why the two-camera interview is the documentary standard. Run an A-cam wide/safe (chest-to-head) and a B-cam tight (close-up on the face). Cutting between them hides every edit and gives the tight angle for emotion.

Two numbers that make cuts look intentional Double the focal length between cameras (e.g. 35–50mm on the A-cam, 85–135mm on the B-cam), and offset the B-cam at least 30° from the A-cam's axis. Under 30° apart reads as a jarring jump; 30°+ reads as a deliberate new angle. Keep both on the same side so the eyeline holds.

Match your two cameras

Same body or same family. If you must mix brands, at least match resolution, frame rate, shutter, white balance in Kelvin, and picture profile, and shoot a gray card on both so the grade has a reference. Mismatched color science is a nightmare in post.

Interviewer placement

Put the interviewer just off one side of the A-cam lens, and outside the spread of the two cameras. If they sit between the lenses, the subject's eyeline lands in dead space and the geography feels wrong.

2 · Framing that flatters

3 · Lens & settings

SettingInterview defaultWhy
Focal length50–85mm equivFlattering compression; under 35mm distorts faces
Aperturef/2.8 (f/1.8–4)Nice separation but enough depth that a lean doesn't go soft. Don't shoot wide open at f/1.2
Shutter2× frame rate180° rule: 1/48 for 24fps, 1/60 for 30fps. Natural motion blur
Frame rate24 or 30fps24 = cinematic, 30 = broadcast feel. Shoot b-roll at 60 for slow-mo
ISOnative / lowestLight the scene instead of cranking ISO and adding grain
White balancemanual KelvinAuto WB drifts shot-to-shot and wrecks camera matching
Log / flat — when NOT to Log preserves range and is worth it on a controlled shoot where you have time to grade. But it looks lifeless straight out of camera and, graded carelessly, adds noise and color casts. For fast turnaround or run-and-gun, shoot a standard Rec.709 profile so it looks right immediately. Log rewards time and skill; if you have neither, don't shoot it.

4 · Lighting

Three lights, three jobs. Key — main source, ~45° off the subject's line of sight and slightly above eye level angled down ~15–30°. Fill — softer, opposite side, to lift the shadow (a bounced reflector does this for free). Back/hair light — behind and above, hitting the shoulders to separate them from the background. That backlight is what makes an interview look "produced."

The one-light look

You don't need three. One big soft key at 45° (softbox, umbrella, or a window), a reflector for fill, and a small backlight if you have one. Bigger and closer = softer = more flattering.

The two lighting mistakes to never make
  • Two shadows — two hard lights of similar strength cross competing shadows on the face. Keep one dominant key, keep fill soft and secondary.
  • Mixed color temperature — a 5,600K window plus a 3,200K lamp makes one side of the face orange and the other blue, and it's brutal to fix. Kill the window, gel your lights to match, or go all-daylight LED on one white balance.
  • Skin tones: darker skin benefits from a touch more key and a defined rim light for separation; lighter skin clips to white faster, so protect highlights. Expose for the face, use soft sources, judge on a real monitor.

5 · On-set audio — the part you can't save later

Bad on-set audio is unrecoverable, so this is where you're most disciplined. The pro standard is two mics: a wireless lav on the subject plus a shotgun on a boom overhead. You pick the better one in the edit, and if one drops out you don't lose the take.

Placement

Lav: 6–8 inches below the chin, centered, clear of scarves, necklaces, and conference lanyards that rub. Shotgun: overhead in front, pointed down at the mouth from 6–12 inches, just out of frame — overhead-and-down rejects room reflections. Close is king.

Levels & monitoring

Peak around −12 dB (loud moments not past ~−6). Clipping at 0 dB is permanent distortion — err low, normalize later. Wear closed-back headphones and listen the whole time; meters won't warn you about a rustling lav or a distant siren, your ears will.

Room tone, 30 seconds, every location

Before you break down, everyone holds still and you record 30–60s of the room's "silence." The editor uses it to patch gaps and smooth cuts so there's no jarring dead air. One minute on set, saves the whole edit.

6 · Directing the guest

Your job is a comfortable person giving self-contained answers, not a great chat that cuts into nothing.

The single most important editing favor Because you'll cut your own voice out, coach them upfront: fold the question into the answer. If you ask "when did you get into crypto?" and they say "2013," it's useless. You want "I got into crypto in 2013." Remind them when they forget.

7 · Run-and-gun at events

Conferences are chaos — PA, crowds, music. A tripod (or a monopod for speed) is still best for the sit-down; a gimbal is for moving b-roll. Solo rig: a 28–75mm zoom (wide-to-tight, no lens swaps), a wireless lav, a small on-camera LED, a monopod, exposure and white balance pre-set so you can grab an interview in under a minute. Win the audio by getting the lav close, not by tightening the camera. Get clear verbal consent on camera. A modern phone with good sound and light beats a cheap camera — audio and framing read as "pro," not the sensor.

8 · Gear — where to spend first

The honest spend order: audio + lens + light before the camera body. Bad sound and bad light can't be fixed in post; a modern-but-modest body is more than enough. A phone with good sound and light beats a cheap camera every time.

TierKitRough spend
1 · Phone-firstPhone (flat/log, locked exposure) · a wireless lav that pairs to it (spend here first) · one small daylight LED or a window · a clamp + mini tripod or compact gimbal$150–500
2 · ProsumerOne mirrorless body · a 24–70 f/2.8 or a 35mm + 85mm pair · wireless lav plus shotgun + boom · a 2-light soft LED kit + reflector · fluid-head tripod + monopod/gimbal$2,000–3,500
3 · Pro two-cameraTwo matched bodies (wide + tight) · 35/50mm + 85/135mm · multitrack field recorder + two lavs + boom + monitoring headphones · 3-point soft LED kit with hair light + gels · two tripods + gimbal + gray card$8,000+
The one-line rule If you can only improve one thing, improve the sound. If two, add light. The camera body comes last.
Coverage checklist — get these every time (tap to check)
  • The wide/safe (A-cam) — runs the whole interview, your fallback
  • The tight (B-cam) — close-up on the face, for emotion and hiding cuts
  • Noddies — the interviewer nodding/listening (no talking), shot after
  • Cutaways / b-roll — hands, the charts/screens they mention, the room
  • Detail inserts — hands, watch, phone, keyboard (60fps if you can)
  • Establishing shots — the building, the room, them walking in / sitting down
  • 30–60s of room tone — every location, before you strike
  • A slate / verbal ID — who this is and the date, so the editor isn't guessing
Answer from memory before you reveal — this is what makes it stick.
Why does a second camera transform the edit, and what two numbers make the cut look intentional?
A second angle hides every jump cut you create when tightening answers. Double the focal length between the two cameras and offset the B-cam at least 30° — under 30° reads as a jarring jump.
What's the one directing instruction that saves the most editing pain?
Get them to fold the question into the answer ("I got into crypto in 2013," not "2013"). You're cutting your own voice out, so every answer has to stand alone.
Which single thing, if you blow it on set, cannot be fixed in post?
Audio. Record two mic sources, get the lav close (6–8" below the chin), monitor on headphones, peak around −12 dB (never clip at 0), and grab room tone. Reverb and clipping are unrecoverable.
Sources
CHAPTER 02

The street interview — the run-and-gun human moment

Most interview advice assumes a scheduled subject, a lit room, and time. The street throws all of that out. You walk up to a stranger who didn't plan to be filmed, you have ten seconds to earn a yes, and you're the director, camera op, sound mixer, and the person making them laugh — all at once, in ninety seconds, before the light changes. It's its own craft. This is the format behind the 360 clips in the road archive.

The spine that changes everything The subject is a real person having a real experience, so the edit must always make them the hero of their thirty seconds. You're filming the moment they learn something and light up — the "wait, that's it?" click — never their prior ignorance. If the only funny version of a clip makes the person small, kill the clip. The dignity is the brand.

— FILMING —

1 · The approach — earning an instant yes

The highest-leverage skill happens before you roll. Pros who do this for a living quote roughly one yes in five, and the rule is: don't take a no personally. You nudge the ratio with speed and warmth.

Approach with the camera down

The most-repeated pro tip: don't have the camera up or pointed as you walk up. Empty hands, a smile, open body language — it signals the connection matters more than the shot, and it's disarming.

Lead with an observation, not "can I interview you?"

A cold "can I interview you?" invites a reflexive no. Open with something human to respond to — "looks like you had a good shopping trip" — then the ask. Give them a half-second of rapport first.

Be transparent, fast

In the first sentence they should know who you are, what this is, and that it's quick and friendly. Transparency removes the "what's the catch?" tension. Speed is kindness — a long hedging approach makes a stranger anxious; a fast, clear, warm one respects their time and lowers the cost of yes.

Before you walk up (tap to check)
  • Camera down / not pointed; hands relatively empty
  • Opening line ready — an observation or compliment, not "can I interview you"
  • One sentence of who / what / why-quick rehearsed
  • Mic on and levels already set — don't fumble gear in front of them
  • Exposure / focus / white balance pre-set for the light you're in
  • An easy first question that makes them smile or feel smart
  • A graceful exit line for a no ("Totally fair — have a great one!")

2 · Consent, ethics & dignity

3 · The one-person-band rig

You're interviewing and operating at the same time, so the rig has to disappear and your attention stays on the human. Keep it minimal — a phone or small camera, one good audio solution, one stabilizer. Every extra device is another thing to monitor while you're supposed to be listening.

Pick one stabilizer

Handheld / phone with IBIS — fastest to deploy, most nimble, best for reaction punch-ins. Gimbal — buttery movement for walking-with-the-subject, but it occupies a hand and costs seconds to balance. Monopod — underrated: height, quick resets, some stability, no gimbal overhead. One, not three.

Phone-first can look pro

Shoot the flattest/highest-quality profile, lock exposure and focus manually (auto hunts and pumps when the crowd moves behind your subject), and let a wireless mic carry the sound. The image rarely gives you away on a street clip — the audio does.

Monitor without losing the person

One earbud in (the subject's mic) keeps you connected to both the sound and the human. If you truly can't monitor, record a safety track onboard the transmitter.

4 · Framing the two-person moment

5 · Audio in chaos — the hardest part

On the street, image forgives and audio does not. Prioritize the subject's voice above everything.

ToolRoleThe catch
Wireless lavThe workhorse — consistent, focused voice, out of frame, moves with themA few seconds to clip on; put it at the collar, not buried under clothing
Handheld interview micThe secret weapon — points a directional capsule at the talker, rejects crowd, looks legit and relaxes a nervous subjectOne hand occupied; you pass it between speakers
On-camera shotgunHands-free — usually your backup/ambience trackFarther from the mouth = more street in the mix
The rules that save street audio Get the mic close — proximity beats every other trick; a mic six inches from the mouth in a crowd beats a distant "better" mic. Use a furry deadcat the moment there's real wind, a low-cut filter for traffic rumble, and a super-cardioid pattern to reject the sides. Record a safety track, and do a 10-second test with playback before the first real interview of a session — verify the environment isn't murdering you before you burn a good subject.

6 · Directing a normie in 60 seconds

Your subject is not an actor and has zero prep, so comfort starts the second you begin interacting, not once the interview begins.

7 · Solo coverage so the edit works

You can't cut energy you didn't shoot. Even solo, grab these every interaction: the reaction (their face at the turn — over-cover it, it's the payoff), the wide (both of you, "this is real, on a real street"), the detail (the phone screen, the gift, the handshake), environmental b-roll (crowd, traffic, signage — your cutaway ammo), and an establishing shot to open on. Batch the b-roll between interviews so it never slows the human moment.

— EDITING —

8 · Why it's a different edit

A sit-down is cut to hide the seams between two angles. A street clip is almost always single camera, which means jump-cut energy is native — own it, don't apologize for it. Audiences expect jump cuts in this format; they read as pace, not error. The whole aesthetic is fast, kinetic, and built around one thing: the reaction is the payoff.

9 · The street-clip structure

HOOK (0–3s)   drop into the reaction / a spicy line — the BEST beat, not the first
DOUBT          their skepticism, in their words ("I always thought crypto was a scam")
TURN           the click — the explanation lands, the laugh hits, the "oh!" face
PAYOFF         the gift, the "wait, that's it?" — held, let it breathe
SOFT CTA       light and human — it's a gift, not a pitch

10 · Cutting the human moment

Don't over-cut a real laugh or pause

The kinetic "cut every 2–4 seconds" rule applies to the setup, not the emotional beat. When the genuine reaction lands, get out of its way and let it play. Hold one beat past comfortable — that's where the audience feels it.

Hide jumps, punch in on the reaction

Cover every seam (a removed "um," a stumble) with environmental b-roll, an insert, or a punch-in — that's what your solo coverage was for. A hard cut to a tighter frame on their face at the turn is the single most powerful move in the format; it points the eye exactly at the emotion.

11 · Energy, captions & the dignity edit

12 · One outing → many clips

Street interviews are a batching machine. Film 6–10 approaches in an afternoon and each becomes its own short — a week of posts from one outing. The shared environmental b-roll covers jumps across every clip from that session. Edit them as a batch with one caption style and the structure template above, varying only the human moment.

Solo run-and-gun gear (minimal, fast — tap to check)
  • Phone or small camera, manual exposure + focus locked
  • Hero audio: wireless lav or a handheld dynamic super-cardioid mic
  • Safety audio: on-camera shotgun (or transmitter onboard recording)
  • Furry deadcat + foam windscreen
  • One earbud for monitoring
  • One stabilizer (IBIS handheld / gimbal / monopod), not three
  • Spare batteries + storage; USB-C top-ups
  • Release forms (paper or phone e-sign) for commercial use
  • Something bright to wear
From memory, then reveal.
What are the two rules of the walk-up that most improve your yes-rate?
Approach with the camera down (empty hands, connection over shot) and lead with an observation, not "can I interview you". Then be transparent and fast — speed is kindness because it lowers the cost of saying yes.
Why is a street edit different from a sit-down edit?
It's single camera, so jump-cut energy is native — own it, don't hide it. The aesthetic is fast and kinetic, and it's built around the reaction as the payoff. Protect the emotional beat; cut fast around it.
What's the one rule that governs both how you film and how you cut this format?
Make the subject the hero of their thirty seconds. Film and cut the moment they learn and light up, never their ignorance. If the funniest version makes the person small, kill it. The dignity is the brand.
Sources
CHAPTER 03

Win on sound — the module that separates pro from amateur

Viewers forgive bad video and leave over bad audio. For interviews, where the entire value is a person talking, the words are the product. If you have limited time to finish, spend a disproportionate share of it here. It's the highest-leverage thing in the whole edit.

1 · The dialogue cleanup chain (do it in this order)

Order matters — each stage assumes the last one cleaned its input. Run it as a fixed recipe.

1. Noise reduction   remove hiss / hum / AC / room (6–10 dB, not 30)
2. High-pass filter  roll off below ~80 Hz — kills rumble
3. EQ cut            −2 to −4 dB around 200–400 Hz — de-mud / de-box
4. Compression       even out loud/soft — ~3–4 dB gain reduction
5. De-ess            tame "sss" at 5–8 kHz — 3–6 dB
6. EQ boost          +1 to +3 dB at 2–5 kHz presence, optional air 10 kHz+
7. Limiter           brick-wall safety ceiling at −1 dBTP

Compression, in plain English

It's an automatic volume-rider: when the voice crosses a line you set (threshold), it turns down by an amount you set (ratio). Loud and quiet get closer, so everything sits steady. Starting point for dialogue: ratio 3:1, attack ~10–15ms, release ~40ms, threshold set for ~3–4 dB gain reduction on the loud words, then add makeup gain. Two gentle passes (3 dB each) sound more natural than one that crushes 8 dB.

De-essing

Compression and presence boosts exaggerate harsh "sss." A de-esser is a compressor that only reacts to that band (5–8 kHz). 3–6 dB of reduction on the peaks. Never a fixed EQ cut there — it dulls the whole voice.

Noise reduction — less is more

Give the tool a half-second of just background (that's what room tone is for) and it learns the fingerprint to subtract. Over-reduce and you get an "underwater," warbly artifact worse than the noise. Pull back until you just stop hearing it, then back off a hair.

2 · Room tone — the secret weapon of smooth edits

Room tone is recorded silence — the faint AC, distant traffic, and electrical hum that's the unique sound of a room. When you cut a filler word out, you create a hole of true digital silence, and the ear is exquisitely sensitive to background suddenly dropping to nothing — it screams "EDIT!" You lay room tone under the whole dialogue track so the background is continuous and every cut disappears. This is why you grab 30–60s of it on set.

3 · Loudness — LUFS and why it matters

Your ears judge loudness, but meters historically showed peaks, and they don't match. Platforms fixed this by normalizing everyone to a loudness target — if you don't master to their number, they change your sound for you. LUFS measures perceived loudness; integrated LUFS is the average of the whole video (the number you master to); true peak is the real max (keep at −1 dBTP). Use the interactive loudness reference in the practice studio to pick your target.

Levels inside the mix Dialogue rides around −12 to −16 dB, rock-steady. Music bed under talking sits 18–30 dB below the dialogue. If your intro, interview, and outro aren't loudness-matched to each other, the viewer rides the volume knob the whole time — the #1 tell of an amateur mix.

4 · Music & ducking

Choose music that doesn't fight the voice: avoid busy mid-range instruments (piano, flutes, lead melodies) that collide with speech; favor beds with energy in the lows and highs that leave the midrange clear. No lyrics ever under dialogue. For crypto/finance, aim for clean, modern, slightly-tense-but-optimistic — not hype-y EDM (dates fast, reads "shitcoin ad"), not sleepy elevator music. And silence is a legitimate, powerful choice — don't score the single most important sentence.

Ducking — how music gets out of the voice's way

Ducking automatically lowers music when someone talks. Auto (sidechain): a compressor on the music track triggered by the dialogue — fast, great for talk-heavy edits (one button in Premiere's Essential Sound or DaVinci Fairlight). Manual keyframes: you draw the music down by hand — more musical, the pro move for intros/outros. For spoken dialogue, duck hard — the music should drop 6–12 dB when the voice comes in (not the 1–3 dB you'd use under a singer). Sidechain start: ratio 4:1–8:1, fast-ish attack (~10–20ms), slow release (300–600ms) so it rises gently between phrases instead of pumping.

SFX — seasoning, not the meal

A tasteful whoosh per transition reads pro; a whoosh on every cut plus impacts on every word reads like a teenager who found the SFX folder. If the viewer notices the sound effects, you used too many. Keep 2–3 you reuse, low in the mix.

5 · Music licensing — the part that gets videos claimed

"Royalty-free" does not mean "free" and does not stop a copyright claim. YouTube's Content ID scans every upload against registered audio; a match fires the rights-holder's policy automatically — they can run ads on your video, mute it, or block it — even when you legitimately licensed the track, because the distributor also registered it. Creative Commons CC-BY requires specific per-track attribution, and credit alone does not stop a claim. See the music & SFX sources in the swipe file for a safe library comparison.

How to actually avoid claims Use a library with channel safelisting (Uppbeat, Epidemic) that auto-clears claims on your uploads, keep your license receipt to dispute anything that slips through, and remember subscription coverage (Epidemic/Artlist) ends when you cancel — old videos get exposed.

6 · Sync & backup

Your good audio usually comes from an external recorder or a mic that isn't the camera, so you marry it to the picture in the edit ("double-system sound"). Waveform auto-sync (Premiere Synchronize, Resolve Auto Sync, PluralEyes) lines them up — which is why you always let the camera record its own scratch audio as the reference. A clap at the top of each take gives a dead-obvious manual sync point. And always record at least two sources — there's no re-shooting a founder who already flew home.

Pre-export audio checklist (tap to check)
  • Room tone under the whole dialogue track; every hard silence patched
  • Noise reduction applied — no "underwater" artifacts
  • High-pass (~80 Hz) on every voice; mud (200–400 Hz) tamed; presence (2–5 kHz) up
  • Compression — steady, ~3–4 dB reduction, natural not squashed
  • De-essing — harsh "sss" tamed
  • Dialogue riding −12 to −16 dB, consistent across the whole piece
  • Music ducked 6–12 dB under dialogue; releases smoothly, no pumping
  • Nothing covers a key word; SFX tasteful, not constant
  • Limiter on the master at −1 dBTP; nothing clips
  • Loudness normalized to target; sections matched to each other
  • Checked on phone speakers and cheap earbuds, not just headphones
From memory, then reveal.
What's the correct order of the dialogue cleanup chain?
Noise reduction → high-pass → EQ cut (mud) → compression → de-ess → EQ boost (presence) → limiter. Each stage assumes the last one cleaned its input.
What loudness should you master an interview to for YouTube + social, and what's the true-peak ceiling?
−14 LUFS integrated, −1 dBTP. That satisfies YouTube and Spotify and sits inside Apple's podcast range. Dialogue rides −12 to −16 dB in the mix.
Why can a track you legally licensed still get claimed, and how do you prevent it?
The distributor also registered it in Content ID, so a match auto-fires. Prevent it with a library that offers channel safelisting (Uppbeat/Epidemic) and keep your license receipt to dispute.
Sources
CHAPTER 04

The cut — editing a conversation into an argument

The camera gave you a rambling forty-minute talk. Your job is to find the ten-minute story hiding inside it, and the six fifteen-second clips hiding inside that. Here's the craft, in the order you actually do it.

1 · Transcribe first, then edit the text

The modern interview edit begins in a transcript, not a timeline. Premiere, DaVinci Resolve, and Descript all support text-based editing — you cut the video by deleting words on the page. Get a clean transcript for every interview and reason about the material as text. It's the difference between a 3.5-hour selection pass and a 40-minute one.

2 · The paper edit, then the radio edit

A paper edit is a document listing, in story order, the soundbites that form the spine of your cut — chosen before you touch the timeline. Pull each usable bite with its timecode, tag it with the story beat it serves (setup, conflict, turn, payoff), then reorder until it works on the page. The test: read it out loud. If it doesn't make sense as text, it won't make sense in the cut.

Find the spine, kill the tangents A crypto interview wanders — regulation, a token launch, a 2021 war story. Pick the one throughline the video is about and treat every soundbite as either serving that spine or being a tangent to cut. Then assemble the paper edit into a radio edit: build the audio story first, eyes closed, ignoring visuals. Fix story problems here where a change costs seconds — only once it plays as a story do you touch picture.

3 · Pacing — making the cut invisible

Don't over-cut the silence

Remove filler clusters and long gaps, not every micro-pause. Aim for a cut that's 80–90% of the original length, not 60–70% — the deepest cuts read as over-edited. Keep the breath at the end of a sentence and the pause between topics that lets the viewer reset. After a big statement, let the moment breathe.

J-cuts and L-cuts — the seamless-conversation trick

Unlink audio from video and offset them. A J-cut: the next audio starts before its picture (you hear the answer begin over the previous shot). An L-cut: the current audio continues after the picture cuts away (you keep hearing them over a reaction or b-roll). The half-second-to-two-second overlap is what makes dialogue feel like a real conversation instead of two people alternating on a stage. Mechanic: unlink, nudge one earlier or later, listen, adjust.

4 · Murch's Rule of Six — how to judge a cut

Walter Murch's priority list for where to cut, in descending importance:

1 Emotion — true to the feeling? 2 Story — advances it? 3 Rhythm — feels right? 4 Eye-trace — respects where we look? 5 Planarity — 2D grammar 6 Spatial continuity — 3D space

Emotion alone is worth more than the other five combined. If a cut forces a sacrifice, sacrifice from the bottom up. For interviews this is freeing: a jump cut that "breaks continuity" (rank 6) is fine if the moment it preserves is emotionally true (rank 1). The founder's voice cracking beats a technically clean cut every time.

5 · Multicam & b-roll — your invisible-mend kit

Sync your two cameras into a multicam sequence and cut angles on a shift — a new question, a punchline, a change of energy — never at random. Every time you tighten an answer you create a jump; cutting to the second angle (or the listener's reaction) at that exact frame hides it completely. That's why interviews are ideal multicam candidates. One camera? B-roll and cutaways hide the same jumps, and Premiere's Morph Cut / Final Cut's Flow blend minor jumps optically.

Time b-roll to the words

When the guest names a token, a chart, an exchange, a hardware wallet — cut to it on the word. Don't blanket the whole interview; place b-roll where it earns its cutaway, and pull back to the face for the emotional beats where it matters most. Build a reusable crypto b-roll library: chart screen-recordings (TradingView, DEX Screener), block-explorer scrolls, wallet UIs, conference-floor establishers.

6 · The cold open — engineering the first 5–10 seconds

The first ten seconds decide whether the next ten minutes get watched. Retention above 75% in the opening 30s is strong; below 60% means the hook is broken. So pull the best 10 seconds of the interview to the very front as a cold-open montage of the sharpest lines, then drop into your title. While cutting, flag great lines in real time — that flag list is your cold open and your clip list.

7 · Color — a simple, repeatable grade

Correct and match first, grade last. Correct each camera to a neutral baseline, then match cameras on the scopes (load the A-cam in a wipe, balance the B-cam to it on the parade and vectorscope; a Color Space Transform matches more accurately than a baked LUT). Fix skin tones with a power window / magic mask onto the vectorscope's skin line. Apply your look last, and lock a base look for a repeatable house style across every interview.

8 · Export & the software call

DeliverableSettings
YouTube 1080p30H.264, ~8–15 Mbps, AAC 384kbps 48kHz
YouTube 4K30H.264, ~35–45 Mbps
Vertical clips9:16, 1080×1920, H.264, high-bitrate master
The honest software call for a solo interview creator Cut long-form and grade in DaVinci Resolve — the best free editor (no watermark, no time limit, 4K export, best-in-class color, text-based editing built in). Knock out vertical clips in CapCut for its caption/AI speed. Premiere ($60/mo) if you're already in Adobe; Final Cut ($299 once) if you're all-Mac.
From memory, then reveal.
What's the difference between a J-cut and an L-cut, and why do they matter for interviews?
J-cut: next audio starts before its picture. L-cut: current audio continues after the picture cuts away. The half-second-to-two-second overlap makes dialogue feel like a real conversation instead of alternating turns.
In Murch's Rule of Six, what wins when a cut forces a trade-off?
Emotion — it's worth more than the other five combined. Sacrifice from the bottom up (give up spatial continuity before you ever compromise emotion). A jump cut is fine if it keeps an emotionally true moment.
You cut a filler word and get a jump. What are your three fixes?
Cut to the second camera angle, cut to b-roll / a cutaway timed to the words, or use Morph Cut / Flow to blend the jump. Interviews are ideal multicam candidates for exactly this reason.
Sources
CHAPTER 05

One interview → a week of content

A 45-minute interview isn't one video, it's a content deposit you draw down for a week. The long-form is the anchor; the clips are how anyone finds it. A one-hour episode reliably yields 10–20 usable clips. The craft is in selection and packaging, not in shooting more.

1 · What to pull

You're hunting self-contained moments that survive without context, roughly in order of performance:

  1. The contrarian take — cuts against consensus ("ETFs are the worst thing that happened to crypto"). Provokes replies.
  2. The number / the stakes — a concrete figure that stops the scroll ("I lost $2M in the Luna collapse").
  3. The story — a 30-second narrative with tension and payoff. These outperform explanations, and they're the ones AI clippers miss.
  4. The disagreement — host and guest openly clash. Conflict is watchable.
  5. The operator insight — "here's exactly how I size a position." Gets saved and shared, which platforms weight heavily.
Mine while you cut, not after Keep a timestamp log open during your first watch. Mark it every time you physically react, or when the guest says "the truth is…", "nobody talks about…", "here's what people get wrong…" — those signposts almost always precede a clip. Leave the session with 8–12 marked moments, then cut to the 5–7 you'll publish. (There's a clip-mining scratchpad in the practice studio.)

2 · Vertical clips & captions

Interviews shoot wide; every clip becomes 9:16. Premiere Auto Reframe and Resolve Smart Reframe track the speaker automatically — genuinely good for a single speaker, but they choke on two people talking over each other (expect manual fixes on ~30% of two-shot clips). For rapid back-and-forth, a split-screen stack (guest top, host bottom) is safest; a manual punch-in on the reactor at the punchline is the high-craft move AI can't do.

Caption style that works

Most short-form is watched on mute, so captions are the delivery. Bold sans-serif (Montserrat/Proxima/Impact-style), 55–75pt, vertically centered in the middle 60–65% safe zone (top 15% is the status bar, bottom 20% is the like/comment UI). 1–2 lines, 3–5 words each. White base + one accent, with karaoke word-highlighting (the active word changes color as spoken) — the single biggest engagement lever. Different caption color per speaker so muted viewers follow the hand-off. And edit the auto-caption errors — the ~8% it misses is disproportionately tickers, protocol names, and dollar figures ("SOL" heard as "soul").

3 · The honest AI-tools verdict

AI does the mechanical 60% and none of the editorial 40%. It transcribes, finds candidates, reframes, and captions fast. It cannot judge what will perform, protect a comedic pause, fix an overlapping two-shot, or catch a caption typo. Every credible run ends with a full manual QA pass.

ToolForTrust it for
Descripttranscript editing, filler removal, Studio SoundThe talking-head master edit and cleanup
Opus Clipauto-finding viral momentsFirst-pass clip candidates — then you curate
Submagic / Captionsanimated captionsFast on-trend English caption styling
Auphonicaudio mastering / levelingFinal audio pass; fixes loud-host/quiet-guest
CapCutfree editor + captionsBudget clips (customize templates so it's not stock)
Resolve / Premierepro finishingHero clips and the long-form master
The division of labor: AI drafts, human decides. Transcription, filler removal, clip candidates, and caption generation save real hours. Auto-selecting final clips, unattended two-shot reframes, and unedited captions create slop.

4 · Thumbnails, titles & native posting

The clips feed the long-form; the long-form lives on thumbnail + title. Thumbnail: a face with legible emotion, ≤3 elements, big 3–5 word text that pops at feed size. The quote format works unusually well for interviews. Title formula: guest name/credential + the bold claim — "Arthur Hayes: The Dollar Ends in 2028." The name is the trust anchor, the claim is the hook, biggest promise in the first ~40 characters.

Post native everywhere, rewrite the wrapper Upload the actual file to each platform (algorithms suppress off-platform links). Keep the clip identical; rewrite the caption/hook/CTA per platform. YouTube = SEO titles + Shorts driving back to the long-form. TikTok = casual + trending sounds. Reels = 7–30s, on-screen text essential. X native = pair the clip with the quote, where contrarian takes spark replies. LinkedIn = the same clip with a professional, value-first caption for the finance room.

5 · Batching, specs & never losing footage

From memory, then reveal.
What kinds of moments make the best clips, and how do you catch them?
Self-contained moments: contrarian takes, concrete numbers, short stories, disagreements, operator insight. Catch them by keeping a timestamp log during your first watch and marking every reaction and verbal signpost ("the truth is…").
Where do captions go on a 9:16 clip, and what's the highest-engagement style?
Vertically centered in the middle 60–65% safe zone (clear of the top status bar and bottom UI). Bold sans-serif, 1–2 lines, 3–5 words, white + one accent, with karaoke word-highlighting. Edit the auto-caption errors on names/tickers/numbers.
What does AI editing genuinely save, and where does it create slop?
Saves: transcription, filler removal, clip candidates, caption generation. Slop: auto-selecting final clips, unattended two-shot reframes, unedited captions. AI drafts, human decides — always a manual QA pass.
Sources
HANDS ON

The practice studio

The tools you'll actually reach for while editing. Everything runs right here in the page — nothing uploads, nothing leaves your browser.

▸ Loudness & levels reference
Pick where the video is going and get the exact numbers to master to. Set your master to the integrated LUFS target and keep true peak at −1 dBTP.
▸ Editor keyboard-shortcut trainer
Speed comes from your hands, not your mouse. Pick your editor, tap the card to flip, and drill the shortcuts that matter most for interview cutting.
tap to reveal · then tap Next
▸ Rate your last edit
Open the last interview you cut and score it honestly on the five things that separate a pro edit from an amateur one. It remembers your first score so you watch it climb over the 30-day plan.
StoryOne clear throughline, tangents cut?
Pacing80–90% length, J/L cuts, breathes?
AudioClean, steady, −14 LUFS, ducked music?
HookBest 10s pulled to the front?
FinishColor matched, captions clean, exported right?
0 / 25
▸ Clip-mining scratchpad
Keep this open on your first watch. Every clip-worthy moment, drop a line: 12:40 — "ETFs were a mistake" (contrarian). It saves as you type, so your clip list is ready before you finish the long-form.
saves in your browser
The loop this whole thing runs on Get clean source → cut for story first → finish the audio → publish, then mine the clips → watch it back and pick ONE thing to do better next time. That single loop, run honestly, is the entire difference between people who get good and people who plateau.
SWIPE FILE

Music, SFX & title formulas

Where to get music that won't get you claimed, and the title/thumbnail patterns that get interviews clicked. Tap any for the detail.

Safe music & SFX sources

FREEYouTube Audio Library · Pixabay · Mixkit +

$0, generally claim-free. YouTube's own library never claims your videos; Pixabay and Mixkit are royalty-free with commercial use and no attribution. Mixkit has 42 free transitions and 20 free whooshes for SFX.

Gotcha: smaller, more generic catalogs, and some YT Audio Library tracks require attribution. Overused tracks can sound stock.

~$7/moUppbeat — budget with safelisting +

Freemium + sub. The safest budget pick: channel safelisting auto-clears claims on your uploads, and downloads stay covered even after you cancel. Free tier needs a credit and has a monthly download cap.

~$17/moEpidemic Sound — depth + safelisting +

Subscription (~$17/mo personal, ~$30 business). Huge, well-organized catalog with SFX; registers in Content ID but auto-clears subscriber claims. Gotcha: coverage ends when you cancel — old videos get exposed.

~$40/moArtlist — universal license +

Subscription, all-in-one music + SFX. Universal license, claim-clearing. Covers content made while subscribed. Deepest cinematic catalog of the three.

READFree Music Archive / Creative Commons +

Mostly CC, $0 — but read each track's exact terms. CC-BY requires specific per-track attribution (artist, track, license), and credit alone does not stop a Content ID claim. An artist can register with Content ID later and claim you months after you published, even when you used it correctly.

Title formulas for interviews

01Name + bold claim +
"[Guest name]: [contrarian claim]"

Example: "Arthur Hayes: The Dollar Ends in 2028." The name is the search-and-trust anchor; the claim is the hook. Keep the biggest promise in the first ~40 characters so it survives truncation.

saves in your browser
02The quote thumbnail +
Face with legible emotion + a curiosity-driving line in quotes

Why it works for interviews: a quote creates a reason to stop, and the guest's face lends credibility. ≤3 elements, big 3–5 word text that pops at feed size on your own screen.

saves in your browser
03The number hook +
"[Guest]: How I [did the thing] with $[number]"

Example: "The trader who turned $5K into $4M — and what he'd never do again." Concrete numbers stop the scroll; the reversal adds curiosity.

saves in your browser
AVOID THESE

The 12 mistakes that kill an interview video

Fixing these is faster than adding anything new. Every one is covered above — this is the checklist version.

Watch for these before you export
  • Bad on-set audio. The one thing post can't fix. Two mics, lav close, monitor on headphones, room tone.
  • No second angle or b-roll. Then every tightened answer is a naked jump cut. Shoot coverage.
  • Editing the timeline before the story. Do the paper edit and radio edit first — fix story where it costs seconds.
  • Over-cutting. Stripping every pause makes people sound robotic. Leave 80–90%, keep the breaths and topic beats.
  • No cold open. Burying the best line at 12:00. Pull the best 10 seconds to the front.
  • Ping-pong cuts. Straight A/B alternation. Use J-cuts and L-cuts so it feels like a conversation.
  • Music fighting the voice. Lyrics or busy mid-range under dialogue, or music not ducked. Duck 6–12 dB, no lyrics.
  • Loudness all over the place. Intro loud, interview quiet. Match sections, normalize to −14 LUFS.
  • Mismatched color. Two cameras that don't match, orange-and-blue skin. Correct and match before you grade.
  • Auto-captions left unedited. "SOL" as "soul," wrong tickers and numbers. Fix the names, they're the whole point.
  • Captions in the UI dead zone. Behind the like/comment buttons. Keep them in the middle 60–65% safe zone.
  • Letting the AI publish for you. Auto-selected clips and unattended reframes are slop. AI drafts, you decide.
THE SYSTEM

Your 30-day editing plan

One technique a day, 20–40 minutes. The goal of month one isn't a masterpiece — it's to make the pro workflow automatic. Tap a day to check it off; your progress saves in this page.

◈ Lock in your trigger
If it's then I

"If [situation], then I [action]" beats "I'll try to edit more." Pin the rep to a moment that already happens, and it saves here so it greets you tomorrow.

0 / 30
Track these, honestly — vibes lie, numbers don't
  • Edit score (out of 25) — from the "rate your last edit" tool. Watch the trend climb.
  • Time-to-cut — how long a long-form + clips takes you. It should drop as templates and shortcuts land.
  • Clips per interview — are you actually mining 5–7, or leaving them on the table?
  • Loudness — did you hit −14 LUFS, or eyeball it? Measure every time.
  • The ONE technique this week — name it, so you know what you were drilling.
ONE PAGE

The cheat sheet

Everything that matters, on one screen. Screenshot it.

The 10 that move the needle most
  • Win on audio first. Two mics, lav close, room tone. It's the one thing viewers won't forgive.
  • Shoot two angles + coverage. A wide, a tight, b-roll, noddies — the edit assembles itself.
  • Get them to answer in full sentences. Fold the question into the answer.
  • Cut the story before the picture. Paper edit → radio edit → then visuals.
  • J-cuts and L-cuts make dialogue feel like a conversation, not a ping-pong match.
  • Emotion beats a clean cut (Murch). Sacrifice from the bottom up.
  • Pull the best 10 seconds to the front. No cold open = buried lede = they leave.
  • Cleanup chain in order, normalize to −14 LUFS, duck music 6–12 dB under the voice.
  • Karaoke captions in the safe zone, and fix the names and tickers.
  • One interview → 5–7 clips. AI drafts, you decide. Watch it back, pick ONE thing to improve.
The whole thing in one sentence Great editing starts on set: get clean sound and real coverage, cut for story first, finish the audio like it's the product, and turn every interview into a week of clips.
← Back to the learn portal