Turn a rambling interview into a video people finish.
Whether you're sitting a founder down or walking up to a stranger on the street, the value is in what you do after: finding the story hiding in the footage, making the sound clean enough that nobody clicks away, and cutting the clips that actually get seen. This is the whole craft, free and in the open — how to film so the edit is easy, the run-and-gun street interview as its own discipline, how to win on audio, how to cut, and how to turn one outing into a week of content. Sourced technique, live tools, and a 30-day plan that saves your progress in this page.
Don't try to learn all of it at once.
This page is deep on purpose. But if you do one thing after reading, do this, because it's the single biggest jump from amateur to pro:
- On your next interview, record two audio sources — a lav clipped 6–8 inches below the chin, plus one backup (camera mic or a second recorder).
- Grab 30 seconds of room tone (everyone silent) before you leave.
- In the edit, run the cleanup chain in order and normalize to the right loudness for your platform.
Clean, consistent, loud-enough sound already puts you ahead of most creators. Everything else on this page builds from there.
How to use this
The chapters run in the order you actually work: shoot → sound → cut → clip. The loop that runs through all of it: get clean source → cut for story first → finish the audio → publish, then mine the clips. Skim it once, then work the 30-day plan — one technique a day. The practice studio has the tools you'll actually reach for while editing.
Shoot it right — the edit is won on set
Every clean cut and every emotional close-up is decided before you press record. The footage you hand the edit is either a gift or a life sentence. Shoot for the cut you don't yet know you'll need.
1 · Two cameras change everything
One camera is viable, but it punishes you: every time you tighten an answer or cut an "um," you get a jump cut, and your only fixes are b-roll, a dissolve, or living with it. That's why the two-camera interview is the documentary standard. Run an A-cam wide/safe (chest-to-head) and a B-cam tight (close-up on the face). Cutting between them hides every edit and gives the tight angle for emotion.
Match your two cameras
Same body or same family. If you must mix brands, at least match resolution, frame rate, shutter, white balance in Kelvin, and picture profile, and shoot a gray card on both so the grade has a reference. Mismatched color science is a nightmare in post.
Interviewer placement
Put the interviewer just off one side of the A-cam lens, and outside the spread of the two cameras. If they sit between the lenses, the subject's eyeline lands in dead space and the geography feels wrong.
2 · Framing that flatters
- Eyeline is the biggest framing decision. Slightly off-lens (subject looks at the interviewer beside the camera) = the documentary look, feels overheard. Straight down the lens = direct address, feels personal. Pick one and hold it on both cameras.
- Rule of thirds: eyes on the upper-third line, subject off-center for storytelling. Centered reads formal/confrontational — save it for testimonials.
- Look room: when they're turned to one side, leave more space in the direction they face, or the shot feels cramped.
- Height: lens at eye level or a hair above. Never shoot up (aggressive) or down (diminishing).
- Separation: pull the subject several feet off the background so it goes soft, and add a backlight. Flat-against-a-wall is the "HR video" look.
3 · Lens & settings
| Setting | Interview default | Why |
|---|---|---|
| Focal length | 50–85mm equiv | Flattering compression; under 35mm distorts faces |
| Aperture | f/2.8 (f/1.8–4) | Nice separation but enough depth that a lean doesn't go soft. Don't shoot wide open at f/1.2 |
| Shutter | 2× frame rate | 180° rule: 1/48 for 24fps, 1/60 for 30fps. Natural motion blur |
| Frame rate | 24 or 30fps | 24 = cinematic, 30 = broadcast feel. Shoot b-roll at 60 for slow-mo |
| ISO | native / lowest | Light the scene instead of cranking ISO and adding grain |
| White balance | manual Kelvin | Auto WB drifts shot-to-shot and wrecks camera matching |
4 · Lighting
Three lights, three jobs. Key — main source, ~45° off the subject's line of sight and slightly above eye level angled down ~15–30°. Fill — softer, opposite side, to lift the shadow (a bounced reflector does this for free). Back/hair light — behind and above, hitting the shoulders to separate them from the background. That backlight is what makes an interview look "produced."
The one-light look
You don't need three. One big soft key at 45° (softbox, umbrella, or a window), a reflector for fill, and a small backlight if you have one. Bigger and closer = softer = more flattering.
- Two shadows — two hard lights of similar strength cross competing shadows on the face. Keep one dominant key, keep fill soft and secondary.
- Mixed color temperature — a 5,600K window plus a 3,200K lamp makes one side of the face orange and the other blue, and it's brutal to fix. Kill the window, gel your lights to match, or go all-daylight LED on one white balance.
- Skin tones: darker skin benefits from a touch more key and a defined rim light for separation; lighter skin clips to white faster, so protect highlights. Expose for the face, use soft sources, judge on a real monitor.
5 · On-set audio — the part you can't save later
Bad on-set audio is unrecoverable, so this is where you're most disciplined. The pro standard is two mics: a wireless lav on the subject plus a shotgun on a boom overhead. You pick the better one in the edit, and if one drops out you don't lose the take.
Placement
Lav: 6–8 inches below the chin, centered, clear of scarves, necklaces, and conference lanyards that rub. Shotgun: overhead in front, pointed down at the mouth from 6–12 inches, just out of frame — overhead-and-down rejects room reflections. Close is king.
Levels & monitoring
Peak around −12 dB (loud moments not past ~−6). Clipping at 0 dB is permanent distortion — err low, normalize later. Wear closed-back headphones and listen the whole time; meters won't warn you about a rustling lav or a distant siren, your ears will.
Room tone, 30 seconds, every location
Before you break down, everyone holds still and you record 30–60s of the room's "silence." The editor uses it to patch gaps and smooth cuts so there's no jarring dead air. One minute on set, saves the whole edit.
6 · Directing the guest
Your job is a comfortable person giving self-contained answers, not a great chat that cuts into nothing.
- Make them comfortable: share the theme and rough questions beforehand, reassure them it's edited and they can retake as much as they want. Chat casually while you set levels.
- Elicit stories, not soundbites: "walk me through the day the fund almost blew up" beats "was that stressful?" Ask for specific moments, then "tell me more about that."
- Let silences sit. When they finish, don't jump in — the beat of silence often pulls the most honest line, and it gives the editor a clean gap with no crosstalk.
- The "ask again": if an answer is gold but rambling, "that was great, give me that once more, a little tighter." A cleaner take is worth the minute.
7 · Run-and-gun at events
Conferences are chaos — PA, crowds, music. A tripod (or a monopod for speed) is still best for the sit-down; a gimbal is for moving b-roll. Solo rig: a 28–75mm zoom (wide-to-tight, no lens swaps), a wireless lav, a small on-camera LED, a monopod, exposure and white balance pre-set so you can grab an interview in under a minute. Win the audio by getting the lav close, not by tightening the camera. Get clear verbal consent on camera. A modern phone with good sound and light beats a cheap camera — audio and framing read as "pro," not the sensor.
8 · Gear — where to spend first
The honest spend order: audio + lens + light before the camera body. Bad sound and bad light can't be fixed in post; a modern-but-modest body is more than enough. A phone with good sound and light beats a cheap camera every time.
| Tier | Kit | Rough spend |
|---|---|---|
| 1 · Phone-first | Phone (flat/log, locked exposure) · a wireless lav that pairs to it (spend here first) · one small daylight LED or a window · a clamp + mini tripod or compact gimbal | $150–500 |
| 2 · Prosumer | One mirrorless body · a 24–70 f/2.8 or a 35mm + 85mm pair · wireless lav plus shotgun + boom · a 2-light soft LED kit + reflector · fluid-head tripod + monopod/gimbal | $2,000–3,500 |
| 3 · Pro two-camera | Two matched bodies (wide + tight) · 35/50mm + 85/135mm · multitrack field recorder + two lavs + boom + monitoring headphones · 3-point soft LED kit with hair light + gels · two tripods + gimbal + gray card | $8,000+ |
- The wide/safe (A-cam) — runs the whole interview, your fallback
- The tight (B-cam) — close-up on the face, for emotion and hiding cuts
- Noddies — the interviewer nodding/listening (no talking), shot after
- Cutaways / b-roll — hands, the charts/screens they mention, the room
- Detail inserts — hands, watch, phone, keyboard (60fps if you can)
- Establishing shots — the building, the room, them walking in / sitting down
- 30–60s of room tone — every location, before you strike
- A slate / verbal ID — who this is and the date, so the editor isn't guessing
- Two-camera / 30° / focal length — Doc Film Academy, Shutterstock
- Framing / eyeline / look room — Wes Jones
- Settings / log — Swole Nerd Productions; 180° rule — Camera Settings
- 3-point lighting — iKan, NBCU Academy
- Mics / on-set audio — Lensrentals, Stillmotion; directing — The Open Notebook
The street interview — the run-and-gun human moment
Most interview advice assumes a scheduled subject, a lit room, and time. The street throws all of that out. You walk up to a stranger who didn't plan to be filmed, you have ten seconds to earn a yes, and you're the director, camera op, sound mixer, and the person making them laugh — all at once, in ninety seconds, before the light changes. It's its own craft. This is the format behind the 360 clips in the road archive.
— FILMING —
1 · The approach — earning an instant yes
The highest-leverage skill happens before you roll. Pros who do this for a living quote roughly one yes in five, and the rule is: don't take a no personally. You nudge the ratio with speed and warmth.
Approach with the camera down
The most-repeated pro tip: don't have the camera up or pointed as you walk up. Empty hands, a smile, open body language — it signals the connection matters more than the shot, and it's disarming.
Lead with an observation, not "can I interview you?"
A cold "can I interview you?" invites a reflexive no. Open with something human to respond to — "looks like you had a good shopping trip" — then the ask. Give them a half-second of rapport first.
Be transparent, fast
In the first sentence they should know who you are, what this is, and that it's quick and friendly. Transparency removes the "what's the catch?" tension. Speed is kindness — a long hedging approach makes a stranger anxious; a fast, clear, warm one respects their time and lowers the cost of yes.
- Camera down / not pointed; hands relatively empty
- Opening line ready — an observation or compliment, not "can I interview you"
- One sentence of who / what / why-quick rehearsed
- Mic on and levels already set — don't fumble gear in front of them
- Exposure / focus / white balance pre-set for the light you're in
- An easy first question that makes them smile or feel smart
- A graceful exit line for a no ("Totally fair — have a great one!")
2 · Consent, ethics & dignity
- Get consent on camera. State who you are and what it's for, then capture a verbal yes on tape — and have them say and spell their name at the top. That doubles as your lower-third source and your clearance record. For anything commercial/brand-attached, back it with a signed release.
- The legal floor (general, not legal advice): recording in public is broadly lawful in the US, but the trap is audio — ten states are all-party-consent (CA, FL, IL, MD, MA, MI, MT, NH, PA, WA), so get everyone's spoken yes before capturing the conversation. Filming children almost always needs a guardian's permission. When in doubt, ask.
- Respect a no gracefully. One yes in five means four graceful exits. "I understand — thank you!" and move on. Never argue, never film someone who declined, never make the refusal into content.
3 · The one-person-band rig
You're interviewing and operating at the same time, so the rig has to disappear and your attention stays on the human. Keep it minimal — a phone or small camera, one good audio solution, one stabilizer. Every extra device is another thing to monitor while you're supposed to be listening.
Pick one stabilizer
Handheld / phone with IBIS — fastest to deploy, most nimble, best for reaction punch-ins. Gimbal — buttery movement for walking-with-the-subject, but it occupies a hand and costs seconds to balance. Monopod — underrated: height, quick resets, some stability, no gimbal overhead. One, not three.
Phone-first can look pro
Shoot the flattest/highest-quality profile, lock exposure and focus manually (auto hunts and pumps when the crowd moves behind your subject), and let a wireless mic carry the sound. The image rarely gives you away on a street clip — the audio does.
Monitor without losing the person
One earbud in (the subject's mic) keeps you connected to both the sound and the human. If you truly can't monitor, record a safety track onboard the transmitter.
4 · Framing the two-person moment
- The arm's-length two-shot (both faces, mic visible between you) — intimate and honest, the native "we're in this together" look.
- Over-the-shoulder on the subject — puts the viewer in your shoes, keeps them the focus, hides you.
- Clean single on the subject (you off-camera) — easiest to make them comfortable; a subject talking to a person rather than into a lens performs far more naturally.
- Leave punch-in room — frame slightly wider than you need so you can crop into the reaction in post without losing sharpness.
- Light the face with what you have — put them facing the open sky, not backlit against it. Moving backgrounds (crowds, traffic) sell "we're really out here" and give the edit cutaway cover.
5 · Audio in chaos — the hardest part
On the street, image forgives and audio does not. Prioritize the subject's voice above everything.
| Tool | Role | The catch |
|---|---|---|
| Wireless lav | The workhorse — consistent, focused voice, out of frame, moves with them | A few seconds to clip on; put it at the collar, not buried under clothing |
| Handheld interview mic | The secret weapon — points a directional capsule at the talker, rejects crowd, looks legit and relaxes a nervous subject | One hand occupied; you pass it between speakers |
| On-camera shotgun | Hands-free — usually your backup/ambience track | Farther from the mouth = more street in the mix |
6 · Directing a normie in 60 seconds
Your subject is not an actor and has zero prep, so comfort starts the second you begin interacting, not once the interview begins.
- Call it a conversation, not an interview — out loud. People relax when they feel heard, not tested.
- Let them look at you, not the lens. A subject talking to a friendly face performs; one staring into glass freezes.
- Don't interrupt to redirect — cutting someone off mid-answer kills their rhythm. Note it, circle back.
- Escalate your questions — general → specific → emotional. When you get specific enough that they have to think, you've hit the spot; that thinking pause is where the genuine reaction lives.
- Engineer the payoff. Ask the doubt out loud ("what did you always think about crypto?") so the turn has something to turn from, then film the click when doubt becomes delight. And listen visibly — your reaction on camera is half the clip.
7 · Solo coverage so the edit works
You can't cut energy you didn't shoot. Even solo, grab these every interaction: the reaction (their face at the turn — over-cover it, it's the payoff), the wide (both of you, "this is real, on a real street"), the detail (the phone screen, the gift, the handshake), environmental b-roll (crowd, traffic, signage — your cutaway ammo), and an establishing shot to open on. Batch the b-roll between interviews so it never slows the human moment.
— EDITING —
8 · Why it's a different edit
A sit-down is cut to hide the seams between two angles. A street clip is almost always single camera, which means jump-cut energy is native — own it, don't apologize for it. Audiences expect jump cuts in this format; they read as pace, not error. The whole aesthetic is fast, kinetic, and built around one thing: the reaction is the payoff.
9 · The street-clip structure
HOOK (0–3s) drop into the reaction / a spicy line — the BEST beat, not the first
DOUBT their skepticism, in their words ("I always thought crypto was a scam")
TURN the click — the explanation lands, the laugh hits, the "oh!" face
PAYOFF the gift, the "wait, that's it?" — held, let it breathe
SOFT CTA light and human — it's a gift, not a pitch
10 · Cutting the human moment
Don't over-cut a real laugh or pause
The kinetic "cut every 2–4 seconds" rule applies to the setup, not the emotional beat. When the genuine reaction lands, get out of its way and let it play. Hold one beat past comfortable — that's where the audience feels it.
Hide jumps, punch in on the reaction
Cover every seam (a removed "um," a stumble) with environmental b-roll, an insert, or a punch-in — that's what your solo coverage was for. A hard cut to a tighter frame on their face at the turn is the single most powerful move in the format; it points the eye exactly at the emotion.
11 · Energy, captions & the dignity edit
- First 1–2 seconds decide everything. Lead with your strongest frame or line — no slow logo, no "hey guys."
- Kinetic connective tissue, one slow beat. Roughly a cut every 2–4 seconds with pattern interrupts to reset attention — except across the payoff, which you protect. Contrast is the tool: fast everywhere makes the one held beat hit harder.
- Captions are non-negotiable — most people watch muted, especially in public. Large, high-contrast, animated word-by-word, and on a noisy clip they rescue any line where the street won.
- Keep some ambience on purpose. Prioritize the subject's voice, but don't scrub it sterile — a bed of real street authenticates the clip. (A car horn over a key word is unrecoverable, which is exactly why you captured a lav and a safety track.)
- The dignity edit: choose the take where they look smart, warm, or delighted — never foolish. Use their discovery as the payoff, never their confusion as the punchline. If the funniest cut makes the person small, don't make it.
12 · One outing → many clips
Street interviews are a batching machine. Film 6–10 approaches in an afternoon and each becomes its own short — a week of posts from one outing. The shared environmental b-roll covers jumps across every clip from that session. Edit them as a batch with one caption style and the structure template above, varying only the human moment.
- Phone or small camera, manual exposure + focus locked
- Hero audio: wireless lav or a handheld dynamic super-cardioid mic
- Safety audio: on-camera shotgun (or transmitter onboard recording)
- Furry deadcat + foam windscreen
- One earbud for monitoring
- One stabilizer (IBIS handheld / gimbal / monopod), not three
- Spare batteries + storage; USB-C top-ups
- Release forms (paper or phone e-sign) for commercial use
- Something bright to wear
- Vox pop / approach — NBCU Academy, Writer's Digest, Benjamin Wiesner
- Run-and-gun / solo — PremiumBeat, Doc Film Academy, Artlist
- Street audio — SYNCO, Hollyland, Desktop Documentaries
- Consent / law — FindLaw, Freedom Forum; comfort — LAI Video
- Short-form retention — OpusClip
Win on sound — the module that separates pro from amateur
Viewers forgive bad video and leave over bad audio. For interviews, where the entire value is a person talking, the words are the product. If you have limited time to finish, spend a disproportionate share of it here. It's the highest-leverage thing in the whole edit.
1 · The dialogue cleanup chain (do it in this order)
Order matters — each stage assumes the last one cleaned its input. Run it as a fixed recipe.
1. Noise reduction remove hiss / hum / AC / room (6–10 dB, not 30)
2. High-pass filter roll off below ~80 Hz — kills rumble
3. EQ cut −2 to −4 dB around 200–400 Hz — de-mud / de-box
4. Compression even out loud/soft — ~3–4 dB gain reduction
5. De-ess tame "sss" at 5–8 kHz — 3–6 dB
6. EQ boost +1 to +3 dB at 2–5 kHz presence, optional air 10 kHz+
7. Limiter brick-wall safety ceiling at −1 dBTP
Compression, in plain English
It's an automatic volume-rider: when the voice crosses a line you set (threshold), it turns down by an amount you set (ratio). Loud and quiet get closer, so everything sits steady. Starting point for dialogue: ratio 3:1, attack ~10–15ms, release ~40ms, threshold set for ~3–4 dB gain reduction on the loud words, then add makeup gain. Two gentle passes (3 dB each) sound more natural than one that crushes 8 dB.
De-essing
Compression and presence boosts exaggerate harsh "sss." A de-esser is a compressor that only reacts to that band (5–8 kHz). 3–6 dB of reduction on the peaks. Never a fixed EQ cut there — it dulls the whole voice.
Noise reduction — less is more
Give the tool a half-second of just background (that's what room tone is for) and it learns the fingerprint to subtract. Over-reduce and you get an "underwater," warbly artifact worse than the noise. Pull back until you just stop hearing it, then back off a hair.
2 · Room tone — the secret weapon of smooth edits
Room tone is recorded silence — the faint AC, distant traffic, and electrical hum that's the unique sound of a room. When you cut a filler word out, you create a hole of true digital silence, and the ear is exquisitely sensitive to background suddenly dropping to nothing — it screams "EDIT!" You lay room tone under the whole dialogue track so the background is continuous and every cut disappears. This is why you grab 30–60s of it on set.
3 · Loudness — LUFS and why it matters
Your ears judge loudness, but meters historically showed peaks, and they don't match. Platforms fixed this by normalizing everyone to a loudness target — if you don't master to their number, they change your sound for you. LUFS measures perceived loudness; integrated LUFS is the average of the whole video (the number you master to); true peak is the real max (keep at −1 dBTP). Use the interactive loudness reference in the practice studio to pick your target.
4 · Music & ducking
Choose music that doesn't fight the voice: avoid busy mid-range instruments (piano, flutes, lead melodies) that collide with speech; favor beds with energy in the lows and highs that leave the midrange clear. No lyrics ever under dialogue. For crypto/finance, aim for clean, modern, slightly-tense-but-optimistic — not hype-y EDM (dates fast, reads "shitcoin ad"), not sleepy elevator music. And silence is a legitimate, powerful choice — don't score the single most important sentence.
Ducking — how music gets out of the voice's way
Ducking automatically lowers music when someone talks. Auto (sidechain): a compressor on the music track triggered by the dialogue — fast, great for talk-heavy edits (one button in Premiere's Essential Sound or DaVinci Fairlight). Manual keyframes: you draw the music down by hand — more musical, the pro move for intros/outros. For spoken dialogue, duck hard — the music should drop 6–12 dB when the voice comes in (not the 1–3 dB you'd use under a singer). Sidechain start: ratio 4:1–8:1, fast-ish attack (~10–20ms), slow release (300–600ms) so it rises gently between phrases instead of pumping.
SFX — seasoning, not the meal
A tasteful whoosh per transition reads pro; a whoosh on every cut plus impacts on every word reads like a teenager who found the SFX folder. If the viewer notices the sound effects, you used too many. Keep 2–3 you reuse, low in the mix.
5 · Music licensing — the part that gets videos claimed
"Royalty-free" does not mean "free" and does not stop a copyright claim. YouTube's Content ID scans every upload against registered audio; a match fires the rights-holder's policy automatically — they can run ads on your video, mute it, or block it — even when you legitimately licensed the track, because the distributor also registered it. Creative Commons CC-BY requires specific per-track attribution, and credit alone does not stop a claim. See the music & SFX sources in the swipe file for a safe library comparison.
6 · Sync & backup
Your good audio usually comes from an external recorder or a mic that isn't the camera, so you marry it to the picture in the edit ("double-system sound"). Waveform auto-sync (Premiere Synchronize, Resolve Auto Sync, PluralEyes) lines them up — which is why you always let the camera record its own scratch audio as the reference. A clap at the top of each take gives a dead-obvious manual sync point. And always record at least two sources — there's no re-shooting a founder who already flew home.
- Room tone under the whole dialogue track; every hard silence patched
- Noise reduction applied — no "underwater" artifacts
- High-pass (~80 Hz) on every voice; mud (200–400 Hz) tamed; presence (2–5 kHz) up
- Compression — steady, ~3–4 dB reduction, natural not squashed
- De-essing — harsh "sss" tamed
- Dialogue riding −12 to −16 dB, consistent across the whole piece
- Music ducked 6–12 dB under dialogue; releases smoothly, no pumping
- Nothing covers a key word; SFX tasteful, not constant
- Limiter on the master at −1 dBTP; nothing clips
- Loudness normalized to target; sections matched to each other
- Checked on phone speakers and cheap earbuds, not just headphones
- Cleanup chain / noise — iZotope; EQ — Audiospectra; compression — Voice123
- Room tone — iZotope, Film Editing Pro
- Loudness / LUFS — LUFS standards, Spotify
- Ducking — iZotope, Pure Audio Insight
- Licensing — Silverman Sound, Uppbeat, Track Club
The cut — editing a conversation into an argument
The camera gave you a rambling forty-minute talk. Your job is to find the ten-minute story hiding inside it, and the six fifteen-second clips hiding inside that. Here's the craft, in the order you actually do it.
1 · Transcribe first, then edit the text
The modern interview edit begins in a transcript, not a timeline. Premiere, DaVinci Resolve, and Descript all support text-based editing — you cut the video by deleting words on the page. Get a clean transcript for every interview and reason about the material as text. It's the difference between a 3.5-hour selection pass and a 40-minute one.
2 · The paper edit, then the radio edit
A paper edit is a document listing, in story order, the soundbites that form the spine of your cut — chosen before you touch the timeline. Pull each usable bite with its timecode, tag it with the story beat it serves (setup, conflict, turn, payoff), then reorder until it works on the page. The test: read it out loud. If it doesn't make sense as text, it won't make sense in the cut.
3 · Pacing — making the cut invisible
Don't over-cut the silence
Remove filler clusters and long gaps, not every micro-pause. Aim for a cut that's 80–90% of the original length, not 60–70% — the deepest cuts read as over-edited. Keep the breath at the end of a sentence and the pause between topics that lets the viewer reset. After a big statement, let the moment breathe.
J-cuts and L-cuts — the seamless-conversation trick
Unlink audio from video and offset them. A J-cut: the next audio starts before its picture (you hear the answer begin over the previous shot). An L-cut: the current audio continues after the picture cuts away (you keep hearing them over a reaction or b-roll). The half-second-to-two-second overlap is what makes dialogue feel like a real conversation instead of two people alternating on a stage. Mechanic: unlink, nudge one earlier or later, listen, adjust.
4 · Murch's Rule of Six — how to judge a cut
Walter Murch's priority list for where to cut, in descending importance:
Emotion alone is worth more than the other five combined. If a cut forces a sacrifice, sacrifice from the bottom up. For interviews this is freeing: a jump cut that "breaks continuity" (rank 6) is fine if the moment it preserves is emotionally true (rank 1). The founder's voice cracking beats a technically clean cut every time.
5 · Multicam & b-roll — your invisible-mend kit
Sync your two cameras into a multicam sequence and cut angles on a shift — a new question, a punchline, a change of energy — never at random. Every time you tighten an answer you create a jump; cutting to the second angle (or the listener's reaction) at that exact frame hides it completely. That's why interviews are ideal multicam candidates. One camera? B-roll and cutaways hide the same jumps, and Premiere's Morph Cut / Final Cut's Flow blend minor jumps optically.
Time b-roll to the words
When the guest names a token, a chart, an exchange, a hardware wallet — cut to it on the word. Don't blanket the whole interview; place b-roll where it earns its cutaway, and pull back to the face for the emotional beats where it matters most. Build a reusable crypto b-roll library: chart screen-recordings (TradingView, DEX Screener), block-explorer scrolls, wallet UIs, conference-floor establishers.
6 · The cold open — engineering the first 5–10 seconds
The first ten seconds decide whether the next ten minutes get watched. Retention above 75% in the opening 30s is strong; below 60% means the hook is broken. So pull the best 10 seconds of the interview to the very front as a cold-open montage of the sharpest lines, then drop into your title. While cutting, flag great lines in real time — that flag list is your cold open and your clip list.
7 · Color — a simple, repeatable grade
Correct and match first, grade last. Correct each camera to a neutral baseline, then match cameras on the scopes (load the A-cam in a wipe, balance the B-cam to it on the parade and vectorscope; a Color Space Transform matches more accurately than a baked LUT). Fix skin tones with a power window / magic mask onto the vectorscope's skin line. Apply your look last, and lock a base look for a repeatable house style across every interview.
8 · Export & the software call
| Deliverable | Settings |
|---|---|
| YouTube 1080p30 | H.264, ~8–15 Mbps, AAC 384kbps 48kHz |
| YouTube 4K30 | H.264, ~35–45 Mbps |
| Vertical clips | 9:16, 1080×1920, H.264, high-bitrate master |
- Paper/radio edit — Eddie AI, Frame.io
- J/L cuts — SpotlightFX; Rule of Six — StudioBinder
- Filler/pacing — ChatCut; multicam — CapCut
- Cold open / retention — Artiphik; color — Pixflow
- Software — Subclip comparison
One interview → a week of content
A 45-minute interview isn't one video, it's a content deposit you draw down for a week. The long-form is the anchor; the clips are how anyone finds it. A one-hour episode reliably yields 10–20 usable clips. The craft is in selection and packaging, not in shooting more.
1 · What to pull
You're hunting self-contained moments that survive without context, roughly in order of performance:
- The contrarian take — cuts against consensus ("ETFs are the worst thing that happened to crypto"). Provokes replies.
- The number / the stakes — a concrete figure that stops the scroll ("I lost $2M in the Luna collapse").
- The story — a 30-second narrative with tension and payoff. These outperform explanations, and they're the ones AI clippers miss.
- The disagreement — host and guest openly clash. Conflict is watchable.
- The operator insight — "here's exactly how I size a position." Gets saved and shared, which platforms weight heavily.
2 · Vertical clips & captions
Interviews shoot wide; every clip becomes 9:16. Premiere Auto Reframe and Resolve Smart Reframe track the speaker automatically — genuinely good for a single speaker, but they choke on two people talking over each other (expect manual fixes on ~30% of two-shot clips). For rapid back-and-forth, a split-screen stack (guest top, host bottom) is safest; a manual punch-in on the reactor at the punchline is the high-craft move AI can't do.
Caption style that works
Most short-form is watched on mute, so captions are the delivery. Bold sans-serif (Montserrat/Proxima/Impact-style), 55–75pt, vertically centered in the middle 60–65% safe zone (top 15% is the status bar, bottom 20% is the like/comment UI). 1–2 lines, 3–5 words each. White base + one accent, with karaoke word-highlighting (the active word changes color as spoken) — the single biggest engagement lever. Different caption color per speaker so muted viewers follow the hand-off. And edit the auto-caption errors — the ~8% it misses is disproportionately tickers, protocol names, and dollar figures ("SOL" heard as "soul").
3 · The honest AI-tools verdict
AI does the mechanical 60% and none of the editorial 40%. It transcribes, finds candidates, reframes, and captions fast. It cannot judge what will perform, protect a comedic pause, fix an overlapping two-shot, or catch a caption typo. Every credible run ends with a full manual QA pass.
| Tool | For | Trust it for |
|---|---|---|
| Descript | transcript editing, filler removal, Studio Sound | The talking-head master edit and cleanup |
| Opus Clip | auto-finding viral moments | First-pass clip candidates — then you curate |
| Submagic / Captions | animated captions | Fast on-trend English caption styling |
| Auphonic | audio mastering / leveling | Final audio pass; fixes loud-host/quiet-guest |
| CapCut | free editor + captions | Budget clips (customize templates so it's not stock) |
| Resolve / Premiere | pro finishing | Hero clips and the long-form master |
4 · Thumbnails, titles & native posting
The clips feed the long-form; the long-form lives on thumbnail + title. Thumbnail: a face with legible emotion, ≤3 elements, big 3–5 word text that pops at feed size. The quote format works unusually well for interviews. Title formula: guest name/credential + the bold claim — "Arthur Hayes: The Dollar Ends in 2028." The name is the trust anchor, the claim is the hook, biggest promise in the first ~40 characters.
5 · Batching, specs & never losing footage
- Batch: build reusable caption + intro/outro + lower-third templates once. Do a whole interview's clips in one session while the material is fresh. Realistic once templated: the long-form master in a few hours, then ~15–20 min per clip.
- Vertical master spec: 1080×1920, 9:16, 30fps, H.264, AAC 128–256kbps, MP4. Below ~8 Mbps shows artifacts; above ~20 gives no visible gain on mobile.
- 3-2-1 backup: 3 copies, 2 media types, 1 off-site. Per interview:
YYYY-MM-DD_Guest/→ 01_Footage (read-only) · 02_Proxies · 03_Project · 04_Assets · 05_Exports. An interview you can't re-shoot is irreplaceable.
- Repurposing / clip selection — Choppity, NPR Training
- Reframe / captions — ELEMENTS, OpusClip
- AI-tools reality — 2hr→20 clips run
- Thumbnails/titles — ampifire; specs — Conbersa; backup — Cut Point
The practice studio
The tools you'll actually reach for while editing. Everything runs right here in the page — nothing uploads, nothing leaves your browser.
12:40 — "ETFs were a mistake" (contrarian). It saves as you type, so your clip list is ready before you finish the long-form.Music, SFX & title formulas
Where to get music that won't get you claimed, and the title/thumbnail patterns that get interviews clicked. Tap any for the detail.
Safe music & SFX sources
FREEYouTube Audio Library · Pixabay · Mixkit +
$0, generally claim-free. YouTube's own library never claims your videos; Pixabay and Mixkit are royalty-free with commercial use and no attribution. Mixkit has 42 free transitions and 20 free whooshes for SFX.
Gotcha: smaller, more generic catalogs, and some YT Audio Library tracks require attribution. Overused tracks can sound stock.
~$7/moUppbeat — budget with safelisting +
Freemium + sub. The safest budget pick: channel safelisting auto-clears claims on your uploads, and downloads stay covered even after you cancel. Free tier needs a credit and has a monthly download cap.
~$17/moEpidemic Sound — depth + safelisting +
Subscription (~$17/mo personal, ~$30 business). Huge, well-organized catalog with SFX; registers in Content ID but auto-clears subscriber claims. Gotcha: coverage ends when you cancel — old videos get exposed.
~$40/moArtlist — universal license +
Subscription, all-in-one music + SFX. Universal license, claim-clearing. Covers content made while subscribed. Deepest cinematic catalog of the three.
READFree Music Archive / Creative Commons +
Mostly CC, $0 — but read each track's exact terms. CC-BY requires specific per-track attribution (artist, track, license), and credit alone does not stop a Content ID claim. An artist can register with Content ID later and claim you months after you published, even when you used it correctly.
Title formulas for interviews
01Name + bold claim +
Example: "Arthur Hayes: The Dollar Ends in 2028." The name is the search-and-trust anchor; the claim is the hook. Keep the biggest promise in the first ~40 characters so it survives truncation.
saves in your browser02The quote thumbnail +
Why it works for interviews: a quote creates a reason to stop, and the guest's face lends credibility. ≤3 elements, big 3–5 word text that pops at feed size on your own screen.
saves in your browser03The number hook +
Example: "The trader who turned $5K into $4M — and what he'd never do again." Concrete numbers stop the scroll; the reversal adds curiosity.
saves in your browserThe 12 mistakes that kill an interview video
Fixing these is faster than adding anything new. Every one is covered above — this is the checklist version.
- Bad on-set audio. The one thing post can't fix. Two mics, lav close, monitor on headphones, room tone.
- No second angle or b-roll. Then every tightened answer is a naked jump cut. Shoot coverage.
- Editing the timeline before the story. Do the paper edit and radio edit first — fix story where it costs seconds.
- Over-cutting. Stripping every pause makes people sound robotic. Leave 80–90%, keep the breaths and topic beats.
- No cold open. Burying the best line at 12:00. Pull the best 10 seconds to the front.
- Ping-pong cuts. Straight A/B alternation. Use J-cuts and L-cuts so it feels like a conversation.
- Music fighting the voice. Lyrics or busy mid-range under dialogue, or music not ducked. Duck 6–12 dB, no lyrics.
- Loudness all over the place. Intro loud, interview quiet. Match sections, normalize to −14 LUFS.
- Mismatched color. Two cameras that don't match, orange-and-blue skin. Correct and match before you grade.
- Auto-captions left unedited. "SOL" as "soul," wrong tickers and numbers. Fix the names, they're the whole point.
- Captions in the UI dead zone. Behind the like/comment buttons. Keep them in the middle 60–65% safe zone.
- Letting the AI publish for you. Auto-selected clips and unattended reframes are slop. AI drafts, you decide.
Your 30-day editing plan
One technique a day, 20–40 minutes. The goal of month one isn't a masterpiece — it's to make the pro workflow automatic. Tap a day to check it off; your progress saves in this page.
"If [situation], then I [action]" beats "I'll try to edit more." Pin the rep to a moment that already happens, and it saves here so it greets you tomorrow.
- Edit score (out of 25) — from the "rate your last edit" tool. Watch the trend climb.
- Time-to-cut — how long a long-form + clips takes you. It should drop as templates and shortcuts land.
- Clips per interview — are you actually mining 5–7, or leaving them on the table?
- Loudness — did you hit −14 LUFS, or eyeball it? Measure every time.
- The ONE technique this week — name it, so you know what you were drilling.
The cheat sheet
Everything that matters, on one screen. Screenshot it.
- Win on audio first. Two mics, lav close, room tone. It's the one thing viewers won't forgive.
- Shoot two angles + coverage. A wide, a tight, b-roll, noddies — the edit assembles itself.
- Get them to answer in full sentences. Fold the question into the answer.
- Cut the story before the picture. Paper edit → radio edit → then visuals.
- J-cuts and L-cuts make dialogue feel like a conversation, not a ping-pong match.
- Emotion beats a clean cut (Murch). Sacrifice from the bottom up.
- Pull the best 10 seconds to the front. No cold open = buried lede = they leave.
- Cleanup chain in order, normalize to −14 LUFS, duck music 6–12 dB under the voice.
- Karaoke captions in the safe zone, and fix the names and tickers.
- One interview → 5–7 clips. AI drafts, you decide. Watch it back, pick ONE thing to improve.