The Vocabulary of Video Production

EN AI · Claude 40 min read

A glossary for video scripting, recording and editing.

Table of Contents

Video Production Vocabulary

I started making videos and quickly found that everyone teaching it assumes you already speak the language. B-roll, negative fill, J-cut, LUFS, thrown around as if the words explained themselves. They don’t. So I wrote down every term I had to stop and look up.

Each entry here carries the problem the word was invented to solve, because that is the only part worth memorizing. None of this vocabulary exists to sound professional. It exists because something kept going wrong on set, in the edit, or at delivery, and the crew needed a short name for the fix. Learn the problem and the jargon dissolves into something usable, to the point you can predict which tool someone will reach for before they do.

1. Writing the script

Almost everything here is YouTube-specific. Narrative screenwriting is a separate craft with its own vocabulary and most of it does not transfer.

The organizing idea is that a video passes through three documents, and they are not drafts of each other. The research document is a private thinking tool with an audience of one. The spoken essay is a linear performance for an ear that cannot rewind. The script is a score for voice, image and pacing at once. Each one is a rewrite from scratch with the previous one open beside it for lookup, and the people who skip the middle document are the ones whose videos come out as a list read aloud. Research runs five to ten times the length of the script it produces, because thinking takes more words than performing does.

Deciding what the video is

Topic vs. thesis. A topic is inert. A thesis is a claim someone could disagree with. The test from The Craft of Research is a sentence you should be able to complete before writing anything, “I am studying X because I want to find out Y in order to help my viewer understand Z.” Stall on the second or third blank and you have a topic. “How Bitcoin Works” is a topic. “How Bitcoin Will Save Your Family” is a thesis, and the difference is the whole video.

Format vs. topic. The topic is what this video is about. The format is the repeatable container it goes in, the thing a viewer recognises before you say a word. “Rust ownership explained” is a topic. “I rewrite a famous program in Rust and count the bugs it kills” is a format, and you can run it thirty times. Channels do not grow on topics, they grow on formats, because a format is what makes the second video predictable enough to click and the tenth one a habit. Pick the format first and the topic becomes an episode.

Baseline and outlier. Your baseline is what a video on your channel normally does. An outlier is one that did some multiple of it, five or ten times. The method Paddy Galloway made standard is that averages teach you nothing and the information is entirely in the anomalies, yours and other people’s. Find a video that beat its channel’s baseline hard, then isolate which variable was different, topic, format, packaging or timing. That is the only cheap way to learn what a niche actually wants.

Swipe file is where you keep the outliers. Titles, thumbnails, cold opens and structures that worked, saved with a note on why. Not to copy, to have something to argue with when your own idea is thin.

Explain vs. argue. Exposition presents information and aims at comprehension. Argument advances a contestable claim and aims at assent. Both show up in any real video, but whichever one dominates decides the structure, because exposition is organized by the logic of the subject and argument by the logic of the claim. The trap is that YouTube’s defaults are all expository, so people trained on explainer conventions keep producing information when what they wanted was persuasion, and the result has nothing anyone can agree or disagree with.

Packaging is the title and the thumbnail treated as one unit. The non-obvious part is the order. Packaging gets decided before the script exists, not after the edit, because the script’s only job is to deliver whatever promise the title and thumbnail made. Write first and package afterwards and you get a video that is about something but promises nothing.

Type 1 vs. type 2 clickbait, Derek Muller’s distinction. Type 1 conveys the actual content in the most curiosity-provoking honest framing available. Type 2 exaggerates or deceives. The rule that separates them is to create curiosity your content actually satisfies. Muller’s own case is the same video retitled from “Strange Applications of the Magnus Effect” to “Backspin Basketball Flies Off Dam”, identical content, fifty million views.

The opening

Hook is the first five to fifteen seconds. It is not an introduction. It has to convince someone who just arrived from a thumbnail that the payoff is real and coming soon. Saying your name, welcoming people back to the channel, and listing what the video will cover are the three most reliable ways to waste that window.

Stakes, promise, anchor are what the rest of the first thirty seconds does. Stakes say why the answer matters. The promise is a specific deliverable, and specific is the operative word, because “some thoughts on retention” leaks attention where “the one line in your hook that decides whether they stay” holds it. The anchor tells the viewer what kind of thing they are watching. Subscribe asks, long sponsor reads and logo bumpers all belong after this, not inside it.

Cold open starts mid-action with no preamble, then explains once the viewer is already committed.

Misconception opening. Derek Muller’s 2008 doctoral work found that students who watched a clear, well-illustrated explanation believed they already knew the material, spent little mental effort and scored the same as before watching. Students who watched a video that first voiced the common wrong answer, then refuted it, reported feeling more confused and scored close to twice as well. His summary is worth keeping, clarity numbs the mind and confusion can crack it open. The practical version is to open on the error rather than the answer.

Do not lead with the thesis. This is the counterintuitive one for anyone with a writing background. The inverted pyramid works in print and kills audio, because stating the conclusion first removes the reason to keep listening. The thesis becomes a reveal you earn, not a headline you announce.

Holding the middle

Bait, open loop, payoff. Ira Glass’s framing is the clearest. You raise a question, it is implied that you will answer it, and the answer arrives just after you have raised the next one. The unit he builds from is an anecdote followed by a moment of reflection that says why you are listening to this at all. A script with no open loops has nothing holding anyone through the middle.

Re-hook is a small new promise dropped every thirty to sixty seconds, so attention resets before it wanders off. Pattern interrupt is the same idea done visually, a cut to a different framing, a graphic, a change of pace, placed where attention would otherwise drift.

Hinge is the structural pivot somewhere between a third and half of the way in, where a video that appeared to be about X turns out to be about Y. Folding Ideas opens In Search of a Flat Earth by actually measuring the curvature of a lake and pivots to QAnon. Line Goes Up promises NFTs and reframes into the whole speculative economy. It reads as generosity, the viewer feels they got more than the title sold them, which is why it survives long runtimes.

Wow factor, MrBeast’s term, is the moment nobody has seen before. His guide treats it as the thing a video is built around rather than decorated with, and it is the honest answer to why a well-made video about nothing underperforms a badly made video about something unprecedented.

Minute bands. The same guide splits a video into the first minute, one to three, three to six, and six onward, and assigns each a different job. It is cruder than the hinge and more usable, because it forces you to name what changes at each boundary instead of assuming momentum carries.

But / therefore. Trey Parker and Matt Stone’s test. Walk the beats of your outline and name the connector between each pair. If it is “and then”, the structure is broken. Every beat should connect by “but” or “therefore”. And-then is a list, but-therefore is an argument.

What YouTube actually measures

Retention graph, AVD, APV. Average view duration is time watched in minutes. Average percentage viewed is that as a share of the video’s length. The graph underneath both is the only honest feedback your writing gets. YouTube names four things on it, the intro (share still watching at thirty seconds), top moments (stretches nobody left), spikes (rewatched, which means either strong or confusing, and you have to go look), and dips (people skipping or leaving). Its own advice on top moments is the useful part, if the best stretch is late, move it earlier, because the audience only ever gets smaller.

CTR and AVD are multiplicative. High click-through with low retention reads as clickbait and gets throttled. Low click-through with high retention reads as an undermarketed good video and just does not get shown. You have to win both, which is the real reason packaging and script are one problem rather than two.

Writing for the ear

The listener has one shot. A reader can backtrack mid-sentence to work out a pronoun. A listener cannot, and when a sentence overflows working memory they do not rewind, they leave. A 2019 study in Science Advances put the transmission rate of human speech at roughly 39 bits a second, near identical across all seventeen languages tested, which is a hard ceiling on how much argument a minute of audio can carry. NPR’s training guide turns it into a rule, the ear handles one fact or idea per sentence.

What breaks in audio. Long dependent clauses before the verb, because they force the listener to hold an unresolved subject. Nested parentheticals, because the ear cannot pop a stack. Dense runs of numbers. Attribution after the quote, which leaves the listener wondering who was talking. The mechanical fixes are dull and they work. Twelve to twenty words a sentence with the verb inside the first eight, wherever you would put a comma try a period instead, round your numbers, and cap any list at three.

The mouth edit. Read the script aloud two lines at a time without looking at the page. Wherever your mouth stumbles, the listener’s attention breaks. This is also the fastest test for a sentence that is grammatically fine and unspeakable.

Words per minute. Conversational delivery lands around 150 words a minute, so ten minutes of video is roughly 1500 words. Useful because it tells you a draft is twice as long as the video you meant to make before you record a single take.

Signposting and repetition. What looks clunky written down is the only thing keeping a listener oriented. State the claim near the top, restate it after your strongest piece of evidence, restate it at the end. Each restatement catches someone who drifted. Repetition in spoken work is structure, not redundancy.

Cutting it down

Reverse outline. After a draft exists, write one sentence per paragraph saying what it claims and one saying what it does for the argument. The instruction that makes it work is to outline what is actually on the page, not what you meant to put there. Paragraphs you cannot summarize have no job. Two claims in one paragraph means split it. Repeated keywords across distant paragraphs means you argued the same thing twice.

Filler is any line that survived because it got spoken, not because it carries something. Restating the last sentence, announcing what you are about to say, hedging. Cutting filler raises retention more reliably than adding B-roll.

Value per minute is the test that decides what stays. If a passage delivers nothing the viewer came for, it does not earn its runtime no matter how well written it is. Zinsser’s benchmark is that most first drafts can lose half their words without losing information or voice. If your final script is less than thirty percent shorter than your first full draft, you have not finished.

Handing it to production

Script callouts are the B-roll, graphics, and on-screen text written into the script while writing, rather than hunted for in the edit. The conventions are small and worth using literally, [B-ROLL: ...], [GRAPHIC: ...], [ON SCREEN TEXT: ...], [ARCHIVE: source, date], marked per line rather than per paragraph so the script and the timeline stay aligned. Deciding a claim needs a visual is a writing decision. Discovering it in the timeline is a rewrite you now have to shoot for.

Dead zone is a stretch where you cannot find a visual without forcing one. Evan Puschak’s read on this is the right one, it means the script is wrong there. Fix the writing rather than decorating it with animation.

Shot list is the script converted into the actual setups to shoot, one line per shot, grouped by location and lighting setup instead of by script order. Shooting in script order is the most common way to waste a day relighting the same room four times.

Storyboard draws the shots before they exist. Worth the effort where framing carries meaning, pointless where a person is just talking to camera.

Chapter titles are written craft, not metadata. Used well they are narrative beats rather than a topical index. Folding Ideas repeats the same thesis across three consecutive chapter names in Line Goes Up. hbomberguy’s chapters escalate, “The Twist You Expected”, “It Got Worse”, “HE’S STILL DOING IT”. Written as an outline before the script, they double as the structure check.

CTA (call to action). The ask. Placement matters more than wording. Put it where the viewer has just received something, right after a payoff, rather than at the start where nothing has been earned or at the very end where most of the audience has already left.

Slugline / scene heading is the INT./EXT., location, and time-of-day line that opens a scene in a screenplay. You only need it once the video has scenes rather than segments.


2. Pre-production

Pre-production / production / post-production. The three phases. Pre-production is everything decided before the camera rolls. Production is the shoot. Post is edit, sound, color and delivery. On a film these are three different teams and three different calendars. Alone they are three different afternoons, and the only reason to keep the words is that skipping the first one is what turns the third one into a rescue operation.

The crew you do not have. The org chart exists and you will hear it constantly in gear reviews and behind-the-scenes footage, so it is worth knowing which of your own tasks each name covers. The producer owns the project and the deals, the line producer owns whether the schedule is affordable. The 1st AD runs the set in the moment, the 2nd AD runs tomorrow. The gaffer owns the light sources, the grip owns everything the light lands on and everything the camera rides on, and the best boy is whichever of them is second in command. Craft services feeds people. You are all of these, which mostly means the failure mode is not incompetence in any one of them but doing them in the wrong order.

Room scout. The real location work for a talking-head channel is acoustic, not visual. A room that photographs well and rings like a bathroom will cost you more in post than a plain room that is dead. Record sixty seconds of nothing in the space and listen on headphones for the fridge, the HVAC, the street, the neighbour’s dog, the buzz off a dimmer. Fix it before the shoot or pick another room, because none of it comes out cleanly afterwards.

Continuity. On a film a script supervisor tracks which hand held the cup and whether the door was open, because scenes are shot out of order. Alone, the version that bites you is across days. Same shirt, same chair position, same lamp, same window light at the same hour. Shoot half a video on Tuesday afternoon and the other half on Thursday morning and the cut announces itself.

Camera marks are the fix. Tape on the floor for the tripod legs, a written note of the lens, the height, the aperture, the white balance and the light positions. It costs two minutes and it is the difference between resuming a shoot and reshooting one.

Pickups are the shots grabbed after the main session, usually because a line came out wrong or an argument needed one more sentence. This is the most useful film term on the list for a solo creator, because the alternative to accepting pickups is rerecording a whole take for one flubbed clause. Write the pickup list while you edit, then shoot them all in one sitting against the marks you saved.

Batching / block shooting is recording several videos, or several segments, in one setup session. The setup is the expensive part, not the recording. Two hours of lighting and framing amortized across four videos is the single biggest schedule win available to a one-person channel.

Wrap and strike mean done and torn down. The solo version worth ritualizing is the offload. Copy the cards to two places before you strike anything, verify the copies play, and only then format. The standard is 3-2-1, three copies, on two kinds of media, one of them somewhere else. Cards get formatted by accident exactly once per career and everybody learns the same way.

Talent release. A signed piece of paper saying you can use someone’s image and voice. Film needs it because distributors demand it. You need it the moment a guest appears on camera or a stranger is recognizable in a shot, because the person who was happy to be filmed can stop being happy about it later, and by then the video has views. Get it in writing, even from a friend, even by email.

Location release is the same thing for a venue. Filming inside a shop, a gym, a conference or an office means someone owns that space and can object. Public street is generally fine, private property that let you in is a conversation you should have before you set up rather than after you publish.

Content ID claim vs. copyright strike. These get confused constantly and they are not the same hazard. A Content ID claim is automated. A rightsholder’s reference file matched something in your video, and they chose to monetize it, track it or block it. It is not a penalty and it does nothing to your channel standing. A copyright strike is a human filing a legal removal request. The video comes down, the strike sits on the channel for 90 days, and three live strikes terminates it. A hundred claims will not close your channel and three strikes will.

Fair use is a defense in United States law, decided by a court weighing four factors, and it is not a setting on YouTube. Commentary, criticism and transformation are the strongest ground, and the practical reality is that a claim gets disputed inside YouTube’s system long before any of that becomes a legal question. Assume every second of licensed music or borrowed footage will be detected, because it will be, and decide in advance whether you would actually defend it.


3. Framing and composition

Shot size is really just distance. The closer the camera sits, the more emotion you get and the less context.

ECU (extreme close-up) shows one detail only, eyes, a ring, a finger on a trigger. Sacrifices all context for intensity.

MS (medium shot) is waist up. The default for conversation because it shows face and body language at once.

WS / LS (wide / long shot) is head to toe or further. The environment does half the storytelling.

Two-shot frames two people, both faces visible. The default for a collab or an interview, because it establishes the relationship before either of them speaks.

OTS (over-the-shoulder) looks past one person’s shoulder at the other, putting the audience physically inside a conversation instead of watching it from outside.

POV shot shows exactly what a character sees, as if the camera were their eyes. Establishing shot is usually a wide shot at the start of a scene that orients the audience in a location before it cuts closer. Insert shot (as a shot type, not just an edit) is a tight, isolated shot of a detail, a ticking clock, a text message, a key turning in a lock, that clarifies something the wider shots can’t. Aerial / bird’s-eye shot looks straight down or from height, which is most of what a drone is actually for, establishing scale or geography.

Dutch angle (canted angle) tilts the camera off the horizon to signal unease or instability. Loud, and it stops working the second time you use it in one video.

Rule of thirds divides the frame into a 3x3 grid, with the subject or their eyes placed on the lines or intersections rather than dead center. Headroom is the space above a subject’s head, too much and they sink and look small, too little and the frame feels claustrophobic (a common rule of thumb is putting the eyes on the upper third line). Leadroom (nose room) is extra space in the direction a subject is looking or moving. Center a person facing straight into the frame edge and it reads as visually wrong, because there’s nowhere for their gaze or motion to go.


4. Camera and optics

Pan / tilt. Camera stays put, pivots side to side or up and down. Surveys a space the audience hasn’t seen yet.

Dolly, truck, pedestal. The three ways the whole camera moves rather than pivots. Dolly is toward or away from the subject, truck is sideways alongside it, pedestal is straight up or down with the angle unchanged. A slider gives you the first two and a gimbal gives you all three badly. The distinction that matters is that a dolly reads as the viewer being pulled in and a zoom does not, which is why the cheap move looks cheap.

Rack focus. The lens shifts focus from one subject to another without cutting, redirecting attention inside a single continuous shot.

Whip pan. A very fast pan that blurs into streaks, almost always used as a disguised transition into the next shot.

Focal length and lens compression. People credit “compression” (background looking closer and larger relative to the subject) to the lens. The cause is camera-to-subject distance. A long lens forces you to stand farther back to keep the same framing, and that distance is what flattens perspective. A wide lens does the opposite: shooting close exaggerates the distance between foreground and background, which is why wide lenses distort faces up close.

Depth of field / aperture (f-stop). Shallow depth of field (low f-number, wide aperture) isolates a subject and reads as dreamy, romantic, or intimate. Deep depth of field (high f-number) keeps foreground to background sharp, used for landscapes or when the environment itself is part of the story. The numbering runs backwards from what you would expect. A small f-number means a big aperture opening, which means shallow depth of field.

ISO and shutter speed. ISO is sensor sensitivity to light. Shutter speed is how long each frame is exposed. A common rule sets shutter speed to double the frame rate (24fps needs 1/48s) for motion blur that reads as natural rather than stuttery or overly smooth.

Shutter angle / the 180-degree shutter rule. This is a completely different “180-degree rule” from the screen-direction rule in the editing section below. The collision of names catches everyone once. Shutter angle comes from old rotary shutters, literally a spinning disc with a pie slice cut out. 180 degrees of the circle open means the shutter is open half the time, which is the natural-looking default. A wider angle (270-360 degrees) produces more blur and a dreamier feel. A narrower angle (90 degrees or less) produces a crisper, stroboscopic look.


5. Lighting

Three-point lighting is the default setup you’re building toward, or deliberately breaking.

  • Key light, the main source, sets the base exposure and mood.
  • Fill light, softer and opposite the key, opens up the shadow the key creates without erasing it.
  • Back light, behind the subject, rims the edges and separates a person from the background so they don’t flatten into it.
  • Hard vs. soft light. Hard is a small or distant source: sharp, defined shadows, unflattering. Soft is a large or close source, or a diffused one: shadow edges disappear. Same bulb, different distance, a completely different feeling.

Bounce light reflects a strong source off a board or reflector to create a softer secondary source, adding light. Negative fill is the inverse: a flag or black solid positioned to block light from bouncing back into the shadow side of the frame, removing light to deepen shadow and add contrast. Modern sensors hold so much dynamic range that cinematographers spend as much effort taking light away as adding it. Negative fill is a real tool, not the absence of one.

Motivated lighting is engineered to look like it’s coming from a real or implied source in the scene (a lamp, a window, a screen), even when the fixture creating it sits somewhere else entirely. It keeps the audience from wondering where the light comes from, so attention stays on the story.

Flags / nets / silks are the grip-department vocabulary for shaping light. If light passes through it, it is a scrim, net or silk. If it blocks light outright, it is a flag. Flags block or create negative fill. Nets and wire scrims cut light by a fixed number of stops without changing its character. Silks (silk, nylon, or muslin) turn hard light soft by diffusing it.

Color temperature (Kelvin) and CRI. Tungsten light sits around 3200K (warm/orange), daylight around 5500-6500K (cool/blue). Mixing them unbalanced is the classic beginner mistake. White-balance for the tungsten lamps and the window reads blue, white-balance for the window and the lamps read orange. Pros either correct one to match the other (gel the window to 3200K, or gel the lamps to daylight) or deliberately split the difference for a stylized look, the tungsten-daylight tension seen in a lot of thrillers is this exact mismatch used on purpose. CRI (color rendering index) measures how accurately a light source renders color against a reference. Cheap lights with low CRI shift toward green or magenta at the extremes of their dimming range, creating color problems that cost hours to fix in post.


6. Sound on the shoot

MOS is a take shot deliberately without sound, which is most B-roll.

Room tone is thirty to sixty seconds of a location’s pure ambient silence, recorded on the same mics as the dialogue. Editors use it to paper over gaps and edits with silence that actually matches, instead of dead digital silence that sounds wrong.

Wild track / wild line is audio recorded with the camera off, a safety net for a line the take botched. The solo version is rerecording one sentence at the desk and dropping it under B-roll so no lip sync is needed.

Lav vs. boom. A lav is a tiny mic clipped to clothing, for hands-free or interview setups. A boom is a mic on a long pole held just outside frame overhead. Lav wins when the camera is wide. Boom wins when clothing rustle would ruin a lav.

Slate / clapperboard announces the take on camera, and its clap gives one sharp spike to line up picture and sound. Immediately useful the moment you record audio on a separate recorder instead of into the camera. A hand clap in frame does the same job.

Take is one attempt at a shot, numbered so nobody loses track of which one worked after take seventeen.

Coverage is shooting one scene from enough sizes and angles, wide, medium, close, reverse, that the editor has real options later instead of one rigid sequence.

Blocking is deciding where people and camera stand and move before rolling, not during.


7. Organizing the footage

Before a cut exists the footage has to be sorted. The order runs string-out, selects, paper edit, assembly cut, rough cut, fine cut, picture lock.

String-out is every usable moment laid end to end, no judgment applied yet, the editor’s raw inventory.

Selects is a filtered shortlist of the best takes pulled from the string-out.

Paper edit builds the intended cut from the transcript, the actual quotes and dialogue on paper, before touching footage. Common in documentary and interview-driven work.

EDL (edit decision list) is a literal instruction list, in and out points per clip, rather than physically cut footage, non-destructive by design and always sourced back to the originals on playback. NLE (non-linear editing) is the model every editor you have used runs on, where nothing you do touches the original file.

Multicam editing syncs feeds from several cameras to one timeline so an editor can switch angles live rather than re-cutting physically. Cameras are lettered by priority, not just order: A-cam gets the primary angle, B-cam captures complementary or reaction coverage, C-cam and beyond add more angles.

Rough cut, fine cut, picture lock. Rough cut is structure only, does the story work. Fine cut is structure locked with execution refined and music in place. Picture lock is the point where visual editing stops entirely. Lock is the point worth respecting even alone, because music, sound mix, captions, graphics and color all key off exact timings. Move a cut afterwards and you redo all of them.


8. Cuts and transitions

A-roll / B-roll. A-roll is the thing actually happening, an interview answer, a scene, a demo. B-roll is whatever gets cut in on top of it. Beginners treat B-roll as garnish and that is backwards. The name is literal, it was a second reel spliced in wherever the join in the first one would show. B-roll exists to fix a problem in the A-roll, an ugly jump, a dead stretch, a claim that has to be shown instead of said.

Hard cut is the default, one shot instantly replaces another. Barely a technique, more the absence of one.

L-cut / J-cut. Sound and picture change at different moments instead of together. In a J-cut, the next shot’s audio arrives before its picture does. In an L-cut, the picture has already moved on while the old audio keeps playing. Staggering sound and image instead of snapping them together is most of what makes an edit feel professional. Nobody in the audience notices it consciously. They just feel one cut as smooth and the other as a jolt.

Jump cut removes a chunk out of one continuous, similar shot. Reads as a lurch forward in time. Usually a mistake to avoid, occasionally used on purpose (vlogs, jitter aesthetic) precisely because it announces itself.

Match cut cuts between two shots that share shape, motion, or composition. Makes the join itself carry meaning (a spinning bone becomes a spinning spacecraft) instead of just moving the story along.

Cutaway cuts away to something else, almost always B-roll, to bridge or hide a jump the main footage can’t cover alone. This is B-roll doing the job it was named for.

Cross-cutting / parallel editing alternates between two things happening at the same time in different places. Builds tension by letting the audience know more than either scene’s characters do.

180-degree rule (the line rule, distinct from the shutter-angle rule above) keeps the camera on one consistent side of an imaginary line between two subjects, so screen left and right stay consistent shot to shot. Break it and the audience loses track of who is looking at whom. Not a style choice, a real perceptual cost.

Cutting on action (match on action) cuts mid-movement, a hand reaching for a door, between two shots of the same action shot at different times, so the motion itself hides the edit and creates false continuity.

Smash cut is an abrupt, jarring cut used deliberately at a moment the audience doesn’t expect one, for comedic or shock effect. A hard cut dropped exactly where the audience expects continuity.

Dissolve / cross-dissolve, wipe, fade, iris. A dissolve signals passage of time, a dream, or a scene change. It is a continuity device, not decoration. A fade to or from black reads as “an act, or the whole story, is ending,” so using it mid-scene misleads the audience. Wipes and iris transitions are stylistically loud and mostly associated with silent-era or vintage film, or fast-cut commercial and social content. Overuse them, wipes especially, and it reads as amateur, because they call attention to themselves instead of serving continuity.

Retention edit is the house style of YouTube, a visual or audio change every five to fifteen seconds so attention never settles. Cut, zoom, graphic, sound effect, anything. It is not the same as pacing, and overdone it reads as anxiety, but the absence of it is what makes a competent talking-head video feel like a lecture.

Punch-in is a cut to a tighter version of the same shot, usually faked in post by scaling a 4K frame inside a 1080 timeline. The cheapest retention edit there is, and the reason to shoot at higher resolution than you deliver.

Speed ramp accelerates or decelerates footage across a transition rather than cutting flat. Usually paired with a whoosh, which is the other half of the trick.

Sound design is the non-music audio layer, whooshes, impacts, clicks, ambience, risers. Roughly half of what people mean when they say a video looks expensive is actually this, which is why sound design is the highest return per hour available in an edit.

Dead air is a stretch with no new information and no visual change. It is where retention graphs dip, and it is almost always fixable by cutting rather than by adding.


9. Audio in post

ADR (automated dialogue replacement) re-records a line after the shoot when the original audio is unusable. Yours will be one sentence redone at the desk, and it will not match the room unless you kept the same mic and distance.

Foley is everyday sound effects, footsteps, a jacket rustling, a door latch, performed by hand in post and synced to picture, because the on-set recording of those sounds is almost never clean enough to use.

Needle drop is one use of one existing song, licensed one time. The part that matters is that it needs two separate licenses, a sync license from whoever owns the composition and a master license from whoever owns that specific recording. Buying one and not the other is the standard way people end up claimed, and a library subscription that covers you on YouTube is not the same as owning either.

Stems are sub-mixes of related elements bounced to one file, a whole drum kit to a single stereo track. Worth exporting your voice, music and effects as three stems if anyone will ever redub or re-cut the video. Mixing balances tracks against each other. Mastering polishes the finished mix as a whole.

Ducking automatically lowers one track’s volume whenever another crosses a threshold, music dipping the instant dialogue starts. Mechanically it is sidechain compression, a compressor on the music track triggered by the dialogue track’s signal instead of its own. Skip it and the music just fights the voice for the whole runtime.

EQ vs. compression. EQ changes what frequencies you hear, tone, by boosting or cutting frequency bands. Compression changes how loud parts are over time, dynamics, acting as an automatic volume rider that quiets loud peaks and can lift quiet parts, tightening the gap between the loudest and quietest moments.

LUFS (Loudness Units Full Scale) is the loudness-normalization standard platforms use, weighted to match human ear sensitivity. YouTube and Spotify both normalize to -14 LUFS, podcasts typically target -16 LUFS, US broadcast is held to -24 LKFS by the FCC’s CALM Act specifically to stop loud commercials, and EU broadcast targets -23 LUFS. The trap is that normalization is not symmetric. YouTube only ever turns audio down to hit its target and never up, so a quiet mix stays quiet there while a hot one just gets attenuated. Spotify does apply positive gain to quiet masters, but caps it to leave 1 dB of headroom, so a track already peaking hard still lands short of the target. Either way, being under is worse than being over.

Stinger is a short musical or sound hit, one to five seconds, used to punctuate a moment or mark a transition. A sitcom sting at a scene’s final beat is the same device.


10. Color and finishing

Color correction is the technical pass first: white balance, exposure, matching shots from different cameras or times of day so they look like they belong together.

Color grading is the creative pass after that, setting mood and style (the blue-orange blockbuster look, a bleached-out flashback). Order matters, grading a shot that hasn’t been corrected first just amplifies whatever was already wrong with it.

LUT is a preset color transform applied during grading. Works properly only on already-corrected footage. A LUT is a style, not a fix.

White balance calibrates what counts as true white so every other color in the shot reads correctly. Get it wrong and the whole scene skews orange or blue no matter what’s actually in frame.

Log / flat color profile vs. Rec.709. Log is a deliberately washed-out, low-contrast recording format that spreads sensor data across a much wider tonal range than a standard viewing monitor can show at once, then compresses it into the recorded file. It looks bad on purpose. The flatness is dynamic range being preserved instead of baked into a finished look too early, so grading later has real room to push shadows and highlights without banding or noise. Rec.709 is the baked-in, ready-to-watch standard that TVs and monitors expect.

Waveform monitor vs. vectorscope. A waveform monitor plots brightness across the frame, left-to-right position matching the actual image, so an operator can see if something is blown out or crushed and exactly where in frame. A vectorscope plots color only, and the one line on it you will actually use is the skin tone indicator, which tells you whether a face is the right colour independent of what you think you see on an uncalibrated monitor.

Chroma subsampling (4:4:4 / 4:2:2 / 4:2:0 ) and bit depth (8 vs. 10-bit). The human eye is far more sensitive to brightness than to color, so video formats throw away color resolution to save space, 4:2:0 (most consumer cameras) keeps a quarter of the color data of 4:4:4 . It shows up worst in chroma-keying, where low color resolution makes a clean green-screen key much harder to pull. Bit depth is a separate axis. 8-bit stores 16.7 million colors, 10-bit stores over a billion, which gives colorists room to push a grade without visible banding. Professionals shoot 10-bit 4:2:2 or better specifically to survive heavy color grading and keying, then usually compress down to 4:2:0 only at final delivery.


11. VFX and compositing

Chroma key (green vs. blue screen). Both colors are chosen because they sit furthest from human skin tones on the color wheel. Green wins for modern digital cameras because a standard camera sensor dedicates twice as many pixels to green as to red or blue, and green also carries higher luminance, needing less lighting on set. Blue is still better for blonde or light hair, since green keying eats into pale edges. Never wear anything close to the key color, it gets keyed out with the background.

Clean plate is the same shot with nothing moving in it, no people, no objects. Shoot ten seconds of one before you step into frame and you can paint out anything later, a light stand, a cable, a logo you did not mean to advertise. Costs nothing at the time and is impossible to get afterwards.

Rotoscoping is hand-tracing a moving subject frame by frame to build a mask that isolates it from the background. What you do when there is no green screen, and the reason masking one arm for three seconds takes an afternoon.

Matte is the mask itself, whether a key, a roto or a shape you drew. Compositing is stacking the layers into one image.

Lower third is a graphic or text overlay in the lower portion of the frame (a name or title caption), also called a chyron. Title card introduces a film or section. End card closes it out, and on YouTube it is where the end screen elements sit, so leave the last fifteen to twenty seconds visually quiet enough to put them somewhere.


12. Delivery and format

Aspect ratio is the frame’s width-to-height shape. 16:9 is the default screen shape now. 2.39:1 ultra-wide reads as prestige because it is associated with expensive anamorphic lenses. 9:16 vertical exists purely because phones are held upright.

Frame rate is how many frames each second of playback contains. 24fps is the old film cadence and still reads as cinematic out of pure habit. 60fps reads as hyper-real or live, sports broadcasts, or the unwanted soap-opera effect when applied to a film. Shoot at a high rate and play back at a normal one, and that’s slow motion.

Codec vs. container. A container (.mp4, .mov) is just the box holding video, audio, and metadata together. A codec (H.264, ProRes) is the actual compression method squeezing the picture into that box. The extension tells you nothing about the codec inside, which catches people constantly.

Proxy is a lightweight stand-in copy of heavy camera-original footage, edited against for smoother performance, then swapped back for the full-resolution original only at final export.

Deliverables are the exact final files a client or platform demands (specific codec, resolution, loudness level), usually spelled out in a formal spec nobody reads until an export gets rejected for missing one line of it.

Resolution vs. bitrate. Resolution (1080p, 4K, 8K) is pixel count. Bitrate is how much data encodes each second, and the two interact. A well-encoded 1080p file can look cleaner than a badly encoded 4K one. Diminishing returns are real and measurable, 1080p30 looks good around 12-16 Mbps, and more bitrate past that plateaus. Resolution also has a ceiling set by viewing distance. Most viewers cannot resolve the difference between 4K and 8K on a normal screen from a normal couch distance. The eye’s angular resolution is the limit, not the display.


13. Discovery and the numbers

Impressions are the times your thumbnail got shown. CTR is the share of those that turned into a click. The number itself is close to meaningless across channels, because a thumbnail shown to your own subscribers converts differently from one shown cold on a stranger’s home page. Most channels live between four and six percent, and the only comparison that means anything is against your own previous videos.

Traffic sources are where the impressions came from, and the split matters more than the total. Browse is the home page and the subscription feed, the algorithm pushing you at someone who was not looking for anything. Suggested is the sidebar and the end of somebody else’s video, the algorithm reading you as the natural next thing. Search is someone typing a question. External is everything off platform. Push traffic and pull traffic reward completely different packaging, and a video that lives on search can carry a title that would die on browse.

Velocity is how fast a video performs in its first day or two, against your own history. It is what people mean when they say a video “took off” or “died”, and it is mostly a verdict on packaging rather than content, because almost nobody has watched it yet.

Session time is how long the viewer stays on YouTube after your video, not just inside it. This is the metric that explains behaviour that otherwise looks irrational, why the platform favours videos that hand the viewer somewhere to go next, and why an end screen and a real playlist are worth more than they look.

CPM vs. RPM. CPM is what an advertiser pays per thousand ad impressions. RPM is what you actually receive per thousand views of your video, after YouTube’s cut and after averaging in every view that carried no ad at all. They are not the same number and they are not close. CPM is the one people quote and RPM is the one that pays rent.

YPP (YouTube Partner Program) is the monetization gate, and it has two steps. Fan funding and shopping open at 500 subscribers, three public uploads in the last 90 days, and either 3,000 valid watch hours in a year or 3 million valid Shorts views in 90 days. Ad revenue needs 1,000 subscribers and either 4,000 qualified watch hours in a year or 10 million qualified Shorts views in 90 days, and the two paths do not add together, Shorts feed views never count toward the watch-hour number. Acceptance also requires no active strikes, which is where section 2 stops being paperwork trivia.

Shorts vocabulary. The hook window is about one second, not three, because the decision to swipe is reflexive and happens before a sentence finishes. Swiped away is the metric for having lost it. A loop is an ending that flows back into the opening so the thing replays without the viewer deciding to, and two loops in twenty seconds beats one clean sixty second watch. Everything in section 1 about earning a slow build is inverted here.


Sources

Scriptwriting and retention

A-roll, B-roll, shot types, camera movement

Camera optics and composition

Editing and cuts

Sound, on set and in post

Lighting

Color and finishing

VFX and compositing

Discovery, analytics and the creator economy

Rights, releases and platform rules

Crew roles, for when you hear them

Delivery and format


More Posts