How do I remove filler words, false starts, and dead air from a raw video recording automatically?
the agent that answers this
Raw recording → clean cut (Claude + Remotion). Take a raw talking-head recording, transcribe it, remove filler words, hesitations and dead air to match a reference "good cut", and render a clean final MP4 with Remotion — replacing the manual editing agency.
This agent takes a raw talking-head recording, transcribes it word-by-word, and builds a cut list that removes filler words, false starts, retakes, and dead air -- while leaving meaningful pauses and restatements intact. It shows you the exact cut list (what's being removed and the resulting runtime) before rendering anything, then delivers a single clean MP4 built by concatenating only the approved segments.
- Cost
- Free
- on your own plan
- Runs
- On demand
- or scheduled
- Built from
- 10 steps
- 4 verified skills
- Runs in
- Claude or Codex
- as you
the steps
10 steps · 4 from verified skills- Step 1tool
DESKTOP MEDIA ADAPTER (mandatory for presenter_video): invoke only the Desktop-installed run-tools/clean_cut_desktop_media_adapter_v1.py wrapper for probe, bounded transcribe, then render. The wrapper sends only a semantic operation to Desktop. Claude/Codex receives no source path, broker URL or bearer secret. Desktop owns direct PyAV range seeking, writes a bounded work/transcript.json through one shared model cache, and writes one bounded normal output file without source staging. Never wrap the source URL in pathlib.Path, resolve it to a filesystem path, download, copy, symlink, stage or cache it. If the wrapper refuses, stop with its typed input error; do not fall back. ffprobe, ffmpeg and Remotion may consume only ordinary derived media produced after the wrapper; OffthreadVideo must not read the broker URL. Then perform the original step: Locate the raw recording (and optional reference): local file path if given, else the founder's Drive editing folder via authenticated Chrome
Your model fills this step- No integration? use a local file path
- No integration? open Drive in Chrome, read file IDs from DOM
- Step 2tool
Preflight: immediately before the first dependency-using step, make sure the remaining work has what it needs, so the run never fails halfway on a missing tool or an unconfigured account. Command-line tools (free, auto-install only if missing; idempotent, later runs find them present): ffmpeg - `ffmpeg -version` || install (`brew install ffmpeg` on macOS, `apt-get install -y ffmpeg` on Linux); Whisper - `python3 -c "import faster_whisper"` || `python3 -m pip install --user faster-whisper` (the video_qa verifier uses faster-whisper); gdown - `gdown --version` || `pipx install gdown`.
Analyze a screen recording of a manual process and produce targeted, working automation scripts. Extracts frames and audio narration from video files,…
Automate This Analyze a screen recording of a manual process and build working automation for it. The user records themselves doing something repetitive or tedious, hands you the video file, and you figure out what they're doing, why, and…
Full skill: automate-this - Step 3tool
Get the raw into the working dir: skip if local, else gdown the public Drive link
Your model fills this step- No integration? use local file in place
- No integration? gdown <fileId>
- Step 4
Extract audio and transcribe to word-level timestamps with faster-whisper distil-large-v3; run ffmpeg silencedetect for pause/sound boundaries
Your model fills this step - Step 5
Establish the target editing style. Default for this user: remove fillers, fumbles, false starts, retakes (including camera-adjust repeats), AND non-speech sounds (coughs/clicks); KEEP all real content AND natural pauses up to 2 seconds Preserve natural breathing room and intentional emphasis up to 2 seconds, all meaningful speech and the later complete take. For each non-speech gap >2 seconds, trim only its excess to <=2 seconds at speech-safe boundaries; a longer intentional pause requires a specific justification in the cut plan and explicit owner approval of that exception at the existing step-7 cut-plan gate. Do not clip consonants. Step 1 source custody and all existing approval, Manager, Judge, render and delivery contracts remain authoritative.
Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel,…
Video Use Principle LLM reasons from raw transcript + on-demand visuals. The only derived artifact that earns its keep is a packed phrase-level transcript (takes_packed.md). Everything else — filler tagging, retake detection, shot…
Full skill: video-use - Step 6
Build the keep/cut edit decision list: (a) remove filler words; (b) remove fumbles/false-starts/retakes including camera-adjust repeats, but do NOT over-cut a meaningful restate (only cut a repeat when the earlier take is abandoned or wrong); for back-to-back verbatim repeats extend the cut past the Whisper word-end since it under-runs; (c) cap dead air by INTER-WORD gaps (not just silencedetect, which leaves room-tone gaps): keep pauses up to 2 seconds, trim longer ones down to 2 seconds, trim silent intro/outro tight; (d) remove non-speech sounds (cough/click; preserve natural breathing room) that show up as non-silent blips inside kept pauses; (e) resume each cut ~0.06s before the next kept word so its onset consonant is not clipped Preserve natural breathing room and intentional emphasis up to 2 seconds, all meaningful speech and the later complete take. For each non-speech gap >2 seconds, trim only its excess to <=2 seconds at speech-safe boundaries; a longer intentional pause requires a specific justification in the cut plan and explicit owner approval of that exception at the existing step-7 cut-plan gate. Do not clip consonants. Step 1 source custody and all existing approval, Manager, Judge, render and delivery contracts remain authoritative. Audit EVERY inter-word gap across the COMPLETE word-level transcript, including room tone that silencedetect misses; silence detection is supplementary, never a coverage filter. Create work/cut/gap-ledger.json with a stable gap ID, source in/out timestamps, original duration, proposed kept duration, reason, intentional exception (null if none), linked cut-plan entries, and the explicit step-7 owner approval reference for any longer exception. Check ledger completeness against every adjacent transcript word pair; resolve missing or uncertain transcript coverage using only the existing authorized media adapter. Include the full gap ledger in the cut plan presented at the existing step-7 gate. Also map newly adjacent kept-word pairs and their planned output gaps to source spans so cuts cannot introduce unaudited gaps. Check both sides of each cut against actual speech; approximate 0.06-second pre-word padding and duplicate word-end extensions are starting points, not permission to clip speech.
End-to-end turn an unedited long-form talking-head / vlog / podcast video into a compact "first cut" (rough cut). Use when asked to edit/剪辑 a raw YouTube (or...
--- name: video-cut description: > End-to-end turn an unedited long-form talking-head / vlog / podcast video into a compact "first cut" (rough cut). Use when asked to edit/剪辑 a raw YouTube (or local) video into a tighter version: download,…
Full skill: Video Cut - Step 7decision
Review gate: present the cut list (timestamps, removals, runtime delta) for approval before rendering
Decision step - Step 8tool
Render frame-accurately by concatenating approved kept segments via ffmpeg trim/concat (hardware h264_videotoolbox), preserving source resolution/fps
Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include…
FFmpeg for Video Production FFmpeg is the essential tool for video/audio processing. This skill covers common operations for Remotion video projects. Quick Reference GIF to MP4 (Remotion-compatible) ffmpeg -i input.gif -movflags faststart…
Full skill: ffmpeg- No integration? ffmpeg trim+concat filter_complex
- No integration? Remotion only for motion-graphics overlays
- Step 9decision
Quality gate: re-transcribe the rendered output and adversarially verify every flub/retake/non-speech sound is gone and NO words were clipped at the seams (spot-check resumed words, AND count speech runs via silencedetect for suspected doublings since fragment transcription hallucinates), intentional content and natural pauses up to 2 seconds are intact, and A/V is synced; fix any problems before delivery Preserve natural breathing room and intentional emphasis up to 2 seconds, all meaningful speech and the later complete take. For each non-speech gap >2 seconds, trim only its excess to <=2 seconds at speech-safe boundaries; a longer intentional pause requires a specific justification in the cut plan and explicit owner approval of that exception at the existing step-7 cut-plan gate. Do not clip consonants. Step 1 source custody and all existing approval, Manager, Judge, render and delivery contracts remain authoritative. Independently audit the ACTUAL rendered MP4 against the owner-approved gap ledger: freshly transcribe/align its complete audio and inspect speech boundaries, enumerate EVERY output inter-word non-speech gap including room tone missed by silencedetect, and report ALL gaps >2 seconds with measured output/source timestamps, duration, ledger ID, intentional-exception justification and explicit step-7 owner-approval evidence. Use the approved source-to-output cut mapping to reconcile actual output with the ledger; planned durations and worker claims are not measurement evidence. Fail the quality gate and repair/re-render/reverify before delivery if any unapproved excess remains, any gap is unresolved, or required ledger/approval evidence is missing. Return to the existing step-7 gate if repair changes approved editorial decisions. Verify the actual rendered tail remains source-bound by comparing decoded terminal frames/timestamps and final duration with the approved source ranges using authorized source-derived evidence; reject synthetic stills, freeze-frame padding, duplicated-frame extension and any tail beyond the approved source segment. Retain existing transcript, audio/video seam and sync QA. Require video_qa in the immutable outcome schema; declare exactly one final_output MP4, a source.txt/.md with the complete expected spoken script, and the three reserved Desktop-authored qa/media.json, qa/transcript.json and qa/video-qa.json reports. Declare the approved gap ledger, cut mapping, exception approvals and actual-output gap/tail QA evidence as additional artifacts, never substitutes for the required video_qa package.
Decision step - Step 10decision
Delivery gate: preview and, only on approval, save to destination (browser upload capped at 10MB; for multi-GB use drag-drop or rclone)
Decision step
common questions
How do I remove filler words, false starts, and dead air from a raw video recording automatically?
The Raw recording → clean cut (Claude + Remotion) agent. A single clean MP4 with fillers/dead-air removed, a reviewable cut list (timestamps + words removed + runtime delta), saved back to the chosen destination after approval.
Is the Raw recording → clean cut (Claude + Remotion) agent free?
Yes. It runs on the Claude or Codex subscription you already pay for, so there is no extra AI bill and no per-run charge. You can build and run unlimited agents on the free plan.
How often does the Raw recording → clean cut (Claude + Remotion) agent run?
You choose: run it on demand, or put it on a schedule (hourly, daily, weekly). Once scheduled it runs unattended, as you, on your own machine.
What does the Raw recording → clean cut (Claude + Remotion) agent need to run?
A Claude or Codex plan capable of running local tools; ffmpeg and faster-whisper (auto-installed if missing); the raw recording as a local file or a shared Drive link.
Does the Raw recording → clean cut (Claude + Remotion) agent use my data? Is it private?
It runs as you, on your own machine, on your real data. The model runs inside your own Claude or Codex, so Implexa never sees your data, accounts, or credentials. Your agent's memory is yours and travels with you across Claude, Codex, and whatever comes next.
How do I build the Raw recording → clean cut (Claude + Remotion) agent?
Install Implexa into your Claude or Codex, then say "build the Raw recording → clean cut (Claude + Remotion) agent" and approve the schedule. Implexa assembles the 10 steps (4 from verified skills) and it runs on its own. About 5 minutes to your first real run.
Can I change what the Raw recording → clean cut (Claude + Remotion) agent does?
Yes. Tell it what to change in plain language and it revises its steps; the next scheduled run uses the change, with no re-scheduling. Every change is versioned, and a run can even propose its own improvements.
changelog
5 proposed improvements- v12Sep 23manual
Remove only v11 Voice Pass; preserve steps 1–9 verbatim, Delivery as step 10, required video_qa, typed inputs and all other authority/settings. Exactly 10 steps; do not reinsert Voice Pass.
- v11Sep 23manual
Steps 5/6/9: preserve 2-second pauses, audit every gap, approve longer exceptions at step 7, verify actual MP4 gaps and source-bound tail, require video_qa. Preserve other contracts.
- v10Aug 21generated
retired duplicate clean-cut setup inputs
- v9Aug 21generated
added range-safe clean-cut input adapter
- v8Aug 20generated
aligned outcome capability with typed presenter input
- v7Aug 20generated
corrected final-output authority for outcome discovery
- v6Aug 20generated
added immutable outcome-search identity
- v5Jul 24manual
rebound step 2 (gap -> skills.sh/automate-this), step 5 (gap -> skills.sh/video-use), step 6 (decision -> clawhub/video-cut), and more
- v4Jun 10manual
Fine-tuned step 5/8 from review feedback: remove non-speech sounds (coughs), cap dead air by inter-word gaps, lead ~0.06s before resumed words to avoid clipping, do not over-cut meaningful restates, support fine word-level cuts
- v3Jun 10manual
added 1 step
- v2Jun 10manual
Refined cut spec: keep natural pauses up to ~3s (only trim longer), remove fillers + fumbles/retakes incl. camera-adjust repeats; support local source files; ffmpeg trim/concat render
- v1Jun 9generated
auto-generated from "Take a raw talking-head recording, transcribe it, remove filler words, hesitatio"
Agents are alive: every change is a version, and a run can propose improvements that get reviewed and applied.
related agents
- on-demandcreator
How do I generate cinematic B-roll clips from a source video and a shot-direction brief using Runway?
cinematic b-roll generator
- on demandmarketing
How do I combine Runway-generated B-roll with a HeyGen talking-head avatar into one finished video?
Cinematic B-Roll Generator with Avatar Presenter
- on demandvideo production
How do I revise a finished video from written feedback without redoing the parts I already approved?
Final Video Editor With Avatar Presenter
- video production
How do I turn a markdown animation brief into working Remotion text-animation components?
Markdown Brief to Remotion Text Animations