An Agent Skill that turns any video into a structured distillation note. Transcription runs locally and for free; on-screen text is recovered by reading deduplicated frames.
Built for coding agents that support Agent Skills (Qoder, Claude Code, Codex, Cursor, Copilot, Gemini CLI, and others).
This skill stands on the shoulders of two prior projects, and it is only fair to say so up front:
- bradautomates/claude-video (MIT) — the original "let the agent watch a video" skill. The overall shape of the pipeline (captions first, extract frames, dedupe near-identical frames, hand frames + transcript to the model) is its idea. No code was copied: the scripts here were written from scratch, and the dedupe step uses a different algorithm (dHash + Hamming distance rather than 16×16 grayscale mean-absolute-difference).
- opencli by jackwener (Apache-2.0) — used as an external command-line dependency to reach sites that require a logged-in session (bilibili and friends). No code borrowed; the skill simply shells out to the installed CLI.
claude-video is excellent, and if your videos have captions you may not need anything else. Two gaps pushed this skill into existence:
- Caption-less videos cost money and leak audio. The original falls back to the Groq or OpenAI Whisper API, which means an API key and shipping your audio to a third party. That is a non-starter for confidential recordings (internal training, client calls, unreleased material). Here, transcription runs on local
faster-whisper(CPU int8): no API key, no network, no cost. A strong desktop CPU reaches roughly 5–6× realtime on thesmallmodel, so a 2h41m lecture transcribes in about 28 minutes. - Sites that require a login are unreachable. Anonymous
yt-dlprequests to bilibili get throttled intoIncompleteReadand timeouts, and its subtitle API returns nothing without authentication. This skill documents a working path throughopencli, including which exit code means "log in" versus "this video genuinely has no subtitles", so the agent stops guessing.
It also leans harder on token economy when reading frames. A 2h41m screen recording produced 2021 scene-change candidates; a dHash pass collapsed them to 69 distinct screens, which were then tiled into 5 contact sheets for triage so only the informative frames were read at full resolution.
A single markdown note: a one-line verdict, the substance, what is reusable, and a dedicated section for facts that only exist on screen (paper titles, URLs, dataset accessions, tool names). Raw transcript and keyframes are kept alongside it.
| Purpose | Dependency | Cost |
|---|---|---|
| Frame extraction, media probing | ffmpeg / ffprobe |
free |
| Downloading URLs, listing captions | yt-dlp |
free |
| Local transcription | faster-whisper (Python) |
free |
| Frame dedupe + contact sheets | pillow (Python) |
free |
| Logged-in sites (bilibili etc.) | opencli |
free, optional |
No API keys. Nothing is uploaded.
# macOS
brew install ffmpeg yt-dlp
# Windows
winget install Gyan.FFmpeg
pip install yt-dlp
# Linux
sudo apt install ffmpeg && pipx install yt-dlp
pip install faster-whisper pillowGPU note: faster-whisper accelerates on NVIDIA CUDA only. On AMD or Intel graphics it runs on CPU, which is fine. With an NVIDIA card, switch device="cpu" to "cuda" and compute_type="int8" to "float16" in scripts/transcribe.py.
Copy the whole folder into your agent's skills directory:
| Host | Path |
|---|---|
| Qoder | ~/.qoder/skills/video-distill/ |
| Claude Code | ~/.claude/skills/video-distill/ |
| Codex | ~/.codex/skills/video-distill/ |
| Cursor | ~/.cursor/skills/video-distill/ |
git clone https://github.com/Seraph310/video-distill.git
cp -r video-distill ~/.qoder/skills/video-distillUse a project-local skills directory instead if you want it scoped to one repository.
Hand the agent a video and ask for what you want:
Distill this video: ~/Downloads/training.mp4
Summarize https://www.youtube.com/watch?v=... into notes
What does this screen recording actually demo? bug-repro.mov
The agent follows SKILL.md. In short:
- Look for subtitles first. Cheapest possible path; a captioned video needs no download and no transcription.
- Transcribe locally only when there are genuinely no subtitles.
- Read the transcript.
- Extract frames only when the picture carries information the audio does not: screen recordings, slide decks, software demos. Skipped for talking-head, podcast, and interview footage.
- Distill audio and screen into one note. Proper nouns are cross-checked against the frames rather than trusted from the audio, timestamps are kept so the note stays an index into the video, and gaps are stated rather than hidden.
Auto-detection is the default. For better accuracy on a known language, pass it explicitly and seed the decoder with domain vocabulary:
python scripts/transcribe.py --input lecture.mp4 --out transcript.txt \
--language zh \
--prompt "A Mandarin lecture on bioinformatics covering Mendelian randomization, eQTL, GWAS, and single-cell analysis."# 1) scene-change candidates (lower the threshold for slide decks on white backgrounds)
ffmpeg -y -hide_banner -i video.mp4 \
-vf "select='gt(scene,0.3)',showinfo,scale=1440:-2" \
-vsync vfr -q:v 3 "frames/f_%04d.jpg" 2> showinfo.log
# 2) collapse near-identical frames; aim for 40-90 distinct screens
python scripts/dedup_frames.py 12 # raise to 20-24 if too many survive
# 3) tile them into contact sheets for triage
python scripts/montage_frames.pyRead the contact sheets first to classify screens, then read only the informative keyframes at full resolution.
Verified against bilibili. Query subtitles before downloading anything:
opencli bilibili subtitle <video-id> -f json # full subtitles
opencli bilibili summary <video-id> -f json # official AI outline with timestamps
opencli bilibili video <video-id> -f json # title, duration, uploaderBranch on the exit code, never on the error text:
| Code | Meaning | Action |
|---|---|---|
0 |
success | use the subtitles |
66 |
no data | this video really has no subtitles, fall back to audio + local transcription |
77 |
auth required | ask the user to run opencli bilibili login and scan the QR code |
69 |
browser bridge down | check the extension |
75 |
timeout | retry |
Before logging in you cannot even determine whether subtitles exist. Do not assume the user is already logged in; verify with the exit code first.
See reference.md for tuning thresholds, model selection, and platform-specific pitfalls.
- Frame reading is the expensive part. Contact-sheet triage exists so you spend image tokens only on screens that carry information.
- Machine-generated subtitles (including bilibili's AI captions) contain homophone errors. Proper nouns need domain knowledge to restore.
- The
smallmodel trades accuracy for speed on purpose, on the assumption that a model will distill the transcript afterwards. Use--model medium --beam_size 5when a human will read the transcript directly.
MIT. See LICENSE.