Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

video-distill

An Agent Skill that turns any video into a structured distillation note. Transcription runs locally and for free; on-screen text is recovered by reading deduplicated frames.

Built for coding agents that support Agent Skills (Qoder, Claude Code, Codex, Cursor, Copilot, Gemini CLI, and others).

Credits

This skill stands on the shoulders of two prior projects, and it is only fair to say so up front:

  • bradautomates/claude-video (MIT) — the original "let the agent watch a video" skill. The overall shape of the pipeline (captions first, extract frames, dedupe near-identical frames, hand frames + transcript to the model) is its idea. No code was copied: the scripts here were written from scratch, and the dedupe step uses a different algorithm (dHash + Hamming distance rather than 16×16 grayscale mean-absolute-difference).
  • opencli by jackwener (Apache-2.0) — used as an external command-line dependency to reach sites that require a logged-in session (bilibili and friends). No code borrowed; the skill simply shells out to the installed CLI.

Why this exists

claude-video is excellent, and if your videos have captions you may not need anything else. Two gaps pushed this skill into existence:

  1. Caption-less videos cost money and leak audio. The original falls back to the Groq or OpenAI Whisper API, which means an API key and shipping your audio to a third party. That is a non-starter for confidential recordings (internal training, client calls, unreleased material). Here, transcription runs on local faster-whisper (CPU int8): no API key, no network, no cost. A strong desktop CPU reaches roughly 5–6× realtime on the small model, so a 2h41m lecture transcribes in about 28 minutes.
  2. Sites that require a login are unreachable. Anonymous yt-dlp requests to bilibili get throttled into IncompleteRead and timeouts, and its subtitle API returns nothing without authentication. This skill documents a working path through opencli, including which exit code means "log in" versus "this video genuinely has no subtitles", so the agent stops guessing.

It also leans harder on token economy when reading frames. A 2h41m screen recording produced 2021 scene-change candidates; a dHash pass collapsed them to 69 distinct screens, which were then tiled into 5 contact sheets for triage so only the informative frames were read at full resolution.

What it produces

A single markdown note: a one-line verdict, the substance, what is reusable, and a dedicated section for facts that only exist on screen (paper titles, URLs, dataset accessions, tool names). Raw transcript and keyframes are kept alongside it.

Requirements

Purpose Dependency Cost
Frame extraction, media probing ffmpeg / ffprobe free
Downloading URLs, listing captions yt-dlp free
Local transcription faster-whisper (Python) free
Frame dedupe + contact sheets pillow (Python) free
Logged-in sites (bilibili etc.) opencli free, optional

No API keys. Nothing is uploaded.

# macOS
brew install ffmpeg yt-dlp
# Windows
winget install Gyan.FFmpeg
pip install yt-dlp
# Linux
sudo apt install ffmpeg && pipx install yt-dlp

pip install faster-whisper pillow

GPU note: faster-whisper accelerates on NVIDIA CUDA only. On AMD or Intel graphics it runs on CPU, which is fine. With an NVIDIA card, switch device="cpu" to "cuda" and compute_type="int8" to "float16" in scripts/transcribe.py.

Install

Copy the whole folder into your agent's skills directory:

Host Path
Qoder ~/.qoder/skills/video-distill/
Claude Code ~/.claude/skills/video-distill/
Codex ~/.codex/skills/video-distill/
Cursor ~/.cursor/skills/video-distill/
git clone https://github.com/Seraph310/video-distill.git
cp -r video-distill ~/.qoder/skills/video-distill

Use a project-local skills directory instead if you want it scoped to one repository.

Usage

Hand the agent a video and ask for what you want:

Distill this video: ~/Downloads/training.mp4
Summarize https://www.youtube.com/watch?v=... into notes
What does this screen recording actually demo? bug-repro.mov

The agent follows SKILL.md. In short:

  1. Look for subtitles first. Cheapest possible path; a captioned video needs no download and no transcription.
  2. Transcribe locally only when there are genuinely no subtitles.
  3. Read the transcript.
  4. Extract frames only when the picture carries information the audio does not: screen recordings, slide decks, software demos. Skipped for talking-head, podcast, and interview footage.
  5. Distill audio and screen into one note. Proper nouns are cross-checked against the frames rather than trusted from the audio, timestamps are kept so the note stays an index into the video, and gaps are stated rather than hidden.

Chinese and other languages

Auto-detection is the default. For better accuracy on a known language, pass it explicitly and seed the decoder with domain vocabulary:

python scripts/transcribe.py --input lecture.mp4 --out transcript.txt \
  --language zh \
  --prompt "A Mandarin lecture on bioinformatics covering Mendelian randomization, eQTL, GWAS, and single-cell analysis."

Frames

# 1) scene-change candidates (lower the threshold for slide decks on white backgrounds)
ffmpeg -y -hide_banner -i video.mp4 \
  -vf "select='gt(scene,0.3)',showinfo,scale=1440:-2" \
  -vsync vfr -q:v 3 "frames/f_%04d.jpg" 2> showinfo.log

# 2) collapse near-identical frames; aim for 40-90 distinct screens
python scripts/dedup_frames.py 12      # raise to 20-24 if too many survive

# 3) tile them into contact sheets for triage
python scripts/montage_frames.py

Read the contact sheets first to classify screens, then read only the informative keyframes at full resolution.

Sites that require a login

Verified against bilibili. Query subtitles before downloading anything:

opencli bilibili subtitle <video-id> -f json   # full subtitles
opencli bilibili summary  <video-id> -f json   # official AI outline with timestamps
opencli bilibili video    <video-id> -f json   # title, duration, uploader

Branch on the exit code, never on the error text:

Code Meaning Action
0 success use the subtitles
66 no data this video really has no subtitles, fall back to audio + local transcription
77 auth required ask the user to run opencli bilibili login and scan the QR code
69 browser bridge down check the extension
75 timeout retry

Before logging in you cannot even determine whether subtitles exist. Do not assume the user is already logged in; verify with the exit code first.

See reference.md for tuning thresholds, model selection, and platform-specific pitfalls.

Limits

  • Frame reading is the expensive part. Contact-sheet triage exists so you spend image tokens only on screens that carry information.
  • Machine-generated subtitles (including bilibili's AI captions) contain homophone errors. Proper nouns need domain knowledge to restore.
  • The small model trades accuracy for speed on purpose, on the assumption that a model will distill the transcript afterwards. Use --model medium --beam_size 5 when a human will read the transcript directly.

License

MIT. See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages