English | 中文
Drop a video link in the chat. Get back a folder with the subtitles, the full text, and a note you can actually review later.
Install · Pipeline · First Run · Limits
I kept saving knowledge videos on Douyin, YouTube, and Bilibili, and never rewatched any of them. This skill turns each one into a folder of text I can search and skim. A 5-minute video becomes a page I can read in 40 seconds.
link or local file
→ captions first (YouTube / Bilibili: official subtitles, no ASR)
→ otherwise download (Douyin direct / yt-dlp) and transcribe (Groq cloud Whisper, free tier)
→ agent writes a structured note with a timestamped outline
→ everything lands in VideoNotes/YYYYMMDD-<topic>/ (video + SRT + full text + note)
A real output, from a 5-minute Douyin video about monetizing a personal brand:
VideoNotes/20260907-personal-brand/
├── jianghushuo.mp4 # the original video
├── jianghushuo_transcript.srt # 201 segments, timestamped
├── jianghushuo_fulltext.txt # plain full text
└── notes.md # timeline table + layered breakdown
One command, powered by the open skills CLI:
npx skills add Djcloveml/video-notesThe CLI walks you through picking agents and scope interactively. To skip the prompts:
Claude Code
npx skills add Djcloveml/video-notes --agent claude-code -yKimi Code
npx skills add Djcloveml/video-notes --agent kimi-code -yCodex / Cursor / others
npx skills add Djcloveml/video-notes --agent codex -y
npx skills add Djcloveml/video-notes --agent cursor -yFull list: 75+ supported agents.
Useful flags:
-g— install globally (~ directory), available in every project. Without it, the skill installs into the current project only.--copy— copy files instead of symlinking (pick this if you plan to hack on the skill).-l— list what's in the repo before installing.
The repo is private while in early development. Until it goes public, installs need GitHub auth on your machine (gh auth login, or set GITHUB_TOKEN) — the CLI picks it up automatically. After that, the plain command works as-is.
To verify the install, send your agent a video link and say "提取知识点" (or "take notes on this video"). First run starts with a short setup check — see below.
The skill checks your machine with scripts/setup.sh and tells you what to install: ffmpeg, a Python venv (it creates ~/video-notes-env itself), a free Groq API key, and yt-dlp. Each missing item comes with the install command.
It also asks you one question, once, and remembers the answer:
- (a) No fallback (default): if a site blocks plain downloads, the skill reports the failure and stops.
- (b) Browser fallback: a CDP-capable browser controller (Playwright, or an agent-browser skill you already have) captures the real network response, which beats most anti-bot walls. You want this if your Douyin direct downloads start failing.
- (c) API-only: captions via APIs for YouTube/Bilibili, nothing else.
Pick (a) if you mostly use YouTube/Bilibili. Pick (b) if you live on Douyin.
- Douyin changes its anti-bot measures every few months. The direct-link route has broken once already; the browser fallback exists for exactly this reason. If neither works on a given day, the skill says so instead of silently failing.
- Groq's free tier caps at about 8 hours of audio a day and 25MB per file. Long recordings need splitting or a paid key.
- ASR makes homophone errors, especially with jargon. The note step fixes the obvious ones; the SRT always keeps the raw transcript.
- Videos behind a login wall need a browser session where you're already logged in.
video-notes/
├── SKILL.md # the pipeline contract
├── reference/ # loaded only when needed
│ ├── acquire.md # per-site routing, caption-first, CDP fallback
│ ├── transcription.md # Groq details, local faster-whisper option
│ ├── note-writing.md # note structure and rules
│ ├── setup.md # what each dependency is for
│ └── evals.md # regression scenarios
└── scripts/
├── setup.sh # env check + first-run onboarding
├── download_douyin.py # no-login Douyin download
└── transcribe_groq.py # audio → SRT + full text
MIT.