Design#Video generation#Motion graphics

srt-whiteboard-animation: subtitle in, hand-drawn animation out

geeklee's MIT agent skill turns SRT subtitles into whiteboard animation: 25–35 s scenes, strokes synced to the narration, built for explainer videos.

Skill details

Install
npx skills add geeklee/srt-whiteboard-animation --skill srt-whiteboard-animation

Whiteboard animation demo of a monkey-mountain story: the hand draws each element as the narration names it

Whiteboard videos — a hand sketching while the narration plays — are painful to make by hand: you draw the artwork, sync it to the audio, then animate each reveal. Generic video models don’t keep the picture on-script either. srt-whiteboard-animation, an AI skill in geeklee’s repo (MIT, created July 27, 2026, about 3.8k stars as of September 30, 2026), installs into Claude Code or Codex: hand it an SRT file and it splits the script into scenes, produces line art, then draws the scene in narrative order — whatever the narration names appears next, so audio and picture stay in sync by construction.

Monkey-mountain demo: elements are drawn stroke by stroke in narration order

The order it works in

Compared with one-shot video generators like html-explainer, this skill runs on confirmation gates. It parses your SRT into 25–35-second scenes with a storyboard strategy (parse_srt.py); only after you approve does it generate line art. Next it reads the subtitles, inspects the image, maps visible subjects to narrative events in the order setup → subject → action → reaction, writes an annotation.json, and opens a browser preview automatically. You drag regions, order, timing and subtitle links; after you confirm, it renders each scene to MP4 and merges them with merge_scenes.py. The SKILL.md explicitly states that silence or blanket approval never counts as confirmation, so a bad storyboard stops at that step instead of poisoning the whole render.

What it hard-constrains

  • Sync by construction: elements appear in subtitle-event order, not screen position — the narration names it, the hand draws it.
  • Nothing leaks early: every region gets an allowed mask; untouched regions can’t show a single line, and overlapping subjects are shielded with protectedRegions.
  • Continuous strokes: no frame-skip reveals — the pen glides along a skeleton or grid, inking first and coloring second (ink/color split 2:1), and finished strokes stay on the canvas.
  • One visual language: 16:9 warm-paper background #F5EBD7, dark-gray sketch lines, sparing red/orange/blue accents, and no text or labels inside the scene.

The render input is line art like this — generated by your image-capable model under the skill’s style rules:

The monkey-mountain scene’s line art: beige paper, gray sketch lines, generous whitespace

Who it’s for

Anyone producing explainers, story narration or course content who already has an SRT file — drop it in and a few scripts later you have a video. Unlike caption-to-video mixers such as story-flicks, the output is an original drawing animated end to end, not stock clips stitched together. Three caveats: your agent must generate images (a WeChat reviewer walked the full flow in Codex); rendering runs locally, with prepare_env.py creating a venv with opencv-python, numpy and PyAV on first run; and it targets short-form — 25–35 seconds per scene, long videos mean stacking scenes. The repo has seen no commits since its creation day, though as a single document plus a handful of scripts there is little to break; the preview’s save button needs Chrome or Edge.