A local studio for code-defined motion templates.
60 templates across maps, text graphics, image motion, newspaper devices and 2.5D depth — each one a pure function of time, exported to real video in the browser.
The UI implements MotionKit.dc.html (kept in design/ for diffing against
future design revisions); the map templates reproduce Canva's Map video
animation.
npm run dev # dev server at localhost:5173
npm run app # run the desktop shell against dist/
npm run dist # build the Windows installernpm test · npm run typecheck · npm run build
npm run dist produces MotionKit Setup <version>.exe — a one-click NSIS
installer, per-user so it needs no admin rights.
The Electron shell (electron/main.cjs) serves dist/ over loopback HTTP
rather than loading it from file://. That is not incidental: export inlines
every referenced asset per frame via fetch, and file:// blocks same-origin
fetches and subresources inside rasterized SVGs, so flags, dropped photos and
webfonts would silently vanish from the output. A 40-line static server on a
random loopback port keeps the renderer byte-identical to the dev experience.
The build output goes outside the project tree (../MotionKit-release).
Vite's file watcher holds handles on anything inside the repo, which makes
electron-builder's staging rename fail with EPERM on Windows.
Data regeneration: npm run gen:iso · npm run gen:cities · npm run gen:urban
Sixty, in eight categories. The gallery rail lists only categories that hold a
template, so a slot declared in CATEGORY_LABELS stays invisible until
something real ships into it.
Nineteen are hand-rolled and described below. The other forty-one come from
families (src/templates/families.tsx): one engine per family, one gallery
entry per kind, with motion, schema and chrome shared — text (16), single image
(11), multi image (8) and newspaper & docs (6). buildFamilies takes the same
define the hand-rolled templates use, so a family entry is not a second class
of template.
map-annotate — the documentary-explainer workhorse. A region held under a
slow drift while actors punch onto it as octagon flag chips, their countries
tint, and links draw between them. The caption plate at the top is the scene
change: the camera never cuts, so a new caption is the edit. Actors are written
Country: Label, links From > To, both as reorderable lists.
Chips start on their country's centroid and are pushed apart by a deterministic relaxation pass — Iraq and Iran are adjacent enough that their chips would otherwise sit on top of each other.
map-route — holds on country 1 while its border strokes on and its flag
fills the shape, flies the great circle to country 2 with the marker rotated to
its bearing and a trail drawing behind it, then lands and repeats the reveal.
map-zoom — pushes from a wide shot onto one country.
map-city — frames the country with its flag, then drops onto one city
(type any of ~9,000 places ≥50k) and gives it the same treatment: its built-up
area outlines and fills. Past country scale there is no offline street or river
data, so the surrounding frame carries a view-fitted graticule, km-labelled
range rings and neighbouring towns rather than a bare colour field.
The camera never pans — it sits on the city for the whole shot and only changes zoom, so the push lands on the city rather than drifting past it. The wide end still shows the whole country because the fit box is the country's bounds mirrored about the city.
City outlines are real: Natural Earth's satellite-derived urban areas,
matched to each city by point-in-polygon (npm run gen:urban). 6,949 of the
9,062 cities have one. Where Natural Earth maps nothing, a seeded blob derived
from population stands in — same shape every render, never Math.random() — and
urbanAreaFor returning null is what selects it.
Framing is a multiple of that footprint, not an absolute kilometre count: a
fixed span that suits a one-million city buries a twelve-million one. cityFill
picks flag, flat colour or outline only, with its own colour in the style group —
by default a flat colour, since a city sitting on top of a matching country flag
reads as one shape rather than two.
Controls mirror Canva's panel — country 1, country 2, marker icon row, flag fill, place names — plus duration, framing, and a full colour set.
lower-third — name plate that slides in over transparency. Alpha-capable,
and the main reason the format picker exists: this is the one you export as
ProRes 4444 and drop onto footage in an NLE. Its plate box is a zRect, so it
positions by dragging on the live preview.
word-card — full-frame statement, revealed line by line behind a per-line
mask so the text wipes up rather than fading.
news-card — the dominant device in the Indian-explainer references: a
pinned clipping on a board, phrases highlighting in sequence, then names ringed
in red. That staggered mark-up is the narration, which is why it reads as
someone reading rather than a slide appearing. Marks are addressed as
line: phrase.
photo-card — a taped photo that drifts while the image pushes inside its
own frame. Two speeds, which is what stops a still reading as a freeze.
doc-reveal — the leaked-document beat: a page pushes in slowly while a
highlighter sweeps one line and a hand-drawn ring closes on another. The marker
uses mix-blend-mode: multiply, so the text stays readable through it the way a
real highlighter behaves.
timeline-scrub — year axis with a date chip riding a playhead.
Alpha-capable, because it is always composited over footage. This appears in
every reference video more often than anything else: it is how these films say
"here is where we are" without cutting away.
The four below come from a minimal astronomy-explainer reference: cold dark
field, a kicker in tiny letterspaced caps top-left, watermark bottom-right, and
one small diagram carrying the whole shot. Its discipline is that nothing is
decorative — every mark is a quantity — so all of them derive their geometry from
the numbers rather than from layout constants. The frame furniture is one shared
Chrome, because the diagram changes and the chrome never does.
stat-compare — bars sized from the largest value in the set. "At true
relative scale" is the entire claim the graphic makes, so a fixed axis would make
it a lie.
unit-chart — one dot per unit, grouped and labelled. Counting beats reading
a number.
measure-span — axis, endpoints, headline value, optional uncertainty
whiskers. At zero spread it draws a point, not a hairline bracket: a measurement
with no stated error should not look like one with a tiny error.
depth-zones — stratified layers scaled to their real extent, with a marker
descending to depth. The abyss dominates the frame because it actually does.
Rows are written Label: value and bands Label: from-to, split on the last
colon so a label may contain one. Anything unparseable is dropped rather than
charted as zero — a silent 0 draws a bar that says something false.
counter — odometer roll on tabular figures, alpha-capable. Every wheel
turns continuously but the tens wheel advances one unit per ten of the ones;
value / 10^i % 10 gives that mechanical behaviour for free, which is what
separates a roll from a number that merely changes.
Both are flat images treated as planes — which is what the reference videos actually do. Neither is a rendered 3D scene.
photo-parallax — 2.5D push on a still. With a depth map this is true
per-pixel parallax via feDisplacementMap: near pixels shift further than far
ones, so the image separates as the camera travels. Without one it falls back to
a two-layer radial split (PRD §13's stated fallback), which reads as depth but is
a guess about where "near" is.
Two things decide whether it works. The displacement scale must be driven by
the camera offset — a constant scale is a permanent warp, where near pixels sit
offset but never move relative to far ones, which is the entire effect. And
feDisplacementMap takes x from the red channel; a greyscale map has R=G=B, so
using green for y smears diagonally. Alpha is flat, giving a constant vertical
shift that the group then translates back out.
Lateral paths (orbit-*, handheld) are where this reads. A pure push-in
needs radial displacement, which this filter cannot express, so it degrades to a
plain push rather than faking it.
photo-reel — a run of related images, each entering and leaving in turn.
Everything else in the kit enters and holds; this is the one that also gets out,
which is what a montage needs — the exit of one card is the entrance of the
next, so slots overlap rather than leaving the frame empty between images. Four
in/out styles: slide, scale, whip (blur peaks at the extremes, not at
rest), mask.
photo-stack — photos as planes at real depths under a pinhole camera. For
a plane parallel to the screen the projection is exactly scale = focal / z, so
genuine perspective falls out of affine transforms — no WebGL, and the parallax
between layers is correct rather than eyeballed. Painter's ordering, and a plane
behind the camera is dropped rather than projected to a mirrored ghost.
The world is projected once into a fixed 4096-unit square and cached as path strings. Mercator is conformal, so camera movement is then a pure affine transform on that group: one attribute changes per frame, not 236 path recomputations. Everything that must stay a constant size — marker, label pill — is placed in screen space from the same camera.
Details that decide whether a map animation reads as real:
- Bearing. Sampled either side of
tand applied as marker rotation. Skipping it is the classic broken-map tell. - Log-space zoom. Linear scale interpolation lurches at world zoom.
- Antimeridian. The arc is unwrapped past ±180° and the basemap is tiled three wide, so Tokyo→Los Angeles flies the short way over the Pacific.
- No
non-scaling-strokeon animated dashes. Chrome measures dashes in screen units under that effect, which silently defeatspathLengthnormalization — the trail renders full-length as micro-dashes. Widths are divided by the camera zoom instead. - Cover, don't fit, the flag. The flag SVGs have a viewBox but no
width/height, so they have no intrinsic size — and for those Chrome ignores
preserveAspectRatioon<image>and letterboxes with the file's own default. That left tall countries part-painted (India's north and southern tip). The image gets a rect that is already 4:3 and covers the bounding box, so every fit mode agrees.coverRectis tested. - Kilometres are latitude-dependent. Mercator stretches by 1/cos(lat), so
zoomForKmdivides by it — otherwise a 50 km view of Oslo and one of Nairobi would be wildly different scales. - Scope every
<defs>id by its content. Several cards render at once and SVG resolves a duplicate id to whichever came first in the document, so a shared literal id meant every card silently borrowed the first card's gradient and clip path.terrainIdkeys on the land colour andcityClipIdon the city, so matching cards share a def and differing ones never collide. - Mix, don't add, when shading by latitude. The terrain bands originally added fixed RGB offsets, which clipped straight to white on a pale land colour — zoomed out far enough to show the ice caps, the whole map went blank. Bands are now a signed mix toward ice or deep cover, so they hold at any base lightness.
- Make the estimate true, not accurate. There is no DOM to measure text
against, and an estimate is not good enough when a highlight must land on one
phrase inside a centred line — the error compounds outward from a mis-placed
centre. Body lines render with
textLengthset to the same figure the mark maths uses, so the two agree by construction rather than by luck. - Store shared geometry once. 6,949 cities resolve to only 4,319 distinct urban polygons — a conurbation is one shape, so Mumbai and every satellite town inside it point at the same record. Keying by city instead of by shape made the file 4.8 MB; deduping brought it to 1.2 MB.
Geo data is world-atlas 50m TopoJSON, flag-icons, a compacted
all-the-cities extract, and Natural Earth 10m urban areas — all generated into
the repo by the gen:* scripts. Nothing is fetched at runtime.
| Path | What |
|---|---|
src/core/geo/ |
world.ts (atlas, projection, country lookup), camera.ts (fit, log zoom, great-circle arc), cities.ts (place lookup, km→zoom), urban.ts (real footprints), iso.ts (generated) |
src/templates/map/ |
templates.tsx (the four templates), scene.tsx (basemap, flag fill, graticule, label, pins), markers.tsx |
src/templates/kit/ |
The family engines — textfx.tsx, imagefx.tsx, paper.tsx, plus science.tsx, depth.tsx and shared common.tsx |
src/templates/families.tsx |
buildFamilies — one gallery entry per kind |
src/core/ |
rect.ts, schema-types.ts, clock.ts, types.ts, tokens.ts, export.ts (WebCodecs render/mux) |
src/panels/ |
resolver.ts — the schema→control walk (PRD §9). ControlPanel.tsx renders whatever it returns. |
src/ui/controls.tsx |
The widgets the resolver can select |
src/screens/ |
Gallery.tsx, Editor.tsx |
Adding a template is one entry in TEMPLATES. Its panel, its control count, the
format gating and the filters all follow from its Zod schema — there is no
per-template UI anywhere.
- Remotion. Templates are already pure functions of normalized time
p, so the swap isp→useCurrentFrame() / durationInFramesand the preview surface →@remotion/player. - The server — partly. Export is real and in-browser (
src/core/export.ts): every template is a pure function ofp, so each frame is serialized to SVG, rasterized, encoded with WebCodecs and muxed to MP4 (H.264) or WebM (VP9), with explicit frame timestamps — a 5s clip is exactly 5s. Assets, flags and webfonts are inlined as data: URIs because SVG-as-image loads no external resources; the missingxmlns(renderToStaticMarkup omits it) is pinned or the blob won't parse. ProRes 4444, PNG sequences, GIF and true alpha remain server work and their picker entries say so. One export runs at a time — PRD's concurrency of 1 by construction. - Preset files. Presets live in memory, not
~/MotionKit/presets/*.json. - Asset persistence.
src/core/assets.tskeeps object URLs for the session, which is enough for templates to render what you dropped on them. The real thing writes content-hashed files into~/MotionKit/assets/via the server — props already store an id rather than a path (PRD §16), so only the store behindassetUrlchanges. Reload and the bytes are gone while the id survives, which is why an unresolved asset draws a placeholder instead of throwing. - Satellite basemap. Raster tiles are out under the offline constraint, so
land is shaded by latitude band instead (
BANDSinscene.tsx). - Streets, rivers, terrain below city scale. Urban footprints are real, but what surrounds them is 50m country geometry plus a graticule. Roads or a rivers layer would be the next step up.
- Fonts load from Google Fonts, not
assets/fonts— see the note inindex.html.