Skip to content

Repository files navigation

MotionKit

MotionKit

A local studio for code-defined motion templates.

60 templates across maps, text graphics, image motion, newspaper devices and 2.5D depth — each one a pure function of time, exported to real video in the browser.


The UI implements MotionKit.dc.html (kept in design/ for diffing against future design revisions); the map templates reproduce Canva's Map video animation.

npm run dev      # dev server at localhost:5173
npm run app      # run the desktop shell against dist/
npm run dist     # build the Windows installer

npm test · npm run typecheck · npm run build

Desktop build

npm run dist produces MotionKit Setup <version>.exe — a one-click NSIS installer, per-user so it needs no admin rights.

The Electron shell (electron/main.cjs) serves dist/ over loopback HTTP rather than loading it from file://. That is not incidental: export inlines every referenced asset per frame via fetch, and file:// blocks same-origin fetches and subresources inside rasterized SVGs, so flags, dropped photos and webfonts would silently vanish from the output. A 40-line static server on a random loopback port keeps the renderer byte-identical to the dev experience.

The build output goes outside the project tree (../MotionKit-release). Vite's file watcher holds handles on anything inside the repo, which makes electron-builder's staging rename fail with EPERM on Windows.

Data regeneration: npm run gen:iso · npm run gen:cities · npm run gen:urban

Templates

Sixty, in eight categories. The gallery rail lists only categories that hold a template, so a slot declared in CATEGORY_LABELS stays invisible until something real ships into it.

Nineteen are hand-rolled and described below. The other forty-one come from families (src/templates/families.tsx): one engine per family, one gallery entry per kind, with motion, schema and chrome shared — text (16), single image (11), multi image (8) and newspaper & docs (6). buildFamilies takes the same define the hand-rolled templates use, so a family entry is not a second class of template.

Map animation

map-annotate — the documentary-explainer workhorse. A region held under a slow drift while actors punch onto it as octagon flag chips, their countries tint, and links draw between them. The caption plate at the top is the scene change: the camera never cuts, so a new caption is the edit. Actors are written Country: Label, links From > To, both as reorderable lists.

Chips start on their country's centroid and are pushed apart by a deterministic relaxation pass — Iraq and Iran are adjacent enough that their chips would otherwise sit on top of each other.

map-route — holds on country 1 while its border strokes on and its flag fills the shape, flies the great circle to country 2 with the marker rotated to its bearing and a trail drawing behind it, then lands and repeats the reveal.

map-zoom — pushes from a wide shot onto one country.

map-city — frames the country with its flag, then drops onto one city (type any of ~9,000 places ≥50k) and gives it the same treatment: its built-up area outlines and fills. Past country scale there is no offline street or river data, so the surrounding frame carries a view-fitted graticule, km-labelled range rings and neighbouring towns rather than a bare colour field.

The camera never pans — it sits on the city for the whole shot and only changes zoom, so the push lands on the city rather than drifting past it. The wide end still shows the whole country because the fit box is the country's bounds mirrored about the city.

City outlines are real: Natural Earth's satellite-derived urban areas, matched to each city by point-in-polygon (npm run gen:urban). 6,949 of the 9,062 cities have one. Where Natural Earth maps nothing, a seeded blob derived from population stands in — same shape every render, never Math.random() — and urbanAreaFor returning null is what selects it.

Framing is a multiple of that footprint, not an absolute kilometre count: a fixed span that suits a one-million city buries a twelve-million one. cityFill picks flag, flat colour or outline only, with its own colour in the style group — by default a flat colour, since a city sitting on top of a matching country flag reads as one shape rather than two.

Controls mirror Canva's panel — country 1, country 2, marker icon row, flag fill, place names — plus duration, framing, and a full colour set.

Reveal

lower-third — name plate that slides in over transparency. Alpha-capable, and the main reason the format picker exists: this is the one you export as ProRes 4444 and drop onto footage in an NLE. Its plate box is a zRect, so it positions by dragging on the live preview.

word-card — full-frame statement, revealed line by line behind a per-line mask so the text wipes up rather than fading.

news-card — the dominant device in the Indian-explainer references: a pinned clipping on a board, phrases highlighting in sequence, then names ringed in red. That staggered mark-up is the narration, which is why it reads as someone reading rather than a slide appearing. Marks are addressed as line: phrase.

photo-card — a taped photo that drifts while the image pushes inside its own frame. Two speeds, which is what stops a still reading as a freeze.

doc-reveal — the leaked-document beat: a page pushes in slowly while a highlighter sweeps one line and a hand-drawn ring closes on another. The marker uses mix-blend-mode: multiply, so the text stays readable through it the way a real highlighter behaves.

Data

timeline-scrub — year axis with a date chip riding a playhead. Alpha-capable, because it is always composited over footage. This appears in every reference video more often than anything else: it is how these films say "here is where we are" without cutting away.

The four below come from a minimal astronomy-explainer reference: cold dark field, a kicker in tiny letterspaced caps top-left, watermark bottom-right, and one small diagram carrying the whole shot. Its discipline is that nothing is decorative — every mark is a quantity — so all of them derive their geometry from the numbers rather than from layout constants. The frame furniture is one shared Chrome, because the diagram changes and the chrome never does.

stat-compare — bars sized from the largest value in the set. "At true relative scale" is the entire claim the graphic makes, so a fixed axis would make it a lie.

unit-chart — one dot per unit, grouped and labelled. Counting beats reading a number.

measure-span — axis, endpoints, headline value, optional uncertainty whiskers. At zero spread it draws a point, not a hairline bracket: a measurement with no stated error should not look like one with a tiny error.

depth-zones — stratified layers scaled to their real extent, with a marker descending to depth. The abyss dominates the frame because it actually does.

Rows are written Label: value and bands Label: from-to, split on the last colon so a label may contain one. Anything unparseable is dropped rather than charted as zero — a silent 0 draws a bar that says something false.

counter — odometer roll on tabular figures, alpha-capable. Every wheel turns continuously but the tens wheel advances one unit per ten of the ones; value / 10^i % 10 gives that mechanical behaviour for free, which is what separates a roll from a number that merely changes.

Depth

Both are flat images treated as planes — which is what the reference videos actually do. Neither is a rendered 3D scene.

photo-parallax — 2.5D push on a still. With a depth map this is true per-pixel parallax via feDisplacementMap: near pixels shift further than far ones, so the image separates as the camera travels. Without one it falls back to a two-layer radial split (PRD §13's stated fallback), which reads as depth but is a guess about where "near" is.

Two things decide whether it works. The displacement scale must be driven by the camera offset — a constant scale is a permanent warp, where near pixels sit offset but never move relative to far ones, which is the entire effect. And feDisplacementMap takes x from the red channel; a greyscale map has R=G=B, so using green for y smears diagonally. Alpha is flat, giving a constant vertical shift that the group then translates back out.

Lateral paths (orbit-*, handheld) are where this reads. A pure push-in needs radial displacement, which this filter cannot express, so it degrades to a plain push rather than faking it.

photo-reel — a run of related images, each entering and leaving in turn. Everything else in the kit enters and holds; this is the one that also gets out, which is what a montage needs — the exit of one card is the entrance of the next, so slots overlap rather than leaving the frame empty between images. Four in/out styles: slide, scale, whip (blur peaks at the extremes, not at rest), mask.

photo-stack — photos as planes at real depths under a pinhole camera. For a plane parallel to the screen the projection is exactly scale = focal / z, so genuine perspective falls out of affine transforms — no WebGL, and the parallax between layers is correct rather than eyeballed. Painter's ordering, and a plane behind the camera is dropped rather than projected to a mirrored ghost.

How the map renders

The world is projected once into a fixed 4096-unit square and cached as path strings. Mercator is conformal, so camera movement is then a pure affine transform on that group: one attribute changes per frame, not 236 path recomputations. Everything that must stay a constant size — marker, label pill — is placed in screen space from the same camera.

Details that decide whether a map animation reads as real:

  • Bearing. Sampled either side of t and applied as marker rotation. Skipping it is the classic broken-map tell.
  • Log-space zoom. Linear scale interpolation lurches at world zoom.
  • Antimeridian. The arc is unwrapped past ±180° and the basemap is tiled three wide, so Tokyo→Los Angeles flies the short way over the Pacific.
  • No non-scaling-stroke on animated dashes. Chrome measures dashes in screen units under that effect, which silently defeats pathLength normalization — the trail renders full-length as micro-dashes. Widths are divided by the camera zoom instead.
  • Cover, don't fit, the flag. The flag SVGs have a viewBox but no width/height, so they have no intrinsic size — and for those Chrome ignores preserveAspectRatio on <image> and letterboxes with the file's own default. That left tall countries part-painted (India's north and southern tip). The image gets a rect that is already 4:3 and covers the bounding box, so every fit mode agrees. coverRect is tested.
  • Kilometres are latitude-dependent. Mercator stretches by 1/cos(lat), so zoomForKm divides by it — otherwise a 50 km view of Oslo and one of Nairobi would be wildly different scales.
  • Scope every <defs> id by its content. Several cards render at once and SVG resolves a duplicate id to whichever came first in the document, so a shared literal id meant every card silently borrowed the first card's gradient and clip path. terrainId keys on the land colour and cityClipId on the city, so matching cards share a def and differing ones never collide.
  • Mix, don't add, when shading by latitude. The terrain bands originally added fixed RGB offsets, which clipped straight to white on a pale land colour — zoomed out far enough to show the ice caps, the whole map went blank. Bands are now a signed mix toward ice or deep cover, so they hold at any base lightness.
  • Make the estimate true, not accurate. There is no DOM to measure text against, and an estimate is not good enough when a highlight must land on one phrase inside a centred line — the error compounds outward from a mis-placed centre. Body lines render with textLength set to the same figure the mark maths uses, so the two agree by construction rather than by luck.
  • Store shared geometry once. 6,949 cities resolve to only 4,319 distinct urban polygons — a conurbation is one shape, so Mumbai and every satellite town inside it point at the same record. Keying by city instead of by shape made the file 4.8 MB; deduping brought it to 1.2 MB.

Geo data is world-atlas 50m TopoJSON, flag-icons, a compacted all-the-cities extract, and Natural Earth 10m urban areas — all generated into the repo by the gen:* scripts. Nothing is fetched at runtime.

Layout

Path What
src/core/geo/ world.ts (atlas, projection, country lookup), camera.ts (fit, log zoom, great-circle arc), cities.ts (place lookup, km→zoom), urban.ts (real footprints), iso.ts (generated)
src/templates/map/ templates.tsx (the four templates), scene.tsx (basemap, flag fill, graticule, label, pins), markers.tsx
src/templates/kit/ The family engines — textfx.tsx, imagefx.tsx, paper.tsx, plus science.tsx, depth.tsx and shared common.tsx
src/templates/families.tsx buildFamilies — one gallery entry per kind
src/core/ rect.ts, schema-types.ts, clock.ts, types.ts, tokens.ts, export.ts (WebCodecs render/mux)
src/panels/ resolver.ts — the schema→control walk (PRD §9). ControlPanel.tsx renders whatever it returns.
src/ui/controls.tsx The widgets the resolver can select
src/screens/ Gallery.tsx, Editor.tsx

Adding a template is one entry in TEMPLATES. Its panel, its control count, the format gating and the filters all follow from its Zod schema — there is no per-template UI anywhere.

Not built here

  • Remotion. Templates are already pure functions of normalized time p, so the swap is p → useCurrentFrame() / durationInFrames and the preview surface → @remotion/player.
  • The server — partly. Export is real and in-browser (src/core/export.ts): every template is a pure function of p, so each frame is serialized to SVG, rasterized, encoded with WebCodecs and muxed to MP4 (H.264) or WebM (VP9), with explicit frame timestamps — a 5s clip is exactly 5s. Assets, flags and webfonts are inlined as data: URIs because SVG-as-image loads no external resources; the missing xmlns (renderToStaticMarkup omits it) is pinned or the blob won't parse. ProRes 4444, PNG sequences, GIF and true alpha remain server work and their picker entries say so. One export runs at a time — PRD's concurrency of 1 by construction.
  • Preset files. Presets live in memory, not ~/MotionKit/presets/*.json.
  • Asset persistence. src/core/assets.ts keeps object URLs for the session, which is enough for templates to render what you dropped on them. The real thing writes content-hashed files into ~/MotionKit/assets/ via the server — props already store an id rather than a path (PRD §16), so only the store behind assetUrl changes. Reload and the bytes are gone while the id survives, which is why an unresolved asset draws a placeholder instead of throwing.
  • Satellite basemap. Raster tiles are out under the offline constraint, so land is shaded by latitude band instead (BANDS in scene.tsx).
  • Streets, rivers, terrain below city scale. Urban footprints are real, but what surrounds them is 50m country geometry plus a graticule. Roads or a rivers layer would be the next step up.
  • Fonts load from Google Fonts, not assets/fonts — see the note in index.html.

About

Local studio for code-defined motion templates — 60 schema-driven React/SVG animations across text, image, montage, map, paper, reveal, data and depth. Panels are generated from each template's Zod schema, export runs in-browser via WebCodecs, and nothing is fetched at runtime.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages