references/cut-catalog.md
# Cut catalog — within-frame seams (worker-built)
> **A worker build-recipe (Step 5) — the sibling of `../hyperframes-animation/rules/`, not a second motion doc.** These are within-frame cuts the **frame worker builds INSIDE its own composition** (Z-scale + blur + opacity tweens, or per-word x-staggers, all on the frame's own paused GSAP timeline). They are **not** the between-frame transition: story owns that via `transition_in`, which the harness's injector stamps from a **separate registry vocabulary** (`crossfade` / `blur-crossfade` / `push-slide` / `zoom-through` / `squeeze`) — the catalog names here (**cut-the-curve / inverse-zoom / waterfall**) are **not** valid `transition_in` values. Use this catalog when a frame's shot sequence has an internal seam — a within-scene text/element swap, a **Scene-to-Scene** cut (a `Scene` is a time window WITHIN one frame, **not** a frame-to-frame boundary), or a text-to-text line change — and you want it to read as one continuous move instead of a hard slideshow cut. (`zoom-through` lives in both worlds: a whole-frame wrapper transition in the registry, an element-level Z-cut here — same idea, different scope.)
Four techniques that create depth and continuity:
1. **Zoom-Through** — within-scene text swaps, Z-axis, moving TOWARD the viewer
2. **Inverse Zoom-Through** — Z-axis swaps moving AWAY from the viewer
3. **Cut the Curve** — between-scene transitions on x/y
4. **Waterfall Cut** — word-by-word cut-the-curve with staggered exits and entries
All four are the same underlying principle: **cut at peak velocity, match direction and
speed on both sides of the cut.** The differences are axis, scope, and granularity.
**Choosing which at a seam:** for an UNFINISHED phrase (building one larger idea across
several visually distinct scenes that still approach the same point — multi-line text, a
run of consecutive cards) use **cut-the-curve** / **waterfall**. For a STATE CHANGE (turning
to a NEW part of the video — most often hook → context, between two distinct chapters) use
**zoom-through**, and **inverse zoom-through** for an arrival / payoff beat. Chain these so
the frame's internal seams feel like one camera moving through the content.
---
## Blur Logic (applies to all Z-axis variants)
Blur sells the speed at the cut, but it must scale with the SUBJECT SIZE:
| Subject | Peak blur | Why |
| ------------------------------------------------------ | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Text-scale (headline, line, word group) | **10px** | At 20px text smears into illegibility — the eye loses the word it was tracking and the cut reads as a glitch, not speed. (Learned 2026-06-10: the until-now inverse zooms shipped at 20px and read mushy; 10px keeps letterforms readable mid-cut.) |
| Full-frame surface (terminal window, card, screenshot) | **18–20px** | Big surfaces have edges and texture that survive heavy blur; lighter blur on a full-frame move reads as a rendering hiccup instead of motion. |
Both sides of a cut use the SAME peak blur — the value must match at the swap frame.
Apply blur to the WRAPPER, never to individual children.
---
## 1. Zoom-Through (forward)
### The Problem
Text enters, holds, exits. Then next text enters, holds, exits. Each text block is
independent — no depth, no continuity. The video feels like a slideshow.
### The Principle
A velocity-matched cut on the Z-axis. You **never see both texts at the same time.** The
outgoing text scales toward the viewer (accelerating), blur and opacity peak at the cut
point hiding a hard swap, and the incoming text continues scaling up from behind
(decelerating into the focal plane). One continuous forward motion, two different texts.
### The Three Phases
**Phase 1: Exit** — text accelerates forward (toward viewer)
- Scale: `1.0 -> 1.2`, Blur: `0px -> 10px` (text-scale; see Blur Logic), Opacity: `1.0 -> 0.15`
- Scale/blur easing: `power3.in` (steep acceleration)
- Opacity easing: `none` (linear — even dimming, separated from scale)
- Duration: 0.2s
**Phase 2: Hard cut** at peak velocity + peak blur
- Outgoing: `opacity: 0` (instant via `tl.set`)
- Incoming: `opacity: 0.15, scale: 0.75, blur: 10px` (instant via `tl.set`)
- All properties match at the cut: blur, opacity, and scale DIRECTION (both scaling up)
**Phase 3: Entry** — text continues forward (growing into focal plane)
- Scale: `0.75 -> 1.0`, Blur: `10px -> 0px`, Opacity: `0.15 -> 1.0`
- Easing: `expo.out` (steep initial burst matching exit velocity, long settle)
- Duration: 0.5s
### Why Opacity Must Be Separate on Exit
Scale uses `power3.in` but that keeps opacity near 1.0 for most of the tween. Splitting
opacity to its own tween with linear ease makes the dimming even. On entry, all properties
can share `expo.out`.
---
## 2. Inverse Zoom-Through (backward)
The mirror: the camera "pulls back" instead of pushing through. The outgoing element
RECEDES away from the viewer; the incoming element arrives OVERSIZED (as if it had been
just behind the camera) and retracts into the focal plane. Both move in the shrinking
direction — same-direction rule preserved, just reversed.
**When to use over the forward variant:** arrival beats. The incoming element lands with
presence because it comes from larger-than-frame — right for a payoff line ("That changes
today."), a giant reply, or a held end-state. Forward zoom-through reads as _progressing
through_ content; inverse reads as _arriving at_ content.
### The Three Phases
**Phase 1: Exit** — element recedes (away from viewer)
- Scale: `1.0 -> 0.8`, Blur: `0px -> 10px` (text-scale)
- Scale/blur easing: `power3.in`; Opacity: `1.0 -> 0.15` on `none` (separate tween)
- Duration: 0.2s
**Phase 2: Hard cut**
- Outgoing: `opacity: 0` via `tl.set`
- Incoming: `opacity: 0.15, scale: 1.25, blur: 10px` via `tl.set`
**Phase 3: Entry** — incoming retracts into place
- Scale: `1.25 -> 1.0`, Blur: `10px -> 0px`, Opacity: `0.15 -> 1.0`
- Easing: `expo.out`, Duration: 0.5s
Shipped examples: `boring → until-now` and `B3 → "That changes today."` (sfx-music-launch);
seams 3/4 in editorial-paper (`follow-up → thinking`, `thinking → compose UI`).
---
## 3. Cut the Curve (Scene Transitions)
### The Principle
Use cut-the-curve for **all scene-to-scene transitions** on x and y axes. The outgoing
scene's hero element accelerates in one direction, the cut lands mid-motion, and the
incoming scene's hero element continues moving in the **same direction** and decelerates.
Nothing exits fully off-screen and nothing enters from fully off-screen — **speed plus
opacity fading trick the eye**; the partial moves are enough.
### Same Path, Same Direction
If Scene A's hero slides left, Scene B's hero enters from the right and continues sliding
left. Both move leftward. One continuous motion.
| Direction | Scene A exit | Scene B entry start | Scene B entry end |
| --------- | -------------- | ------------------- | ----------------- |
| Leftward | `x: 0 -> -230` | `x: +230` | `x: 0` |
| Rightward | `x: 0 -> +230` | `x: -230` | `x: 0` |
| Upward | `y: 0 -> -230` | `y: +230` | `y: 0` |
| Downward | `y: 0 -> +230` | `y: -230` | `y: 0` |
### Velocity matching via mirrored eases
The cleanest match: exit `power4.in` and entry `power4.out` with the SAME distance and
duration — mathematically the two halves of one `power4.inOut` composite, so the entering
element picks up at exactly the 50% point of the notional path at identical velocity
(e.g. 230px / 0.3s ≈ 3,070 px/s at the cut on both sides).
The fade trick: the exit's opacity completes at ~25–30% of its travel (fade duration
≈ 0.18–0.3s vs motion 0.3–0.34s) — the element vanishes while still visibly accelerating,
and nothing has to reach the frame edge. Entry fades IN fast from ~0.35 under its
deceleration. Time the LAST fading element to die right at the hard cut — gaps where
nothing is moving read as awkward dead air.
### Rules
- Use cut-the-curve for all scene transitions — it's the default, not an accent
- Same direction on both sides; mirrored `.in`/`.out` eases, same distance + duration
- Exit duration short (0.2–0.4s), entry duration >= exit duration
- Partial travel + fade, never full off-screen moves
---
## 4. Waterfall Cut (word-by-word cut-the-curve)
Cut-the-curve at WORD granularity — the strongest version of the leftward cut for
text-to-text seams. Each word of the outgoing line ramps out on its own pronounced curve;
each word of the incoming line cascades in mid-flight. The stagger turns the cut into a
wave the eye rides across the seam.
### Exit (per word)
- Motion: `x: 0 -> -230` over 0.34s on **power4.in** — a much more pronounced ramp than
the usual power2: the word barely creeps, then RIPS
- Fade: `opacity -> 0` over 0.18s (separate tween, `power1.in`) — completes when the word
is only ~25–30% into its travel
- Stagger: reading order, ~0.022s per word, timed so the LAST word finishes fading right
at the hard cut
### Entry (per word)
- `fromTo x: +230 -> 0, opacity: 0.35 -> 1` over 0.3s on **power4.out** — the mirrored
back half of the composite; every word ignites already moving at matched velocity
- Waterfall stagger with SHRINKING gaps (start 0.05s, multiply by ~0.84 per word) so the
cascade accelerates across the line — the cascade should speed up word over word, not run
at a flat per-word delay
- Pre-set all words to `x: +230, opacity: 0` at build time — `immediateRender: false`
alone leaves un-started words sitting visible at rest during the stagger window
### Whole-line variant
A single-line beat (e.g. a big intro line) exits as one group with the same pronounced
ramp, but stretch its fade to ~0.3s ending ~0.02s before the cut — a lone element that
fades early leaves dead air that a word cascade would have covered.
Shipped example: `until-now.html` B1→B2→B3 (sfx-music-launch).
---
## Choosing a Variant
| | Zoom-Through | Inverse Zoom | Cut the Curve | Waterfall Cut |
| -------------- | --------------------------- | --------------------------- | ----------------- | ------------------------- |
| Scope | Within-scene text swap | Arrival/payoff beat | Between scenes | Text-to-text seam |
| Axis | Z, toward viewer | Z, away from viewer | X / Y | X, per-word |
| Peak blur | 10px text / 20px full-frame | 10px text / 20px full-frame | none required | none (fade does the work) |
| Opacity at cut | 0.15 | 0.15 | exit faded by cut | last word dies at cut |
| Feel | progressing through | arriving at | carried sideways | a wave across the seam |
---
## Anti-Patterns
| Don't | Why | Instead |
| ---------------------------------------- | ------------------------------------------- | ------------------------------------------------------ |
| Two texts visible during a zoom-through | Overlapping text breaks the Z-axis illusion | Hard cut at blur peak, one text at a time |
| 20px blur on text-scale subjects | Letterforms smear; reads as a glitch | 10px for text, 18–20px only full-frame |
| Elements on different paths across a cut | Eye tracks one direction, cut goes another | Same property, same direction |
| Mismatched blur/opacity at the swap | Visible flash or brightness jump | Identical values at the cut frame |
| Gentle easing on entry (`power2.out`) | Entry velocity feels slower than exit | Mirror the exit: `power4.out` / `expo.out` |
| Full off-screen exits / entries | Wastes time and breaks the speed illusion | Partial travel + early fade |
| Lone element fading long before its cut | Dead air at the seam | Fade ends ~0.02s before the cut, or use a word cascade |
| Zoom-through on body text | Small text at 0.75 scale is unreadable | Only headlines and short phrases |
| Scene cuts without cut-the-curve | Static cuts feel like a slideshow | Cut-the-curve is the default |
references/motion-language.md
# Motion language — the move vocabulary + the motion doctrine + the seek-safe core
> The motion layer for **Step 4 (Visual design)**. When you write a frame's **time-coded shot sequence**, you name each scene's move **inline from the vocabulary below** — a named palette of the moves the golden corpus actually uses. Each move carries the **backing rule id** in this skill's local `../hyperframes-animation/rules/`; cite that id so the move resolves to a real recipe when a **frame worker** implements it in Step 5 (the worker reads the rule body in `../hyperframes-animation/rules/<id>.md` — it reproduces the move, it does not guess from the name). You name motion by **role / move name**, never by raw GSAP curve, ms, or stagger formula — the worker maps the curve. Between-frame **transitions are not yours**: story names `transition_in`, the harness injects it; that injected transition **is** the frame's exit. For cuts a worker builds INSIDE a frame (within-scene swaps, scene-to-scene seams), see the catalog in `cut-catalog.md`.
A good promo feels like one continuous film — one camera, one motion feel, **smooth and timed to the voiceover** — not a pile of slides that animate once and freeze. The doctrine in Part 2 is load-bearing: when in doubt, do what it says.
---
# Part 1 — the move vocabulary
Reach into this palette when naming a scene's motion. Pick the move that matches the beat, name it in the shot sequence, and cite the rule id after `→`. The blueprints (`../hyperframes-animation/blueprints/`) name these same moves in their `rule mapping`; you're drawing from one shared palette. Compose 2–4 across a shot's scenes (entrance → sequential reveal → settle), not all at once.
## Kinetic type
- **hard-cut / flash word-swap** — a word or line replaces the previous one on an instant cut (no fade/roll); the swap itself is the beat. → `discrete-text-sequence`
- **in-place token cycle** — a fixed line holds and only its variable slot changes, token → token → token. → `discrete-text-sequence`
- **per-word staggered reveal** — a phrase assembles word-by-word (or chunk-by-chunk), each landing on its own beat. → `dynamic-content-sequencing`
- **kinetic beat-slam** — short phrases slam in on a shared percussive beat array, each with a distinct entrance, resolving on a locked finale; the recipe for "punchy / rhythmic" taglines. → `kinetic-beat-slam`
## Typewriter
- **type-on with caret** — text types in character-by-character behind a blinking caret. → `discrete-text-sequence` (+ `context-sensitive-cursor` for the caret blink / color)
- **backspace-and-retype** — the line types, deletes the last word(s), and retypes a new one (typo-correction, reframe). → `discrete-text-sequence` (+ `context-sensitive-cursor`)
## Count-up / data
- **value-scaled counter** — a number counts up and its font size grows with the value, so the climb itself escalates. → `counting-dynamic-scale`
- **bars / progress / star wipe** — a number paired with a graphic that fills: bar-height stagger, a progress bar / ring filling, a fractional star-rating wipe. → `stat-bars-and-fills`
## Reveal / decode
- **3D char flip-decode** — characters flip in 3D and resolve from scrambled glyphs to the real text (decryption feel). → `hacker-flip-3d`
- **SVG self-draw** — an outline / icon / ring draws itself stroke-by-stroke. → `svg-path-draw`
## Camera
- **push / focus / drift** — a sequential camera move on the frame root (pull-back → focus → push) plus continuous micro-drift; the cinematic baseline. → `multi-phase-camera`
- **zoom-to-target** — zoom into a non-centered element (scale + counter-translate to keep it framed). → `coordinate-target-zoom`
- **pan / focus-lock** — a virtual camera transforming one `.world` wrapper to pan / zoom / lock onto a region. → `viewport-change`
- **camera-cursor-tracking** — the viewport locks to a moving focal point (a typing cursor), static framing then focal-locked tracking. → `camera-cursor-tracking`
## Layout motion
- **cluster→outward expansion** — elements start clustered at center and expand outward to their final positions in lockstep. → `center-outward-expansion`
- **orbit** — elements flip in from 3D space and settle into a continuous elliptical orbit (entry flips in-place at the orbital position). → `orbit-3d-entry`
- **split-tilt cards** — two cards side-by-side with opposing rotationY tilts, entering from their respective sides (comparison / before-after). → `split-tilt-cards`
- **logo/avatar ring + connectors** — avatars or logos on an elliptical ring with SVG connection lines to a center point, staggered entry. → `avatar-cloud-network`
## Surface / UI
- **3D page-scroll reveal** — a full webpage as a tilted 3D card whose internal content scrolls to reveal specific sections. → `3d-page-scroll`
- **cursor click + ripple** — a cursor moves to a target, depresses with it on click, and emits an expanding ripple. → `cursor-click-ripple`
- **button press** — a tactile press: compression then spring recovery, optional release burst / glow. → `press-release-spring` (or `physics-press-reaction` for a click that compresses cursor + target together)
- **keyword glow** — keywords light up with glow + scale + color on an attack-decay-rest envelope, synced to a word rail. → `asr-keyword-glow`
## Morph / handoff
- **scale-swap** — two elements at the same screen center hand off: the outgoing cluster shrinks + fades as the incoming one arrives. → `scale-swap-transition`
- **card morph-anchor** — a container morphs apparent size + corner radius + surface between two shots, then fades to reveal the real target beneath (HyperFrames uses uniform `scale`, not `width`/`height`). → `card-morph-anchor`
## Seam cuts (worker-built, inside a frame)
The velocity-matched cuts a worker authors between a frame's own Scenes. Name the seam in the shot sequence; the recipe is in the catalog, not a single `../hyperframes-animation/rules/` id.
- **zoom-through / inverse zoom-through** — a within-scene swap on the Z-axis; forward reads "progressing through", inverse reads "arriving at" (payoff). → `cut-catalog.md`
- **cut-the-curve** — a scene-to-scene cut where both sides move the same direction at matched velocity. → `cut-catalog.md`
- **waterfall cut** — cut-the-curve at word granularity, a wave across a text-to-text seam. → `cut-catalog.md`
## Emphasis / marker
- **highlight / circle / burst / scribble** — a marker-drawn emphasis on a word or element: yellow highlight sweep, hand-drawn circle, radiating burst, scribble, or rough sketch-outline. → `css-marker-patterns`
## Aliveness during a hold (use sparingly — see Part 2)
- **subtle jitter** — the sanctioned way to keep a settled frame alive: a small, low-amplitude positional/scale jitter on the held element. The motion-graphics trick that reads "alive" without reading "weak." → `sine-wave-loop` (low-amplitude register)
- **live SVG internals** — internal SVG parts move so an icon feels alive (rotating hands, oscillating blades, pulsing dots, dash-flow); fine because it's the subject doing something, not a card breathing. → `svg-icon-enrichment`
- **finite bounded ambient** — a single bounded breathe/drift on ONE held hero, only when genuinely needed; de-emphasized — prefer sequential reveal or jitter first. → `sine-wave-loop`
## The added moves — now backed by local rules
Five moves the golden corpus needs were added to this skill's `../hyperframes-animation/rules/`, rounding out the vocabulary above:
- **depth-of-field / selective-blur** — blur the off-focus subset to spotlight the focal element → `depth-of-field-blur`
- **motion-blur streak** — directional velocity blur on a fast fly-in / camera push-through → `motion-blur-streak`
- **3D depth scatter-assemble** — glyphs/elements scatter into a tumbling 3D cloud, then reassemble → `depth-scatter-assemble`
- **spring-pop entrance** — the canonical entrance pop; default to a smooth long-tail settle, overshoot only when explicitly playful → `spring-pop-entrance`
- **ambient glow / bloom** — un-triggered soft glow blooming behind a static hero → `ambient-glow-bloom`
---
# Part 2 — the motion doctrine (load-bearing)
These four rules are the difference between a clip that reads as a serious launch video and one that reads as an agent-made PowerPoint. Follow them as written.
## 1. Smooth beats bouncy — `power3` is the default
Elements should use **long-tail decel curves that let them settle smoothly. `power3` is enough in most cases.** No bouncy, no overshoot, no `back.out` / `bounce.out` / `elastic.out` as a default.
Bouncy is the **#1 instant turn-off** in user-made Remotion / HyperFrames videos, and the agent almost never gets it right — it thinks bouncy adds emphasis, but it buys that emphasis at the cost of cleanliness. The serious launch-video shops feel the same. **Smooth always wins.** Overshoot is demoted to a **rare, explicitly-playful exception** (a consumer/fun logo slam, a deliberate bell-hit) — never the house style. Name the intent as a long-tail settle; the worker maps `power3` (or `expo.out` on a fast arrival). See `../hyperframes-animation/rules/spring-pop-entrance.md` — it now leads with the smooth settle. (The exact form of that settle is a critically-damped spring; the worker has a baked, seek-safe `springEase` — ζ=1 — in `../hyperframes-animation/adapters/gsap-easing-and-stagger.md` → Spring Eases for when the settle is the hero. Real physics, same doctrine — not a license for bounce.)
## 2. Sequential reveal in the back ~50%, timed to the voiceover
This is the anti-PowerPoint mechanism — sharper than "put development in the middle."
- **Don't dump everything on screen in the first ~25%** of the scene. Rushing all content in up front is exactly what forces the slideshow feel.
- **Reveal each piece — a line, a card, even an h1 — when the voiceover mentions it**, sequencing reveals across the **later ~50%** of the scene. Same amount of agent work, but the cut becomes coherent and gains rhythm.
- **Less is more.** Fewer things on screen, each arriving on its VO beat, beats a full canvas that animated once and froze.
Practically: a frame's shot sequence front-loads almost nothing — the entrance carries only what the VO is saying at t=0, and the rest of the elements wait in the timeline for their spoken cue. A reveal maps onto a development-class move from Part 1 (`per-word staggered reveal`, `cluster→outward expansion`, a `count-up`, an `asr-keyword-glow` synced to the word rail).
## 3. No lazy breathing, no bad pan/push — "no motion over bad motion"
The agent's two reflexive ways to fake "aliveness" both read cheap:
- **No lazy breathing.** Scaling cards/text up and down in a circular loop to look "alive" is the cheap tell. Don't reach for it.
- **No bad slow pan / push in the back half.** A slow pan or push on elements in the later ~50% of a scene **disrupts the viewer's sightline and causes eye discomfort** — it actively makes the frame worse, not better.
The fix for both is the same: **stagger element reveals in time with the script** (rule 2). And the governing principle: **"I'd rather have NO motion than BAD motion."** A held, still frame is better than a frame kept "alive" by breathing or a drifting camera. The **only sanctioned aliveness** during a hold is **subtle jitter** — a small low-amplitude jitter that keeps a frame from feeling dead without looking weak (it's in current video work now). Everything else holds.
## 4. Internal seams are velocity-matched cuts
When a frame has an internal seam — a within-scene swap, a Scene-to-Scene cut, a text-to-text line change — make it a **velocity-matched cut**, not a hard slideshow cut: cut at peak velocity, match direction and speed on both sides. The catalog (the four techniques, the blur logic, and which to use when) is `cut-catalog.md`; the moves are listed under **Seam cuts** in Part 1.
## One-line summary
Smooth long-tail (`power3`) over bouncy; reveal sequentially in the back ~50% timed to the VO (not dumped in the first 25%); no lazy breathing and no bad slow pan/push — prefer stillness, with subtle jitter as the only aliveness; cut at peak velocity with matched direction/speed (→ `cut-catalog.md`).
---
# Part 3 — the seek-safe core (hard rules)
The frame is a **paused GSAP timeline seeked frame-by-frame**, so some "continuous" intents from a real-time engine can't render — don't name them. These are non-negotiable regardless of doctrine.
- **No infinite / forever motion** — "particles loop endlessly," "logo rotates forever," "marquee on repeat." Any aliveness (the subtle jitter, a live SVG internal, a needed bounded ambient) is a **finite tween over the hold**, never `repeat` / `yoyo`.
- **No randomness or wall-clock** — no `Math.random` particle fields, no `Date.now` drift. Every render must be identical; name deterministic motion only (stagger and any variation derive from the element index).
- **Entrances use `fromTo`** — state the from-state explicitly so a seek to `t=0` lands the element correctly; never rely on a CSS-hidden start (it renders visible before the tween claims it, and flickers under seek).
- **No CSS `transition` / `@keyframes` for motion** — CSS animation runs on the browser clock, independent of the HF seek clock; it desyncs and flickers. Drive all motion inside the paused GSAP timeline.
- **Entrance + sequential reveal only — no mid-video exit.** The frame unmounts via the harness transition; that injected `transition_in` **is** the exit. Exit motion belongs only to the final frame. (Worker-built seam cuts in `cut-catalog.md` are within-frame, not the frame's exit.)
## Forbidden — the failure modes
**Slideshow (the primary failure):** everything dumped on screen in the first ~25%; content enters then freezes; nothing revealed on its VO cue. Fix with rule 2 (sequential reveal timed to the VO).
**Cheap aliveness:** circular breathing as "life"; a slow pan/push in the back half disrupting the eye; many elements floating independently as "motion." Fix with rule 3 (stillness + subtle jitter only).
**Bouncy:** `back.out` / `bounce.out` / `elastic.out` as the default entrance; hand-keyed overshoot. Fix with rule 1 (`power3` long-tail; overshoot only when explicitly playful).
**Always:** no `repeat` / `yoyo`; no `Math.random` / `Date.now`; no all-elements-entering-simultaneously (sequence or stagger).
## Naming motion in a shot — example
> Scene 1 (0.0–1.0s): solid field; hero headline enters via **per-word staggered reveal** (`dynamic-content-sequencing`) on a smooth long-tail settle (`power3`); slow **push** on the root (`multi-phase-camera`) holds steady — no back-half re-push.
> Scene 2 (1.0–3.0s): as the VO names each capability, five feature icons reveal **sequentially** via **cluster→outward expansion** (`center-outward-expansion`), then a **value-scaled counter** (`counting-dynamic-scale`) ticks up beneath them — the back-half reveal, timed to the script, not dumped at t=0.
> Scene 3 (3.0–4.2s): hold on the result; **keyword glow** (`asr-keyword-glow`) lands on the payoff word as the VO says it; settles and holds still — at most **subtle jitter** (`sine-wave-loop`, low amplitude) keeps it alive; no breathing, no drift.
Name the move + its rule id (or `cut-catalog.md` for a seam cut) per scene; let the worker pick curves, ms, and stagger — defaulting to `power3`.
references/story-design.md
# Story design — product launch video
Step 3 of the product-launch flow. Output: `STORYBOARD.md` (the narrative plan, one frame per beat) and `SCRIPT.md` (the locked spoken narration).
This step decides **what the video says, in what order, and how each beat is said** — and it says each beat in the SHAPE of a proven script. It does not design layout, composition, or motion (that is Step 4). For exact file syntax follow `../hyperframes-core/references/storyboard-format.md` and `../hyperframes-core/references/script-format.md`.
## What story design produces
For every beat, four things:
1. **Position in the SEQUENCE** — the shot order. Story truth decides which beats exist and in what order (the arc).
2. **Voiceover written in a blueprint's script shape** — the spoken line, drafted to sound like the proven script for the shape this beat is reaching toward (see the script bank below).
3. **A candidate `blueprint:` id** — the proven shot SHAPE this beat leans toward (a tag, not a commitment; Step 4 confirms or overrides).
4. **`transition_in`** — how this beat enters from the one before it.
The big idea: **the blueprint is applied from the very first step.** The blueprints were reverse-engineered from 50 golden clips; each one implies a proven script. So we write the VO in that script's shape from the start — the voiceover is blueprint-shaped before Step 4 ever runs, which makes the blueprint's hit-rate downstream high.
This is still a SOFT discipline. Story truth comes first: **never invent, bend, drop, or reorder a beat to fit a blueprint.** The script patterns only shape HOW a beat is said and which proven shape it leans toward — they never decide which beats exist.
## Read first
1. `hyperframes.json` — locked brief: angle, length, aspect ratio, language.
2. `frame.md` — tone, mood, design system, brand register.
3. `capture/extracted/visible-text.txt` — product facts, page copy, positioning, proof, CTA.
4. `capture/extracted/asset-descriptions.md` — the ONLY source for the captured asset inventory.
5. `user_script.txt` and `VO_MODE`, when present.
Do not inspect `capture/assets/`, contact sheets, screenshots, or raw files in Step 3. Treat `asset-descriptions.md` as the canonical asset list. Never invent asset filenames.
## Method
### 1. Extract the product truth
From the brief and captured text, name:
- **Audience** — who the video speaks to.
- **Pain / desire** — what they already want fixed or achieved.
- **Promise** — the one-line thesis of the whole video.
- **Product role** — what the product does in the story.
- **Proof** — features, UI moments, metrics, logos, demos.
- **CTA** — what the viewer should do next.
Build the sequence around the **promise**, not a feature list. A website is an information layout; a video is an emotional sequence. Reorder, merge, and omit captured content freely — do not follow page order.
### 2. Choose the arc (the sequence backbone)
Pick ONE arc — it fixes the beat order. Compound only when useful (e.g. `PAS with feature-benefit progression`).
| Arc | Use when | Beat order |
| ------------------------- | ------------------------------------------------ | --------------------------------------------------------------------------- |
| `PAS` | Pain is known and urgent (broken B2B workflows). | hook → pain → agitation → solution tease → product intro → proof/demo → CTA |
| `Future Pacing` | Sells a new future / category / paradigm. | imagine → name product → remove pain → mechanism → outcome → CTA |
| `Demo Loop` | UI is self-explanatory; best shown working. | question → product intro → demo cycle 1 → demo cycle 2 → trust → CTA |
| `BAB` | Bridges an old workflow to a better one. | before → after tease → bridge/product → step 1 → step 2 → wow → CTA |
| `Feature-Benefit Cascade` | Feature-rich or desire/status-driven. | category hook → feature → benefit → feature → benefit → climax → CTA |
Use feature→benefit rhythm inside any arc when there are many capabilities — always translate a feature into viewer value, never stack raw features.
`frame.md` tunes the VOICE, not the arc: restrained/B2B → plain, low-hype; bold/launch → short, punchy; warm/human → friendly direct address; premium/cinematic → aspirational, fewer words.
### 3. Lay out the beats, each with a role
One clear job per beat — never "more benefits" or "another feature." Beat `type` (= blueprint **role**):
`hook | pain_point | product_intro | feature_showcase | benefit_highlight | social_proof | branding | cta`
The opening 3–5s needs ONE hook that creates tension, curiosity, or desire — a shocking stat, pain validation, a rhetorical question, direct address, an imagine/future-pace, a category announcement, or visual spectacle. Never open with generic company description. Per `../hyperframes-creative/references/story-spine.md`: the hook speaks the viewer's outcome language (what they gain, never a feature list), and the promise (`message`) lands by beat 2 — features after that are its evidence.
A UI demo is usually a SEQUENCE of 3+ consecutive `feature_showcase` / `benefit_highlight` beats on the same surface (input → response → result → benefit), not one isolated frame.
### 4. Write each beat's VO in its blueprint's script shape
For each beat, look up its **role** in the script bank below, find the blueprint whose SHAPE fits the beat you already chose, and **draft the voiceover to sound like that blueprint's pattern.** Tag the candidate `blueprint:` id on the frame.
- The bank is the heart of this step: proven product-launch clips reversed into the one VO line each implies, grouped by role → blueprint, each with a **pattern** to imitate.
- If two blueprints fit the beat, prefer the one whose script shape matches the line you'd naturally write.
- If NO shape fits the beat, omit `blueprint:` and write the VO plainly — Step 4 composes that frame freely. Do not force a wrong shape.
- **Vary the shapes across the video.** Reaching for the same blueprint every beat re-creates the sameness this exists to avoid. `kinetic-type-beats` is the workhorse (6 roles) — lean on it, but not everywhere.
- **Write each VO as discrete cues, not one run-on breath.** Step 5 reveals each on-screen piece _when the voiceover names it_ (the anti-PowerPoint mechanism — `motion-language.md` Part 2). A line with clear phrase boundaries — "Content, sentiment, engagement — in one place" — hands the shot its reveal cadence for free; a single long clause leaves the frame nothing to pace to. The bank patterns are already cue-segmented — keep that rhythm.
Step 3 only TAGS the candidate id and writes the shaped VO. Step 4 (visual design) picks and instantiates the blueprint into a time-coded shot; it may override or drop a Step 3 candidate. The full menu with picking guidance lives in `../hyperframes-animation/blueprints-index.md`.
---
## The script bank — what each beat's VO sounds like
> Proven product-launch clips, each reversed into the one spoken line it implies. Grouped by **role → blueprint**. Real product names kept (swap in your own). Draft your beat's VO in the SHAPE of the matching pattern. Kept 1:1 with the role declarations in `../hyperframes-animation/blueprints-index.md` — when a blueprint gains a role there, add its script shape here.
### HOOK
**kinetic-type-beats** — the words ARE the motion
- Mailoji — "Still using a @gmail address? Or @outlook, or @hotmail, or @yahoo?"
- Outrank — "Getting traffic is hard. Insanely hard."
- AiAgent — "An AI agent that's easy to use — and optimised for you."
- Uizard — "Transform your sketches into prototypes — automatically."
- _Pattern:_ a punchy claim or rhetorical jab whose KEY WORD swaps in place (or escalates beat by beat) — the swap/escalation is the joke.
**typewriter-reveal** — someone is typing this
- "Need answers about your audience — right now?"
- Contra — "You are more than your job title. You are more than your resume."
- _Pattern:_ a relatable line typed live and edited mid-stream (a word backspaces and retypes) — the everyday thought, in your own words.
**spatial-pan-stations** — a panned timeline
- Rows — "From VisiCalc to Excel to Google Sheets — the spreadsheet has barely changed since 1979."
- _Pattern:_ a march of named milestones across time, landing on "...until now / ...up to us."
**constellation-hub** — nodes ring a center, camera pushes in
- "Content, sentiment, engagement, analysis — every platform you're on, in one place."
- _Pattern:_ a spread of tools/channels collapsing onto one center — "it connects everything."
**ticker-takeover** — options cycle, then a hero crashes in
- Notion — "A doc? A database? A wiki? — no, it's all of them, in one place."
- _Pattern:_ a "could be X, or Y, or Z?" cycle on one swapping word, then a hero claim crashes in and replaces it — "actually, this is what it is."
**prompt-type-submit-generate** — watch me ask
- "Build me a landing page for my coffee brand — dark, minimal, launch-ready."
- "What if you could just… ask?"
- _Pattern:_ the VO speaks (or frames) the ask itself while the prompt types live — the question is the whole hook; the answer stays off-screen, or a second ask starts before the cut.
**zoom-out-workspace-reveal** — the mystery re-scopes
- "This isn't a finished animation — it's your canvas."
- _Pattern:_ near-silence or one slow tease over the close-up mystery, with the landing line timed to the zoom-out — the words answer "what am I looking at?" exactly when the workspace does.
**fixed-anchor-cycle** — a roll-call around a pinned line
- "For founders, for designers, for marketers, for teams — for anyone who ships."
- _Pattern:_ one static claim holds while its audience/option list cycles fast beneath it — the list is the sentence's swapping object, and the brand line lands after the cycle clears.
**cursor-ui-demo** — a canvas already alive
- "Your whole team is already in here — designing, commenting, shipping."
- _Pattern:_ a "we're all here working" line over an ambient multi-cursor canvas — presence, not features; the live activity is the proof.
**dataviz-countup** — cold-open on one exploding stat
- "One million users. Ninety days."
- _Pattern:_ ONE statistic counts up as the VO speaks it — scale alone carries the tension; no product, no context yet.
**cta-morph-press** — a lone widget does its trick
- "It starts as a search bar. It becomes your whole workflow."
- _Pattern:_ one small widget on a bare field morphs in place and performs its payload while the VO names the transformation — small thing, big claim, hand off to the title.
### PROBLEM
**kinetic-type-beats** — pain lands alone on a bare canvas
- Butter — "What if your sessions didn't have to be boring and unstructured — or buried under a dozen tabs?"
- SmartCue — "You asked for better leads. We were the cure — MQLs that actually convert, a sales team that becomes your ally."
- _Pattern:_ 3–5 short pain statements (or a "what if?" framing), each landing solo before the next — no product yet.
**spatial-pan-stations** — a panned web of pain
- Vauban — "Coordinating legal documents, signatures, and cross-border transactions — it's a tangled mess."
- _Pattern:_ pain "stations" traversed one by one, ending on a knot — "too many disconnected steps."
**dataviz-countup** — the data IS the argument
- "67% of professionals say leadership is disconnected — and it's costing them a 65% boost in profitability."
- _Pattern:_ a count-up / chart / stat the camera pushes through to dramatize a worsening or large-scale problem.
**overwhelm-surround** — buried by your own tools
- "Slack, email, docs, tickets, three more tabs — and somehow it all lands on you."
- _Pattern:_ recognizable tools pile in until they surround and bury the viewer — the pain is being swamped, not one bad number.
### PRODUCT_INTRO
**kinetic-type-beats** — "Introducing…" name-drop
- "Elevating experiences, removing manual touchpoints, automating processes — so you can focus on the customer."
- Uizard — "Introducing Uizard — the design tool for everyone."
- _Pattern:_ hard-cut through "Introducing…" / tagline / value beats and resolve on the brand name or logo.
**logo-assemble-lockup** — wordless premium sting
- Manifold — "Manifold." _(wordless mark assembles; VO optional — just the name)_
- _Pattern:_ an abstract system assembles around a fixed mark — no copy, or just the product name landing.
**cursor-ui-demo** — first cursor-led look
- ClickUp — "This is ClickUp — click through and watch your whole workspace change."
- "Pull up any contact, find the right advisor, and you're matched in seconds."
- PaLM 2 — "Meet PaLM 2 — what is it, what can it do, and how was it built?"
- _Pattern:_ a cursor sweeps in to introduce the surface — a light first look landing on a hovered hero element or fresh result.
**dataviz-countup** — open on the result
- SuperX — "X growth — discover what really works: 19.6 million impressions."
- _Pattern:_ a confident "look at the data / the result" open — scroll a tilted card grid to one glowing hero metric, tagline assembling word by word.
**video-text-pivot** — show it work, then the number
- "Watch it run — then look at what it saved: 14 hours, every week."
- _Pattern:_ the product video plays, then slides aside to hand the frame to one impact stat — "see it work, now see what it's worth."
**prompt-type-submit-generate** — something you talk to
- "Meet Ada — describe what you need, attach what you have, and let it run."
- _Pattern:_ the first look IS the composer — the VO introduces the product as a conversation partner while a long prompt types with attachments and option picks, ending on (or just after) the submit.
**spatial-pan-stations** — decode the idea, then land on it live
- "Capture, structure, publish — that's the idea. Here it is running."
- _Pattern:_ a labeled concept strip read station by station, the final pan bridging into the live demo — "here's the idea → here it is working."
**device-surface-showcase** — introduced by completing its core loop
- "Open a space, drop in your notes, hit share — that's it, that's the product."
- _Pattern:_ the product introduced by DOING its core loop once inside its real interface, stepwise and cursorless — the VO narrates the steps plainly and lands on "that's it."
**titlecard-reveal** — a near-still title prelude
- "Relay. Docs that write themselves. Let's look inside."
- _Pattern:_ 2–3 near-still cards seamed by blur-snap handoffs — name, one-line claim, then into the product; each card carries one short spoken phrase.
### KEY_FEATURE
**grid-card-assemble** — enumerate breadth at once
- Postcards — "Want more? Unlimited exports, 1,400 fonts, an AI assistant, version history — and it's free to try."
- ClearVPN — "ClearVPN is built to help you: streaming access, secure browsing, location changing — only the essentials."
- Copilot — "Command bar, Zapier, white-labeling, API, SOC2 — and a whole lot more."
- _Pattern:_ a tile/pill/card grid self-assembles to show many capabilities at once — "look how much it does."
**cursor-ui-demo** — one workflow, end to end
- "Scroll your feed, then jump straight to your notifications — it's all one click away."
- Flowrite — "Pick your recipient, set your intention, choose a tone — and Flowrite writes it for you."
- Descript — "Tune your edits, add a crossfade, automate the volume, normalize loudness — then export."
- _Pattern:_ one specific multi-step workflow shown end-to-end across 2–4 real edits, landing on the action button or result.
**device-surface-showcase** — experienced inside its real interface
- HRS — "Your flight's cancelled — so book a hotel, a taxi, and get reimbursed, all in one digital journey."
- HelpKit — "A dynamic table of contents that follows your readers as they scroll."
- Graphite — "Use unique themes, recolor one element or several, and configure it all in a handy window."
- Contra — "Pick a template, make it yours, and launch a portfolio that's unmistakably you."
- _Pattern:_ a feature shown being USED inside its real surface — the device/window is the hero and its screens advance through a flow.
**comparison-split** — two paired capabilities, side by side
- "Design on the left, code on the right — always one source of truth."
- _Pattern:_ two complementary capabilities of equal weight shown together — "X and Y, in lockstep."
**video-text-pivot** — the feature, then its result
- "Here's the editor in action — and the result: a publish-ready cut in minutes."
- _Pattern:_ a feature clip runs, then yields the frame to a metric / impact line — the clip proves it, the number lands it.
**prompt-type-submit-generate** — one ask, one answer
- "Ask for revenue by region — and the chart draws itself."
- _Pattern:_ ONE prompt→response round trip — the VO names the ask, lets the status theater breathe a beat, then calls the result as it streams in.
**agent-progress-theater** — the machine visibly works
- "Kick off the scan — it checks every file, flags what's broken, and fixes what it can. Watch."
- _Pattern:_ trigger → working theater → receipt: the VO hands the work over, then reads the findings as rows land and check off — present tense, the machine is the subject.
**panel-edit-live-sync** — one gesture, two surfaces
- "Drag the value — the button follows. Pick a unit — the code converts. Live."
- _Pattern:_ 2–4 short cause→effect couplets, each pairing a gesture with its live mirror ("do X — Y answers"), ending held on the last edit.
**transcript-scroll-artifact-reveal** — the work, then the deliverable
- "It planned, researched, built, and tested — and left you the spreadsheet to prove it."
- _Pattern:_ a "look how much happened" line read down the transcript at traversal pace, then one pivot line cashes it in on the artifact — evidence first, payoff last.
**camera-journey** — a cinematic flight over the result
- "A month of content — planned, scheduled, and ready before you sat down."
- _Pattern:_ a flying camera explores the generated artifact while the VO makes one calm, sweeping claim; the content acts by itself — no hands, no clicks.
**dataviz-countup** — the numbers prove the feature
- "Response time, down 40%. Coverage, 3×. Every claim, a number."
- _Pattern:_ the feature proven by its metrics — a stat montage scrubbed / counted up while the VO reads each number as it lands.
### BENEFITS
**kinetic-type-beats** — rapid-fire value montage
- AiAgent — "No API keys, GPT-4 access, simple UX, clean UI — moving fast."
- Uizard — "Export to Sketch, create style guides, share, collaborate — and code less."
- _Pattern:_ a staccato montage of 8–12 short value phrases, each flashing and clearing at high tempo.
**grid-card-assemble** — an accumulating value list
- Lineicons — "Consistent and clean, tons of free icons, a Figma plugin, a powerful editor, every format you need."
- Plutio — "Manage projects, track time, send invoices, write proposals — all deeply customisable."
- _Pattern:_ short value phrases populate a vertical list ~1/sec, co-resident and accumulating, each popping into its slot.
**titlecard-reveal** — the calm value beat
- CSS Scan Pro — "A smart color picker — with instant tints and shades."
- _Pattern:_ one clean two-line value title, one slide-up crossfade, then held still. Low motion is the point.
**camera-journey** — travel the value chain
- "Leave one comment here — and the forecast updates over there."
- _Pattern:_ a cause→effect round trip: the VO's first half lands on the action, its second half on the payoff the camera travels to — "do this small thing here, get this big thing there."
**zoom-out-workspace-reveal** — this is just one corner
- "That one file it fixed? A corner of everything it already did."
- _Pattern:_ the VO tells the micro-story during the close-up dwell, then the scale line lands with the zoom-out — the benefit is breadth, revealed in one move.
**fixed-anchor-cycle** — everything changes, this stays
- "Same prompt. Every tool you own."
- _Pattern:_ one pinned claim while the entire surface re-skins around it — a short "works everywhere" line, then let the cycling themes speak.
**cursor-ui-demo** — demo, value line, demo
- "Watch it draft the reply — that's an hour back — now watch it file the ticket."
- _Pattern:_ a demo|text|demo sandwich — the VO alternates showing and telling, naming the value in the text beat between two live interactions.
### SOCIAL_PROOF
**constellation-hub** — the hub at the center of your stack
- kyvos — "On any BI tool — Tableau, Looker, Power BI — Kyvos sits at the center of your stack."
- "Connects to thousands of apps — including every one you already use."
- _Pattern:_ the product mark is the hub and partner logos orbit it — "sits at the center of everything you use." In the scatter-drift end-card variant there is no hub and no ring: a serif claim holds center while app icons pop in scattered and drift slowly outward — breadth said with count and spread, not geometry.
**grid-card-assemble** — a logo wall pulling back to a vast ecosystem
- Lineicons — "Used by more than 100,000 designers, developers, and companies — including these."
- ClickUp — "Connect Google Drive, Slack, GitHub, Stripe — your whole ecosystem in one place."
- Copilot — "With thousands of partner apps — Airtable, Calendly, Jira, Asana — you can embed anything."
- _Pattern:_ a wall of partner/app logos builds, then a camera zoom-out reveals a vast ecosystem.
**titlecard-reveal** — busy → clean proof card
- Trumpet — "We supply the trumpet, you bring the band — loved by 1,000+ sales, success, and marketing teams."
- _Pattern:_ wipe a busy open away to a clean lockup plus a "loved by N+ teams" line that settles and holds.
**dataviz-countup** — the numbers vouch
- "Twelve thousand teams. 4.9 stars. 99.98% uptime."
- _Pattern:_ proof by count-up — adoption, rating, and scale metrics tick up as the VO reads them; the biggest number lands last.
### CTA
**kinetic-type-beats** — punchy closing line beat-by-beat
- Stylebit — "Go pro, connect up to five Figma accounts — more features coming. Join now."
- revid.ai — "Boost your engagement and turbocharge your social media."
- _Pattern:_ a closing line (or short value stack) that snaps in beat by beat and lands on the logo or URL.
**logo-assemble-lockup** — logo build → push-through into the URL
- Strapi Cloud — "Deploy on Strapi Cloud — no server hassle, same flexibility. Try it now at strapi.io/cloud."
- Highlander — "Highlander is ready for the future you're building. Let's raise — at highlander.ai."
- STUDIO AI — "Get early access to STUDIO."
- Glorify — "Get started for free now — no credit card required."
- _Pattern:_ the logo builds, a fast camera push-through streaks giant CTA letters past the lens, resolving on the URL or action verb.
**cta-morph-press** — identity condenses into one click
- Linear — "This is Linear. Start building — it's free."
- _Pattern:_ the brand mark condenses straight into the single thing you click — "here's us → click here," no spatial set.
**prompt-type-submit-generate** — the command is the ask
- "One command — npm install relay — and you're live."
- _Pattern:_ the closing invitation IS a typed install command — the headline demotes, the terminal pill types it out, and the VO speaks the command (or one short line over it); the card holds with only the caret blinking.
**constellation-hub** — the orbit collapses into action
- "Everything you use, one hub — click, and it's yours."
- _Pattern:_ the ecosystem ring collapses onto the core on one click that springs the product open — the VO turns "it connects everything" into the action line.
**titlecard-reveal** — a calm end-card stack
- "Try it free. No credit card. relay.app."
- _Pattern:_ 2–3 near-still cards hard-cut in sequence, one short closing phrase each, terminating on the held logo/URL — the calm is the confidence.
### BRAND_OUTRO
**kinetic-type-beats** — a verb barrage resolving on one word
- Phantom — "Buy, store, stake, swap, send, connect, explore — multichain."
- _Pattern:_ a rapid center-channel barrage of single-word verbs asserting breadth, resolving on the brand's one defining word. No logo lockup needed.
**typewriter-reveal** — persistent mark + typed CTA rail
- Collato — "Next time, just Collato it — sign up for free today."
- _Pattern:_ the mark holds dead-center the whole time while a sub-line types or swaps into the final CTA.
**logo-assemble-lockup** — elements clear, the lockup draws itself in
- Copilot — "Copilot." _(pills disperse off-frame, the mark draws on; VO optional — just the name)_
- Dora AI — "Dora AI — join the waitlist."
- _Pattern:_ feature/UI elements clear the stage off all four edges, then the logo mark draws itself on and the wordmark completes the lockup.
**ticker-takeover** — use-cases cycle, the brand takes the frame
- "For notes, for tasks, for plans, for teams — [brand] holds it all."
- _Pattern:_ closing use-cases/verbs cycle through one slot, then the brand mark crashes in and owns the final frame.
**fixed-anchor-cycle** — the anchor holds, the words cycle
- bolt.new — "Prompt, run, edit, deploy — enjoy."
- Anthropic — "Opus 4.6 — by Anthropic."
- _Pattern:_ the brand name sits immovable while tagline words or praise quotes cycle beside it — per-word highlight stepping, or an accelerating flurry — landing on the completed lockup.
---
## VO_MODE handling
**No pasted script** — write the VO yourself, in the matching blueprint's script shape:
- 1–2 sentences per spoken beat, usually 6–20 words.
- Concrete and human; active verbs; say what the product does for a person.
- Avoid: "seamless experience," "unlock the power of," "streamline your workflow," long noun-phrase lists, a whole beat that is just "Or…".
- Silent beats are allowed when the visual proves the point — leave them out of `SCRIPT.md`.
**`VO_MODE = restructure`** — treat `user_script.txt` as source material. Rewrite, reorder, merge, or omit to fit the arc and target length. You may still shape each segment toward its beat's blueprint pattern.
**`VO_MODE = verbatim`** — do NOT change the user's words. Segment the script into beat-sized chunks at sentence/paragraph boundaries (split a long sentence only at a natural clause boundary). Final duration follows the provided script. Blueprint shaping does not apply to wording — only to which shape each verbatim chunk is paired with.
## Asset candidates
`asset_candidates` is the Step-3 → Step-4 handoff. Rules:
1. Read only `capture/extracted/asset-descriptions.md` to know what exists.
2. Use only filenames listed there; write as `assets/<basename>`.
3. One line, candidates separated by semicolons, a short description after `—`.
4. Prefer `[video]` assets when motion proves the product better than a still.
5. Use content assets (UI, screenshots, product photos, charts, demos). Skip tiny icons, favicons, badges, decorative chrome, repeated logo variants — unless the beat needs them. Partner / third-party logos come from `/media-use` (`resolve --type logo --entity <brand>`) — never redrawn by hand.
6. Pure-typography beats may use an empty asset list. Do not use nested lists.
Example:
```md
- asset_candidates: assets/dashboard-hero.png — dark analytics dashboard, wide screenshot; assets/demo-loop.mp4 — query-to-result interaction clip
```
## transition_in
Between-frame transition — how each frame ENTERS from the one before it. The harness's injector stamps it onto the two whole-frame clips (opacity / transform / filter on the frame wrappers). Name a **registry type** directly; optionally add a direction and/or a duration (`push-slide LEFT`, `crossfade 0.4s`). `cut` / `none` / empty = a hard cut.
The five registry types:
- **`crossfade`** — a plain opacity dissolve; the neutral choice when two frames sit in the same visual world.
- **`blur-crossfade`** — dissolve through a soft blur + slight scale; use when the two frames' backgrounds differ a lot, so the blur masks the color clash a plain crossfade would expose.
- **`push-slide`** `[LEFT|RIGHT|UP|DOWN]` — outgoing slides off, incoming pushes in from the opposite edge; a lateral "next beat" feel for a run of consecutive cards / feature beats.
- **`zoom-through`** — outgoing scales up + blurs out, incoming scales up from small into focus; for a STATE CHANGE / turning to a new section (hook → context).
- **`squeeze`** — outgoing compresses to a line on one edge as incoming expands from the other; a snappy, mechanical beat change.
Pick a small set and repeat them: default to `crossfade` (or `blur-crossfade` when the backgrounds clash), and reach for `zoom-through` at section boundaries. Frame 1's `transition_in` is a placeholder.
## Music & silence
The storyboard's top YAML block carries a `music:` field — the BGM mood the audio step retrieves against (e.g. `music: confident minimal tech underscore`). Omitting it falls back to `message:` → `arc:` → a neutral default, so BGM plays unless turned off explicitly.
- **`music: none`** — BGM off (narration, if any, still runs).
- **`music: none` + no `SCRIPT.md`** — the canonical **fully-silent marker**: no narration, no BGM, no SFX. `audio.mjs` generates nothing and Step 3.1 is a clean skip. Use exactly this spelling when the user asks for a silent / music-free video.
## Frame template
Use the exact fields required by the core storyboard format. The narrative shape each frame satisfies:
```md
## Frame N — Short name
- scene: one clear visual idea
- voiceover: "spoken line, written in the candidate blueprint's script shape, or empty"
- duration: rough estimate in seconds
- transition_in: crossfade
- status: outline
- src: compositions/frames/NN-short-name.html
- type: hook
- persuasion: Pain validation
- beat: urgency
- blueprint: kinetic-type-beats — candidate shape from the role→blueprint menu; omit when none fits
- asset_candidates: assets/example.png — short asset description
narrativeRole: what this beat does in the viewer journey.
keyMessage: the one idea the viewer should remember.
```
- `persuasion` — a concrete move (Pain agitation, Negative contrast, Friction reduction, Show-don't-tell proof, Feature-to-benefit translation, Statistical proof, Authority by association, Social proof, Risk reversal, Future pacing, Value stacking, Rule of three, Scarcity/urgency, Status seeking…). Never "show benefit." Invent one if none fits and explain it in the prose.
- `beat` — a specific emotion (anxiety, frustration, overwhelm, tension, urgency, skepticism, FOMO → relief, curiosity, clarity, intrigue, aspiration → trust, confidence, control, ease, power, awe, excitement, belonging → triumph, motivation, urgency-to-act, peace of mind, inevitability). Compound allowed (e.g. `relief + control`).
## Final checklist
- The arc is named; the sequence is narrative-driven, not page-order-driven.
- The opening uses one clear hook strategy that creates tension/curiosity/desire.
- Each beat has one job; every beat has `type`, `persuasion`, `beat`.
- Each beat's `voiceover` is written in its candidate blueprint's script shape (from the bank), with the candidate `blueprint:` tagged wherever a shape fits — and omitted where none does.
- Each `voiceover` is phrase-segmented into cues (each cue a piece Step 5 can reveal on) — not one long run-on clause.
- Shapes vary across the video; no single blueprint on every beat.
- Story truth was never bent to fit a blueprint — no beat invented/dropped/reordered for a shape.
- Every visual beat has suitable `asset_candidates` (filenames only from `asset-descriptions.md`), unless intentionally typography-only.
- UI/product demos use a multi-beat sequence when the value depends on workflow.
- `transition_in` is a registry type (`crossfade` / `blur-crossfade` / `push-slide` / `zoom-through` / `squeeze`) — default `crossfade` (`blur-crossfade` on a background clash), `zoom-through` at section boundaries, repeated across the video.
- `SCRIPT.md` contains only locked spoken narration; silent beats are intentional and omitted.
references/visual-design.md
# Visual design — product-launch per-frame shot method
> The method behind **Step 4 (Frame visual design)**. You (the orchestrator) read it to **enrich `STORYBOARD.md` frames in place** — story-design wrote the skeleton (each frame's `scene`, `voiceover`, `transition_in`, the narrative fields, its `asset_candidates`, and optionally a candidate blueprint id); you add how each frame **looks and moves**. The unit you write per frame is a **time-coded shot sequence** — a shot directed across its whole duration, not a static slide. You write **no HTML** (that's the frame workers), you **never read `capture/`** (story already chose the assets), and you do **not** select assets or name transitions (story owns both). `frame.md` is your palette/type truth by role. Layout is a compact vocabulary in this file (the **Layout** section below), stated inline per Scene; motion vocabulary + the motion doctrine + the seek-safe core → `motion-language.md`; the proven shapes → `../hyperframes-animation/blueprints-index.md` + `blueprints/<id>.md`; concrete rules resolve in Step 5 from this skill's local `../hyperframes-animation/rules/`. Adding palette theory or a generic font rule here? Wrong home — `frame.md` + `hyperframes-creative`.
## The unit is a time-coded shot sequence
A frame's visual layer is **a sequence of time windows paced to the voiceover**, not a bag of effect tags. The failure that reads as PowerPoint is **front-loading**: the agent rushes the whole canvas on screen in the first ~25%, and then it just sits (the old representation — a flat set of effect names + a prose note — fired everything at entrance and left the rest empty). A time-coded shot sequence written **against the VO** makes that impossible: each window states what is on screen and what is moving, and **nothing appears before the voiceover reaches it.**
Write each frame as a handful of windows cued by the spoken line:
```
Scene 1 (0.0–Xs): only what the VO is saying at t=0 enters — never the whole canvas
Scene 2 (Xs–Ys): the next piece reveals as the VO names it (a line / card / stat / icon)
… one window per spoken cue — as many or as few as the line calls for
Scene N (…–end): content has resolved; hold the read (stillness; subtle jitter at most)
```
- Each `Scene` line names **what's on screen**, **what moves in this window**, and **where it sits** (layout, inline). Times are real seconds across the frame's `duration`.
- **Pace reveals to the voiceover; never front-load.** This is the core anti-PowerPoint mechanism (→ `motion-language.md` Part 2 Rule 2). At t=0 show only what the VO is saying then; reveal each further piece — a line, a card, even an h1 — **when the VO names it**, spreading reveals across the shot and especially the **back ~50%**. **The window count = the number of spoken cues the line calls for** — a two-beat line is two windows, a five-feature list is five or six. There is **no fixed count and no mandatory "middle" act**; the only sin is dumping everything up front.
- **End on a held read.** Once the content has resolved it holds and reads — **prefer stillness to bad motion**: no forced camera drift, no lazy breathing, no back-half pan/push; at most a subtle jitter keeps it alive (→ `motion-language.md`). On a short shot the final reveal and the hold are the **same window** — the hold is not a separate mandatory act. Only the final frame has a real exit; every other frame's exit is the harness transition (story's `transition_in`).
- A **deliberately held** frame — content already revealed, now reading still — is legitimate and often right (a climax, a breather). The failure is never "too still"; it is **front-loaded-then-frozen** (everything dumped by ~25%, nothing cued to the VO). Place held beats deliberately for rhythm so the video isn't uniformly busy (allocate them in `## Video direction`). Reveal pacing + holds + the idle budget → `motion-language.md`.
## Pick the shape — instantiate a blueprint
Don't invent each shot from scratch. The frame's **role** (its `type` / `beat`) points to a proven shape:
1. **Match the role to a blueprint.** Open `../hyperframes-animation/blueprints-index.md`, find the frame's role in the **role→blueprint menu**, and pick the blueprint whose intent fits this beat (story may already have named a candidate id — confirm or override it). Read that `blueprints/<id>.md`: it is a short, product-agnostic, **time-coded shot template with `[slots]`** and a named **signature move** (the thing that makes the shape itself — the SVG ring, the push-THROUGH, the in-place token swap).
2. **Instantiate its `[slots]` with THIS product's content** — three postures:
- **Reproduce** — the blueprint fits the beat and your content maps onto its slots cleanly. Fill every `[slot]` with this product's word / asset / stat and follow its Scene timing. Write the resulting Scene lines.
- **Adapt** — the _structure_ fits but the content / asset-count / surface doesn't (or you want a fresher surface to avoid templating). State **what you keep / what you change** in one line, then write the adapted Scene lines. You may extend or vary; you may **never** drop the **signature move** (drop it and you picked the wrong blueprint), and you keep the reveals **paced to the VO** — never collapse the shape to a single front-loaded dump.
- **Compose** — no blueprint fits the beat. Build the shot from the **motion vocabulary** in `motion-language.md`: still pace the reveals to the VO across the shot, never fire everything at t=0. Mark it `blueprint: compose`.
3. **Keep the signature move.** Whichever posture, the blueprint's signature move (named in its file) is the spine of the shot — it usually lands on the shot's key reveal. Carry it through.
The blueprint's own Scene lines, motion vocabulary, and `rule mapping` are your raw material; you are choosing a shape and casting this product into it, not copying an engineering spec.
## What you add to each frame
Story-design's `## Frame N` block already carries the narrative + `asset_candidates`. You append the shot. Story's `scene` / `voiceover` / `transition_in` / role fields stay untouched.
```
## Frame 3 — The problem
- scene: a 20-minute timer over a stack of rejected takes ← refine only if it could read sharper
- voiceover: "…" ← story's; leave it
- transition_in: crossfade ← story's; leave it
- type: pain_point ← story's
- persuasion: Pain agitation
- beat: frustration
- blueprint: dataviz-countup (Adapt) ← you add: the id you instantiated (or "compose")
- focal: assets/reject-stat.png ← you add: the hero asset for this beat
- roles: reject-stat = cutout · timer = supporting · backdrop = background (dim ~40%) ← you add: role per candidate
- sfx: impact-soft, riser ← you add: the sound the beat wants (fetched + mounted at root; never yours to embed)
Adapt: keep the count-up-ring signature; one stat not three, and the trend chart becomes the rejected-takes count climbing.
Scene 1 (0.0–1.2s): solid backdrop (dim ~40%); a circular progress ring + bold center number seat dead-center, ring sweeps and number counts 0→20 on one heavy ease — Centered template, ~50% of frame. Slow push-in runs underneath.
Scene 2 (1.2–3.4s): as the VO names the count, the camera pushes THROUGH the ring into the rejected-takes stack lower-left; the stack grows beat-by-beat as a reject counter ticks up beside it (the count-up reveals on its spoken cue, not at t=0). Asymmetric 60/40, 3 depth layers.
Scene 3 (3.4–5.0s): land the hero stat card dead-center, accent glow blooms behind it and holds; the stat reads clean and STILL — no continuing push, no breathing (a held beat beats bad motion). The stillness reads against the prior motion.
```
The lightweight tags:
- **`blueprint:`** — the id you instantiated (with `(Reproduce)` / `(Adapt)`), or `compose`. One id per frame.
- **`focal:`** — which existing candidate is the hero of this beat.
- **`roles:`** — each candidate's role: `cutout` (foreground subject, lay text around it) · `background` (full-bleed, dim 30–50%) · `supporting` (secondary). You **consume** the candidates story chose — never add, swap, or drop one (coverage is story's call; if a frame truly has the wrong candidates, flag it back, don't reach into `capture/`).
- **`sfx:`** — name the sound the beat wants (an impact for a slam, a whoosh for a push, a riser into a reveal). The audio script's `fetch-sfx` pass retrieves it and the assembler mounts it at the root — you only **name** it, never embed an `<audio>` element.
**Layout is stated INLINE in each Scene line** — name the template, density, depth, and hierarchy as part of "where it sits" (`Centered, ~50% of frame`, `asymmetric 60/40, 3 depth layers`), drawing on the **Layout** vocabulary below; never write px / scale / shadow recipes (the worker writes those).
**Motion is named INLINE in each Scene line** — name the move from `motion-language.md`'s vocabulary (`ring sweeps`, `pushes THROUGH`, `count-up`, `glow blooms`) and let it settle on a long-tail curve (`power3` default — smooth beats bouncy; see `motion-language.md`). Never write ease curves / ms / stagger (those resolve in Step 5 from this skill's local `../hyperframes-animation/rules/`).
## Layout — named inline per Scene
State each Scene's layout as part of "where it sits." **If the blueprint already implies a composition** (a ring around a center, stations on a wide canvas, two cards from opposite wings), that wins — describe it directly; the vocabulary below is for **composing freely** or a generic beat, not a menu you must pick from. Never write px / scale / shadow (the worker does). One frame's layout can EVOLVE across its Scenes (Scene 1 centered hero → Scene 2 rearranges to a grid).
- **Framing vocabulary** — centered (hero / climax) · rule-of-thirds · split-screen (comparison) · layered-depth (immersive) · asymmetric 60/40 or 70/30 (editorial) · triptych (three panels) · full-width strip. Vary the framing across the video so it doesn't read as one repeated template — let the beat decide, not a quota.
- **Density** — primary visual ≥ 40% of canvas; ≥ 3 depth layers (background + midground + foreground); never a lone small cluster floating in empty space. Squint test: after blur you can still pick out the #1 element.
- **Hierarchy** — combine ≥ 2 of size (3:1) / weight (800 vs 400) / contrast / position (upper-third is golden) / motion, so one element clearly dominates.
- **Depth** — layer 2–3 of: size, blur, opacity gradient, overlap, shadow-stack.
- **Don't show**: nav bars, footers, scrollbars, real cursors / browser chrome, generic decorative shapes standing in for a real asset, floating bokeh / purple-blue "AI" gradients — unless it's an intentional UI-demo reconstruction.
## `## Video direction` — write the invariants ONCE
The whole video shares one look and one motion grammar. Write a **`## Video direction`** block ONCE at the top of `STORYBOARD.md` so every frame inherits it and per-frame Scene lines carry only the **delta**. This block is load-bearing — it is what binds many independent shots into one film. **Keep it.**
- **palette system** — from `frame.md`: which roles map to which hues. Never invent.
- **motion grammar + reveal model** — long-tail eases (`power3` default, smooth over bouncy) + the **VO-paced reveal** model every frame follows (reveal each piece on its spoken cue; never front-load) + what may stay alive during a hold (subtle jitter at most; no lazy breathing) (→ `motion-language.md`).
- **rhythm / held-frame allocation** — name the **held / breather frames** (often before a climax) so the video varies its energy: most frames reveal to the VO, a few hold still (a held read beats bad motion; the anti-monotony discipline; → `motion-language.md`).
- **negative list** — what never appears: off-brand textures / effects the pack forbids, **plus both motion failure modes** — slideshow (front-load then freeze) and screensaver (everything floating independently) (→ `motion-language.md`).
Do **not** repeat these per frame — restating video-level rules in every frame is exactly the bloat this layer prevents.
## Palette & type — from `frame.md`, never invented
- **Palette** — `frame.md` (the adopted pack) is the color truth; apply its roles per frame. Generic basics (one accent, tint neutrals, avoid pure `#000`/`#fff`) → `hyperframes-creative/references/house-style.md`.
- **Type** — fonts resolve via `frame.md`'s type tokens; reference them **by role** (display / body / mono / the pack's ramp), never by raw family or px. Generic typography craft (embedded fonts, dark-bg optical compensation, `tabular-nums`) → `hyperframes-creative/references/typography.md`.
## Caption-band keep-out (plan side)
The bottom ~17% of the canvas is reserved for the caption pill. Plan every frame's content into the **top ~83%** so nothing important lands in the band (the worker enforces the pixel cutoff; you plan the layout). Holds even when captions are disabled — bottom-edge consistency.
## Where the detail lives
| For… | Read |
| ------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| the proven shapes + role→blueprint menu + how to pick | `../hyperframes-animation/blueprints-index.md` → `blueprints/<id>.md` (local) |
| motion — shot model, vocabulary, holds, idle budget, stillness, seek-safe | `motion-language.md` (local) |
| layout — templates, density, depth, hierarchy, caption band | the **Layout** vocabulary in this file |
| concrete eases / ms / stagger + rule recipe bodies (Step 5) | local `../hyperframes-animation/rules/` (the frame worker reads it; you don't) |
| palette + type tokens | the project's `frame.md`; basics → `hyperframes-creative` `house-style.md` / `typography.md` |
| "produced, not generated" foreground density | `hyperframes-creative/references/video-composition.md` |
| within-frame cuts / seams (zoom-through · cut-the-curve · waterfall) | `cut-catalog.md` (the worker builds them inside the composition) |
| transitions | story-design owns `transition_in`; you don't touch it |
## Before you finish — checklist
- **`## Video direction`** written once at the top (palette · motion grammar + shot model + idle budget · stillness allocation · negative list incl. both failure modes); per-frame entries are deltas, not restatements.
- Every frame is a **time-coded shot sequence** with real second windows across its `duration` — not a tag bag.
- **No frame front-loads** — at t=0 only what the VO is saying enters; each further piece reveals on its spoken cue, across the back ~50%. Window count follows the VO, not a fixed number.
- Every frame names an **`blueprint:`** id (Reproduce / Adapt) or `compose`; an Adapt states keep/change and **keeps the signature move**; nothing collapses to a single front-loaded dump — reveals stay paced to the VO.
- **Held frames are deliberate** — allocated in Video direction for rhythm; a held read is fine (prefer stillness to bad motion), but no frame may be front-loaded-then-frozen.
- Each frame's `asset_candidates` have a `focal` + per-candidate `roles`; none added, swapped, or dropped.
- Layout named **inline** per Scene (template / density / depth / hierarchy — the **Layout** vocabulary here); motion named **inline** per Scene from the vocabulary (`motion-language.md`). No px / ease curves / ms / JS.
- Content planned into the top ~83% (caption band clear).
- Palette / type pulled from `frame.md` by role — nothing invented.
- You wrote no HTML and never read `capture/`.
scripts/assemble-index.mjs
#!/usr/bin/env node
// assemble-index.mjs — deterministic top-level index.html assembly for a
// product-launch project. No subagent, no judgment: turns STORYBOARD.md + the
// built frame files (+ optional audio_meta.json) into the standalone index.html
// the renderer consumes, and stages the frame-named capture assets into assets/.
//
// index.html is a *standalone* composition (root <div id="root"> directly in
// <body>, no <template> wrapper — template is for sub-comps). Structure is
// modeled on the canonical fixture packages/studio/fixtures/storyboard-sample/
// index.html and the authoritative head/audio template in
// packages/core/docs/quickstart-template.html. Frame mount order = STORYBOARD
// document order. Transitions are NOT written here — the transitions injector
// mutates this file afterward (data-start/duration/track-index + GSAP).
//
// Track lanes. Same-track time-overlap is this workflow's own assembly convention,
// not a framework rule: the render never reads data-track-index, and no lint rule
// checks overlap (timeline_track_too_dense counts elements per lane for readability).
// The convention exists because the frame injector below ping-pongs 0/1 for overlaps:
// 1 frame sub-comp clips (sequential; the injector 0/1-ping-pongs for overlaps)
// 2 captions sub-comp clip (full-duration overlay, on top of frames)
// 10 per-frame voice <audio>
// 11 BGM <audio>
// 20+i SFX <audio> (one lane each)
//
// audio_meta.json contract (produced by audio.mjs; OPTIONAL — absent ⇒ silent
// video, frames only). Durations come from STORYBOARD (audio sync-durations
// writes them), NOT from here; this file carries only media PATHS, keyed by
// frame number:
// { "bgm": { "path": "assets/bgm/x.mp3", "volume": 0.12 } | null,
// "voices":[ { "frame": 3, "path": "assets/voice/03.wav" } ],
// "sfx": [ { "frame": 3, "file": "assets/sfx/x.mp3", "offset_s": 0,
// "duration_s": 1.0, "volume": 0.35 } ] }
//
// Reads: --storyboard STORYBOARD.md, --hyperframes <project root>,
// [--audio-meta audio_meta.json]. On disk: each built frame's src html,
// capture/{assets,assets/videos,screenshots}/<basename> for staging, compositions/captions.html.
// Writes: <project>/index.html + stages assets/<basename> + (guard ① below)
// repairs a frame file in place when its root is missing data-width/height.
//
// Pre-assembly frame guards (run in the same pass that reads each frame, so common
// `lint` failures surface HERE instead of after assembly + a wasted render):
// ① AUTO-REPAIR — a sub-comp root missing data-width/data-height: inject the canvas
// dims (the renderer needs them on the cloned root; else lint root_missing_dimensions).
// ② APPROVED VIDEO HOIST — an explicitly marked frame video is moved to the host root;
// audio remains orchestrator-owned and unmarked media is still a hard failure.
// ③ HARD FAIL — a timed element (data-start+duration+track-index) that is not the root
// and lacks class="clip" (shows the whole frame), or two same-track clips that overlap.
//
// Exit 0 = index.html written + summary. Exit 1 = fatal contract break (no
// frames, a built/animated frame missing its src/file, a frame with no
// duration, an inner data-composition-id mismatch, or a guard ②/③ violation).
// No backstop: fix upstream.
import { existsSync, readFileSync, writeFileSync } from "node:fs";
import { spawnSync } from "node:child_process";
import { basename, join, resolve } from "node:path";
import { parseStoryboard } from "./lib/storyboard.mjs";
import { parseFormat } from "./lib/dimensions.mjs";
import { stageAssets } from "./lib/assets.mjs";
import { parseColors, semanticColors } from "./lib/tokens.mjs";
import { bgmDefaultVolume } from "../../media-use/audio/scripts/lib/bgm.mjs";
// ---------- argv ----------
const argv = process.argv.slice(2);
const flag = (name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
// Deliberate escape from the bgm_pending refusal below — for previewing while a detached
// generate is still running. Off by default so a silent film can't ship by accident.
const allowPendingBgm = argv.includes("--allow-pending-bgm");
function die(msg) {
console.error(`✗ assemble-index.mjs: ${msg}`);
process.exit(1);
}
// Ensure the BGM track is at least `total` seconds long. HeyGen (and most music
// libraries) return a short loopable clip (~15–30s); mounting it at data-duration=total
// would leave the video's TAIL SILENT. If the file is short, loop-extend it to `total`
// (with a 0.4s fade-in + 1.5s fade-out) into a sibling *.loop.mp3 and return that path.
// Needs ffprobe+ffmpeg (present in the render env); degrades to the original + a warning
// when they're absent, so assembly never hard-fails on audio tooling.
function ensureBgmCovers(relPath, hyperframesDir, total) {
const abs = join(hyperframesDir, relPath);
const probe = spawnSync(
"ffprobe",
["-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", "--", abs],
{ encoding: "utf8" },
);
if (probe.status !== 0) return { looped: false, short: false, reason: "ffprobe unavailable" };
const dur = parseFloat(String(probe.stdout || "").trim());
if (!Number.isFinite(dur) || dur <= 0)
return { looped: false, short: false, reason: "unreadable duration" };
if (dur >= total - 0.1) return { looped: false, short: false, dur }; // already covers
const relOut = relPath.replace(/\.([^./]+)$/, ".loop.$1");
const absOut = join(hyperframesDir, relOut);
const fadeOut = Math.max(0, total - 1.5);
const ff = spawnSync(
"ffmpeg",
[
"-y",
"-stream_loop",
"-1",
"-i",
abs,
"-t",
String(total),
"-af",
`afade=t=in:st=0:d=0.4,afade=t=out:st=${fadeOut}:d=1.5`,
"-c:a",
"libmp3lame",
"-q:a",
"2",
absOut,
],
{ encoding: "utf8" },
);
if (ff.status !== 0 || !existsSync(absOut))
return { looped: false, short: true, dur, reason: "ffmpeg unavailable" };
return { looped: true, rel: relOut, from: dur };
}
const hyperframesDir = resolve(flag("hyperframes", "."));
const storyboardPath = resolve(flag("storyboard", join(hyperframesDir, "STORYBOARD.md")));
const audioMetaPath = resolve(flag("audio-meta", join(hyperframesDir, "audio_meta.json")));
const outPath = resolve(flag("out", join(hyperframesDir, "index.html")));
const r3 = (x) => Math.round(x * 1000) / 1000;
const anomalies = [];
const frameErrors = []; // fatal per-frame composition violations (guards ②/③) — reported together
const repairs = []; // auto-repairs applied to frame files in place (guard ①)
// ---------- parse storyboard ----------
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const { width: WIDTH, height: HEIGHT } = parseFormat(manifest.globals.format);
// ---------- per-frame composition guards (see header ①②③) ----------
// String-level checks on each frame's HTML — no DOM parse, deterministic, run in
// the same pass that already reads the file. OPEN_TAG matches one opening tag while
// tolerating quoted attribute values that contain ">" (e.g. inline styles).
const OPEN_TAG = "<([a-zA-Z][a-zA-Z0-9-]*)((?:[^>\"']|\"[^\"]*\"|'[^']*')*)>";
const attrPresent = (attrs, name) => new RegExp(`(?:^|\\s)${name}(?:[\\s=]|$)`).test(attrs);
const attrValue = (attrs, name) => {
const m = attrs.match(new RegExp(`(?:^|\\s)${name}\\s*=\\s*(?:"([^"]*)"|'([^']*)')`));
return m ? (m[1] ?? m[2]) : null;
};
// The root (or a nested-comp mount) legitimately carries timing without class="clip".
const isRootish = (attrs) =>
/(?:^|\s)id\s*=\s*["']root["']/.test(attrs) ||
attrPresent(attrs, "data-composition-id") ||
attrPresent(attrs, "data-composition-src");
// Locate the composition root opening tag: prefer id="root", else the first element
// carrying data-composition-id. Returns { start, end, full, attrs } or null.
function findRootTag(html) {
const re = new RegExp(OPEN_TAG, "g");
let m;
let firstCompId = null;
while ((m = re.exec(html))) {
const attrs = m[2];
if (/(?:^|\s)id\s*=\s*["']root["']/.test(attrs))
return { start: m.index, end: m.index + m[0].length, full: m[0], attrs };
if (attrPresent(attrs, "data-composition-id") && !firstCompId)
firstCompId = { start: m.index, end: m.index + m[0].length, full: m[0], attrs };
}
return firstCompId;
}
function attrValueFrom(attrs, name) {
const escaped = name.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
const match = attrs.match(new RegExp(`(?:^|\\s)${escaped}\\s*=\\s*(?:"([^"]*)"|'([^']*)')`));
return match ? (match[1] ?? match[2]) : null;
}
function escapeHtmlAttr(value) {
return value
.replaceAll("&", "&")
.replaceAll('"', """)
.replaceAll("<", "<")
.replaceAll(">", ">");
}
function approvedVideoAttrs(attrs) {
const forwarded = [];
for (const name of ["id", "src", "poster", "preload", "aria-label", "data-media-start"]) {
const value = attrValueFrom(attrs, name);
if (value !== null) forwarded.push(`${name}="${escapeHtmlAttr(value)}"`);
}
for (const name of ["muted", "playsinline", "loop"]) {
if (attrPresent(attrs, name)) forwarded.push(name);
}
return forwarded.join(" ");
}
function approvedVideoLayout(attrs) {
const names = ["x", "y", "width", "height"];
const raw = Object.fromEntries(
names.map((name) => [name, attrValueFrom(attrs, `data-frame-video-${name}`)]),
);
const rawFit = attrValueFrom(attrs, "data-frame-video-fit");
const values = Object.fromEntries(names.map((name) => [name, Number(raw[name])]));
if (
names.some(
(name) => raw[name] === null || raw[name].trim() === "" || !Number.isFinite(values[name]),
) ||
values.width <= 0 ||
values.height <= 0
) {
return {
style: null,
error:
"approved frame video layout data-frame-video-x/y/width/height must all be finite numeric values, with positive width and height",
};
}
const fit = rawFit ?? "cover";
if (!["cover", "contain", "fill", "none", "scale-down"].includes(fit)) {
return {
style: null,
error:
'approved frame video layout data-frame-video-fit must be "cover", "contain", "fill", "none", or "scale-down"',
};
}
return {
style: `position:absolute;left:${values.x}px;top:${values.y}px;width:${values.width}px;height:${values.height}px;object-fit:${fit}`,
error: null,
};
}
function hoistApprovedVideos(html, label) {
const videos = [];
const errors = [];
const scan = html
.replace(/<!--[\s\S]*?-->/g, (match) => " ".repeat(match.length))
.replace(/<script\b[\s\S]*?<\/script[^>]*>/gi, (match) => " ".repeat(match.length))
.replace(/<style\b[\s\S]*?<\/style[^>]*>/gi, (match) => " ".repeat(match.length));
const re = /<video\b((?:[^>"']|"[^"]*"|'[^']*')*)>([\s\S]*?)<\/video\s*>/gi;
const repaired = html.replace(re, (full, attrs, inner, offset) => {
if (scan[offset] !== "<") return full;
if (attrValueFrom(attrs, "data-frame-video") !== "approved") return full;
const rawStart = attrValueFrom(attrs, "data-start");
const rawDuration = attrValueFrom(attrs, "data-duration");
const rawTrack = attrValueFrom(attrs, "data-track-index");
if (rawStart === null || rawDuration === null || rawTrack === null) {
errors.push(
`${label}: approved frame video must declare quoted data-start, data-duration, and data-track-index`,
);
return full;
}
const start = Number(rawStart);
const duration = Number(rawDuration);
const track = Number(rawTrack);
if (
!Number.isFinite(start) ||
!Number.isFinite(duration) ||
duration <= 0 ||
!Number.isFinite(track)
) {
errors.push(
`${label}: approved frame video must declare finite data-start, positive data-duration, and data-track-index`,
);
return full;
}
const layout = approvedVideoLayout(attrs);
if (layout.error) {
errors.push(`${label}: ${layout.error}`);
return full;
}
videos.push({
attrs: approvedVideoAttrs(attrs),
inner,
start,
duration,
track,
layoutStyle: layout.style,
});
return "<!-- approved frame video hoisted by assemble-index -->";
});
return { html: repaired, videos, errors };
}
// Returns { errors: string[], repairedHtml: string|null, repairNote: string|null }.
function guardFrame(html, label) {
const errors = [];
const originalHtml = html;
const approved = hoistApprovedVideos(html, label);
html = approved.html;
errors.push(...approved.errors);
// Scan a copy with comments + <script>/<style> bodies blanked, so a tag-like string
// in a comment (e.g. "<!-- match the host <video> coords -->") or in GSAP code can't
// trip ②/③. ① still splices into the ORIGINAL html, so its offsets stay correct.
const scan = html
.replace(/<!--[\s\S]*?-->/g, " ")
.replace(/<script\b[\s\S]*?<\/script[^>]*>/gi, " ")
.replace(/<style\b[\s\S]*?<\/style[^>]*>/gi, " ");
// ② media inside a sub-comp — never driven by the runtime (renders blank/black).
const media = scan.match(/<(video|audio)(?=[\s/>])/i);
if (media) {
errors.push(
`${label}: has a <${media[1].toLowerCase()}> inside the sub-composition. This workflow hoists media to index.html so the frame injector owns it; the framework itself renders media inside a sub-composition identically to media at the host root (verified by render, and pinned by packages/producer/tests/sub-composition-video), so this is an assembly convention, not a runtime limit. Move the clip to index.html as a root-level <video>/<audio> and drive any per-scene motion on the main timeline (composition-patterns.md archetype B).`,
);
}
// ③ timed-element checks: missing class="clip", and same-track window overlap.
const re = new RegExp(OPEN_TAG, "g");
const clips = [];
let m;
while ((m = re.exec(scan))) {
const attrs = m[2];
if (
!attrPresent(attrs, "data-start") ||
!attrPresent(attrs, "data-duration") ||
!attrPresent(attrs, "data-track-index")
)
continue;
if (isRootish(attrs)) continue;
if (!/(?:^|\s)class\s*=\s*["'][^"']*\bclip\b[^"']*["']/.test(attrs)) {
errors.push(
`${label}: a timed <${m[1]}> (data-start/duration/track-index) has no class="clip" — it renders for the whole frame instead of only its window. Add class="clip", or remove the timing attrs if it is a GSAP-animated element meant to be present throughout.`,
);
}
const track = attrValue(attrs, "data-track-index");
const start = parseFloat(attrValue(attrs, "data-start"));
const dur = parseFloat(attrValue(attrs, "data-duration"));
if (track != null && Number.isFinite(start) && Number.isFinite(dur))
clips.push({ track, start, end: start + dur });
}
const EPS = 1e-3; // adjacent clips that merely touch are legal
const byTrack = new Map();
for (const c of clips) {
const arr = byTrack.get(c.track);
if (arr) arr.push(c);
else byTrack.set(c.track, [c]);
}
for (const [track, list] of byTrack) {
list.sort((a, b) => a.start - b.start);
for (let i = 1; i < list.length; i++) {
if (list[i].start < list[i - 1].end - EPS) {
errors.push(
`${label}: clips on track ${track} overlap (one ends at ${r3(list[i - 1].end)}s, the next starts at ${r3(list[i].start)}s). This workflow's injector assumes one clip per lane at a time. The render itself tolerates the overlap; put them on distinct data-track-index lanes or fix their windows.`,
);
break; // one report per track is enough
}
}
}
// ① auto-repair: ensure the root carries data-width / data-height.
let repairedHtml = approved.html !== originalHtml ? approved.html : null;
let repairNote = null;
const root = findRootTag(html);
if (root) {
const needW = !attrPresent(root.attrs, "data-width");
const needH = !attrPresent(root.attrs, "data-height");
if (needW || needH) {
const inject =
(needW ? ` data-width="${WIDTH}"` : "") + (needH ? ` data-height="${HEIGHT}"` : "");
const newTag = root.full.replace(/(\/?>)$/, `${inject}$1`);
repairedHtml = html.slice(0, root.start) + newTag + html.slice(root.end);
repairNote = `${label}: injected${needW ? " data-width" : ""}${needH ? " data-height" : ""} (${WIDTH}×${HEIGHT}) on the root — was missing (would lint root_missing_dimensions)`;
}
}
return { errors, repairedHtml, repairNote, hoistedVideos: approved.videos };
}
// ---------- resolve mountable frames in document order ----------
// A frame mounts when its src html exists on disk. A built/animated frame
// missing its src/file is a contract break (die). An outline frame with no
// file is skipped (still a placeholder) with an anomaly note.
const mounted = [];
for (const f of manifest.frames) {
const label = `frame ${f.number ?? f.index}${f.title ? ` (${f.title})` : ""}`;
const built = f.status === "built" || f.status === "animated";
if (!f.src) {
if (built) die(`${label} is ${f.status} but has no \`src\` — the orchestrator must write it`);
anomalies.push(`${label}: status ${f.status}, no src — skipped`);
continue;
}
const compAbs = join(hyperframesDir, f.src);
// Read directly and handle ENOENT here rather than an existsSync precheck — the
// check→read/write pair is a TOCTOU race CodeQL flags (js/file-system-race).
let html;
try {
html = readFileSync(compAbs, "utf8");
} catch {
if (built)
die(`${label} is ${f.status} but its src ${f.src} is not on disk — re-dispatch the worker`);
anomalies.push(`${label}: src ${f.src} not on disk (status ${f.status}) — skipped`);
continue;
}
if (!Number.isFinite(f.durationSeconds) || f.durationSeconds <= 0) {
die(
`${label}: no positive duration (got ${JSON.stringify(f.duration)}) — run audio sync-durations`,
);
}
// Host data-composition-id MUST equal the inner file's, or the runtime never
// finds the timeline. frame_id = src basename (frame-worker contract); verify
// the inner html actually declares it.
const compId = basename(f.src).replace(/\.html?$/i, "");
// Guard against blank/partial scene files: a worker that errors or is
// interrupted mid-write leaves an empty (or markup-less) file that exists but
// fails at render with "Composition HTML is empty or could not be parsed".
// Catch it here — before emitting data-composition-src — and re-dispatch.
if (!html.trim() || !/<\w/.test(html)) {
die(
`${label}: ${f.src} is empty or has no HTML — the worker wrote a blank/partial file. Re-dispatch that worker before assembling.`,
);
}
// pre-assembly guards: ① repair missing root dims in place, ②/③ collect fatal violations.
const guard = guardFrame(html, label);
if (guard.repairedHtml) {
writeFileSync(compAbs, guard.repairedHtml);
html = guard.repairedHtml;
repairs.push(guard.repairNote);
}
for (const e of guard.errors) frameErrors.push(e);
if (
!html.includes(`data-composition-id="${compId}"`) &&
!html.includes(`data-composition-id='${compId}'`)
) {
die(`${label}: ${f.src} has no data-composition-id="${compId}" (host/inner id must match)`);
}
mounted.push({
frame: f,
compId,
durationSeconds: r3(f.durationSeconds),
hoistedVideos: guard.hoistedVideos,
});
}
if (frameErrors.length) {
die(
`${frameErrors.length} frame composition violation(s) — fix the worker output and re-assemble:\n` +
frameErrors.map((e) => ` • ${e}`).join("\n"),
);
}
if (mounted.length === 0) die("no mountable frames (none built with an on-disk src)");
// cumulative starts — emitted data-start[i] + data-duration[i] == start[i+1] by
// construction (renderer computes end the same way), so adjacent clips touch
// exactly with no float-overlap.
let acc = 0;
for (const m of mounted) {
m.start = acc;
acc += m.durationSeconds;
}
const TOTAL = r3(acc);
// ---------- duration expectation (advisory) ----------
// Frontmatter `duration:` carries the brief's rough length expectation
// (storyboard-format.md § Frontmatter). Never blocks the build: report where
// the cut lands, and flag a large gap so the agent judges whether the drift
// serves the piece.
let durationNote = "";
const rawTarget = manifest.globals.extra?.duration;
if (rawTarget != null && String(rawTarget).trim() !== "") {
const targetMatch = String(rawTarget).match(/(\d+(?:\.\d+)?)/);
const target = targetMatch ? parseFloat(targetMatch[1]) : NaN;
if (!Number.isFinite(target) || target <= 0) {
anomalies.push(
`frontmatter duration "${rawTarget}" is not parseable (e.g. "22s") — skipped the expectation check`,
);
} else {
const diff = r3(TOTAL - target);
durationNote = ` (expected ~${target}s, ${diff >= 0 ? "+" : ""}${diff}s)`;
const pct = Math.abs((diff / target) * 100);
if (pct > 10) {
anomalies.push(
`total ${TOTAL}s lands ${Math.round(pct)}% ${diff > 0 ? "over" : "under"} the brief's ~${target}s expectation — ` +
`judge whether the drift serves the piece (pacing, narration fit); re-pace, or update \`duration:\` if the new length is intended`,
);
}
}
}
const startOfFrameNumber = new Map();
for (const m of mounted) if (m.frame.number != null) startOfFrameNumber.set(m.frame.number, m);
// ---------- audio_meta (optional) ----------
let audio = { bgm: null, voices: [], sfx: [] };
if (existsSync(audioMetaPath)) {
try {
const parsed = JSON.parse(readFileSync(audioMetaPath, "utf8"));
// bgm_pending rides along: without it this step cannot tell a detached generate that has
// not landed yet from a film that is silent by design, and it would build the silent one.
audio = {
bgm: parsed.bgm ?? null,
bgm_pending: !!parsed.bgm_pending,
voices: parsed.voices ?? [],
sfx: parsed.sfx ?? [],
};
} catch (e) {
die(`audio_meta.json parse: ${e.message}`);
}
}
const voiceByFrame = new Map();
for (const v of audio.voices) if (v.frame != null) voiceByFrame.set(v.frame, v);
// ---------- build <body> in track order ----------
const body = [];
let voiceCount = 0;
for (const m of mounted) {
// (track 1) frame sub-comp clip — no class="clip" semantics needed; .scene CSS sizes it.
body.push(
` <div`,
` id="el-${m.compId}"`,
` class="scene"`,
` data-composition-id="${m.compId}"`,
` data-composition-src="${m.frame.src}"`,
` data-start="${m.start}"`,
` data-duration="${m.durationSeconds}"`,
` data-track-index="1"`,
` ></div>`,
);
// (track 10) voice — only when the file is actually on disk.
const v = m.frame.number != null ? voiceByFrame.get(m.frame.number) : undefined;
if (v?.path) {
if (existsSync(join(hyperframesDir, v.path))) {
body.push(
` <audio`,
` id="el-${m.compId}-voice"`,
` src="${v.path}"`,
` data-start="${m.start}"`,
` data-duration="${m.durationSeconds}"`,
` data-track-index="10"`,
` data-volume="1"`,
` ></audio>`,
);
voiceCount++;
} else {
anomalies.push(`${m.compId}: voice ${v.path} not on disk — skipped`);
}
}
body.push("");
}
// Approved frame videos are mounted at the host root after frame clips. Translate
// frame-relative timing to the global timeline and keep them off audio/frame lanes.
for (const [frameIndex, m] of mounted.entries()) {
for (const video of m.hoistedVideos ?? []) {
const globalStart = r3(m.start + video.start);
const track = 1000 + frameIndex * 1000 + video.track;
const id = /(?:^|\s)id\s*=/.test(video.attrs) ? "" : ` id="el-${m.compId}-video-${frameIndex}"`;
body.push(
` <video${id} ${video.attrs}`,
` class="clip"`,
...(video.layoutStyle ? [` style="${video.layoutStyle}"`] : []),
` data-start="${globalStart}"`,
` data-duration="${r3(video.duration)}"`,
` data-track-index="${track}"`,
` >${video.inner}</video>`,
"",
);
}
}
// (track 11) BGM — duck under narration when any voice is present. Loop-extend a short
// track to the full video length so the tail isn't silent (libraries return ~15–30s clips).
let bgmEmitted = false;
let bgmNote = "";
if (audio.bgm?.path) {
if (existsSync(join(hyperframesDir, audio.bgm.path))) {
let bgmSrc = audio.bgm.path;
const cov = ensureBgmCovers(audio.bgm.path, hyperframesDir, TOTAL);
if (cov.looped) {
bgmSrc = cov.rel;
bgmNote = ` (looped ${cov.from.toFixed(1)}s→${TOTAL}s)`;
} else if (cov.short) {
anomalies.push(
`bgm is ${cov.dur?.toFixed?.(1) ?? "?"}s (< ${TOTAL}s) and could not be extended (${cov.reason}) — the tail will be silent; install ffmpeg`,
);
}
// An explicit volume from audio_meta always wins; otherwise the shared
// media-use default (bed ~ -18 dB under narration, forward for a silent film).
const vol = audio.bgm.volume != null ? audio.bgm.volume : bgmDefaultVolume(voiceCount > 0);
body.push(
` <!-- BGM -->`,
` <audio`,
` id="el-bgm"`,
` src="${bgmSrc}"`,
` data-start="0"`,
` data-duration="${TOTAL}"`,
` data-track-index="11"`,
` data-volume="${vol}"`,
` ></audio>`,
"",
);
bgmEmitted = true;
} else {
anomalies.push(`bgm ${audio.bgm.path} not on disk — skipped`);
}
} else if (audio.bgm_pending) {
// The distinction the flag exists to make. A warning is not enough here: assemble is re-run
// on Step 6 rework, long after the audio step's own warning scrolled past, and it would
// happily build a silent film from a snapshot whose JSON says the bed is still generating.
// Refuse by default; --allow-pending-bgm is the deliberate escape for previewing mid-generate.
if (!allowPendingBgm) {
die(
"audio_meta.json says bgm_pending — the music bed is still generating and is NOT in this " +
"assembly. Wait for the track, re-run the audio step, then assemble again. To assemble a " +
"deliberately silent preview anyway, pass --allow-pending-bgm.",
);
}
anomalies.push(
"bgm still generating (bgm_pending) — assembled without a bed per --allow-pending-bgm",
);
}
// (track 2) captions — captions.mjs writes this or legally skips; key off existence.
let captionsEmitted = false;
if (existsSync(join(hyperframesDir, "compositions/captions.html"))) {
body.push(
` <!-- captions -->`,
` <div`,
` id="el-captions"`,
` class="scene"`,
` data-composition-id="captions"`,
` data-composition-src="compositions/captions.html"`,
` data-start="0"`,
` data-duration="${TOTAL}"`,
` data-track-index="2"`,
` ></div>`,
"",
);
captionsEmitted = true;
}
// (track 20+i) SFX — placed at its frame's start + offset.
let sfxEmitted = 0;
audio.sfx.forEach((cue, i) => {
const host = cue.frame != null ? startOfFrameNumber.get(cue.frame) : undefined;
if (!host) {
anomalies.push(`sfx ${cue.file}: frame ${cue.frame} not mounted — skipped`);
return;
}
const rel = cue.file;
if (!existsSync(join(hyperframesDir, rel))) {
anomalies.push(`sfx ${rel} not on disk — skipped`);
return;
}
const t = r3(host.start + (cue.offset_s ?? 0));
const dur = r3(cue.duration_s ?? 1);
const vol = cue.volume != null ? cue.volume : 0.35;
if (sfxEmitted === 0) body.push(` <!-- SFX -->`);
body.push(
` <audio`,
` id="el-sfx-${i}"`,
` src="${rel}"`,
` data-start="${t}"`,
` data-duration="${dur}"`,
` data-track-index="${20 + i}"`,
` data-volume="${vol}"`,
` ></audio>`,
);
sfxEmitted++;
});
// ---------- stage frame-named assets: capture/ → assets/ (idempotent backstop) ----------
// Frame workers + the live preview reference assets/<basename>; stage-assets.mjs
// already ran this at Step 4 close. Re-run as a backstop so a late-named asset
// still lands. Shared logic: lib/assets.mjs (first-wins, safe to call twice).
const {
staged,
wanted,
anomalies: assetAnomalies,
} = stageAssets({
hyperframesDir,
frames: manifest.frames,
});
for (const a of assetAnomalies) anomalies.push(a);
// ---------- <head> ----------
// ---------- ground color ----------
// Per-frame roots carry data-start/data-duration and get clip-gated against the
// global timeline in render (only the first frame's [0,dur] window overlaps global
// 0), so a frame's own full-bleed background can't be relied on as the video ground —
// every frame after the first would render on the bare body color (black). Paint the
// ground on the always-present root composition instead, using the project's frame.md
// canvas color (the same ground role the caption skin maps to --cap-canvas). Falls
// back to the body letterbox color when frame.md is absent or has no resolvable ground.
const framePath = join(hyperframesDir, "frame.md");
let groundColor = null;
if (existsSync(framePath)) {
try {
const roles = semanticColors(parseColors(readFileSync(framePath, "utf8")));
if (roles && roles.canvas) groundColor = roles.canvas;
} catch {
/* leave groundColor null — #root stays transparent over the body letterbox */
}
}
const headStyle = [
" * {",
" margin: 0;",
" padding: 0;",
" box-sizing: border-box;",
" }",
" html,",
" body {",
` width: ${WIDTH}px;`,
` height: ${HEIGHT}px;`,
" overflow: hidden;",
" background: #000;",
" }",
" #root {",
" position: relative;",
` width: ${WIDTH}px;`,
` height: ${HEIGHT}px;`,
" overflow: hidden;",
...(groundColor ? [` background: ${groundColor};`] : []),
" }",
" .scene {",
" position: absolute;",
" inset: 0;",
" width: 100%;",
" height: 100%;",
" }",
].join("\n");
const html = `<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=${WIDTH}, height=${HEIGHT}" />
<script src="https://cdn.jsdelivr.net/npm/gsap@3.14.2/dist/gsap.min.js" integrity="sha384-sG0Hv1tP1lZCk9KQmrIbY/XNwi+OY84GQqhMscbnsoBFqAz8KNCil1kvfL3Hbbk2" crossorigin="anonymous"></script>
<style>
${headStyle}
</style>
</head>
<body>
<div
id="root"
data-composition-id="main"
data-start="0"
data-duration="${TOTAL}"
data-width="${WIDTH}"
data-height="${HEIGHT}"
>
${body.join("\n")}
</div>
<script>
window.__timelines = window.__timelines || {};
window.__timelines["main"] = gsap.timeline({ paused: true });
</script>
</body>
</html>
`;
writeFileSync(outPath, html);
// ---------- summary ----------
console.log(`✓ wrote ${outPath}`);
console.log(` canvas: ${WIDTH}×${HEIGHT}`);
console.log(` frames (track 1): ${mounted.length}`);
console.log(` voice (track 10): ${voiceCount}`);
console.log(` bgm (track 11): ${bgmEmitted ? "yes" + bgmNote : "no"}`);
console.log(` captions (track 2): ${captionsEmitted ? "yes" : "no"}`);
console.log(` sfx (track 20+): ${sfxEmitted}`);
console.log(` assets staged: ${staged}/${wanted.size}`);
console.log(` total duration: ${TOTAL}s${durationNote}`);
if (repairs.length) {
console.log(`\nrepaired (frame files updated in place):`);
for (const rp of repairs) console.log(` - ${rp}`);
}
if (anomalies.length) {
console.log(`\nanomalies (non-fatal):`);
for (const a of anomalies) console.log(` - ${a}`);
}
scripts/assemble-index.test.mjs
import assert from "node:assert/strict";
import { existsSync, mkdirSync, mkdtempSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { spawnSync } from "node:child_process";
import test from "node:test";
const assembleScript = new URL("./assemble-index.mjs", import.meta.url).pathname;
// ── bgm_pending at the assembly boundary ─────────────────────────────────────
// Regression: the flag survived into audio_meta.json but assemble rebuilt its audio object
// from three named keys and dropped it, so the step that actually builds the film could not
// tell "not ready yet" from "silent by design" and would ship the silent one.
function assembleWith({ audioMeta, extraArgs = [] }) {
const dir = mkdtempSync(join(tmpdir(), "product-launch-assemble-"));
writeFileSync(
join(dir, "STORYBOARD.md"),
"---\nformat: 1920x1080\nmessage: T\n---\n\n## Frame 1 — A\n- duration: 3s\n- src: compositions/frames/01-a.html\n",
);
mkdirSync(join(dir, "compositions", "frames"), { recursive: true });
writeFileSync(
join(dir, "compositions", "frames", "01-a.html"),
'<div data-composition-id="01-a" data-width="1920" data-height="1080">' +
'<section class="clip" data-start="0" data-duration="3"></section></div>',
);
if (audioMeta) writeFileSync(join(dir, "audio_meta.json"), JSON.stringify(audioMeta));
const r = spawnSync(
process.execPath,
[
assembleScript,
"--storyboard",
join(dir, "STORYBOARD.md"),
"--hyperframes",
dir,
...extraArgs,
],
{ encoding: "utf8" },
);
return { dir, r };
}
test("assemble REFUSES while bgm_pending and no bed on disk", () => {
const { dir, r } = assembleWith({
audioMeta: { bgm: null, bgm_pending: true, voices: [], sfx: [] },
});
assert.notEqual(r.status, 0, "should not assemble a silent film over a pending bed");
assert.match(r.stderr, /bgm_pending/);
// Refusing means producing nothing, not a half-built index.
assert.equal(existsSync(join(dir, "index.html")), false);
});
test("--allow-pending-bgm assembles anyway, and says so", () => {
const { dir, r } = assembleWith({
audioMeta: { bgm: null, bgm_pending: true, voices: [], sfx: [] },
extraArgs: ["--allow-pending-bgm"],
});
assert.equal(r.status, 0, r.stderr);
assert.equal(existsSync(join(dir, "index.html")), true);
assert.match(r.stdout + r.stderr, /pending/i);
});
test("a film that is silent BY DESIGN still assembles untouched", () => {
// The whole point of carrying the flag: this case must stay distinguishable from the above.
const { dir, r } = assembleWith({ audioMeta: { bgm: null, voices: [], sfx: [] } });
assert.equal(r.status, 0, r.stderr);
assert.equal(existsSync(join(dir, "index.html")), true);
assert.doesNotMatch(r.stderr, /bgm_pending/);
});
scripts/audio.mjs
#!/usr/bin/env node
// audio.mjs — product-launch audio ADAPTER. The TTS / BGM / SFX implementation
// no longer lives here: it is the shared engine at
// ../../media-use/audio/scripts/audio.mjs. This file only (a) maps the
// product-launch model (SCRIPT.md frames + STORYBOARD.md music/sfx) into the
// engine's neutral audio_request.json, (b) converts the engine's id-keyed
// audio_meta back into the frame-keyed shape captions.mjs / assemble-index.mjs
// already consume, and (c) keeps the local `sync-durations` pass (it rewrites
// STORYBOARD.md, which is product-launch-specific).
//
// Three modes (unchanged CLI surface):
// (default) generate — engine --only tts,bgm. BGM mode is "retrieve" (strict:
// no HeyGen credential ⇒ skip, never a detached generate, since this
// workflow has no wait-bgm step). Runs in the background during Step 4.
// sync-durations — write real voice durations into STORYBOARD.md (local).
// fetch-sfx — engine --only sfx, merged into the existing meta (Step 5,
// after the frames' `sfx:` cues exist).
//
// node audio.mjs --script ./SCRIPT.md --storyboard ./STORYBOARD.md --hyperframes . --out ./audio_meta.json
// node audio.mjs sync-durations --audio-meta ./audio_meta.json --storyboard ./STORYBOARD.md
// node audio.mjs fetch-sfx --storyboard ./STORYBOARD.md --hyperframes .
import { spawnSync } from "node:child_process";
import { existsSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { dirname, join, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { parseStoryboard } from "./lib/storyboard.mjs";
const HERE = dirname(fileURLToPath(import.meta.url));
const DEFAULT_ENGINE = join(HERE, "..", "..", "media-use", "audio", "scripts", "audio.mjs");
const flag = (argv, name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
const pad2 = (n) => String(n).padStart(2, "0");
// SCRIPT.md → [{ frame, text }]. `## … (Frame N)` opens a line; `**key:**` rows
// are metadata; the indented block is the spoken text (the only TTS input).
function parseScript(md) {
const out = [];
let cur = null;
const flush = () => {
if (cur && cur.text.trim()) out.push({ frame: cur.frame, text: cur.text.trim() });
cur = null;
};
for (const line of md.split(/\r?\n/)) {
const h = line.match(/^#{2,3}\s+.*?\(frame\s+(\d+)\)/i);
if (h) {
flush();
cur = { frame: Number(h[1]), text: "" };
continue;
}
if (!cur) continue;
if (/^\s*\*\*/.test(line)) continue;
const m = line.match(/^(?: {4,}|\t)(.+)$/);
if (m) cur.text += (cur.text ? " " : "") + m[1].trim();
}
flush();
return out;
}
// Path of the engine's neutral meta — a stable sidecar so `--only` merges
// (generate then fetch-sfx) accumulate, while audio_meta.json holds the PL shape.
const neutralPath = (plOutPath) => join(dirname(plOutPath), "audio_engine_meta.json");
// Run the shared engine. Returns nothing; dies on a non-zero exit.
function runEngine({ request, hyperframesDir, neutral, only, extra = [] }, die) {
const reqPath = join(hyperframesDir, "audio_request.json");
writeFileSync(reqPath, JSON.stringify(request, null, 2));
const engine = process.env.HF_MEDIA_ENGINE || DEFAULT_ENGINE;
if (!existsSync(engine)) die(`media audio engine not found at ${engine} (set $HF_MEDIA_ENGINE)`);
const args = [
engine,
"--request",
reqPath,
"--hyperframes",
hyperframesDir,
"--out",
neutral,
"--only",
only,
...extra,
];
const r = spawnSync("node", args, { stdio: "inherit" });
if (r.status !== 0) die(`media audio engine exited ${r.status}`);
}
// Engine neutral meta (id-keyed) → product-launch meta (frame-keyed) consumed by
// captions.mjs / assemble-index.mjs. id is the zero-padded frame number.
function toProductLaunchMeta(neutral) {
const voices = (neutral.voices ?? []).map((v) => ({
frame: Number(v.id),
path: v.path,
duration_s: v.duration_s,
words: (v.words ?? []).map((w) => ({ id: w.id, text: w.text, start: w.start, end: w.end })),
}));
const bgm = neutral.bgm
? {
path: neutral.bgm.path,
volume: neutral.bgm.volume,
query: neutral.bgm.query ?? null,
duration_s: neutral.bgm.duration_s ?? null,
}
: null;
// bgm_pending must survive the neutral → PL translation. A detached generate (Lyria/MusicGen)
// leaves `bgm: null, bgm_pending: true` until the track lands; dropping the flag made
// "not ready yet" indistinguishable from "silent by design", so a later `fetch-sfx` snapshot
// turned a still-generating bed into no music at all with nothing to signal it.
const bgmPending = !!neutral.bgm_pending;
const sfx = (neutral.sfx ?? []).map((s) => ({
frame: Number(s.id),
file: s.file,
offset_s: s.offset_s ?? 0,
duration_s: s.duration_s ?? 1,
volume: s.volume ?? 0.35,
}));
return { bgm, bgm_pending: bgmPending, voices, sfx };
}
// ── generate (TTS + BGM) ────────────────────────────────────────────────────
function runGenerate(argv) {
const die = (m) => {
console.error(`✗ audio generate: ${m}`);
process.exit(1);
};
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
const scriptPath = resolve(flag(argv, "script", join(hyperframesDir, "SCRIPT.md")));
const outPath = resolve(flag(argv, "out", join(hyperframesDir, "audio_meta.json")));
const userVoice = flag(argv, "voice", null);
const provider = flag(argv, "provider", process.env.HF_TTS_PROVIDER || "auto");
const speed = Number(flag(argv, "speed", "1.0")) || 1.0;
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const g = manifest.globals;
const lines = existsSync(scriptPath)
? parseScript(readFileSync(scriptPath, "utf8")).map((l) => ({
id: pad2(l.frame),
text: l.text,
}))
: [];
// The canonical fully-silent marker (SKILL.md Step 3.1): `music: none` in
// the storyboard's top YAML block turns BGM off; combined with no SCRIPT.md
// the project is fully silent — generate nothing and remove any stale meta
// from a previous run (assemble treats an absent audio_meta.json as silent).
const bgmOff =
String(g.extra?.music ?? "")
.trim()
.toLowerCase() === "none";
if (bgmOff && !lines.length) {
rmSync(outPath, { force: true });
rmSync(neutralPath(outPath), { force: true });
console.log(
"✓ audio generate: project marked silent (music: none, no SCRIPT.md) — nothing to generate",
);
return;
}
if (!lines.length) console.error("· no SCRIPT.md — silent film (BGM only)");
// BGM mood: storyboard `music:` → message → arc → default. `mode: retrieve` is
// strict here (no wait-bgm step downstream).
const query = (g.extra && g.extra.music) || g.message || g.arc || "calm cinematic underscore";
const request = {
provider,
speed,
lines,
bgm: bgmOff
? { mode: "none" }
: { mode: "retrieve", query, blob: g.message || "", arc: g.arc || "" },
};
if (userVoice) request.voice = userVoice;
const neutral = neutralPath(outPath);
runEngine({ request, hyperframesDir, neutral, only: "tts,bgm" }, die);
const meta = toProductLaunchMeta(JSON.parse(readFileSync(neutral, "utf8")));
writeFileSync(outPath, JSON.stringify(meta, null, 2));
console.log(
`✓ audio generate: ${meta.voices.length} voice + ${meta.bgm ? "1 bgm" : "no bgm"} → ${outPath}`,
);
}
// ── fetch-sfx ────────────────────────────────────────────────────────────────
function runFetchSfx(argv) {
const die = (m) => {
console.error(`✗ audio fetch-sfx: ${m}`);
process.exit(1);
};
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
const outPath = resolve(flag(argv, "audio-meta", join(hyperframesDir, "audio_meta.json")));
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
// Per-frame `sfx:` cues (comma-separated) → engine lines carrying only sfx.
// `filter(Boolean)` alone is not enough: a storyboard that spells "no SFX here" as
// `sfx: none` used to reach the engine as a cue literally NAMED "none", which then failed
// to resolve. The absence sentinels are part of the storyboard vocabulary, so drop them.
const SFX_NONE = new Set(["none", "no", "n/a", "na", "skip", "-", "—", "–"]);
const lines = [];
for (const f of manifest.frames) {
const names = (f.extra?.sfx ?? "")
.split(",")
.map((s) => s.trim())
.filter((s) => s && !SFX_NONE.has(s.toLowerCase()));
if (names.length && f.number != null) lines.push({ id: pad2(f.number), sfx: names });
}
const neutral = neutralPath(outPath);
const request = { lines, bgm: { mode: "none" } };
// --only sfx is a MERGE, not an overwrite: the engine reads the existing neutral
// sidecar (audio_engine_meta.json) and recomputes only the sfx section, so the
// voices/bgm written by the earlier generate (--only tts,bgm) pass are preserved.
runEngine({ request, hyperframesDir, neutral, only: "sfx" }, die);
const meta = toProductLaunchMeta(JSON.parse(readFileSync(neutral, "utf8")));
writeFileSync(outPath, JSON.stringify(meta, null, 2));
console.log(`✓ audio fetch-sfx: ${meta.sfx.length} SFX cue(s) → ${outPath}`);
// This pass rewrites audio_meta.json from the neutral sidecar. If a detached BGM generate is
// still running, the bed it eventually writes is NOT folded back in — the snapshot we just
// took has no music. Say so instead of leaving a silent film that the storyboard claims has a
// bed (observed live: the caller had to notice on its own and rebuild the entry).
if (meta.bgm_pending && !meta.bgm) {
console.warn(
"⚠ audio fetch-sfx: a detached BGM generate is still pending, so this snapshot has no bed. " +
"Re-run `fetch-sfx` (or re-point audio_meta.json at the track) once it lands, before assembling.",
);
}
}
// ── sync-durations (local; rewrites STORYBOARD.md) ────────────────────────────
function runSyncDurations(argv) {
const die = (m) => {
console.error(`✗ audio sync-durations: ${m}`);
process.exit(1);
};
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const audioMetaPath = resolve(flag(argv, "audio-meta", join(hyperframesDir, "audio_meta.json")));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
if (!existsSync(audioMetaPath)) die(`audio_meta.json not found at ${audioMetaPath}`);
const meta = JSON.parse(readFileSync(audioMetaPath, "utf8"));
const durByFrame = new Map();
for (const v of meta.voices ?? []) {
if (v.frame != null && v.duration_s) durByFrame.set(v.frame, v.duration_s);
}
// Read directly and handle ENOENT here, rather than an existsSync precheck —
// the check→write pair (write-back below) is a TOCTOU race CodeQL flags.
let storyboardRaw = "";
try {
storyboardRaw = readFileSync(storyboardPath, "utf8");
} catch {
die(`STORYBOARD.md not found at ${storyboardPath}`);
}
const lines = storyboardRaw.split(/\r?\n/);
const FRAME_RE = /^#{2,3}\s+(?:frame|beat|scene)\b.*?(\d+)/i;
let curFrame = null;
let updated = 0;
for (let i = 0; i < lines.length; i++) {
const h = lines[i].match(FRAME_RE);
if (h) {
curFrame = Number(h[1]);
continue;
}
if (curFrame != null && durByFrame.has(curFrame)) {
const m = lines[i].match(/^(\s*[-*]\s+duration\s*:\s*).*/i);
if (m) {
lines[i] = `${m[1]}${durByFrame.get(curFrame)}s`;
durByFrame.delete(curFrame);
updated++;
}
}
}
writeFileSync(storyboardPath, lines.join("\n"));
const missing = [...durByFrame.keys()];
console.log(
`✓ audio sync-durations: ${updated} frame duration(s) updated` +
(missing.length ? ` · no \`- duration:\` line for frame(s) ${missing.join(", ")}` : ""),
);
}
// ── dispatch ──────────────────────────────────────────────────────────────────
const sub = process.argv[2];
if (sub === "sync-durations") runSyncDurations(process.argv.slice(3));
else if (sub === "fetch-sfx") runFetchSfx(process.argv.slice(3));
else runGenerate(process.argv.slice(2)); // default: generate
scripts/audio.test.mjs
import assert from "node:assert/strict";
import { existsSync, mkdtempSync, readFileSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { spawnSync } from "node:child_process";
import test from "node:test";
const script = new URL("./audio.mjs", import.meta.url).pathname;
function runAudio({ args = [], env = {} } = {}) {
const dir = mkdtempSync(join(tmpdir(), "product-launch-audio-"));
const engine = join(dir, "engine.mjs");
writeFileSync(join(dir, "STORYBOARD.md"), "message: Test\n");
writeFileSync(
engine,
`import { readFileSync, writeFileSync } from "node:fs";
const argv = process.argv.slice(2);
const flag = (name) => argv[argv.indexOf(name) + 1];
const request = JSON.parse(readFileSync(flag("--request"), "utf8"));
writeFileSync(new URL("request.json", import.meta.url), JSON.stringify(request));
writeFileSync(flag("--out"), JSON.stringify({ voices: [], bgm: null, sfx: [] }));
`,
);
const result = spawnSync(
process.execPath,
[script, "--hyperframes", dir, "--storyboard", join(dir, "STORYBOARD.md"), ...args],
{ encoding: "utf8", env: { ...process.env, HF_MEDIA_ENGINE: engine, ...env } },
);
assert.equal(result.status, 0, result.stderr);
return JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
}
test("passes --provider to the shared audio engine", () => {
assert.equal(runAudio({ args: ["--provider", "kokoro"] }).provider, "kokoro");
});
test("uses HF_TTS_PROVIDER when --provider is omitted", () => {
assert.equal(runAudio({ env: { HF_TTS_PROVIDER: "elevenlabs" } }).provider, "elevenlabs");
});
test("--provider takes precedence over HF_TTS_PROVIDER", () => {
assert.equal(
runAudio({ args: ["--provider", "kokoro"], env: { HF_TTS_PROVIDER: "elevenlabs" } }).provider,
"kokoro",
);
});
// ── the canonical fully-silent marker (SKILL.md Step 3.1) ────────────────────
// `music: none` in the storyboard's top YAML block + no SCRIPT.md marks the
// project fully silent: generate must produce nothing (an absent
// audio_meta.json is what assemble treats as silent) and clear stale meta.
/** Like runAudio, but with a caller-controlled storyboard and no request assertion. */
function runAudioRaw({ storyboard, scriptMd = null, preexistingMeta = null }) {
const dir = mkdtempSync(join(tmpdir(), "product-launch-audio-"));
const engine = join(dir, "engine.mjs");
writeFileSync(join(dir, "STORYBOARD.md"), storyboard);
if (scriptMd != null) writeFileSync(join(dir, "SCRIPT.md"), scriptMd);
if (preexistingMeta != null) writeFileSync(join(dir, "audio_meta.json"), preexistingMeta);
writeFileSync(
engine,
`import { readFileSync, writeFileSync } from "node:fs";
const argv = process.argv.slice(2);
const flag = (name) => argv[argv.indexOf(name) + 1];
const request = JSON.parse(readFileSync(flag("--request"), "utf8"));
writeFileSync(new URL("request.json", import.meta.url), JSON.stringify(request));
writeFileSync(flag("--out"), JSON.stringify({ voices: [], bgm: null, sfx: [] }));
`,
);
const result = spawnSync(
process.execPath,
[script, "--hyperframes", dir, "--storyboard", join(dir, "STORYBOARD.md")],
{ encoding: "utf8", env: { ...process.env, HF_MEDIA_ENGINE: engine } },
);
return { dir, result };
}
test("music: none + no SCRIPT.md = fully silent: no engine run, no audio_meta.json", () => {
const { dir, result } = runAudioRaw({ storyboard: "---\nmessage: Test\nmusic: none\n---\n" });
assert.equal(result.status, 0, result.stderr);
assert.match(result.stdout, /marked silent/);
// The engine was never invoked...
assert.equal(existsSync(join(dir, "request.json")), false);
// ...and no meta exists (absence ⇒ assemble treats the film as silent).
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
});
test("fully-silent run removes stale audio_meta.json from a previous non-silent run", () => {
const { dir, result } = runAudioRaw({
storyboard: "---\nmessage: Test\nmusic: none\n---\n",
preexistingMeta: JSON.stringify({ bgm: { path: "old.mp3" }, voices: [], sfx: [] }),
});
assert.equal(result.status, 0, result.stderr);
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
});
test("music: none with narration keeps TTS but turns BGM off (not fully silent)", () => {
const { dir, result } = runAudioRaw({
storyboard: "---\nmessage: Test\nmusic: none\n---\n",
scriptMd: "## Hook (Frame 1)\n\n Spoken line for frame one.\n",
});
assert.equal(result.status, 0, result.stderr);
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
assert.equal(request.bgm.mode, "none");
assert.equal(request.lines.length, 1);
});
test('a quoted music: "none" is still the silent marker (frontmatter stripQuotes)', () => {
// YAML authors quote scalars freely; the vendored storyboard parser strips
// matching quotes at parse time (storyboard.mjs stripQuotes), so the marker
// must not depend on the unquoted spelling.
const { dir, result } = runAudioRaw({
storyboard: '---\nmessage: Test\nmusic: "none"\n---\n',
});
assert.equal(result.status, 0, result.stderr);
assert.match(result.stdout, /marked silent/);
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
});
test("a storyboard music mood still retrieves BGM (marker is exact, not fuzzy)", () => {
const { dir, result } = runAudioRaw({
storyboard: "---\nmessage: Test\nmusic: upbeat synthwave with heavy drums\n---\n",
});
assert.equal(result.status, 0, result.stderr);
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
assert.equal(request.bgm.mode, "retrieve");
assert.equal(request.bgm.query, "upbeat synthwave with heavy drums");
});
// ── fetch-sfx ────────────────────────────────────────────────────────────────
// Regressions found while running the full product-launch workflow end to end
// (linear.app site showcase, 2026-07-30).
/** Runs the fetch-sfx subcommand. `neutralOut` is what the stub engine writes to --out. */
function runFetchSfx({ storyboard, neutralOut = { voices: [], bgm: null, sfx: [] } }) {
const dir = mkdtempSync(join(tmpdir(), "product-launch-sfx-"));
const engine = join(dir, "engine.mjs");
writeFileSync(join(dir, "STORYBOARD.md"), storyboard);
writeFileSync(
engine,
`import { readFileSync, writeFileSync } from "node:fs";
const argv = process.argv.slice(2);
const flag = (name) => argv[argv.indexOf(name) + 1];
const request = JSON.parse(readFileSync(flag("--request"), "utf8"));
writeFileSync(new URL("request.json", import.meta.url), JSON.stringify(request));
writeFileSync(flag("--out"), ${JSON.stringify(JSON.stringify(neutralOut))});
`,
);
const result = spawnSync(
process.execPath,
[script, "fetch-sfx", "--hyperframes", dir, "--storyboard", join(dir, "STORYBOARD.md")],
{ encoding: "utf8", env: { ...process.env, HF_MEDIA_ENGINE: engine } },
);
return { dir, result };
}
const FRAME_WITH_SFX = (sfx) =>
`---\nmessage: Test\n---\n\n## Frame 1 — Hook\n- duration: 3s\n- sfx: ${sfx}\n`;
test("fetch-sfx: `sfx: none` is an absence marker, not a cue named none", () => {
const { dir, result } = runFetchSfx({ storyboard: FRAME_WITH_SFX("none") });
assert.equal(result.status, 0, result.stderr);
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
// Used to reach the engine as { sfx: ["none"] } — a cue that cannot resolve.
assert.deepEqual(request.lines, []);
});
test("fetch-sfx: the other absence spellings are markers too", () => {
for (const spelling of ["None", "n/a", "NA", "skip", "-", "—"]) {
const { dir, result } = runFetchSfx({ storyboard: FRAME_WITH_SFX(spelling) });
assert.equal(result.status, 0, result.stderr);
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
assert.deepEqual(request.lines, [], `spelling: ${spelling}`);
}
});
test("fetch-sfx: a real cue still reaches the engine, and mixed lists drop only the marker", () => {
const { dir, result } = runFetchSfx({ storyboard: FRAME_WITH_SFX("whoosh, none, click") });
assert.equal(result.status, 0, result.stderr);
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
assert.deepEqual(request.lines, [{ id: "01", sfx: ["whoosh", "click"] }]);
});
test("fetch-sfx: carries bgm_pending through and warns that the snapshot has no bed", () => {
const { dir, result } = runFetchSfx({
storyboard: FRAME_WITH_SFX("whoosh"),
// A detached Lyria/MusicGen generate that has not landed yet.
neutralOut: { voices: [], bgm: null, bgm_pending: true, sfx: [] },
});
assert.equal(result.status, 0, result.stderr);
const meta = JSON.parse(readFileSync(join(dir, "audio_meta.json"), "utf8"));
// The flag used to be dropped in the neutral → PL translation, making "not ready yet"
// indistinguishable from "silent by design".
assert.equal(meta.bgm_pending, true);
assert.equal(meta.bgm, null);
assert.match(result.stderr + result.stdout, /pending/i);
});
test("fetch-sfx: a resolved bed reports bgm_pending false and no warning", () => {
const { dir, result } = runFetchSfx({
storyboard: FRAME_WITH_SFX("whoosh"),
neutralOut: { voices: [], bgm: { path: "assets/bgm/track.mp3", volume: 0.12 }, sfx: [] },
});
assert.equal(result.status, 0, result.stderr);
const meta = JSON.parse(readFileSync(join(dir, "audio_meta.json"), "utf8"));
assert.equal(meta.bgm_pending, false);
assert.equal(meta.bgm.path, "assets/bgm/track.mp3");
assert.doesNotMatch(result.stderr, /pending/i);
});
scripts/build-frame.mjs
#!/usr/bin/env node
// build-frame.mjs — Step 2 design system in ONE command. The LLM only chooses a
// preset; this does the deterministic rest: copy the preset's FRAME.md → frame.md,
// remix its colors/typography onto the project's brand tokens, copy the preset's
// caption-skin.html, and self-validate. "Strict on brand" is deterministic, so it's
// a script, not LLM hand-editing (which mis-copies hex / breaks keys).
//
// node build-frame.mjs --preset capsule --hyperframes .
// [--tokens capture/extracted/tokens.json] [--preset-dir <abs path to frame-presets>]
//
// Remix rule — ONLY `colors:` values and `typography:` fontFamily change; keys,
// structure, geometry, and components are untouched:
// colors — map brand tokens onto the preset's keys BY ROLE: the ink-role key takes
// the brand ink (darkest/ink-named), the canvas-role key takes the brand
// canvas (lightest), and every other color is repainted with the nearest
// brand accent's hue+saturation while KEEPING its own lightness, so tint
// families (sun / sun-soft / haze) stay a family. Empty brand colors → the
// preset palette is kept (it is already a complete, good design).
// fonts — the preset's display family → the brand display font, its body family →
// the brand body font, wherever they appear. Empty brand fonts → kept.
import {
copyFileSync,
existsSync,
mkdirSync,
readdirSync,
readFileSync,
writeFileSync,
} from "node:fs";
import { dirname, join, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import {
brandRolesFromStats,
chroma,
isIconFont,
lum,
parseColors,
parseFonts,
pickAccent,
semanticColors,
STATUS_ROLE_KEY,
UA_DEFAULT_COLORS,
} from "./lib/tokens.mjs";
const __dirname = dirname(fileURLToPath(import.meta.url));
const argv = process.argv.slice(2);
const flag = (name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
const die = (m) => {
console.error(`✗ build-frame: ${m}`);
process.exit(1);
};
const presetName = flag("preset", null);
const hyperframesDir = resolve(flag("hyperframes", "."));
const presetDir = resolve(
flag("preset-dir", join(__dirname, "../../hyperframes-creative/frame-presets")),
);
const tokensPath = resolve(flag("tokens", join(hyperframesDir, "capture/extracted/tokens.json")));
if (!presetName) die("--preset <name> is required");
const presetFrame = join(presetDir, presetName, "FRAME.md");
if (!existsSync(presetFrame)) {
const avail = existsSync(presetDir)
? readdirSync(presetDir, { withFileTypes: true })
.filter((d) => d.isDirectory())
.map((d) => d.name)
: [];
die(
`no FRAME.md for preset "${presetName}" under ${presetDir}\n available: ${avail.join(", ")}`,
);
}
// ── HSL helpers (recolor = brand hue+sat, original lightness) ──────────────────
function hexToHsl(hex) {
const m = /^#?([0-9a-fA-F]{6})$/.exec(String(hex).trim());
if (!m) return null;
const n = parseInt(m[1], 16);
const r = ((n >> 16) & 255) / 255,
g = ((n >> 8) & 255) / 255,
b = (n & 255) / 255;
const max = Math.max(r, g, b),
min = Math.min(r, g, b),
d = max - min;
let h = 0;
const l = (max + min) / 2;
const s = d === 0 ? 0 : l > 0.5 ? d / (2 - max - min) : d / (max + min);
if (d !== 0) {
h = max === r ? (g - b) / d + (g < b ? 6 : 0) : max === g ? (b - r) / d + 2 : (r - g) / d + 4;
h *= 60;
}
return { h, s, l };
}
function hslToHex(h, s, l) {
h = (((h % 360) + 360) % 360) / 360;
const hue = (p, q, t) => {
t = (t + 1) % 1;
if (t < 1 / 6) return p + (q - p) * 6 * t;
if (t < 1 / 2) return q;
if (t < 2 / 3) return p + (q - p) * (2 / 3 - t) * 6;
return p;
};
let r, g, b;
if (s === 0) {
r = g = b = l;
} else {
const q = l < 0.5 ? l * (1 + s) : l + s - l * s;
const p = 2 * l - q;
r = hue(p, q, h + 1 / 3);
g = hue(p, q, h);
b = hue(p, q, h - 1 / 3);
}
const to = (x) =>
Math.round(x * 255)
.toString(16)
.padStart(2, "0")
.toUpperCase();
return `#${to(r)}${to(g)}${to(b)}`;
}
const hueDist = (a, b) => {
const d = Math.abs(a - b) % 360;
return d > 180 ? 360 - d : d;
};
function hexToRgb(hex) {
const m = /^#?([0-9a-fA-F]{6})$/.exec(String(hex).trim());
if (!m) return null;
const n = parseInt(m[1], 16);
return [(n >> 16) & 255, (n >> 8) & 255, n & 255];
}
const rgbToHsl = (r, g, b) =>
hexToHsl("#" + [r, g, b].map((x) => Math.round(x).toString(16).padStart(2, "0")).join(""));
// Repaint a chromatic rgba()/rgb() tint with the brand accent's RGB, keeping its alpha.
// A near-neutral rgb (shadow / scrim overlay) is left untouched; a non-rgba string → null.
function remapRgbaToAccent(val, brAccent, brAccent2, prAccentHsl, prAccent2Hsl) {
const m = /^rgba?\(\s*([\d.]+)[\s,]+([\d.]+)[\s,]+([\d.]+)\s*(?:[,/]\s*([\d.]+%?))?\s*\)$/i.exec(
String(val).trim(),
);
if (!m) return null;
const r = +m[1],
g = +m[2],
b = +m[3],
a = m[4];
if (Math.max(r, g, b) - Math.min(r, g, b) < 16) return null; // neutral overlay — keep as-is
const src = rgbToHsl(r, g, b);
const useSecond =
brAccent2 &&
prAccentHsl &&
prAccent2Hsl &&
src &&
hueDist(src.h, prAccent2Hsl.h) < hueDist(src.h, prAccentHsl.h);
const t = hexToRgb(useSecond ? brAccent2 : brAccent);
if (!t) return null;
return a !== undefined
? `rgba(${t[0]}, ${t[1]}, ${t[2]}, ${a})`
: `rgb(${t[0]}, ${t[1]}, ${t[2]})`;
}
// ── brand tokens ──────────────────────────────────────────────────────────────
let brandColors = [];
let brandFonts = [];
let brandFontWeights = []; // weights the brand text font actually ships (tokens fonts[].weights)
let brandColorStats = []; // rich per-color usage stats (areaBg / interactiveBg / textCount …)
// Icon/glyph fonts capture surfaces as "fonts" — they are never the brand text face
// (webflow-icons, Font Awesome, icomoon …) and must not become display/body or contribute weights.
if (existsSync(tokensPath)) {
try {
const t = JSON.parse(readFileSync(tokensPath, "utf8"));
brandColors = (t.colors ?? [])
.map((c) => (typeof c === "string" ? c : (c?.hex ?? c?.value ?? "")))
.map((c) => String(c).trim())
.filter((c) => /^#?[0-9a-fA-F]{6}$/.test(c))
.map((c) => (c.startsWith("#") ? c : `#${c}`));
brandFonts = (t.fonts ?? [])
.map((f) => (typeof f === "string" ? f : (f?.family ?? f?.name ?? "")))
.map((f) => String(f).split(",")[0].replace(/['"]/g, "").trim())
.filter(Boolean)
.filter((f) => !isIconFont(f));
// Union of the (non-icon) brand fonts' available weights — used to clamp the preset's
// type ramp so a font shipping only 400/500 never faux-bolds a 600/700 heading.
brandFontWeights = [
...new Set(
(t.fonts ?? [])
.filter((f) => f && typeof f === "object" && !isIconFont(f.family ?? f.name ?? ""))
.flatMap((f) => (Array.isArray(f.weights) ? f.weights : []))
.map((w) => parseInt(w, 10))
.filter((w) => Number.isFinite(w)),
),
].sort((a, b) => a - b);
brandColorStats = Array.isArray(t.colorStats) ? t.colorStats : [];
} catch (e) {
die(`tokens.json parse: ${e.message}`);
}
}
let md = readFileSync(presetFrame, "utf8");
const presetColors = parseColors(md);
const summary = [];
// ── color remix ───────────────────────────────────────────────────────────────
if (brandColors.length && presetColors.length) {
const pr = semanticColors(presetColors);
// Brand roles: prefer the function-based reading of capture colorStats (canvas =
// largest background, accent = top interactive bg, ink = dominant contrasting text).
// Fall back to the legacy luminance/chroma heuristic only when stats are absent —
// but pick the accent via pickAccent either way so a UA-default link color never wins.
const br =
brandRolesFromStats(brandColorStats, brandColors) ??
(() => {
// strip UA-default link colors so a stray <a> color can't become ink/canvas/accent
const clean = brandColors.filter((h) => !UA_DEFAULT_COLORS.has(h.toUpperCase()));
const s = semanticColors(clean.map((h, i) => [`c${i}`, h]));
return {
ink: s.ink,
canvas: s.canvas,
accent: pickAccent(brandColorStats, clean, [s.ink, s.canvas]) ?? s.accent,
accent2: s.accent2,
};
})();
if (!br.accent) die("accent 选取失败:品牌色里没有可用的强调色");
if (chroma(br.accent) <= 40) {
console.warn(
` ⚠ accent ${br.accent} 彩度很低 (${chroma(br.accent)}) — 确认这是品牌色而非中性/默认色`,
);
}
// Map by LUMINANCE POLARITY. The preset's darker value takes the brand's darker value and
// the lighter takes the lighter — UNLESS the brand's GROUND polarity differs from the
// preset's. Every shipped preset is light-ground; a dark-mode brand (Linear, Vercel,
// Raycast…) has its canvas darker than its ink (colorStats already resolved the real
// ground as the largest-area background). On a polarity MISMATCH we INVERT the mapping so a
// light preset becomes the dark brand (canvas↔ink swap) instead of forcing the brand onto
// an off-brand light video; neutral/tint lightness is then mirrored (L→1−L) so the whole
// palette flips to the brand's ground. Same-polarity (the common case) is unchanged.
const darker = (a, b) => ((lum(a) ?? 0) <= (lum(b) ?? 0) ? a : b);
const prDark = darker(pr.ink, pr.canvas);
const prLight = prDark === pr.ink ? pr.canvas : pr.ink;
const brDark = darker(br.ink, br.canvas);
const brLight = brDark === br.ink ? br.canvas : br.ink;
const presetGroundDark = (lum(pr.canvas) ?? 255) < (lum(pr.ink) ?? 0);
const brandGroundDark = (lum(br.canvas) ?? 255) < (lum(br.ink) ?? 0);
const invert = presetGroundDark !== brandGroundDark;
const mapDark = invert ? brLight : brDark; // preset's dark value → this brand value
const mapLight = invert ? brDark : brLight; // preset's light value → this brand value
const flipL = (l) => (invert ? 1 - l : l); // mirror tint/neutral lightness when flipping
const prAccentHsl = hexToHsl(pr.accent);
const prAccent2Hsl = hexToHsl(pr.accent2);
const newByKey = new Map();
for (const [key, val] of presetColors) {
const ph = hexToHsl(val);
let next;
if (val === prDark) next = mapDark;
else if (val === prLight) next = mapLight;
else if (STATUS_ROLE_KEY.test(key))
// semantic status colors (green/red …) — the HUE carries the meaning; never repaint.
// MUST precede the accent checks: a preset's red "negative" is often its 2nd-most-chromatic
// color and would otherwise be claimed as accent2 and recolored to the brand hue.
next = val;
else if (val === pr.accent)
next = br.accent; // primary accent → the EXACT brand color
else if (pr.accent2 !== pr.accent && val === pr.accent2)
next = br.accent2; // exact 2nd accent
else if (!ph) {
// rgba()/rgb() tint → repaint its rgb with the brand accent, keep alpha (a neutral
// overlay is kept). A non-color non-hex value (var(), named) falls through unchanged.
next = remapRgbaToAccent(val, br.accent, br.accent2, prAccentHsl, prAccent2Hsl) ?? val;
} else if (chroma(val) < 16) {
// NEUTRAL source (grey text-ladder, hairline borders) → keep it NEUTRAL. Apply at most a
// whisper of the brand hue (sat ≤ 0.06); never the accent's full saturation — that is what
// turned the grey ladder into saturated blue.
const bh = hexToHsl(br.accent);
next = bh ? hslToHex(bh.h, Math.min(ph.s, 0.06), flipL(ph.l)) : val;
} else {
// chromatic tint → repaint with the nearest brand accent's hue+sat, keep THIS color's
// lightness so tint families stay families.
const useSecond =
pr.accent !== pr.accent2 &&
prAccentHsl &&
prAccent2Hsl &&
hueDist(ph.h, prAccent2Hsl.h) < hueDist(ph.h, prAccentHsl.h);
const bh = hexToHsl(useSecond ? br.accent2 : br.accent);
next = bh ? hslToHex(bh.h, bh.s, flipL(ph.l)) : val;
}
if (next !== val) newByKey.set(key, next);
}
// rewrite only the value of each colors: line; everything else byte-identical.
let inBlock = false;
md = md
.split(/\r?\n/)
.map((line) => {
if (/^colors:\s*$/.test(line)) {
inBlock = true;
return line;
}
if (inBlock && /^\S/.test(line)) inBlock = false;
if (!inBlock) return line;
const m = line.match(
/^(\s+)([\w-]+):\s*(?:"[^"]*"|'[^']*'|#[0-9a-fA-F]{3,8}|rgba?\([^)]*\)|[^#\n]*?)(\s+#.*)?$/,
);
if (m && newByKey.has(m[2])) return `${m[1]}${m[2]}: "${newByKey.get(m[2])}"${m[3] ?? ""}`;
return line;
})
.join("\n");
summary.push(
`colors: ${invert ? "INVERTED (dark-mode brand on light preset) · " : ""}dark ${prDark}→${mapDark}, light ${prLight}→${mapLight}, accent ${pr.accent}→${br.accent}` +
` (${newByKey.size}/${presetColors.length} keys repainted${brandColorStats.length ? ", via colorStats" : ""})`,
);
} else {
summary.push(
brandColors.length
? "colors: preset has no parseable colors — kept"
: "colors: no brand colors — preset palette kept",
);
}
// ── font remix ────────────────────────────────────────────────────────────────
if (brandFonts.length) {
const pf = parseFonts(md);
const strip = (q) => (q ? q.replace(/^"|"$/g, "") : null);
const pDisplay = strip(pf.display);
const pBody = strip(pf.body);
const pMono = strip(pf.mono);
// A monospace brand face is for code / labels / chrome — never the reading display or body.
// Split the brand fonts: the primary NON-mono family carries display AND body (the common
// single-sans case, e.g. Inter for everything), and a captured mono (Berkeley Mono,
// JetBrains Mono…) is routed onto the preset's mono role instead of turning the body
// monospace. (Distinct display/body brands still resolve to a clean sans; hand-tune the
// display in frame.md if a separate display face is wanted.)
const isMonoFont = (n) =>
/(?:^|[\s_-])mono(?:[\s_-]|$)|monospace|consol|courier|menlo|monaco|jetbrains|berkeley|space\s*mono|ibm\s*plex\s*mono|sf\s*mono|roboto\s*mono|source\s*code|fira\s*code|geist\s*mono|dm\s*mono/i.test(
String(n),
);
const nonMono = brandFonts.filter((f) => !isMonoFont(f));
const monoFonts = brandFonts.filter(isMonoFont);
const bDisplay = nonMono[0] ?? brandFonts[0];
const bBody = nonMono[0] ?? brandFonts[0];
const bMono = monoFonts[0] ?? null;
const escRe = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
// Replace the preset family as a WHOLE WORD/PHRASE everywhere — frontmatter values,
// component strings like "Space Grotesk 600", AND prose — case-sensitive with word
// boundaries so a single-word family ("Inter") can never corrupt a substring
// ("interactive"). Quote-exact replace alone missed names baked into longer strings + prose.
const swapFamily = (from, to) => {
if (from && to && from !== to) md = md.replace(new RegExp(`\\b${escRe(from)}\\b`, "g"), to);
};
swapFamily(pDisplay, bDisplay);
if (pBody !== pDisplay) swapFamily(pBody, bBody);
// route the brand mono onto the preset's mono role (only if the preset has a DISTINCT mono
// family — never collapse body/display into mono)
if (bMono && pMono && pMono !== pBody && pMono !== pDisplay) swapFamily(pMono, bMono);
summary.push(
`fonts: display ${pDisplay}→${bDisplay}, body ${pBody}→${bBody}` +
(bMono && pMono && pMono !== pBody && pMono !== pDisplay ? `, mono ${pMono}→${bMono}` : ""),
);
} else {
summary.push("fonts: no brand fonts — preset fonts kept");
}
// ── cap type weights to the brand font's available faces ──────────────────────
// The remix swaps the font FAMILY but keeps the preset's weights; a brand font that ships
// only e.g. 400/500 would faux-bold every 600/700 heading. Clamp each `typography:` weight
// to the NEAREST weight the brand font actually provides (tokens.json fonts[].weights).
if (brandFonts.length && brandFontWeights.length) {
const avail = brandFontWeights;
const nearest = (n) =>
avail.reduce((best, w) => {
const dw = Math.abs(w - n),
db = Math.abs(best - n);
return dw < db || (dw === db && w > best) ? w : best;
}, avail[0]);
let capped = 0;
const cap = (num) => {
const n = parseInt(num, 10);
if (avail.includes(n)) return String(n);
const c = nearest(n);
if (c !== n) capped++;
return String(c);
};
let inType = false;
md = md
.split(/\r?\n/)
.map((line) => {
if (/^typography:\s*$/.test(line)) {
inType = true;
return line;
}
if (inType && /^\S/.test(line)) inType = false;
let out = line;
// (a) structured `weight: NNN` in the typography ramp
if (inType) out = out.replace(/(\bweight:\s*)(\d{3})\b/g, (m, pfx, num) => pfx + cap(num));
// (b) a weight baked into a quoted `typography:` component value, e.g.
// cta-button → typography: "Basier Square 600" (NNN not followed by a unit like px)
out = out.replace(
/(typography:\s*"[^"]*?\b)(\d{3})\b(?![a-z%])/gi,
(m, pfx, num) => pfx + cap(num),
);
return out;
})
.join("\n");
if (capped)
summary.push(`fonts: capped ${capped} type weight(s) to brand faces {${avail.join(", ")}}`);
}
// ── brand-adaptation note ─────────────────────────────────────────────────────
// The remix fixes the NORMATIVE frontmatter, but the preset's PROSE still carries its
// original weight ranges / color-names. Prepend a short "frontmatter is truth" header so a
// reader (or frame worker) interprets any lingering preset prose THROUGH the brand values —
// instead of fragile per-sentence prose surgery.
if (brandFonts.length || (brandColors.length && presetColors.length)) {
const bD = brandFonts[0];
const bB = brandFonts[1] ?? brandFonts[0];
const note =
`## Brand adaptation (READ FIRST — the frontmatter is the source of truth)\n\n` +
`This is the **${presetName}** preset remixed onto the captured brand. The YAML frontmatter above ` +
`(colors · typography · components) is **normative and already correct — use it verbatim.** The prose ` +
`below is the ORIGINAL preset's intent; read it THROUGH the frontmatter:\n\n` +
(brandFonts.length
? `- **Fonts** — already set to **${bD}** (display) / **${bB}** (body); ignore any preset font name lingering in prose.\n`
: "") +
(brandFontWeights.length
? `- **Weights** — the brand font ships \`{${brandFontWeights.join(", ")}}\` only; every weight is clamped to these — ignore higher preset weights (e.g. 600/700) in prose.\n`
: "") +
`- **Colors** — use the frontmatter hex; preset color NAMES in prose (e.g. "cobalt", "cream") mean the remapped brand values.\n`;
if (/^# .*$/m.test(md)) md = md.replace(/^# .*$/m, (m) => `${m}\n\n${note}`);
else md = `${note}\n${md}`;
summary.push("brand-adaptation note prepended");
}
// ── stage brand font files + emit @font-face ──────────────────────────────────
// A brand font is rarely a Google font, so renaming the family in frame.md is not enough:
// nothing loads the actual face. If the capture downloaded font files, copy them to
// assets/fonts/ under CLEAN, face-named names (so captions.mjs' family-prefix matcher
// finds them too) and append a ready-to-paste, ROOT-RELATIVE @font-face block to frame.md.
//
// The staged NAME is a contract, not cosmetics: captions.mjs derives each face's weight and
// style back out of it. So the name has to carry every axis that distinguishes one face from
// another, and the dedup key has to be the whole face. Naming on weight alone made Google's
// two-file Newsreader download (upright + italic, both scoring "Regular") collide on one
// slot: the italic sorts first, took the name, the upright was never staged, and the block
// below then asserted font-style:normal over italic bytes.
if (brandFonts.length) {
const norm = (s) =>
String(s)
.toLowerCase()
.replace(/[^a-z0-9]/g, "");
const extOf = (f) => (f.match(/\.(woff2|woff|ttf|otf)$/i)?.[1] ?? "").toLowerCase();
const FMT = { woff2: "woff2", woff: "woff", ttf: "truetype", otf: "opentype" };
const weightInfo = (name) => {
const s = name.toLowerCase();
// A numeric axis is the font's own answer, so it beats the word heuristic. Fontsource
// names every face that way and carries no weight WORD at all, so word-only parsing
// scored a whole family "Regular" and staged exactly one of its faces.
//
// A weight token must not be buried inside a longer run: this reads capture files,
// which are commonly hash-named, and "Newsreader-a1b200c3.woff2" is not a 200-weight
// face. Hence a non-digit before (which also stops "2100" reading as 100) and no
// alphanumeric after. "Roboto900.ttf" still parses.
const numeric = /(?:^|[^0-9])([1-9]00)(?![0-9a-z])/.exec(s);
if (numeric) return { n: Number(numeric[1]), w: numeric[1] };
if (/black|heavy|ultra|extrabold/.test(s)) return { n: 800, w: "ExtraBold" };
if (/semibold|demibold/.test(s)) return { n: 600, w: "SemiBold" };
if (/bold/.test(s)) return { n: 700, w: "Bold" };
if (/medium/.test(s)) return { n: 500, w: "Medium" };
if (/light|thin/.test(s)) return { n: 300, w: "Light" };
return { n: 400, w: "Regular" };
};
const styleOf = (name) => (/italic|oblique/i.test(name) ? "italic" : "normal");
const fams = [...new Set(brandFonts)];
const srcDirs = [
join(hyperframesDir, "capture/assets/fonts"),
join(hyperframesDir, "assets/fonts"),
].filter((d) => existsSync(d));
const files = [];
for (const d of srcDirs)
for (const f of readdirSync(d).sort()) if (extOf(f)) files.push({ d, f });
// Single family → all font files belong to it (the common captured case, hash-named files
// included). Multiple families → assign each file to the longest family key its name contains.
const ranked = [...fams].sort((a, b) => norm(b).length - norm(a).length);
const famOf = (f) =>
fams.length === 1 ? fams[0] : ranked.find((x) => norm(f).includes(norm(x)));
const outDir = join(hyperframesDir, "assets/fonts");
const faces = [];
const stagedNames = new Set();
for (const { d, f } of files) {
const fam = famOf(f);
if (!fam) continue;
const { n, w } = weightInfo(f);
const style = styleOf(f);
const clean = `${fam.replace(/[^A-Za-z0-9]/g, "")}-${w}${style === "italic" ? "-Italic" : ""}.${extOf(f)}`;
if (stagedNames.has(clean)) continue;
mkdirSync(outDir, { recursive: true });
if (!existsSync(join(outDir, clean))) copyFileSync(join(d, f), join(outDir, clean));
stagedNames.add(clean);
faces.push(
`@font-face{font-family:"${fam}";font-weight:${n};font-style:${style};font-display:block;src:url("assets/fonts/${clean}") format("${FMT[extOf(f)]}");}`,
);
}
if (faces.length) {
md +=
`\n\n## Font loading (auto-generated)\n\n` +
`The brand font ships as local files in \`assets/fonts/\` — do NOT link Google Fonts for it. ` +
`Paste this \`<style>\` into every frame's \`<head>\`/\`<template>\` (captions use the same files) ` +
`so \`font-family\` resolves in preview, snapshot, and render alike:\n\n` +
"```html\n<style>\n" +
faces.join("\n") +
"\n</style>\n```\n";
summary.push(
`fonts: staged ${stagedNames.size} face(s) → assets/fonts/ + @font-face in frame.md`,
);
}
}
// ── write frame.md ────────────────────────────────────────────────────────────
const framePath = join(hyperframesDir, "frame.md");
writeFileSync(framePath, md);
// ── copy caption-skin.html ────────────────────────────────────────────────────
const presetSkin = join(presetDir, presetName, "caption-skin.html");
let skinCopied = false;
if (existsSync(presetSkin)) {
const skinDir = join(hyperframesDir, ".hyperframes");
mkdirSync(skinDir, { recursive: true });
copyFileSync(presetSkin, join(skinDir, "caption-skin.html"));
skinCopied = true;
}
// ── self-validate ─────────────────────────────────────────────────────────────
const outColors = parseColors(md);
if (outColors.length !== presetColors.length) {
die(`color keys changed (${presetColors.length}→${outColors.length}) — keys must be preserved`);
}
const outRoles = semanticColors(outColors);
const li = lum(outRoles.ink),
lc = lum(outRoles.canvas);
// ink (type) and canvas (ground) must differ enough to READ — in EITHER direction. A
// light-mode spec has ink darker than canvas; a dark-mode spec (the polarity flip above)
// the reverse. Assert luminance SEPARATION, not a fixed polarity.
if (li != null && lc != null && Math.abs(li - lc) < 40) {
die(
`ink (${outRoles.ink}, lum ${li.toFixed(0)}) and canvas (${outRoles.canvas}, lum ${lc.toFixed(0)}) lack contrast — bad brand mapping`,
);
}
console.log(`✓ build-frame: ${presetName} → ${framePath}`);
for (const s of summary) console.log(` ${s}`);
console.log(
` .hyperframes/caption-skin.html: ${skinCopied ? "copied" : "preset ships none — captions will use the default pill"}`,
);
console.log(` self-check: keys preserved, ink/canvas contrast ok ✓`);
scripts/captions.mjs
#!/usr/bin/env node
// captions.mjs — build the captions sub-composition from STORYBOARD + audio_meta.
//
// One mode: `build`. Reads STORYBOARD.md (frame order + durations → cumulative
// frame starts) + audio_meta.json (voices[].words, frame-relative) → absolute-
// timed caption groups → writes:
// compositions/captions.html — a self-contained sub-composition the index
// assembler mounts on its captions track (data-composition-id="captions").
// caption_groups.json — the computed groups (debug / inspection / --out).
// caption-overrides.json — an empty `[]` shim (silences the captions runtime's
// validate-time fetch; only written when captions.html is).
// No narration / no words → legal skip: nothing written, assemble-index then omits
// the captions track (it keys off compositions/captions.html existence).
//
// node captions.mjs build --storyboard ./STORYBOARD.md --audio-meta ./audio_meta.json --hyperframes . --out ./caption_groups.json
//
// CAPTION LOOK — two sources, picked automatically:
// 1. PRESET SKIN (preferred). If a project-local `.hyperframes/caption-skin.html`
// exists (Step 2 copies the chosen frame-preset's skin into the project), it is
// the caption look.
// It is a brand-token-strict skin with three reserved holes; this script fills them
// and wraps the result in a <template> for the engine:
// - `var GROUPS = [];` → the computed caption groups
// - `var DURATION = 0;` + data-duration="0" (and data-width/height="0") → real values
// - `<style data-brand-tokens></style>` → :root tokens derived from the project's
// frame.md (colors + fonts), mapped to a fixed semantic vocab every skin shares:
// --cap-ink / --cap-canvas / --cap-accent / --cap-accent-2 / --font-display /
// --font-body, plus --cap-band-top / --cap-band-height (the keep-out band).
// So the brand-token overlay from Step 2 flows into the captions automatically.
// 2. DEFAULT (fallback). No skin file → the built-in Roboto/black pill (buildCaptionsHtml).
//
// Grouping mirrors the proven heuristics (frame boundary · sentence-end punct ·
// silence gap · density-aware word cap); word timings come inline from audio_meta.
import { existsSync, mkdirSync, readdirSync, readFileSync, writeFileSync } from "node:fs";
import { dirname, join, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { parseStoryboard } from "./lib/storyboard.mjs";
import { captionBand, parseFormat } from "./lib/dimensions.mjs";
import { parseColors, parseFonts, semanticColors } from "./lib/tokens.mjs";
const flag = (argv, name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
const r3 = (x) => Number(x.toFixed(3));
// ── grouping params ───────────────────────────────────────────────────────────
const SILENCE_GAP = 0.18; // s of silence between words → split
const TAIL_PAD = 0.12; // s the group lingers after its last word
const SENT_END = /[.?!,;:—]$/;
const DENSITY_WINDOW = 1.0; // s window for words/sec density
function wordCap(density) {
return density > 3.5 ? 2 : density > 2.5 ? 3 : 4;
}
function runBuild(argv) {
const skip = (reason) => {
console.log(`captions: skipped (${reason})`);
process.exit(0);
};
const die = (m) => {
console.error(`✗ captions build: ${m}`);
process.exit(1);
};
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
const audioMetaPath = resolve(flag(argv, "audio-meta", join(hyperframesDir, "audio_meta.json")));
const outPath = resolve(flag(argv, "out", join(hyperframesDir, "caption_groups.json")));
const htmlPath = join(hyperframesDir, "compositions/captions.html");
const overridesPath = join(hyperframesDir, "caption-overrides.json");
const skinArg = flag(argv, "skin", null);
const hiddenSkinPath = join(hyperframesDir, ".hyperframes", "caption-skin.html");
const legacySkinPath = join(hyperframesDir, "caption-skin.html");
const skinPath = resolve(
skinArg ?? (existsSync(hiddenSkinPath) ? hiddenSkinPath : legacySkinPath),
);
const framePath = resolve(flag(argv, "frame", join(hyperframesDir, "frame.md")));
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const { width: W, height: H } = parseFormat(manifest.globals.format);
if (!existsSync(audioMetaPath)) skip("no audio_meta.json (silent film)");
const meta = JSON.parse(readFileSync(audioMetaPath, "utf8"));
if (!Array.isArray(meta.voices) || meta.voices.length === 0) skip("no narration");
// cumulative frame starts (by frame number) + total duration, from STORYBOARD.
const startByFrame = new Map();
let acc = 0;
for (const f of manifest.frames) {
if (f.number != null) startByFrame.set(f.number, acc);
acc += Number.isFinite(f.durationSeconds) ? f.durationSeconds : 0;
}
const total = r3(acc);
// absolute word stream: frame start + frame-relative word timing.
const words = [];
for (const v of meta.voices) {
const base = startByFrame.get(v.frame);
if (base == null || !Array.isArray(v.words)) continue;
for (const w of v.words) {
const text = String(w.text ?? "").trim();
if (!text || /^[.?!,;:—–-]+$/.test(text)) continue; // drop empties + bare punctuation
if (!isFinite(w.start) || !isFinite(w.end)) continue;
words.push({ text, start: r3(base + w.start), end: r3(base + w.end), frame: v.frame });
}
}
words.sort((a, b) => a.start - b.start);
if (words.length === 0) skip("no usable words");
// density at i = words whose start falls within [w.start, w.start + WINDOW).
const densityAt = (i) => {
const t0 = words[i].start;
let n = 0;
for (let j = i; j < words.length && words[j].start < t0 + DENSITY_WINDOW; j++) n++;
return n / DENSITY_WINDOW;
};
// group: split on frame change / silence gap / word cap; always flush after a
// sentence-ending word.
const groups = [];
let cur = null;
for (let i = 0; i < words.length; i++) {
const w = words[i];
const prev = cur && cur.words[cur.words.length - 1];
const crossFrame = cur && w.frame !== cur.frame;
const gap = prev && w.start - prev.end > SILENCE_GAP;
const full = cur && cur.words.length >= cur.cap;
if (!cur || crossFrame || gap || full) {
if (cur) groups.push(cur);
cur = { frame: w.frame, cap: wordCap(densityAt(i)), words: [] };
}
cur.words.push(w);
if (SENT_END.test(w.text)) {
groups.push(cur);
cur = null;
}
}
if (cur) groups.push(cur);
// finalize: ids, start/end (tail-padded, clamped < next group's start), text.
const finalized = groups.map((g, gi) => {
const first = g.words[0];
const last = g.words[g.words.length - 1];
const next = groups[gi + 1];
let end = r3(last.end + TAIL_PAD);
if (next && next.words[0].start < end) end = r3(next.words[0].start);
return {
id: `caption-group-${gi}`,
frame: g.frame,
start: r3(first.start),
end,
text: g.words.map((w) => w.text).join(" "),
words: g.words.map((w, wi) => ({
id: `caption-word-${gi}-${wi}`,
text: w.text,
start: r3(w.start),
end: r3(w.end),
})),
};
});
// ── write caption_groups.json ──
mkdirSync(dirname(outPath), { recursive: true });
writeFileSync(
outPath,
JSON.stringify({ total_duration_s: total, width: W, height: H, groups: finalized }, null, 2),
);
// ── write compositions/captions.html (preset skin if present, else default) ──
mkdirSync(dirname(htmlPath), { recursive: true });
let source;
if (existsSync(skinPath)) {
const tokens = frameTokensCss(framePath, H);
const faces = brandFontFaces(framePath, hyperframesDir);
const fonts = existsSync(framePath) ? parseFonts(readFileSync(framePath, "utf8")) : {};
writeFileSync(
htmlPath,
buildFromSkin(
readFileSync(skinPath, "utf8"),
finalized,
total,
W,
H,
tokens,
die,
faces,
fonts,
),
);
source = `preset skin (${skinPath.replace(hyperframesDir + "/", "")})`;
} else {
writeFileSync(htmlPath, buildCaptionsHtml(finalized, total, W, H));
source = "default (built-in pill)";
}
// ── write caption-overrides.json shim ──
// Atomic create-if-absent: `wx` throws if the file already exists (which we
// ignore) — no existsSync→writeFileSync TOCTOU gap.
try {
writeFileSync(overridesPath, "[]\n", { flag: "wx" });
} catch {
/* overrides shim already present */
}
console.log(
`✓ captions build: ${finalized.length} group(s) from ${words.length} words → compositions/captions.html (total ${total}s) · skin: ${source}`,
);
}
// ── preset-skin path ────────────────────────────────────────────────────────
// Fill the skin's three reserved holes + the root's 0-placeholders, then wrap the
// fragment in a <template> (the engine clones template contents only). One generic
// fill works for every preset's skin — no per-skin transform.
//
// Every preset's skin is authored against ITS OWN fonts/metrics (broadside→Barlow @
// line-height 1.02, capsule→Bodoni, …). When the project's brand font differs (it
// almost always does), three things must be reconciled so ANY skin renders correctly
// for ANY brand — done here generically, not per-project:
// · @font-face for the brand fonts (else the renderer can't supply them → fallback)
// · the skin's preset-font FALLBACK literals (var(--font-x, "Barlow")) repointed to
// the brand family, so no undeclared font name trips font_family_without_font_face
// · a metric safety net: a heavier brand font overflows a tight preset line-height,
// so the active-word highlight clips — a line-height floor + word padding fixes it
// · data-composition-id + dimensions on the <template> root (skins lead with
// <script>/<style>, so the root element must carry the id, not the first child)
function buildFromSkin(skin, groups, total, W, H, tokens, die, faces = "", fonts = {}) {
const fillOnce = (src, re, repl, label) => {
const n = (src.match(re) || []).length;
if (n !== 1) die(`caption-skin.html: expected exactly one ${label}, found ${n}`);
return src.replace(re, () => repl);
};
let out = skin;
// Strip HTML doc-comments first. A skin's authoring comment can contain tag-like text
// (broadside's literally says "<template>"), which the linter's tag scanner then picks
// up as the root element → false root_missing_composition_id / root_missing_dimensions.
// The comments are preview/authoring docs, not needed in the generated composition.
// Strip in a fixpoint loop, not a single global pass: removing one comment can
// re-form a marker from a nested/partial pair (e.g. <!--<!---->-->), which one
// pass misses — CodeQL flags the single replace as incomplete sanitization.
for (let prev = ""; prev !== out; ) {
prev = out;
out = out.replace(/<!--[\s\S]*?-->/g, "");
}
// brand :root tokens + @font-face for the brand fonts, both into the reserved hole
out = fillOnce(
out,
/<style data-brand-tokens>\s*<\/style>/,
`<style data-brand-tokens>\n${faces ? faces + "\n" : ""}${tokens}\n </style>`,
"<style data-brand-tokens></style> hole",
);
// Resolve the skin's font-family var()s to the brand family LITERAL. Two reasons:
// (1) the linter's used-font scanner naively comma-splits, so var(--x, "Brand") yields
// junk tokens ('var(--x', 'brand")') that never match the @font-face → a false
// font_family_without_font_face; a plain "Brand" literal matches the @font-face.
// (2) it drops the preset's own fallback name (Barlow / IBM Plex Mono / …), which has
// no @font-face in this project. The :root token stays for any other consumer.
if (fonts.display)
out = out.replace(/var\(\s*--font-display\s*(?:,\s*"[^"]*"\s*)?\)/g, fonts.display);
if (fonts.body) out = out.replace(/var\(\s*--font-body\s*(?:,\s*"[^"]*"\s*)?\)/g, fonts.body);
out = fillOnce(
out,
/var GROUPS = \[\];/,
`var GROUPS = ${JSON.stringify(groups)};`,
"`var GROUPS = [];` hole",
);
out = fillOnce(out, /var DURATION = 0;/, `var DURATION = ${total};`, "`var DURATION = 0;` hole");
out = fillOnce(out, /data-duration="0"/, `data-duration="${total}"`, '`data-duration="0"` hole');
out = fillOnce(out, /data-width="0"/, `data-width="${W}"`, '`data-width="0"` hole');
out = fillOnce(out, /data-height="0"/, `data-height="${H}"`, '`data-height="0"` hole');
// font-robust safety net — appended last so it wins the cascade over the skin's own
// (preset-font-tuned) line-height. Kept SNUG (1.1) so the plate hugs the text. NO extra
// word/pill padding: inspect's `text_box_overflow` on the highlight words is a cosmetic
// false-positive here (heavy-glyph ink slightly exceeds the line box, but there's no
// overflow:hidden — nothing is clipped); zeroing it would need an airy line-height that
// balloons the pill, which is worse. Override only if a brand font genuinely clips.
out += "\n<style>\n .caption-line { line-height: 1.1 !important; }\n</style>";
return `<template id="captions-template" data-composition-id="captions" data-width="${W}" data-height="${H}">\n${out.trim()}\n</template>\n`;
}
export { buildFromSkin };
// @font-face for the brand display/body fonts, matched from the project's font dirs
// (staged assets/fonts first, else capture/assets/fonts) by family-name prefix, with
// weight parsed from the filename. Paths are relative to compositions/captions.html.
// Returns "" when frame.md or font files are absent (then the skin's fallback applies).
function brandFontFaces(framePath, hyperframesDir) {
if (!existsSync(framePath)) return "";
const { display, body } = parseFonts(readFileSync(framePath, "utf8"));
const families = [
...new Set([display, body].filter(Boolean).map((f) => f.replace(/^"|"$/g, ""))),
];
if (!families.length) return "";
const dirs = [
// ROOT-RELATIVE — compositions are served with the project root as their base URL, so a
// "../" prefix escapes the root (lint: invalid_parent_traversal_in_asset_path) and 404s in
// Studio/preview. Mirror what the frame workers use for images.
{ abs: join(hyperframesDir, "assets/fonts"), rel: "assets/fonts" },
{ abs: join(hyperframesDir, "capture/assets/fonts"), rel: "capture/assets/fonts" },
].filter((d) => existsSync(d.abs));
const weightOf = (n) => {
const s = n.toLowerCase();
// A numeric axis is the font's own answer, so it beats the word heuristic. Fontsource
// names every face this way ("inter-latin-500-normal.woff2") and carries no weight
// WORD at all, so word-only parsing collapsed a whole family onto 400 and shipped
// exactly one of its faces.
//
// A weight token must not be buried inside a longer run: capture/assets/fonts holds
// hash-named files, and "Newsreader-a1b200c3.woff2" is not a 200-weight face. Hence a
// non-digit before (which also stops "2100" reading as 100) and no alphanumeric after.
// "Roboto900.ttf" still parses — requiring separators on both sides would have lost it.
const numeric = /(?:^|[^0-9])([1-9]00)(?![0-9a-z])/.exec(s);
if (numeric) return Number(numeric[1]);
if (/black|heavy|ultra|extrabold/.test(s)) return 800;
if (/semibold|demibold/.test(s)) return 600; // before /bold/ — "demibold" contains "bold"
if (/bold/.test(s)) return 700;
if (/medium/.test(s)) return 500;
if (/light|thin/.test(s)) return 300;
return 400; // book / regular / roman
};
// Weight is not the only axis in a filename. Google Fonts ships Newsreader as
// "Newsreader-Italic-VariableFont_opsz,wght.ttf" + "Newsreader-VariableFont_opsz,wght.ttf",
// and the italic sorts first — so without a style axis the italic file claimed the
// family's ONLY 400 slot, the upright file was dropped as a duplicate, and the face
// was declared with no `font-style`. @font-face is deliberately global (the composition
// CSS scoper exempts it, and it has to be), so the whole document then rendered that
// family in italics — captions italicizing every sibling composition.
const styleOf = (n) => (/italic|oblique/i.test(n) ? "italic" : "normal");
const fmtOf = (f) =>
/\.woff2$/i.test(f)
? "woff2"
: /\.woff$/i.test(f)
? "woff"
: /\.ttf$/i.test(f)
? "truetype"
: "opentype";
// Normalize away ALL non-alphanumerics (spaces, underscores, hyphens) on BOTH the
// family name and the filename. Real font files use "_" / "-" as word separators
// ("TT_Norms_Pro_Bold.woff2"), so stripping only whitespace never matched them — the
// family key "ttnormspro" failed `startsWith` against "tt_norms_pro_bold", and the
// function silently returned "" → captions shipped with NO @font-face for any
// underscore/hyphen-named brand font (e.g. TT Norms Pro), which is exactly the
// font_family_without_font_face bug.
const norm = (s) => s.toLowerCase().replace(/[^a-z0-9]/g, "");
const faces = [];
const seen = new Set();
const claimed = new Set(); // each file is claimed by the MOST SPECIFIC family only
// Match the longest family key first so "TT Norms Pro" can't swallow the files that
// belong to "TT Norms Pro Mono" (its key is a prefix of the longer one's).
const ranked = [...families].sort((a, b) => norm(b).length - norm(a).length);
for (const fam of ranked) {
const key = norm(fam);
for (const d of dirs) {
let files = [];
try {
files = readdirSync(d.abs);
} catch {
continue;
}
for (const f of files.sort()) {
if (!/\.(woff2|woff|ttf|otf)$/i.test(f)) continue;
if (claimed.has(f)) continue; // a more specific family already took this file
if (!norm(f.replace(/\.(woff2|woff|ttf|otf)$/i, "")).startsWith(key)) continue;
const w = weightOf(f);
const style = styleOf(f);
const dedup = `${fam}-${w}-${style}`;
if (seen.has(dedup)) continue; // one src per face; assets/fonts wins over capture
seen.add(dedup);
claimed.add(f);
faces.push(
` @font-face { font-family: '${fam}'; src: url('${d.rel}/${f}') format('${fmtOf(f)}'); font-weight: ${w}; font-style: ${style}; font-display: block; }`,
);
}
}
}
// Loud signal instead of a silent "". If frame.md named a brand font but no file
// matched, the caption text WILL fall back to a generic font in the render — surface
// the cause here (at build time) rather than letting it surface 2 steps later as a
// font_family_without_font_face lint error disconnected from its root cause.
if (!faces.length) {
const where = dirs.length
? dirs.map((d) => d.rel).join(" / ")
: "assets/fonts or capture/assets/fonts (neither exists)";
console.warn(
` ⚠ captions: frame.md names font ${families.map((f) => `"${f}"`).join(", ")} ` +
`but no matching .woff2/.woff/.ttf/.otf was found in ${where} — captions will fall back ` +
`(text may render in the wrong font). Stage a font file whose name starts with the family ` +
`(e.g. "TT Norms Pro" → TT_Norms_Pro_Bold.woff2) so it ships with the project.`,
);
}
return faces.join("\n");
}
export { brandFontFaces }; // exported as a seam for unit testing
// frame.md colors:/typography: → a :root token block, mapped to the fixed semantic
// vocab every preset skin references. Robust to per-preset key names: colors are
// matched by name, then by luminance. Brand-token overlay (Step 2) flows through
// because the values come from the project's frame.md. No frame.md → band vars only.
function frameTokensCss(framePath, H) {
const band = captionBand(H);
const out = [];
if (existsSync(framePath)) {
const md = readFileSync(framePath, "utf8");
const colors = parseColors(md);
for (const [k, v] of colors) out.push(` --${k}: ${v};`); // raw, for completeness
const sem = semanticColors(colors);
if (sem.ink) out.push(` --cap-ink: ${sem.ink};`);
if (sem.canvas) out.push(` --cap-canvas: ${sem.canvas};`);
if (sem.accent) out.push(` --cap-accent: ${sem.accent};`);
if (sem.accent2) out.push(` --cap-accent-2: ${sem.accent2};`);
const { display, body } = parseFonts(md);
if (display) out.push(` --font-display: ${display}, system-ui, serif;`);
if (body) out.push(` --font-body: ${body}, system-ui, sans-serif;`);
}
out.push(` --cap-band-top: ${band.bandTopY}px;`);
out.push(` --cap-band-height: ${band.bandHeight}px;`);
return ` :root {\n${out.join("\n")}\n }`;
}
// ── default path (no preset skin) ─────────────────────────────────────────────
// Self-contained captions sub-composition. The <template> holds the band container
// + style AND the <script> (the HyperFrames loader only executes scripts INSIDE the
// cloned template — a sibling <script> after </template> never runs, so the timeline
// never registers and captions render blank). The script builds per-word spans and a
// paused, seek-safe GSAP timeline (opacity for group show/hide, a quick color tween
// per word for the karaoke highlight — no className flips, no JS state) and ends each
// group with a hard tl.set kill so an exit can't get stuck. gsap is loaded via CDN
// inside the template (matching the frame compositions). Band = captionBand(H).
function buildCaptionsHtml(groups, total, W, H) {
const band = captionBand(H);
const fs = Math.round(H * 0.038);
const pad = Math.round(fs * 0.4);
return `<template id="captions-template">
<div
data-composition-id="captions"
data-width="${W}"
data-height="${H}"
data-duration="${total}"
id="captions-root"
>
<div id="cap"></div>
</div>
<style>
#captions-root {
position: absolute;
inset: 0;
pointer-events: none;
}
#cap {
position: absolute;
left: 0;
right: 0;
top: ${band.bandTopY}px;
height: ${band.bandHeight}px;
display: flex;
align-items: center;
justify-content: center;
}
.caption-group {
position: absolute;
max-width: 80%;
padding: ${pad}px ${Math.round(pad * 1.8)}px;
background: rgba(0, 0, 0, 0.72);
border-radius: ${Math.round(fs * 0.3)}px;
font-family: Roboto, sans-serif;
font-weight: 700;
font-size: ${fs}px;
line-height: 1.25;
text-align: center;
color: #fff;
opacity: 0;
}
.caption-word {
color: rgba(255, 255, 255, 0.55);
}
</style>
<script src="https://cdn.jsdelivr.net/npm/gsap@3.14.2/dist/gsap.min.js" integrity="sha384-sG0Hv1tP1lZCk9KQmrIbY/XNwi+OY84GQqhMscbnsoBFqAz8KNCil1kvfL3Hbbk2" crossorigin="anonymous"></script>
<script>
(function () {
var GROUPS = ${JSON.stringify(groups)};
var cap = document.getElementById("cap");
var tl = gsap.timeline({ paused: true });
GROUPS.forEach(function (g) {
var el = document.createElement("div");
el.className = "caption-group";
g.words.forEach(function (w) {
var s = document.createElement("span");
s.className = "caption-word";
s.textContent = w.text + " ";
el.appendChild(s);
});
cap.appendChild(el);
tl.fromTo(el, { opacity: 0 }, { opacity: 1, duration: 0.18, overwrite: "auto" }, g.start);
tl.to(el, { opacity: 0, duration: 0.12, overwrite: "auto" }, g.end);
tl.set(el, { opacity: 0, visibility: "hidden" }, g.end + 0.12); // deterministic hard kill
g.words.forEach(function (w, i) {
tl.to(el.children[i], { color: "#ffffff", duration: 0.06 }, w.start);
});
});
tl.to({}, { duration: ${total} }, 0); // full-span anchor
window.__timelines = window.__timelines || {};
window.__timelines["captions"] = tl;
})();
</script>
</template>
`;
}
if (process.argv[1] && resolve(process.argv[1]) === fileURLToPath(import.meta.url)) {
const sub = process.argv[2];
if (sub === "build" || sub === undefined) runBuild(process.argv.slice(sub === "build" ? 3 : 2));
else {
console.error(
"usage: node captions.mjs build [--storyboard …] [--audio-meta …] [--hyperframes .]",
);
process.exit(2);
}
}
scripts/captions.test.mjs
import assert from "node:assert/strict";
import { mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import test from "node:test";
import { fileURLToPath } from "node:url";
import { brandFontFaces, buildFromSkin } from "./captions.mjs";
const presetsDir = fileURLToPath(
new URL("../../hyperframes-creative/frame-presets/", import.meta.url),
);
const skins = readdirSync(presetsDir, { withFileTypes: true })
.filter((entry) => entry.isDirectory())
.map((entry) => ({
name: entry.name,
source: readFileSync(
new URL(
`../../hyperframes-creative/frame-presets/${entry.name}/caption-skin.html`,
import.meta.url,
),
"utf8",
),
}))
.filter(({ source }) => source.includes(".caption-word.is-active"));
for (const canvas of ["#f7f3e8", "#111827"]) {
for (const skin of skins) {
test(`${skin.name} preserves word-state rules on ${canvas}`, () => {
const active = skin.source.match(/\.caption-word\.is-active\s*\{[^}]*\}/s)?.[0];
const spoken = skin.source.match(/\.caption-word\.is-spoken\s*\{[^}]*\}/s)?.[0];
assert.ok(active, "skin must define an active-word rule");
assert.ok(spoken, "skin must define a spoken-word rule");
const output = buildFromSkin(
skin.source,
[],
1,
1920,
1080,
`:root { --cap-canvas: ${canvas}; --cap-ink: #111111; --cap-accent: #ffcc00; }`,
(message) => {
throw new Error(message);
},
);
assert.ok(output.includes(active));
assert.ok(output.includes(spoken));
assert.equal(output.match(/\.caption-word\.is-active\s*\{/g)?.length, 1);
assert.equal(output.match(/\.caption-word\.is-spoken\s*\{/g)?.length, 1);
});
}
}
// @font-face is document-global on purpose — the composition CSS scoper exempts it, and it
// has to, since a face declaration cannot be scoped. That makes brandFontFaces the one part
// of the captions sub-composition whose output reaches every sibling composition, so it has
// to describe each face exactly: get an axis wrong and the whole document renders the brand
// family wrong.
function withFontProject(files, run) {
const dir = mkdtempSync(join(tmpdir(), "hf-captions-fonts-"));
try {
mkdirSync(join(dir, "assets/fonts"), { recursive: true });
for (const name of files) writeFileSync(join(dir, "assets/fonts", name), "");
writeFileSync(
join(dir, "frame.md"),
'typography:\n display: { fontFamily: "Newsreader", weight: 400 }\n body: { fontFamily: "Inter", weight: 400 }\n',
);
return run(join(dir, "frame.md"), dir);
} finally {
rmSync(dir, { recursive: true, force: true });
}
}
test("an italic file is declared italic, and never squats the family's upright slot", () => {
// Google Fonts' own Newsreader download. The italic sorts first, so before the style axis
// existed it claimed the family's only 400 slot, the upright was dropped as a duplicate,
// and the face shipped with no font-style — italicizing every sibling composition that
// used Newsreader.
const faces = withFontProject(
["Newsreader-Italic-VariableFont_opsz,wght.ttf", "Newsreader-VariableFont_opsz,wght.ttf"],
brandFontFaces,
);
const lines = faces.split("\n").filter((line) => line.includes("Newsreader"));
assert.equal(lines.length, 2, "both the upright and the italic file must be declared");
const upright = lines.find((line) => line.includes("Newsreader-VariableFont"));
const italic = lines.find((line) => line.includes("Newsreader-Italic-VariableFont"));
assert.ok(upright, "the upright file must survive");
assert.match(upright, /font-style: normal/);
assert.ok(italic, "the italic file must survive");
assert.match(italic, /font-style: italic/);
});
test("a numeric weight in the filename is read as the weight", () => {
// Fontsource names every face numerically and carries no weight WORD, so word-only
// parsing scored the whole family 400 and shipped exactly one of its four faces.
const faces = withFontProject(
[
"inter-latin-400-normal.woff2",
"inter-latin-500-normal.woff2",
"inter-latin-600-italic.woff2",
"inter-latin-700-normal.woff2",
],
brandFontFaces,
);
const lines = faces.split("\n").filter((line) => line.includes("Inter"));
assert.equal(lines.length, 4, "each face is a distinct weight/style pair");
for (const [file, weight, style] of [
["inter-latin-400-normal", 400, "normal"],
["inter-latin-500-normal", 500, "normal"],
["inter-latin-600-italic", 600, "italic"],
["inter-latin-700-normal", 700, "normal"],
]) {
const line = lines.find((candidate) => candidate.includes(file));
assert.ok(line, `${file} must be declared`);
assert.match(line, new RegExp(`font-weight: ${weight};`));
assert.match(line, new RegExp(`font-style: ${style};`));
}
});
test("word-named weights still parse when the filename carries no numeric axis", () => {
const faces = withFontProject(
["Newsreader_Bold.woff2", "Newsreader_Regular.woff2"],
brandFontFaces,
);
const lines = faces.split("\n").filter((line) => line.includes("Newsreader"));
assert.equal(lines.length, 2);
assert.match(
lines.find((line) => line.includes("Bold")),
/font-weight: 700; font-style: normal/,
);
assert.match(
lines.find((line) => line.includes("Regular")),
/font-weight: 400; font-style: normal/,
);
});
// build-frame.mjs stages captured brand fonts under a REWRITTEN name, and brandFontFaces
// derives the face's axes back out of that name. The two are a contract, and it is easy to
// break silently from either side: build-frame used to drop the style token while renaming,
// so an italic file arrived as "Newsreader-Regular.ttf" and was declared upright — leaving
// the document-global normal slot pointing at italic bytes even once brandFontFaces learned
// about styles. These two tests pin both ends of that contract.
test("the names build-frame.mjs stages round-trip back to the right face", () => {
const faces = withFontProject(
[
"Newsreader-Regular.ttf",
"Newsreader-Regular-Italic.ttf",
"Inter-400.woff2",
"Inter-600-Italic.woff2",
],
brandFontFaces,
);
for (const [file, weight, style] of [
["Newsreader-Regular.ttf", 400, "normal"],
["Newsreader-Regular-Italic.ttf", 400, "italic"],
["Inter-400.woff2", 400, "normal"],
["Inter-600-Italic.woff2", 600, "italic"],
]) {
const line = faces.split("\n").find((candidate) => candidate.includes(`/${file}'`));
assert.ok(line, `${file} must be declared`);
assert.match(line, new RegExp(`font-weight: ${weight};`));
assert.match(line, new RegExp(`font-style: ${style};`));
}
});
test("a weight token buried in a longer run is not read as a weight", () => {
// capture/assets/fonts commonly holds hash-named files, and a hash is not a weight.
const faces = withFontProject(
["Newsreader-a1b200c3.woff2", "Inter-2100.woff2", "Inter900.woff2"],
brandFontFaces,
);
const weightOf = (file) =>
Number(/font-weight: (\d+);/.exec(faces.split("\n").find((l) => l.includes(file)))?.[1]);
// "200" sits mid-run (…b200c3), so the word path decides: Regular.
assert.equal(weightOf("Newsreader-a1b200c3.woff2"), 400);
// 4-digit guard: "2100" must not read as 100.
assert.equal(weightOf("Inter-2100.woff2"), 400);
// ...but a trailing weight with no separator is still a weight.
assert.equal(weightOf("Inter900.woff2"), 900);
});
// captions.mjs ships once per creation workflow because each skill installs standalone,
// and the three copies are meant to be byte-identical. This PR alone had to land the same
// two-axis fix in all three; a future one that lands in only one drifts silently.
test("captions.mjs is byte-identical across the three workflows that ship it", () => {
const [first, ...rest] = ["product-launch-video", "faceless-explainer", "pr-to-video"].map(
(skill) => ({
skill,
source: readFileSync(new URL(`../../${skill}/scripts/captions.mjs`, import.meta.url), "utf8"),
}),
);
for (const other of rest) {
assert.equal(other.source, first.source, `${other.skill} drifted from ${first.skill}`);
}
});
test("every build-frame.mjs copy stages the style axis it promises", () => {
for (const skill of ["product-launch-video", "faceless-explainer", "pr-to-video"]) {
const source = readFileSync(
new URL(`../../${skill}/scripts/build-frame.mjs`, import.meta.url),
"utf8",
);
// The staged filename must carry the style, or the italic and upright faces of one
// weight collide on a single name and only whichever sorts first survives.
assert.match(
source,
/const clean = `\$\{fam\.replace\(\/\[\^A-Za-z0-9\]\/g, ""\)\}-\$\{w\}\$\{style === "italic" \? "-Italic" : ""\}\./,
`${skill}/build-frame.mjs must keep the style token in the staged name`,
);
// ...and the emitted descriptor must report the real style, not a hardcoded normal.
assert.doesNotMatch(
source,
/font-weight:\$\{n\};font-style:normal/,
`${skill}/build-frame.mjs must not assert font-style:normal over captured bytes`,
);
}
});
scripts/capture-skill-guardrails.test.mjs
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import test from "node:test";
import { fileURLToPath } from "node:url";
function read(relativePath) {
return readFileSync(fileURLToPath(new URL(relativePath, import.meta.url)), "utf8");
}
test("product launch capture treats blocked output as a hard gate", () => {
const skill = read("../SKILL.md");
assert.match(skill, /hyperframes capture[^\n]+--json/);
assert.match(skill, /capture\/BLOCKED\.md[^\n]+hard stop/i);
assert.match(skill, /do not[\s\S]{0,120}synthetic[\s\S]{0,120}fallback/i);
assert.match(skill, /very little text[\s\S]{0,180}empty asset/i);
assert.match(skill, /show-it-as-is[\s\S]{0,240}provided screenshot/i);
});
test("CLI capture reference documents the two budgets and machine diagnostics", () => {
const reference = read("../../hyperframes-cli/references/init-and-scaffold.md");
assert.match(reference, /--capture-budget/);
assert.match(reference, /--skip-vision/);
assert.match(reference, /HYPERFRAMES_CAPTURE_PHASE/);
assert.match(reference, /BLOCKED\.md[^\n]+hard stop/i);
assert.match(reference, /--timeout[^\n]+navigation[^\n]+--capture-budget/i);
assert.match(reference, /fresh output directory/i);
assert.match(reference, /outer caller timeout[\s\S]{0,180}does not prove/i);
});
test("webpage motion workflow refuses blocked capture artifacts", () => {
const module = read("../../motion-graphics/categories/webpage/module.md");
assert.match(module, /BLOCKED\.md[\s\S]{0,120}stop/i);
assert.match(module, /provided[\s\S]{0,120}screenshot[\s\S]{0,120}explicit/i);
assert.match(module, /do not[^\n]+blocked[^\n]+capture/i);
assert.match(module, /asset-free fallback/i);
});
scripts/frame-packets.mjs
#!/usr/bin/env node
// Thin wrapper over the shared packet builder in hyperframes-core — this file only
// pins the paths that are specific to this workflow skill. The logic (frame
// splitting, rule citation, packet bounds, `_role.md` assembly) has one owner:
// ../../hyperframes-core/scripts/lib/frame-packets-core.mjs
import { dirname, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import * as core from "../../hyperframes-core/scripts/lib/frame-packets-core.mjs";
const SKILL_DIR = resolve(dirname(fileURLToPath(import.meta.url)), "..");
const CONFIG = {
animationDir: resolve(SKILL_DIR, "../hyperframes-animation"),
corePath: resolve(SKILL_DIR, "../hyperframes-core/references/frame-worker-core.md"),
deltaPath: resolve(SKILL_DIR, "sub-agents/frame-worker.md"),
};
export function buildRolePayload({ outDir }) {
return core.buildRolePayload({ ...CONFIG, outDir });
}
export function buildFramePackets(options) {
return core.buildFramePackets({ ...CONFIG, ...options });
}
if (core.isMainModule(import.meta.url)) core.runCli({ buildFramePackets, buildRolePayload });
scripts/frame-packets.test.mjs
import assert from "node:assert/strict";
import { existsSync, mkdirSync, mkdtempSync, readFileSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { dirname, join } from "node:path";
import test from "node:test";
import { buildFramePackets } from "./frame-packets.mjs";
function write(path, contents) {
mkdirSync(dirname(path), { recursive: true });
writeFileSync(path, contents);
}
test("packets inline the blueprint body and the Scene-cited rule recipes", () => {
const project = mkdtempSync(join(tmpdir(), "plv-packets-"));
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Hook\n\n- duration: 3s\n- src: compositions/frames/01-hook.html\n- blueprint: dataviz-countup\n- scene: hero stat punches in\n\nScene 1 (0.0–1.5s): the stat enters via spring-pop-entrance, then counting-dynamic-scale runs the tally.\n\n## Frame 2 — Freeform\n\n- duration: 4s\n- src: compositions/frames/02-freeform.html\n- blueprint: compose\n\nScene 1 (0.0–4.0s): a quiet hold, no named motion.\n`,
);
const result = buildFramePackets({ projectDir: project });
assert.equal(result.length, 2);
const hook = readFileSync(result[0].path, "utf8");
assert.match(hook, /## Selected blueprint: dataviz-countup/);
assert.match(hook, /## Selected motion rule: spring-pop-entrance/);
assert.match(hook, /## Selected motion rule: counting-dynamic-scale/);
assert.match(hook, /RULES_DIR: /);
const freeform = readFileSync(result[1].path, "utf8");
assert.doesNotMatch(freeform, /## Selected blueprint/);
assert.doesNotMatch(freeform, /## Selected motion rule/);
});
test("_role.md is the core contract + this workflow's delta, verbatim", () => {
const project = mkdtempSync(join(tmpdir(), "plv-role-"));
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Hook\n\n- duration: 3s\n- src: compositions/frames/01-hook.html\n`,
);
buildFramePackets({ projectDir: project });
const rolePath = join(project, ".hyperframes", "frame-packets", "_role.md");
assert.ok(existsSync(rolePath));
const role = readFileSync(rolePath, "utf8");
assert.match(role, /# Frame worker — core contract/);
assert.match(role, /# Frame worker — product-launch delta/);
});
test("packet validation is atomic and leaves no partial output on overflow", () => {
const project = mkdtempSync(join(tmpdir(), "plv-atomic-"));
const outDir = join(project, ".hyperframes", "frame-packets");
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Big\n\n- duration: 3s\n- src: compositions/frames/01-big.html\n\n${"padding line\n".repeat(300)}`,
);
assert.throws(
() => buildFramePackets({ projectDir: project, outDir, maxPacketBytes: 2_000 }),
/limit 2000/,
);
assert.equal(existsSync(outDir), false);
});
// ── the blueprint qualifier ──────────────────────────────────────────────────
// Regression: visual-design.md documents `blueprint:` as the id plus a
// `(Reproduce)` / `(Adapt)` qualifier, and prints `dataviz-countup (Adapt)` as
// its worked example. The resolver used the raw field as the filename, so every
// qualified blueprint looked for a file that cannot exist and inlined "" —
// packets shipped without the document the frame was designed against, and the
// run still reported success. The cases above only ever used bare ids.
test("a qualified blueprint resolves to the same body as the bare id", () => {
const project = mkdtempSync(join(tmpdir(), "plv-blueprint-qualified-"));
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Adapted\n\n- duration: 3s\n- src: compositions/frames/01-adapted.html\n- blueprint: device-surface-showcase (Adapt)\n\n## Frame 2 — Reproduced\n\n- duration: 3s\n- src: compositions/frames/02-reproduced.html\n- blueprint: device-surface-showcase (Reproduce)\n\n## Frame 3 — Bare\n\n- duration: 3s\n- src: compositions/frames/03-bare.html\n- blueprint: device-surface-showcase\n`,
);
const packets = buildFramePackets({ projectDir: project });
const blueprintSections = packets.map((packet) => {
const body = readFileSync(packet.path, "utf8");
const start = body.indexOf("## Selected blueprint:");
assert.notEqual(start, -1, `${packet.frameId} inlined no blueprint`);
return body.slice(start);
});
assert.match(blueprintSections[0], /## Selected blueprint: device-surface-showcase\n/);
// The qualifier is direction for the worker, not a different document: all
// three frames must inline byte-identical blueprint bodies.
assert.equal(new Set(blueprintSections).size, 1);
});
test("a qualified `compose` still selects no blueprint", () => {
const project = mkdtempSync(join(tmpdir(), "plv-blueprint-compose-"));
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Freeform\n\n- duration: 3s\n- src: compositions/frames/01-freeform.html\n- blueprint: compose (Adapt)\n`,
);
const [packet] = buildFramePackets({ projectDir: project });
assert.doesNotMatch(readFileSync(packet.path, "utf8"), /## Selected blueprint/);
});
test("a blueprint with no file fails the run instead of shipping an empty section", () => {
const project = mkdtempSync(join(tmpdir(), "plv-blueprint-missing-"));
const outDir = join(project, ".hyperframes", "frame-packets");
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Typo\n\n- duration: 3s\n- src: compositions/frames/01-typo.html\n- blueprint: device-surface-showcses\n`,
);
assert.throws(
() => buildFramePackets({ projectDir: project, outDir }),
/01-typo: blueprint "device-surface-showcses" has no file/,
);
assert.equal(existsSync(outDir), false);
});
test("an uninstalled animation skill degrades with a warning, it does not fail the run", () => {
// hyperframes-animation installs on demand, so an absent blueprints/ means the
// library isn't there yet — not that the frame named a bad id. Matches how an
// absent rules/ already behaves.
const project = mkdtempSync(join(tmpdir(), "plv-blueprint-uninstalled-"));
write(join(project, "frame.md"), "# tokens\n");
write(
join(project, "STORYBOARD.md"),
`---\nformat: 1920x1080\n---\n\n## Frame 1 — Hook\n\n- duration: 3s\n- src: compositions/frames/01-hook.html\n- blueprint: dataviz-countup (Adapt)\n`,
);
const [packet] = buildFramePackets({
projectDir: project,
animationDir: join(project, "absent-animation-skill"),
});
assert.doesNotMatch(readFileSync(packet.path, "utf8"), /## Selected blueprint/);
});
scripts/lib/assets.mjs
// assets.mjs — stage frame-named capture assets into assets/.
// Shared by stage-assets.mjs (Step 4 close, BEFORE the frame workers run) and
// assemble-index.mjs (Step 5, idempotent backstop). Only assets a frame names
// in `asset_candidates` are staged; unnamed assets never reach the project.
// asset_candidates value form: "assets/<basename> — desc; assets/… — …".
import { copyFileSync, existsSync, mkdirSync } from "node:fs";
import { basename, join } from "node:path";
export function basenamesFromCandidates(value) {
if (typeof value !== "string") return [];
return value
.split(";")
.map((seg) => seg.split(/\s+[—–-]\s+/)[0].trim()) // strip the " — description"
.filter(Boolean)
.map((p) => basename(p.replace(/^assets\//, "")));
}
// Copy each frame's asset_candidates from capture/{assets,assets/videos,
// assets/svgs, screenshots} into assets/. Already-staged files are left as is
// (first-wins), so calling this twice is safe. Returns { staged, wanted, anomalies }.
export function stageAssets({ hyperframesDir, frames }) {
const wanted = new Set();
for (const f of frames) {
for (const b of basenamesFromCandidates(f.extra?.asset_candidates)) wanted.add(b);
}
const captureDirs = [
join(hyperframesDir, "capture/assets"),
join(hyperframesDir, "capture/assets/videos"), // videos download into a subdir
join(hyperframesDir, "capture/assets/svgs"), // inline SVGs extract into a subdir
join(hyperframesDir, "capture/screenshots"),
];
const assetsDir = join(hyperframesDir, "assets");
const anomalies = [];
let staged = 0;
if (wanted.size > 0) {
mkdirSync(assetsDir, { recursive: true });
for (const b of wanted) {
const dest = join(assetsDir, b);
if (existsSync(dest)) {
staged++;
continue;
} // first-wins / already staged
const src = captureDirs.map((d) => join(d, b)).find((p) => existsSync(p));
if (src) {
copyFileSync(src, dest);
staged++;
} else {
anomalies.push(
`asset "${b}" named by a frame but not found under capture/ — frame will 404 it`,
);
}
}
}
return { staged, wanted, anomalies };
}
scripts/lib/dimensions.mjs
// dimensions.mjs — canvas size + caption-band geometry for the product-launch
// pipeline. Single source of truth = the STORYBOARD frontmatter `format` global
// ("1920x1080" / "1080x1920" / "1080x1080", or a named orientation). Every
// script and the index assembler reads the size from here; none hardcodes it.
// Named orientation presets. Square/portrait are 1080-based so they share the
// long-edge pixel budget with landscape (same render-cost ballpark).
export const ORIENTATION_PRESETS = {
landscape: { width: 1920, height: 1080 }, // 16:9 — default
portrait: { width: 1080, height: 1920 }, // 9:16 — reels / shorts / TikTok
square: { width: 1080, height: 1080 }, // 1:1 — feed
};
export const DEFAULT_DIMENSIONS = ORIENTATION_PRESETS.landscape;
function sane(w, h) {
return Number.isFinite(w) && Number.isFinite(h) && w >= 240 && h >= 240 && w <= 8192 && h <= 8192;
}
// Parse a STORYBOARD `format` global into { width, height, source }. Accepts
// "WxH" (e.g. "1920x1080"; `x` or `×`, any inner spacing) or a named orientation;
// falls back to landscape so a storyboard with a missing/garbled format still
// renders (no behavior change vs the old landscape lock).
export function parseFormat(format) {
const s = typeof format === "string" ? format.trim().toLowerCase() : "";
if (ORIENTATION_PRESETS[s]) return { ...ORIENTATION_PRESETS[s], source: `orientation=${s}` };
const m = s.match(/^(\d+)\s*[x×]\s*(\d+)$/);
if (m) {
const w = parseInt(m[1] ?? "", 10);
const h = parseInt(m[2] ?? "", 10);
if (sane(w, h)) return { width: w, height: h, source: "format" };
}
return { ...DEFAULT_DIMENSIONS, source: "default(landscape)" };
}
// Caption band geometry, derived from canvas height: the bottom ~16.67% (180px
// at h=1080). Frame content must end `safetyPx` above the band top. Holds even
// when captions are disabled (bottom-edge consistency).
export const CAPTION_BAND_FRACTION = 0.1667;
export function captionBand(height, safetyPx = 20) {
const h = Number.isFinite(height) ? height : DEFAULT_DIMENSIONS.height;
const bandHeight = Math.round(h * CAPTION_BAND_FRACTION);
const bandTopY = h - bandHeight; // foreground must end at/above this y
return { bandHeight, bandTopY, foregroundMaxY: bandTopY - safetyPx };
}
scripts/lib/pad-frame-duration.mjs
// pad-frame-duration.mjs — keeps a frame's own #root/clip data-duration in
// sync with the padded index.html wrapper duration transitions.mjs computes.
//
// The frame's OWN internal file declares its #root/clip data-duration to the
// STORYBOARD's content-only length (frame-worker.md: duration is "fixed
// upstream"). When an outgoing transition pads the index.html WRAPPER's
// data-duration to cover the transition tail, the frame's own internal
// duration is left short — the render engine clip-gates the sub-composition's
// visible content at that shorter value, so content vanishes abruptly at
// content-end instead of fading gracefully through the wrapper's extended
// fade-out tween. Pad the frame's own file to match so both durations agree.
import { readFileSync, writeFileSync } from "node:fs";
import { resolve } from "node:path";
export function padFrameInternalDuration(hyperframesDir, frameSrc, frameId, newDuration) {
const framePath = resolve(hyperframesDir, frameSrc);
let html;
try {
html = readFileSync(framePath, "utf8");
} catch (err) {
if (err?.code === "ENOENT") return;
throw err;
}
const tagRe = /<[a-z][\w:-]*\s[^<>]*?>/gi;
let m;
while ((m = tagRe.exec(html)) !== null) {
const tag = m[0];
if (!tag.includes(`data-composition-id="${frameId}"`)) continue;
if (!/data-duration="[\d.]+"/.test(tag)) continue;
const newTag = tag.replace(/data-duration="[\d.]+"/, `data-duration="${newDuration}"`);
if (newTag === tag) return;
writeFileSync(framePath, html.slice(0, m.index) + newTag + html.slice(m.index + tag.length));
return;
}
}
scripts/lib/pad-frame-duration.test.mjs
import { test } from "node:test";
import assert from "node:assert/strict";
import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { padFrameInternalDuration } from "./pad-frame-duration.mjs";
// Regression: an outgoing transition pads the index.html WRAPPER's
// data-duration to cover the transition tail, but the frame's own internal
// file kept its shorter content-only duration, so the render engine
// clip-gated the sub-composition's visible content at the shorter value —
// content vanished abruptly instead of fading through the wrapper's
// extended fade-out tween. A user diagnosed and verified this fix
// themselves: pad the frame's own #root/clip data-duration to match.
test("padFrameInternalDuration pads the matching frame's own data-duration", () => {
const dir = mkdtempSync(join(tmpdir(), "transitions-pad-"));
const framesDir = join(dir, "compositions", "frames");
mkdirSync(framesDir, { recursive: true });
const frameSrc = "compositions/frames/scene-1.html";
const framePath = join(dir, frameSrc);
writeFileSync(
framePath,
`<template>
<div
id="root"
data-composition-id="scene-1"
data-width="1920"
data-height="1080"
data-duration="4.2"
></div>
</template>`,
);
try {
padFrameInternalDuration(dir, frameSrc, "scene-1", 4.7);
const updated = readFileSync(framePath, "utf8");
assert.match(updated, /data-duration="4\.7"/);
assert.doesNotMatch(updated, /data-duration="4\.2"/);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
test("padFrameInternalDuration only touches the tag matching the given frame id", () => {
const dir = mkdtempSync(join(tmpdir(), "transitions-pad-scope-"));
const framesDir = join(dir, "compositions", "frames");
mkdirSync(framesDir, { recursive: true });
const frameSrc = "compositions/frames/scene-2.html";
const framePath = join(dir, frameSrc);
const original = `<template>
<div id="root" data-composition-id="scene-2" data-duration="3.0">
<div data-composition-id="unrelated-child" data-duration="1.0"></div>
</div>
</template>`;
writeFileSync(framePath, original);
try {
padFrameInternalDuration(dir, frameSrc, "scene-2", 3.5);
const updated = readFileSync(framePath, "utf8");
assert.match(updated, /data-composition-id="scene-2" data-duration="3\.5"/);
assert.match(updated, /data-composition-id="unrelated-child" data-duration="1\.0"/);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
test("padFrameInternalDuration is a no-op when the frame file does not exist", () => {
const dir = mkdtempSync(join(tmpdir(), "transitions-pad-missing-"));
try {
assert.doesNotThrow(() =>
padFrameInternalDuration(dir, "compositions/frames/missing.html", "missing", 5),
);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
scripts/lib/storyboard.mjs
// storyboard.mjs — vendored lenient parser for STORYBOARD.md.
//
// Faithful plain-JS port of @hyperframes/core/storyboard
// (packages/core/src/storyboard/parseStoryboard.ts). Vendored because skills
// ship standalone: installed via `npx skills add`, a skill's scripts can't reach
// the monorepo's core package, and the core export points at .ts source that
// `node` (which runs these scripts) can't load. CANONICAL contract = the core
// parser + skills/hyperframes-core/references/storyboard-format.md; keep this in
// lockstep. Behavior: never throws, accepts freeform narrative, recognizes
// Frame/Beat/Scene headings at H2/H3, preserves unknown keys verbatim under
// `extra` (keys lowercased). Pure node — no deps.
export const FRAME_STATUSES = ["outline", "built", "animated"];
export const DEFAULT_FRAME_STATUS = "outline";
// Detection-only frame heading (ends at the keyword); ReDoS-hardened — keep as-is.
const FRAME_HEADING_RE = /^(#{2,3})[ \t]+(?:frame|beat|scene)\b/i;
const FRAME_TITLE_SEP_RE = /^[\s.:—-]+/;
const HEADING_LEVEL_RE = /^(#{1,6})\s+/;
const META_RE = /^\s*[-*]\s+([A-Za-z_][\w-]*)\s*:\s*(.+?)\s*$/;
const LEADING_INT_RE = /^(\d+)/;
const DURATION_NUM_RE = /(\d+(?:\.\d+)?)/;
const TRANSITION_KEYS = new Set(["transition_in", "transitionin", "transition"]);
const SCENE_KEYS = new Set(["scene", "description", "summary", "caption"]);
export const VOICEOVER_ALIASES = ["voiceover", "vo", "voice_over", "narration"];
const VOICEOVER_KEYS = new Set(VOICEOVER_ALIASES);
export function parseStoryboard(source) {
const warnings = [];
const { globals, bodyStartLine, body } = parseFrontmatter(source, warnings);
const frames = parseFrames(body, bodyStartLine, warnings);
return { globals, frames, warnings };
}
function emptyGlobals() {
return { extra: {} };
}
function isFrameStatus(value) {
return FRAME_STATUSES.includes(value);
}
// ── Frontmatter ─────────────────────────────────────────────────────────────
function findFrontmatterRange(lines, warnings) {
let start = 0;
while (start < lines.length && (lines[start] ?? "").trim() === "") start++;
if ((lines[start] ?? "").trim() !== "---") return null;
for (let i = start + 1; i < lines.length; i++) {
if ((lines[i] ?? "").trim() === "---") return { start, end: i };
}
warnings.push({
message: "Frontmatter opening '---' has no closing '---'; treating whole file as body.",
line: start + 1,
});
return null;
}
function parseFrontmatterEntries(lines, start, end, warnings) {
const globals = emptyGlobals();
for (let i = start + 1; i < end; i++) {
const raw = lines[i] ?? "";
if (raw.trim() === "") continue;
const colon = raw.indexOf(":");
if (colon === -1) {
warnings.push({
message: `Ignored non key:value frontmatter line: "${raw.trim()}"`,
line: i + 1,
});
continue;
}
const key = raw.slice(0, colon).trim().toLowerCase();
assignGlobal(globals, key, stripQuotes(raw.slice(colon + 1).trim()));
}
return globals;
}
function parseFrontmatter(source, warnings) {
const lines = source.split(/\r?\n/);
const range = findFrontmatterRange(lines, warnings);
if (!range) return { globals: emptyGlobals(), bodyStartLine: 1, body: source };
const globals = parseFrontmatterEntries(lines, range.start, range.end, warnings);
const body = lines.slice(range.end + 1).join("\n");
return { globals, bodyStartLine: range.end + 2, body };
}
function assignGlobal(globals, key, value) {
switch (key) {
case "format":
globals.format = value;
break;
case "message":
globals.message = value;
break;
case "arc":
globals.arc = value;
break;
case "audience":
globals.audience = value;
break;
default:
globals.extra[key] = value;
}
}
// ── Frames ──────────────────────────────────────────────────────────────────
function openFrameSection(line, headingLine) {
const match = FRAME_HEADING_RE.exec(line);
if (!match) return null;
const headingText = line.slice(match[0].length).replace(FRAME_TITLE_SEP_RE, "").trim();
return { headingText, headingLine, level: (match[1] ?? "##").length, lines: [] };
}
function endsFrameSection(line, current) {
if (!current) return false;
const heading = HEADING_LEVEL_RE.exec(line);
return heading !== null && (heading[1] ?? "").length <= current.level;
}
function parseFrames(body, bodyStartLine, warnings) {
const lines = body.split(/\r?\n/);
const sections = [];
let current = null;
for (let i = 0; i < lines.length; i++) {
const line = lines[i] ?? "";
const opened = openFrameSection(line, bodyStartLine + i);
if (opened) {
sections.push(opened);
current = opened;
} else if (endsFrameSection(line, current)) {
current = null;
} else if (current) {
current.lines.push(line);
}
}
return sections.map((section, idx) => buildFrame(section, idx + 1, warnings));
}
function buildFrame(section, index, warnings) {
const frame = { index, status: DEFAULT_FRAME_STATUS, narrative: "", extra: {} };
const { number, title } = parseHeading(section.headingText);
if (number !== undefined) frame.number = number;
if (title) frame.title = title;
const narrativeLines = [];
for (const line of section.lines) {
const meta = META_RE.exec(line);
if (meta) {
applyMeta(
frame,
(meta[1] ?? "").toLowerCase(),
(meta[2] ?? "").trim(),
section.headingLine,
warnings,
);
} else {
narrativeLines.push(line);
}
}
frame.narrative = narrativeLines.join("\n").trim();
return frame;
}
function parseHeading(text) {
if (!text) return {};
const intMatch = LEADING_INT_RE.exec(text);
if (!intMatch) return { title: text };
const number = Number.parseInt(intMatch[1] ?? "", 10);
const rest = text
.slice((intMatch[0] ?? "").length)
.replace(/^[\s.:—-]+/, "")
.trim();
return { number, title: rest || undefined };
}
// Dispatch a recognized metadata key to its field, else stash under `extra`.
// Mirrors core's META_SETTERS map exactly (direct keys + alias sets).
function applyMeta(frame, key, value, headingLine, warnings) {
switch (key) {
case "duration":
applyDuration(frame, value, headingLine, warnings);
return;
case "status":
applyStatus(frame, value, headingLine, warnings);
return;
case "poster":
applyPoster(frame, value);
return;
case "src":
frame.src = value;
return;
}
if (TRANSITION_KEYS.has(key)) {
frame.transitionIn = value;
return;
}
if (SCENE_KEYS.has(key)) {
frame.scene = value;
return;
}
if (VOICEOVER_KEYS.has(key)) {
frame.voiceover = stripQuotes(value);
return;
}
frame.extra[key] = value;
}
function applyPoster(frame, value) {
const num = DURATION_NUM_RE.exec(value);
if (num) frame.poster = Number.parseFloat(num[1] ?? "");
}
function applyDuration(frame, value, headingLine, warnings) {
frame.duration = value;
const num = DURATION_NUM_RE.exec(value);
if (num) {
frame.durationSeconds = Number.parseFloat(num[1] ?? "");
return;
}
warnings.push({
message: `Frame ${frame.index}: could not parse duration "${value}".`,
line: headingLine,
frameIndex: frame.index,
});
}
function applyStatus(frame, value, headingLine, warnings) {
const normalized = value.toLowerCase();
if (isFrameStatus(normalized)) {
frame.status = normalized;
return;
}
frame.extra.status = value;
warnings.push({
message: `Frame ${frame.index}: unknown status "${value}"; defaulting to "${DEFAULT_FRAME_STATUS}".`,
line: headingLine,
frameIndex: frame.index,
});
}
function stripQuotes(value) {
if (value.length >= 2) {
const first = value[0];
const last = value[value.length - 1];
if ((first === '"' && last === '"') || (first === "'" && last === "'")) {
return value.slice(1, -1);
}
}
return value;
}
scripts/lib/tokens.mjs
// tokens.mjs — shared brand-token parsing + semantic role mapping for frame.md /
// FRAME.md. Used by build-frame.mjs (remix a preset onto brand tokens) and
// captions.mjs (derive caption colors from frame.md). One mapping → frames and
// captions stay consistent. Pure node.
// Collect `key: value` pairs under the top-level `colors:` block (until dedent).
export function parseColors(md) {
const out = [];
let inBlock = false;
for (const line of md.split(/\r?\n/)) {
if (/^colors:\s*$/.test(line)) {
inBlock = true;
continue;
}
if (!inBlock) continue;
if (/^\S/.test(line)) break; // dedent to a top-level key → end of block
const m = line.match(
/^\s+([\w-]+):\s*(?:"([^"]+)"|'([^']+)'|(#[0-9a-fA-F]{3,8}|rgba?\([^)]*\)|[^#\s][^#\n]*?))\s*(?:#.*)?$/,
);
if (m) out.push([m[1], (m[2] ?? m[3] ?? m[4]).trim()]);
}
return out;
}
// relative luminance of a #rrggbb (null for non-hex like rgba()).
export function lum(v) {
const m = /^#?([0-9a-fA-F]{6})$/.exec(String(v).trim());
if (!m) return null;
const n = parseInt(m[1], 16);
return 0.2126 * ((n >> 16) & 255) + 0.7152 * ((n >> 8) & 255) + 0.0722 * (n & 255);
}
// chroma (max−min channel) of a #rrggbb — a cheap "how colorful" proxy; −1 for non-hex.
export function chroma(v) {
const m = /^#?([0-9a-fA-F]{6})$/.exec(String(v).trim());
if (!m) return -1;
const n = parseInt(m[1], 16);
const r = (n >> 16) & 255,
g = (n >> 8) & 255,
b = n & 255;
return Math.max(r, g, b) - Math.min(r, g, b);
}
// Browser user-agent default colors for links / visited links. These leak into a
// capture from any UNSTYLED <a> and are NOT brand colors — but being pure & saturated
// they beat a real accent on chroma alone. Never let one become the accent.
export const UA_DEFAULT_COLORS = new Set(
["#0000EE", "#0000FF", "#0000CC", "#1A0DAB", "#551A8B", "#EE0000"].map((c) => c.toUpperCase()),
);
export const ICON_FONT_PATTERN =
/(?:^|[\s_-])icons?(?:[\s_-]|$)|icomoon|font\s*-?awesome|glyphicons?|material\s*icons|feather\s*icons|(?:icon|glyph).*font|font.*(?:icon|glyph)|^vidaxlfont$/i;
export function isIconFont(name) {
return ICON_FONT_PATTERN.test(String(name));
}
// Semantic STATUS roles (green "positive", red "negative"/"error", amber "warning" …). Their HUE
// carries the meaning, so they are never a brand ACCENT — a status red is frequently the most
// chromatic color in a palette (e.g. #dc2626 chroma 182 beats a deep-blue accent #1E40AF chroma
// 145) and would otherwise win a pure chroma ranking, painting captions/highlights the error red.
// build-frame.mjs uses this same key set to protect status colors during the preset→brand remix.
export const STATUS_ROLE_KEY =
/(?:^|[-_])(?:positive|negative|success|error|warning|danger|good|bad|up|down|info|neutral|alert|caution|critical)(?:[-_]|$)/i;
// Pick the brand ACCENT — never by raw chroma alone, never a UA-default link color.
// Priority:
// 1) with capture colorStats → the colorful color that RECURS across the UI. The brand
// accent shows up in MANY roles (link text + icon + button + badge), whereas a one-off
// CTA fill appears in just one. So rank chromatic (chroma>40) candidates by role
// diversity first, then total prevalence, then interactive use, then chroma. This keeps
// a pervasive brand color (e.g. an indigo used everywhere) ahead of a single bright
// button fill (e.g. a lime used once) — the old "top interactiveBg" rule picked the
// latter. Requiring interactiveBg>0 is dropped so a text/icon-only accent can still win.
// 2) no stats → most chromatic color AFTER removing UA defaults + `exclude`.
// A stray default link color (e.g. #0000EE) can win under neither path.
export function pickAccent(stats, colors, exclude = []) {
const ban = new Set([...exclude, ...UA_DEFAULT_COLORS].map((c) => String(c).toUpperCase()));
const ok = (h) => /^#[0-9a-fA-F]{6}$/.test(String(h)) && !ban.has(String(h).toUpperCase());
// Prominence rank from the (frequency-ordered) `colors` palette: index 0 = most used.
// A saturated color sitting at the TAIL is almost always a one-off (a single CTA fill),
// not the brand accent — capture colorStats counts are too sparse to tell these apart
// (e.g. Linear's indigo and a lime CTA both register count≈1), but palette ORDER does.
const rank = new Map((colors ?? []).map((h, i) => [String(h).toUpperCase(), i]));
const prom = (h) => (rank.has(String(h).toUpperCase()) ? rank.get(String(h).toUpperCase()) : 1e9);
if (Array.isArray(stats) && stats.length) {
const roles = (s) =>
((s.interactiveBg || 0) > 0 ? 1 : 0) +
((s.textCount || 0) > 0 ? 1 : 0) +
((s.bgCount || 0) > 0 ? 1 : 0);
const a = stats
.filter((s) => ok(s?.hex) && chroma(s.hex) > 40)
.sort(
(x, y) =>
roles(y) - roles(x) || // used in MORE roles (link+icon+button) = the brand accent
prom(x.hex) - prom(y.hex) || // earlier in the palette = more prominent
(y.count || 0) - (x.count || 0) ||
(y.interactiveBg || 0) - (x.interactiveBg || 0) ||
chroma(y.hex) - chroma(x.hex),
);
if (a.length) return a[0].hex;
}
const c = (colors ?? [])
.map(String)
.filter(ok)
.sort((x, y) => chroma(y) - chroma(x));
return c[0];
}
// Derive brand roles from rich capture colorStats (areaBg / interactiveBg / textCount /
// maxArea) — by semantic FUNCTION, not luminance/chroma proxies. Returns null when stats
// are unusable, so the caller can fall back. canvas = the color painting the most real
// background area (the page ground, dark or light); ink = the dominant text color that
// actually contrasts with the canvas; accent via pickAccent.
export function brandRolesFromStats(stats, colorsInOrder) {
if (!Array.isArray(stats) || !stats.length) return null;
const v = stats.filter((s) => /^#[0-9a-fA-F]{6}$/.test(s?.hex || ""));
if (!v.length) return null;
const canvas = [...v].sort(
(a, b) =>
(b.areaBg || 0) - (a.areaBg || 0) ||
(b.maxArea || 0) - (a.maxArea || 0) ||
(b.bgCount || 0) - (a.bgCount || 0),
)[0]?.hex;
// pass the frequency-ordered palette (tokens.colors) so pickAccent can use palette
// PROMINENCE — colorStats counts alone are too sparse to rank rare accents.
const accent = pickAccent(v, colorsInOrder ?? v.map((s) => s.hex), [canvas]);
if (!canvas || !accent) return null;
const cl = lum(canvas) ?? 0;
const ink =
[...v]
.filter((s) => s.hex !== canvas && s.hex !== accent)
.sort((a, b) => (b.textCount || 0) - (a.textCount || 0))
.find((s) => Math.abs((lum(s.hex) ?? 0) - cl) > 64)?.hex ??
(cl > 128 ? "#000000" : "#FFFFFF");
const accent2 =
pickAccent(v, colorsInOrder ?? v.map((s) => s.hex), [canvas, ink, accent]) ?? accent;
return { ink, canvas, accent, accent2 };
}
// Map a list of [key, value] colors to semantic roles. ink = a dark/ink-named
// color (else darkest); canvas = a paper/cream/white-named color (else lightest);
// accents = whatever's left, ranked by chroma (the loudest color is almost always
// the brand accent) — UA-default link colors AND semantic status colors (positive/
// negative/error…) excluded so neither a stray <a> color nor a status red ever wins.
// For an unkeyed brand list, pass synthetic keys — name matching simply no-ops and it
// falls back to luminance/chroma, which is what we want. NOTE: when capture colorStats
// exist, prefer brandRolesFromStats() — it picks by function, not these proxies.
export function semanticColors(colors) {
if (!colors.length) return {};
const named = (re) => colors.find(([k]) => re.test(k));
const hexes = colors.filter(([, v]) => lum(v) != null);
const byLum = [...hexes].sort((a, b) => (lum(a[1]) ?? 1e9) - (lum(b[1]) ?? 1e9));
const pick = (m, fallback) => (m ? m[1] : fallback ? fallback[1] : undefined);
// "ink" must be a whole word-segment so "soft-pink"/"pink" don't match it.
const ink = pick(
named(/(?:^|[-_])ink(?:[-_]|$)|black|charcoal|^text(?:-dark)?$|outline|noir/i),
byLum[0] ?? colors[0],
);
const canvas = pick(
named(/cream|paper|canvas|white|bg|ground|surface|base|sand|parchment|off-?white|bone/i),
byLum[byLum.length - 1] ?? colors[colors.length - 1],
);
const accents = colors
.filter(
([k, v]) =>
v !== ink &&
v !== canvas &&
!UA_DEFAULT_COLORS.has(String(v).toUpperCase()) &&
!STATUS_ROLE_KEY.test(k), // a status red/green carries meaning by hue — never an accent
)
.sort((a, b) => chroma(b[1]) - chroma(a[1]))
.map(([, v]) => v);
return { ink, canvas, accent: accents[0] ?? ink, accent2: accents[1] ?? accents[0] ?? ink };
}
// Collect role→fontFamily under the top-level `typography:` block; pick a display
// + body family from the usual role names. Returns quoted families (or null).
export function parseFonts(md) {
const roles = {};
let inBlock = false;
for (const line of md.split(/\r?\n/)) {
if (/^typography:\s*$/.test(line)) {
inBlock = true;
continue;
}
if (!inBlock) continue;
if (/^\S/.test(line)) break;
const m = line.match(/^\s+([\w-]+):\s*\{[^}]*fontFamily:\s*"([^"]+)"/);
if (m) roles[m[1]] = m[2];
}
const q = (s) => (s ? `"${s}"` : null);
const body = roles.body ?? roles.subtitle ?? Object.values(roles)[0];
const display =
roles.display ??
roles.headline ??
roles["card-headline"] ??
roles["section-headline"] ??
roles["quote-display"] ??
roles.h1 ??
roles.h2 ??
roles.title ??
roles.hero ??
body;
// the monospace / chrome family (code, tags, ticks, page numbers) — so the remix can
// route a captured brand mono (Berkeley Mono, JetBrains Mono…) onto this role instead
// of the reading body. null when the preset has no distinct mono role.
const mono =
roles.mono ??
roles["mono-tag"] ??
roles["mono-chrome"] ??
roles["mono-tick"] ??
roles.code ??
roles.data ??
roles.pagenum ??
null;
return { display: q(display), body: q(body), mono: q(mono) };
}
scripts/lib/tokens.test.mjs
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import { dirname, join } from "node:path";
import test from "node:test";
import { fileURLToPath } from "node:url";
import { brandRolesFromStats, ICON_FONT_PATTERN, isIconFont } from "./tokens.mjs";
import { brandRolesFromStats as facelessBrandRolesFromStats } from "../../../faceless-explainer/scripts/lib/tokens.mjs";
import { brandRolesFromStats as prBrandRolesFromStats } from "../../../pr-to-video/scripts/lib/tokens.mjs";
const scriptsDir = join(dirname(fileURLToPath(import.meta.url)), "..", "..", "..");
test("recognizes brand-specific icon font names", () => {
assert.equal(isIconFont("vidaXLfont"), true);
assert.equal(isIconFont("BrandGlyphFont"), true);
assert.equal(isIconFont("Poppins"), false);
assert.equal(isIconFont("HelveticaFont"), false);
assert.equal(isIconFont("Airbnb Cereal Font"), false);
assert.equal(isIconFont("SF Pro Text Font"), false);
assert.equal(isIconFont("Uber Move Font"), false);
assert.equal(isIconFont("Circular Std font"), false);
assert.equal(isIconFont("brand-font"), false);
});
test("keeps sibling skill icon-font classifiers aligned", () => {
for (const skill of ["faceless-explainer", "pr-to-video"]) {
const source = readFileSync(join(scriptsDir, skill, "scripts", "build-frame.mjs"), "utf8");
assert.match(
source,
new RegExp(String.raw`ICON_FONT_PATTERN\s*=\s*${escapeRegExp(ICON_FONT_PATTERN)}`),
);
}
});
function escapeRegExp(value) {
return String(value).replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
}
test("preserves a prominent second accent used outside interactive backgrounds", () => {
const colors = ["#FFFFFF", "#2D1238", "#F3E62B", "#111111"];
const stats = [
{ hex: "#FFFFFF", areaBg: 1000, maxArea: 1000 },
{ hex: "#2D1238", textCount: 4, interactiveBg: 3 },
{ hex: "#F3E62B", textCount: 3, interactiveBg: 0 },
{ hex: "#111111", textCount: 20 },
];
assert.deepEqual(brandRolesFromStats(stats, colors), {
canvas: "#FFFFFF",
ink: "#111111",
accent: "#F3E62B",
accent2: "#2D1238",
});
for (const sibling of [facelessBrandRolesFromStats, prBrandRolesFromStats]) {
assert.deepEqual(sibling(stats, colors), brandRolesFromStats(stats, colors));
}
});
scripts/lib/transition-registry.mjs
// transition-registry.mjs — loader for this skill's vendored transition registry
// (./transitions.json). The registry is the curated Tier-B subset (transform /
// opacity / filter on the two frame clip wrappers `#el-<id>`, no overlay DOM) +
// each type's GSAP template. Vendored into the skill so it ships standalone; the
// recipes originate from the shared catalog skills/hyperframes-animation/
// transitions/ (css-*.md) — keep them in step if those shared recipes change.
import { readFileSync } from "node:fs";
import { resolve, dirname } from "node:path";
import { fileURLToPath } from "node:url";
const here = dirname(fileURLToPath(import.meta.url));
export const DEFAULT_REGISTRY_PATH = resolve(here, "./transitions.json");
let _cache = null;
export function loadTransitionRegistry(registryPath = DEFAULT_REGISTRY_PATH) {
if (_cache && _cache.path === registryPath) return _cache.data;
let data;
try {
data = JSON.parse(readFileSync(registryPath, "utf8"));
} catch (e) {
throw new Error(`transition registry not loadable at ${registryPath}: ${e.message}`);
}
if (!Array.isArray(data.transitions) || data.transitions.length === 0) {
throw new Error(`transition registry ${registryPath} has no transitions[]`);
}
_cache = { path: registryPath, data };
return data;
}
// Convenience: a Map name -> transition record.
export function transitionsByName(registryPath = DEFAULT_REGISTRY_PATH) {
const data = loadTransitionRegistry(registryPath);
const map = new Map();
for (const t of data.transitions) map.set(t.name, t);
return map;
}
scripts/lib/transitions.json
{
"_comment": "Vendored transition registry for the product-launch workflow — the curated Tier-B subset (transform/opacity/filter on the two frame clip wrappers #el-<id>, no overlay DOM, no per-frame cooperation). Each type carries its GSAP template; the transitions.mjs injector stamps it onto window.__timelines[\"main\"]. Recipes originate from the shared catalog skills/hyperframes-animation/transitions/ (css-*.md) — keep in step if those change. Token placeholders the injector substitutes: __OLD__ (#el-<from>), __NEW__ (#el-<to>), __T__ (overlap-start s), __DUR__ (this boundary's duration), __DX__/__DXIN__ (horizontal travel + incoming offset), __DY__/__DYIN__ (vertical).",
"transitions": [
{
"name": "crossfade",
"energy": "any",
"default_duration_s": 0.5,
"directions": [],
"source": "css-dissolve.md",
"gsap_template": [
"tl.to(__OLD__, { opacity: 0, duration: __DUR__, ease: \"power2.inOut\" }, __T__);",
"tl.fromTo(__NEW__, { opacity: 0 }, { opacity: 1, duration: __DUR__, ease: \"power2.inOut\" }, __T__);"
]
},
{
"name": "blur-crossfade",
"energy": "calm",
"default_duration_s": 0.6,
"directions": [],
"source": "css-dissolve.md",
"note": "Default when the two frames' #root backgrounds differ a lot — the blur masks the background-color clash a plain crossfade would expose.",
"gsap_template": [
"tl.to(__OLD__, { filter: \"blur(10px)\", scale: 1.03, opacity: 0, duration: __DUR__, ease: \"power2.inOut\" }, __T__);",
"tl.fromTo(__NEW__, { filter: \"blur(10px)\", scale: 0.97, opacity: 0 }, { filter: \"blur(0px)\", scale: 1, opacity: 1, duration: __DUR__, ease: \"power2.inOut\" }, __T__);"
]
},
{
"name": "push-slide",
"energy": "medium",
"default_duration_s": 0.5,
"directions": ["LEFT", "RIGHT", "UP", "DOWN"],
"default_direction": "LEFT",
"source": "css-push.md",
"note": "Directional. The injector picks __DX__/__DY__ from the direction and emits the horizontal OR vertical pair (not both).",
"gsap_template_horizontal": [
"tl.to(__OLD__, { x: __DX__, duration: __DUR__, ease: \"power3.inOut\" }, __T__);",
"tl.fromTo(__NEW__, { x: __DXIN__, opacity: 1 }, { x: 0, duration: __DUR__, ease: \"power3.inOut\" }, __T__);"
],
"gsap_template_vertical": [
"tl.to(__OLD__, { y: __DY__, duration: __DUR__, ease: \"power3.inOut\" }, __T__);",
"tl.fromTo(__NEW__, { y: __DYIN__, opacity: 1 }, { y: 0, duration: __DUR__, ease: \"power3.inOut\" }, __T__);"
]
},
{
"name": "zoom-through",
"energy": "high",
"default_duration_s": 0.4,
"directions": [],
"source": "css-scale.md",
"gsap_template": [
"tl.to(__OLD__, { scale: 2.5, opacity: 0, filter: \"blur(8px)\", duration: __DUR__, ease: \"power3.in\" }, __T__);",
"tl.fromTo(__NEW__, { scale: 0.5, opacity: 0, filter: \"blur(8px)\" }, { scale: 1, opacity: 1, filter: \"blur(0px)\", duration: __DUR__, ease: \"power3.out\" }, __T__);"
]
},
{
"name": "squeeze",
"energy": "medium",
"default_duration_s": 0.4,
"directions": [],
"source": "css-push.md",
"note": "Old compresses to a vertical line on the left edge; new expands from the right edge. Incoming starts off (scaleX 0) so its higher-track stacking is harmless.",
"gsap_template": [
"tl.to(__OLD__, { scaleX: 0, transformOrigin: \"left center\", duration: __DUR__, ease: \"power3.inOut\" }, __T__);",
"tl.fromTo(__NEW__, { scaleX: 0, transformOrigin: \"right center\", opacity: 1 }, { scaleX: 1, transformOrigin: \"right center\", duration: __DUR__, ease: \"power3.inOut\" }, __T__);"
]
}
],
"default_high_energy": "zoom-through",
"default_calm": "blur-crossfade",
"max_duration_s": 2.0
}
scripts/media-contract.test.mjs
import { test } from "node:test";
import assert from "node:assert/strict";
import { existsSync, mkdirSync, mkdtempSync, readFileSync, writeFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { dirname, join } from "node:path";
import { tmpdir } from "node:os";
import { spawnSync } from "node:child_process";
const skillDir = join(dirname(fileURLToPath(import.meta.url)), "..");
test("frame worker documents the approved video-hoist contract", () => {
const instructions = readFileSync(join(skillDir, "sub-agents", "frame-worker.md"), "utf8");
assert.match(instructions, /data-frame-video="approved"/);
assert.match(instructions, /assemble-index\.mjs.*hoists it to the host root/i);
assert.match(instructions, /Audio remains orchestrator-owned/i);
});
test("assemble hoists an approved timed frame video to the host root", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-"));
mkdirSync(join(project, "compositions"));
const framePath = join(project, "compositions", "frame-1.html");
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1 — Demo\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
framePath,
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video data-frame-video="approved" data-frame-video-x="0" data-frame-video-y="0" data-frame-video-width="1920" data-frame-video-height="1080" src="https://cdn.example/clip.mp4" poster="poster.png" preload="auto" muted playsinline loop style="background:url(https://evil.example/x)" nonce="unsafe" onerror="alert(1)" srcdoc="<script>alert(2)</script>" data-start="0.25" data-duration="1.5" data-media-start="12.75" data-track-index="7"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.equal(result.status, 0, result.stderr);
const index = readFileSync(join(project, "index.html"), "utf8");
const frame = readFileSync(framePath, "utf8");
assert.match(index, /data-start="0\.25"/);
assert.match(index, /data-duration="1\.5"/);
assert.match(index, /data-media-start="12\.75"/);
assert.match(index, /data-track-index="1007"/);
assert.match(index, /src="https:\/\/cdn\.example\/clip\.mp4"/);
assert.match(index, /poster="poster\.png"/);
assert.match(index, /preload="auto"/);
assert.match(index, /\smuted(?:\s|>)/);
assert.match(index, /\splaysinline(?:\s|>)/);
assert.match(index, /\sloop(?:\s|>)/);
assert.doesNotMatch(index, /onerror=/i);
assert.doesNotMatch(index, /srcdoc=/i);
assert.doesNotMatch(index, /nonce=/i);
assert.match(
index,
/style="position:absolute;left:0px;top:0px;width:1920px;height:1080px;object-fit:cover"/,
);
assert.doesNotMatch(frame, /<video\b/i);
});
test("assemble preserves approved-video geometry through sanitized host CSS", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-layout-"));
mkdirSync(join(project, "compositions"));
const framePath = join(project, "compositions", "frame-1.html");
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
framePath,
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video class="clip approved-demo" data-frame-video="approved" data-frame-video-x="120" data-frame-video-y="240" data-frame-video-width="960" data-frame-video-height="540" data-frame-video-fit="cover" src="clip.mp4" style="position:fixed;background:url(https://evil.example/x)" nonce="unsafe" onerror="alert(1)" srcdoc="<script>alert(2)</script>" data-start="0" data-duration="1" data-track-index="7"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.equal(result.status, 0, result.stderr);
const index = readFileSync(join(project, "index.html"), "utf8");
assert.match(
index,
/style="position:absolute;left:120px;top:240px;width:960px;height:540px;object-fit:cover"/,
);
assert.doesNotMatch(index, /approved-demo/);
assert.doesNotMatch(index, /evil\.example/);
assert.doesNotMatch(index, /position:fixed/);
assert.doesNotMatch(index, /data-frame-video-(?:x|y|width|height|fit)=/);
assert.doesNotMatch(index, /onerror=|srcdoc=|nonce=/i);
assert.doesNotMatch(readFileSync(framePath, "utf8"), /<video\b/i);
});
test("rejects partial or unsafe approved-video layout geometry", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-layout-invalid-"));
mkdirSync(join(project, "compositions"));
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
join(project, "compositions", "frame-1.html"),
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video data-frame-video="approved" data-frame-video-x="0" data-frame-video-y="0" data-frame-video-width="calc(100% + 1px)" data-frame-video-height="1080" data-frame-video-fit="cover" src="clip.mp4" data-start="0" data-duration="1" data-track-index="0"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.notEqual(result.status, 0);
assert.match(result.stderr, /approved frame video layout.*finite numeric/i);
assert.equal(existsSync(join(project, "index.html")), false);
});
test("rejects an approved video without mandatory layout geometry", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-layout-missing-"));
mkdirSync(join(project, "compositions"));
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
join(project, "compositions", "frame-1.html"),
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video data-frame-video="approved" src="clip.mp4" data-start="0" data-duration="1" data-track-index="0"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.notEqual(result.status, 0);
assert.match(result.stderr, /approved frame video layout.*finite numeric/i);
assert.equal(existsSync(join(project, "index.html")), false);
});
test("rejects empty approved-video layout coordinates", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-layout-empty-"));
mkdirSync(join(project, "compositions"));
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
join(project, "compositions", "frame-1.html"),
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video data-frame-video="approved" data-frame-video-x="" data-frame-video-y="0" data-frame-video-width="1920" data-frame-video-height="1080" src="clip.mp4" data-start="0" data-duration="1" data-track-index="0"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.notEqual(result.status, 0);
assert.match(result.stderr, /approved frame video layout.*finite numeric/i);
assert.equal(existsSync(join(project, "index.html")), false);
});
test("rejects an approved video with missing admission timing", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-missing-"));
mkdirSync(join(project, "compositions"));
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
join(project, "compositions", "frame-1.html"),
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"><video data-frame-video="approved" src="clip.mp4" data-duration="1" data-track-index="0"></video></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline();</script></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.notEqual(result.status, 0);
assert.match(result.stderr, /must declare quoted data-start/i);
});
test("does not hoist declarations hidden in comments or scripts", () => {
const project = mkdtempSync(join(tmpdir(), "hf-frame-video-hidden-"));
mkdirSync(join(project, "compositions"));
writeFileSync(
join(project, "STORYBOARD.md"),
"---\nformat: 16:9\n---\n\n## Frame 1\n- status: built\n- duration: 2s\n- src: compositions/frame-1.html\n",
);
writeFileSync(
join(project, "compositions", "frame-1.html"),
`<html><body><div id="root" data-composition-id="frame-1" data-width="1920" data-height="1080"></div><script>window.__timelines = {}; window.__timelines["frame-1"] = gsap.timeline(); const s = '<video data-frame-video="approved" data-start="0" data-duration="1" data-track-index="1"></video>';</script><!-- <video data-frame-video="approved" data-start="0" data-duration="1" data-track-index="2"></video> --></body></html>`,
);
const result = spawnSync(
process.execPath,
[join(skillDir, "scripts", "assemble-index.mjs"), "--hyperframes", project],
{ encoding: "utf8" },
);
assert.equal(result.status, 0, result.stderr);
assert.doesNotMatch(readFileSync(join(project, "index.html"), "utf8"), /data-track-index="1001"/);
});
scripts/stage-assets.mjs
#!/usr/bin/env node
// stage-assets.mjs — copy each frame's named asset_candidates from capture/ into
// assets/ so Step 5 frame workers reference real files and the live preview
// (Step 3 / Step 6) shows them. Runs at Step 4 close, once visual design is
// locked. assemble-index.mjs re-runs the same staging idempotently as a backstop.
//
// Reads: --storyboard STORYBOARD.md, --hyperframes <project root>.
// Writes: assets/<basename> for each named, found asset.
// Exit 0 always once the storyboard parses — a missing asset is a non-fatal
// anomaly (the frame would 404 it), not a contract break.
import { existsSync, readFileSync } from "node:fs";
import { join, resolve } from "node:path";
import { parseStoryboard } from "./lib/storyboard.mjs";
import { stageAssets } from "./lib/assets.mjs";
const argv = process.argv.slice(2);
const flag = (name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
function die(msg) {
console.error(`✗ stage-assets.mjs: ${msg}`);
process.exit(1);
}
const hyperframesDir = resolve(flag("hyperframes", "."));
const storyboardPath = resolve(flag("storyboard", join(hyperframesDir, "STORYBOARD.md")));
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const { staged, wanted, anomalies } = stageAssets({ hyperframesDir, frames: manifest.frames });
console.log(`✓ staged ${staged}/${wanted.size} asset(s) into assets/`);
if (anomalies.length) {
console.log(`\nanomalies (non-fatal):`);
for (const a of anomalies) console.log(` - ${a}`);
}
scripts/stage-assets.test.mjs
import assert from "node:assert/strict";
import { existsSync, mkdirSync, mkdtempSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import test from "node:test";
import { stageAssets } from "./lib/assets.mjs";
// ── captured SVGs are stageable ──────────────────────────────────────────────
// Regression: `hyperframes capture` extracts inline SVGs into capture/assets/svgs/
// (assetDownloader.ts), and the capture manifest advertises them to the agent as
// `assets/svgs/<name>.svg`, so a frame names one in `asset_candidates` exactly as
// it names a screenshot. stageAssets only searched capture/{assets,assets/videos,
// screenshots}, so every captured SVG resolved to nothing — reported as a
// non-fatal anomaly, and the frame 404'd the logo it had been told to use.
function projectWithCapturedSvg() {
const dir = mkdtempSync(join(tmpdir(), "product-launch-stage-assets-"));
mkdirSync(join(dir, "capture/assets/svgs"), { recursive: true });
mkdirSync(join(dir, "capture/screenshots"), { recursive: true });
writeFileSync(join(dir, "capture/assets/svgs/brand-mark.svg"), "<svg/>");
writeFileSync(join(dir, "capture/screenshots/hero.png"), "png");
return dir;
}
const frames = [
{ extra: { asset_candidates: "assets/svgs/brand-mark.svg — the mark; assets/hero.png — hero" } },
];
test("stages an SVG that capture wrote into capture/assets/svgs/", () => {
const dir = projectWithCapturedSvg();
const { staged, wanted, anomalies } = stageAssets({ hyperframesDir: dir, frames });
assert.equal(wanted.size, 2);
assert.equal(staged, 2, `expected both assets staged, got anomalies: ${anomalies.join("; ")}`);
assert.deepEqual(anomalies, []);
assert.ok(existsSync(join(dir, "assets/brand-mark.svg")));
assert.ok(existsSync(join(dir, "assets/hero.png")));
});
test("still reports an asset that exists nowhere under capture/", () => {
const dir = projectWithCapturedSvg();
const { staged, anomalies } = stageAssets({
hyperframesDir: dir,
frames: [{ extra: { asset_candidates: "assets/svgs/absent.svg — never captured" } }],
});
assert.equal(staged, 0);
assert.equal(anomalies.length, 1);
assert.match(anomalies[0], /absent\.svg/);
});
scripts/transitions.mjs
#!/usr/bin/env node
// transitions.mjs — inter-frame transition injector + verifier for the video workflow.
//
// inject — read STORYBOARD frame order + each frame's transition_in, overlap
// the frame clip wrappers in index.html, and stamp the GSAP template.
// verify — deterministic gate over the injector's output.
//
// transition_in (written by story-design on the INCOMING frame) names a registry
// type directly: crossfade | blur-crossfade | push-slide | zoom-through | squeeze,
// optionally "<type> <DIR>" / "<type> <N>s" (e.g. "push-slide LEFT", "crossfade
// 0.4s"). `cut` / `none` / empty ⇒ hard cut (no overlap, no stamp).
//
// Mechanics — EXTEND-OUTGOING-ONLY (keeps voice/SFX/captions synced; their timing
// is keyed to the original frame start). At boundary i→i+1 (type = the incoming
// frame's transition_in): extend ONLY the outgoing wrapper's data-duration by
// `dur` so it holds its final frame across the window; do NOT move any data-start;
// the incoming — already present from the cut on a higher track — fades/pushes in
// over it. Then 0/1-ping-pong ALL frame clips' data-track-index (adjacent
// overlapping wrappers never share a track — lint timeline_track_too_dense) and
// stamp the token-substituted GSAP template into __timelines["main"] at T =
// incoming start. captions(2)/voice(10)/bgm(11)/sfx(20+) are never touched.
//
// node transitions.mjs inject --storyboard ./STORYBOARD.md --hyperframes .
// node transitions.mjs verify --storyboard ./STORYBOARD.md --index ./index.html
import { existsSync, readFileSync, writeFileSync } from "node:fs";
import { join, resolve } from "node:path";
import { parseStoryboard } from "./lib/storyboard.mjs";
import { parseFormat } from "./lib/dimensions.mjs";
import { loadTransitionRegistry, transitionsByName } from "./lib/transition-registry.mjs";
import { padFrameInternalDuration } from "./lib/pad-frame-duration.mjs";
const flag = (argv, name, def) => {
const i = argv.indexOf(`--${name}`);
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : def;
};
const NO_TRANSITION = new Set(["cut", "none", ""]);
const r3 = (x) => Number(x.toFixed(3));
// transition_in → { type, direction?, dur? } | null (hard cut).
function parseTransitionIn(raw) {
const s = (raw ?? "").trim();
if (NO_TRANSITION.has(s.toLowerCase())) return null;
const parts = s.split(/\s+/);
const spec = { type: parts[0].toLowerCase() };
for (const p of parts.slice(1)) {
const m = p.match(/^(\d+(?:\.\d+)?)s?$/);
if (m) spec.dur = Number(m[1]);
else spec.direction = p.toUpperCase();
}
return spec;
}
// Mounted STORYBOARD frames present in index.html, in document order: { id, frame }.
function mountedFramesInOrder(manifest, html) {
const out = [];
for (const f of manifest.frames) {
if (!f.src) continue;
const id = f.src
.split("/")
.pop()
.replace(/\.html?$/i, "");
if (html.includes(`id="el-${id}"`)) out.push({ id, frame: f });
}
return out;
}
// Frame clip wrappers parsed out of index.html (ids carry hyphens; excludes
// el-captions and audio by keying off the known frame-id set). The id is matched
// from anywhere in the tag's attribute list — never assume it is the first attribute
// (the index assembler emits data-hf-id before id, so an id-first regex finds nothing
// and inject crashes on the empty clip map).
function parseFrameClips(html, frameIds) {
const clipRe = /<div\b([^>]*)><\/div>/g;
const clips = new Map();
let m;
while ((m = clipRe.exec(html)) !== null) {
const attrs = m[1];
const idm = attrs.match(/\bid="el-([A-Za-z0-9_-]+)"/);
if (!idm || !frameIds.has(idm[1])) continue;
const num = (re) => {
const x = attrs.match(re);
return x ? Number(x[1]) : null;
};
clips.set(idm[1], {
id: idm[1],
block: m[0],
start: num(/data-start="([\d.]+)"/),
duration: num(/data-duration="([\d.]+)"/),
track: num(/data-track-index="(\d+)"/) ?? 0,
});
}
return clips;
}
// The host wrapper is extended across an outgoing transition, so the mounted
// frame must remain visually populated for the same local-time window. Extend
// the frame root and every non-audio timed element that reached the original
// storyboard boundary. This also repairs worker files whose root was already
// inflated while their ground/content clips still ended at the synced duration.
function extendFrameTail(hyperframesDir, frame, baseDuration, targetDuration, die) {
if (!frame?.src || targetDuration <= baseDuration) return;
const framePath = join(hyperframesDir, frame.src);
let html;
try {
html = readFileSync(framePath, "utf8");
} catch {
die(`outgoing frame file not found at ${framePath}`);
}
const compId = frame.src
.split("/")
.pop()
.replace(/\.html?$/i, "");
const EPS = 0.011;
let foundRoot = false;
let extended = 0;
const rewritten = html.replace(/<([A-Za-z][\w:-]*)\b([^>]*)>/g, (tag, name, attrs) => {
const durationMatch = attrs.match(/\bdata-duration="([\d.]+)"/);
if (!durationMatch) return tag;
const duration = Number(durationMatch[1]);
if (!Number.isFinite(duration)) return tag;
const compositionMatch = attrs.match(/\bdata-composition-id="([^"]+)"/);
if (compositionMatch?.[1] === compId && !foundRoot) {
foundRoot = true;
return tag.replace(/\bdata-duration="[\d.]+"/, `data-duration="${targetDuration}"`);
}
if (name.toLowerCase() === "audio") return tag;
const startMatch = attrs.match(/\bdata-start="([\d.]+)"/);
if (!startMatch) return tag;
const start = Number(startMatch[1]);
const end = start + duration;
if (end < baseDuration - EPS || end >= targetDuration - EPS) return tag;
extended++;
return tag.replace(/\bdata-duration="[\d.]+"/, `data-duration="${r3(targetDuration - start)}"`);
});
if (!foundRoot) die(`${frame.src} has no data-composition-id="${compId}" root`);
writeFileSync(framePath, rewritten);
console.log(
` ${compId}: extended root + ${extended} tail clip(s) ${baseDuration}s→${targetDuration}s`,
);
}
// Resolve a transition_in spec to a registry record (calm default on unknown).
function resolveRecord(spec, byName, reg, warn) {
let rec = byName.get(spec.type);
if (!rec) {
rec = byName.get(reg.default_calm);
warn(`transition_in "${spec.type}" not in registry — using ${reg.default_calm}`);
}
return rec;
}
function resolveDur(spec, rec, reg) {
let dur = spec.dur ?? rec.default_duration_s ?? 0.5;
return Math.min(dur, reg.max_duration_s ?? 2.0);
}
// GSAP lines for one transition record (token substitution).
function buildGsap(rec, fromId, toId, dur, T, direction, canvasW, canvasH, die) {
const subs = {
__OLD__: `"#el-${fromId}"`,
__NEW__: `"#el-${toId}"`,
__T__: String(T),
__DUR__: String(dur),
};
let template;
if (rec.directions && rec.directions.length > 0) {
const dir = (direction || rec.default_direction || rec.directions[0]).toUpperCase();
const vertical = dir === "UP" || dir === "DOWN";
template = vertical ? rec.gsap_template_vertical : rec.gsap_template_horizontal;
if (!template)
die(`transition ${rec.name}: missing ${vertical ? "vertical" : "horizontal"} template`);
if (vertical) {
const dy = dir === "UP" ? -canvasH : canvasH;
subs.__DY__ = String(dy);
subs.__DYIN__ = String(-dy);
} else {
const dx = dir === "LEFT" ? -canvasW : canvasW;
subs.__DX__ = String(dx);
subs.__DXIN__ = String(-dx);
}
} else {
template = rec.gsap_template;
if (!template) die(`transition ${rec.name}: missing gsap_template`);
}
return template.map((line) => {
let out = line;
for (const [k, v] of Object.entries(subs)) out = out.split(k).join(v);
return out;
});
}
function runInject(argv) {
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
const indexPath = join(hyperframesDir, "index.html");
const die = (msg) => {
console.error(`✗ transitions inject: ${msg}`);
process.exit(1);
};
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const { width: CW, height: CH } = parseFormat(manifest.globals.format);
const reg = loadTransitionRegistry();
const byName = transitionsByName();
// Read directly and handle ENOENT here, rather than an existsSync precheck —
// the check→write pair (write-back below) is a TOCTOU race CodeQL flags.
let html = "";
try {
html = readFileSync(indexPath, "utf8");
} catch {
die(`index.html not found at ${indexPath} — run assemble-index.mjs first`);
}
const order = mountedFramesInOrder(manifest, html);
if (order.length === 0) die("no frame clips found in index.html");
const frameIds = new Set(order.map((x) => x.id));
const clips = parseFrameClips(html, frameIds);
const gsapLines = [];
const applied = [];
for (let i = 1; i < order.length; i++) {
const spec = parseTransitionIn(order[i].frame.transitionIn);
if (!spec) continue; // hard cut
const incoming = clips.get(order[i].id);
const outgoing = clips.get(order[i - 1].id);
const rec = resolveRecord(spec, byName, reg, (m) =>
console.error(` ! frame ${order[i].id}: ${m}`),
);
const dur = resolveDur(spec, rec, reg);
const T = r3(incoming.start); // cut = incoming start (frames tile)
const baseDuration = outgoing.duration;
outgoing.duration = r3(baseDuration + dur); // extend outgoing only
extendFrameTail(hyperframesDir, order[i - 1].frame, baseDuration, outgoing.duration, die);
padFrameInternalDuration(
hyperframesDir,
order[i - 1].frame.src,
outgoing.id,
outgoing.duration,
);
gsapLines.push(
...buildGsap(rec, outgoing.id, incoming.id, dur, T, spec.direction, CW, CH, die),
);
applied.push({ from: outgoing.id, to: incoming.id, type: rec.name, dur, T });
}
if (applied.length === 0) {
console.log(`✓ transitions inject: 0 transitions (all cuts) — index.html unchanged`);
return;
}
// 0/1 ping-pong all frame clips in play order.
const ordered = [...clips.values()].sort((a, b) => a.start - b.start || a.id.localeCompare(b.id));
ordered.forEach((c, i) => {
c.track = i % 2;
});
// rewrite each clip block: start unchanged; duration possibly extended; track ping-ponged.
for (const c of clips.values()) {
const nb = c.block
.replace(/data-duration="[\d.]+"/, `data-duration="${c.duration}"`)
.replace(/data-track-index="\d+"/, `data-track-index="${c.track}"`);
html = html.replace(c.block, nb);
}
// stamp the GSAP after the master timeline anchor.
const anchor = 'window.__timelines["main"] = gsap.timeline({ paused: true });';
if (!html.includes(anchor)) die("master timeline anchor not found in index.html");
// The transition tweens alone leave window.__timelines["main"] spanning only the
// last transition (e.g. 24.7s), shorter than the real composition. The Studio
// reads main.duration() as its master duration and parses clips against it, so a
// short master collapses its timeline (clips dropped, duration wrong, blank stage)
// — the render engine is unaffected (it trusts the root data-duration attr). Stamp
// a full-span anchor so main.duration() == composition total. Mirrors the
// `tl.to({}, { duration })` anchor captions.html already uses.
const rootDurMatch = html.match(/data-composition-id="main"[^>]*?data-duration="([\d.]+)"/);
const totalDur = rootDurMatch ? Number(rootDurMatch[1]) : null;
const block = [
anchor,
" // ── frame transitions (injected by transitions.mjs) ──",
' (function () { var tl = window.__timelines["main"];',
...gsapLines.map((l) => " " + l),
...(totalDur
? [
` tl.to({}, { duration: ${totalDur} }, 0); // full-span anchor — main.duration() == composition total (Studio master duration)`,
]
: []),
" })();",
].join("\n");
html = html.replace(anchor, block);
writeFileSync(indexPath, html);
console.log(`✓ transitions inject: ${applied.length} transition(s) stamped into index.html`);
for (const a of applied) console.log(` ${a.from}→${a.to}: ${a.type} ${a.dur}s @ T=${a.T}s`);
const tracks = ordered
.map((c) => `${c.id}[t${c.track} ${c.start}→${r3(c.start + c.duration)}]`)
.join(" ");
console.log(` tracks: ${tracks}`);
}
function runVerify(argv) {
const hyperframesDir = resolve(flag(argv, "hyperframes", "."));
const storyboardPath = resolve(flag(argv, "storyboard", join(hyperframesDir, "STORYBOARD.md")));
const indexPath = resolve(flag(argv, "index", join(hyperframesDir, "index.html")));
const bail = (msg) => {
console.error(`✗ transitions verify: ${msg}`);
process.exit(1);
};
if (!existsSync(storyboardPath)) bail("STORYBOARD.md not found");
if (!existsSync(indexPath)) bail("index.html not found");
const manifest = parseStoryboard(readFileSync(storyboardPath, "utf8"));
const html = readFileSync(indexPath, "utf8");
const order = mountedFramesInOrder(manifest, html);
const frameIds = new Set(order.map((x) => x.id));
const clips = parseFrameClips(html, frameIds);
const EPS = 0.011;
const overlaps = (a, b) =>
a.start < b.start + b.duration - EPS && b.start < a.start + a.duration - EPS;
const fail = [];
// (4) global no same-track overlap. This is this workflow's own lane convention,
// not a lint rule: nothing in the framework rejects same-track overlap.
const all = [...clips.values()];
for (let i = 0; i < all.length; i++)
for (let j = i + 1; j < all.length; j++) {
const a = all[i];
const b = all[j];
if (a.track === b.track && overlaps(a, b))
fail.push(`same-track overlap: ${a.id}[t${a.track}] & ${b.id}[t${b.track}]`);
}
const bm = html.match(/frame transitions \(injected[\s\S]*?\}\)\(\);/);
const txBlock = bm ? bm[0] : "";
let expected = 0;
for (let i = 1; i < order.length; i++) {
const spec = parseTransitionIn(order[i].frame.transitionIn);
if (!spec) continue;
expected++;
const to = clips.get(order[i].id);
const from = clips.get(order[i - 1].id);
if (!to || !from) {
fail.push(`boundary ${order[i - 1].id}→${order[i].id}: wrapper missing`);
continue;
}
if (!txBlock.includes(`"#el-${from.id}"`) || !txBlock.includes(`"#el-${to.id}"`))
fail.push(`boundary ${from.id}→${to.id}: injected block does not reference both ids`);
const overlapAmt = r3(from.start + from.duration - to.start);
if (overlapAmt <= 0) fail.push(`boundary ${from.id}→${to.id}: no overlap (${overlapAmt}s)`);
if (from.track === to.track)
fail.push(`boundary ${from.id}→${to.id}: both on track ${from.track}`);
}
if (expected > 0 && !txBlock)
fail.push(`${expected} transition(s) expected but no injected block found`);
if (fail.length) {
console.error(`✗ transitions verify: ${fail.length} failure(s):`);
for (const f of fail) console.error(` - ${f}`);
process.exit(1);
}
console.log(
`✓ transitions verify: ${expected} transition(s) verified (cross-track, overlap>0, both ids referenced, no same-track overlap)`,
);
}
const sub = process.argv[2];
const rest = process.argv.slice(3);
if (sub === "inject") runInject(rest);
else if (sub === "verify") runVerify(rest);
else {
console.error("usage: node transitions.mjs <inject|verify> [args...]");
process.exit(2);
}
SKILL.md
---
name: product-launch-video
description: "Turn a product or marketing URL, pasted script, or brief into a product launch / promo video — SaaS promos, feature reveals, product demos, app and company launches. Use when the user wants to market, launch, promote, or reveal a product; the default for any commercial URL. Site tours / showcases of a website route here too — the brief carries the show-it-as-is intent. Unclear → /hyperframes."
---
> **First, keep this skill fresh — confirm with the user before running:** `npx hyperframes skills update product-launch-video`. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them.
> **media-use**: Before sourcing audio/images/logos, call `/media-use` to resolve BGM/SFX/images from the HeyGen catalog and brand logos from their official sources. Run `--adopt` first to register existing assets. See `/media-use` skill.
> **figma source**: If the source is a figma.com URL, run `/figma` first — asset export, brand tokens, and components/storyboard reconstruction if needed — then build this workflow from its output. Don't drive Figma via raw MCP tools directly: that skips SVG sanitization, `.media/manifest.jsonl` provenance, and brand-token `var()` binding, so a later brand change can't propagate without a full re-import.
# Product Launch to HyperFrames
Use this skill to capture a product, understand its brand, plan a launch video, and build it frame by frame in HyperFrames.
> **The front door is `/hyperframes`.** You are the orchestrator. Run each step, verify its gate, and only then continue to the next step. This skill is for a **product being marketed, launched, promoted, or revealed**, including requests such as "promo for our site" when the purpose is promotional. A site tour / showcase ask stays here too: `BRIEF.md` carries the show-it-as-is intent, and the captured screens become the assets the video features. Any other intent, a bare "make a video", or any uncertainty → read `/hyperframes` first — the intent layer owns every route decision, and a fresh creation arriving here without `BRIEF.md` goes through it anyway (Setup's opening rule).
You are the orchestrator. Work in `videos/<project>/`. Run steps in order and pass each gate before continuing. User-gated steps are Step 0, Step 3, and Step 6. Read `../hyperframes-core/references/brief-contract.md` before Step 0 — it defines the gate types and how `BRIEF.md`'s `flow`/`storyboard` derive the mode that governs the Step 3/4/6 gates. Do every step yourself except Step 5, where you dispatch one sub-agent per frame. Do not put design or motion rules here; those live in the frame-worker sub-agent, this skill's local `../hyperframes-animation/rules/` + `../hyperframes-animation/blueprints/`, and `hyperframes-creative`.
Workflow: Step 0 setup -> `hyperframes.json`; Step 1 capture -> `capture/`; Step 2 design system -> `frame.md`; Step 3 storyboard/script -> `STORYBOARD.md` and `SCRIPT.md`; Step 3.1 audio -> `audio_meta.json`; Step 4 visual design -> enriched `STORYBOARD.md`; Step 5 frames -> `compositions/frames/NN-*.html` and `index.html`; Step 6 final render -> `renders/video.mp4`.
---
## Step 0: Setup
Goal: Enter with a confirmed brief, create the HyperFrames project, and make the brief durable.
**The brief is confirmed by the intent layer, not by questions asked here.** Opening rule, in order: **(1)** `BRIEF.md` exists → read it and ask nothing — the brief is settled, and its `flow`/`storyboard` derive the mode (brief contract § 1). **(2)** No `BRIEF.md` but the project exists (`hyperframes.json` / `STORYBOARD.md` on disk) → resume from the storyboard's frontmatter and the recorded preferences; never re-interrogate a half-built project. **(3)** Neither — a fresh creation request that arrived here directly → read `/hyperframes` and run its intent layer (`references/intent-interview.md`): it checks recipes and remembered defaults, conducts this route's questions (`../hyperframes/references/routes/product-launch-video.md`), and hands back the locked brief. Edit requests skip all of this — go do the edit.
Initialize only if `hyperframes.json` is missing. Name `<project>` from the brand or domain in kebab-case, such as `acme-promo`; never use workspace name or timestamp.
`npx hyperframes init "videos/<project>" --non-interactive --example=blank --skill=product-launch-video` — `init` checks the installed skills against the latest on GitHub and updates the global set if any are out of date.
After init, let `<PROJECT_ROOT>` be `videos/<project>` and run every subsequent relative-path command with that directory as its working directory. In the commands below, `.` means `<PROJECT_ROOT>`; never write `.media`, `capture`, or output files in the caller directory.
**Write `BRIEF.md` immediately after init** (never before — `init` refuses a non-empty directory): the intent layer's locked brief, shape per `../hyperframes-core/references/brief-format.md`. Resolve `<MEDIA_DIR>` as the installed `/media-use` skill directory. Then record each preference-backed answer with `node <MEDIA_DIR>/scripts/prefs.mjs record --hyperframes .` (`brief-format.md` names the subset). If the intent layer adopted a recipe, run `node <MEDIA_DIR>/scripts/recipe.mjs use --hyperframes . --name <name>`; it copies its `frame.md` into the project (Step 2 is then skipped) and returns the skeletons Step 3 drafts from. A recipe fills answers, not approvals; the review gates still run.
**Show sign-in status before proceeding past Setup** — run `npx hyperframes auth status` and relay its output verbatim. It reports whether voice/BGM will use HeyGen or local engines and, when signed out, how to sign in. Note the exit code contract: `auth status` **exits 1 when not signed in** (and when the stored credential is rejected) — that non-zero exit is the normal signed-out state, not a command failure, so don't treat it as an error, don't retry it, and don't chain it with `&&`/`set -e` in a way that would abort the workflow. Apply one branch:
- **Collaborative:** wait for the user to sign in or explicitly choose `offline` / `go`.
- **Autonomous:** state the status and continue through the available local engines.
Do not silently omit a required capability when no offline provider exists; surface the blocker. Do not fold this decision into another question or write keys into a per-repo `.env`. Auth ownership and offline fallbacks: `/media-use` `references/setup-providers.md` § Providers.
**Gate:** `hyperframes.json` and `BRIEF.md` exist; the preference-backed answers were recorded (brief contract § 2); sign-in status was shown (signed in, or continuing offline).
---
## Step 1: Capture assets
Goal: Collect the source material, brand signals, and usable assets for the video.
Classify the input and choose the path. Explicit URL -> capture it and use the site for narration and assets. Pasted script/brief -> save verbatim as `user_script.txt`; `VO_MODE` (verbatim or restructured) comes from `BRIEF.md` — the intent layer asks it when a script arrives (ask once here only if the brief somehow lacks it). Then resolve capture target: URL in text -> use it; brand name only -> `WebSearch`, confirm URL in one line, then crawl; no URL/site (or the brief says don't scrape) -> no-capture path.
Run capture with: `npx hyperframes capture "<URL>" -o ./capture --json`. Keep the default
post-navigation budget unless the caller owns a smaller deadline; then pass a positive
`--capture-budget <milliseconds>` that leaves time for downstream work. `--timeout` controls page
navigation only. Use `--skip-vision` only when optional image captioning is intentionally disabled.
Inspect the command result and output directory immediately. A non-zero exit, JSON `ok: false`, or
`capture/BLOCKED.md` is a **hard stop** for the capture path: report the recorded reason and do not
consume partial screenshots, DOM, tokens, or assets. Do not manufacture a synthetic no-capture
fallback after a failed URL capture. Continue through the no-capture path only when the original
brief supplied the source material, or when the user explicitly switches to a provided screenshot
or brief after the failure.
Warnings such as `very little text content` together with an empty asset catalog are not proof of a
usable page. For a site tour or show-it-as-is brief, require trustworthy captured structure or a
provided screenshot; if neither exists, stop. Do not invent or rebuild the page merely because the
capture is unusable.
For a site tour or show-it-as-is brief, the captured page is the visual source of truth. Use the real screenshot instead of rebuilding the full website in HTML. If the shot needs internal movement, keep the screenshot as the base and overlay real captured assets at measured positions, or rebuild only the one component that moves. For a scroll shot, animate the viewport over `capture/screenshots/full-page.png` — the 1x plate of the whole document, pixel-exact for a 1920-wide viewport travelling down it. It is absent when the page was too tall to capture in one piece; fall back to the overlapping scroll-position shots in the same directory. Pushing in past 1:1 wants its own 2x capture of that region instead, since the plate has no headroom above 1x. Recreate the whole page only when the user explicitly asks for a stylized interpretation; an unusable capture alone is not authorization.
If `GEMINI_API_KEY`, `GOOGLE_API_KEY`, or an OpenRouter key exists, capture auto-captions assets into `capture/extracted/asset-descriptions.md`. This is not a review gate. Without a vision key, use DOM context and continue.
No-capture path: create `capture/extracted/tokens.json`, `capture/extracted/visible-text.txt`, `capture/extracted/asset-descriptions.md`, and `capture/assets/` by hand. `tokens.json` should be `{ "title": "", "description": "", "colors": [], "fonts": [] }`; fill title/description from the brief when possible. `visible-text.txt` contains the full brief or script. `asset-descriptions.md` should say no assets were captured unless the user gave asset notes.
**Gate:** capture JSON reported `ok: true`; `capture/BLOCKED.md` does not exist;
`capture/extracted/tokens.json`, `capture/extracted/visible-text.txt`,
`capture/extracted/asset-descriptions.md`, and `capture/assets/` exist; and you can state the brand in
one clear sentence. Treat `asset-descriptions.md` as the main asset inventory. If it is missing after
real capture, stop and report capture incomplete. Warnings about a degraded optional phase are
acceptable only when this structural gate still passes.
---
## Step 2: Design System
Goal: Choose one shipped frame preset; a script turns it into this video's `frame.md` + caption skin.
When `BRIEF.md` names a `style_preset` — the user picked it by eye from the showcases at the intent layer — use it; the judgment call is yours only when the brief is silent. Then you make the one call — **which preset**: read `../hyperframes-creative/references/design-spec.md` and pick the preset whose look best fits the brand and brief. Then run:
```bash
node <SKILL_DIR>/scripts/build-frame.mjs --preset <name> --hyperframes .
```
The script does the rest deterministically: copies the preset's `FRAME.md` → `frame.md` and **remixes** it onto the brand tokens in `capture/extracted/tokens.json` (brand colors mapped onto the preset's color keys by role — ink, canvas, accents — keeping keys/structure/components; the preset's display + body fonts swapped for the brand's), copies the preset's caption skin to `.hyperframes/caption-skin.html`, and self-validates (exits 1 on a broken mapping). Proceed to the next step as soon as it exits 0 — no hand-editing of the spec.
`tokens.json` with no brand colors/fonts (e.g. no capture) → the script keeps the preset's own palette, a complete shippable design. If the brief names brand colors/fonts the capture missed, add them to `capture/extracted/tokens.json` before running (or use the user's `design.md` to populate it); only adjust `frame.md` by hand afterward if a mapping truly needs it.
**Gate:** `build-frame.mjs` exited 0 — `frame.md` exists from a named preset, and (when the preset ships one) `.hyperframes/caption-skin.html` exists as the caption skin source; the chosen preset was recorded as a preference (`--key style_preset --workflow <this workflow>`, brief contract § 2).
---
## Step 3: Storyboard and Script
Goal: Turn the brief and captured material into an approved frame-by-frame story plan.
Read `../hyperframes-creative/references/story-spine.md` (hook language, value-before-evidence, storyboard-as-proposal, source-traceable visuals), `references/story-design.md`, `../hyperframes-animation/blueprints-index.md`, `../hyperframes-core/references/storyboard-format.md`, and `../hyperframes-core/references/script-format.md`. Use them to write `STORYBOARD.md` and, when narration is needed, `SCRIPT.md`. Set the frontmatter `duration:` from the brief's `length` — a rough expectation; assembly reports where the cut lands against it.
Use `story-design.md` for story blueprint, hook, persuasion logic, beats, `VO_MODE`, and asset choices. As a **soft guide**, consult the role→blueprint menu in `../hyperframes-animation/blueprints-index.md`: for each beat, note a candidate blueprint id when one fits. Story truth still decides which beats exist — never force a beat to fit a blueprint, and never invent a beat just because a proven shape is available. Choose each visual frame's `asset_candidates` from `capture/extracted/asset-descriptions.md` (the canonical inventory) — don't browse raw `capture/assets/`. Do not ask the user to pick assets unless that inventory is missing or unusable. Use the exact required fields from the storyboard and script references.
After drafting, run the review loop's plan pass — `../hyperframes-core/references/review-loop.md` § 1: open the board (don't ask whether to), present the plan as a proposal, and ask the two questions — approve or change, and **sketches first** (recommended) or skip. Feedback loops through chat or the board's comments file until approved. This is a **checkpoint gate** (brief contract § 1): in autonomous mode there is no board and nothing to ask — post the same summary as a heads-up and proceed; sketches collapse into the build, and the one preview question comes at Step 6.
**Gate:** `STORYBOARD.md` exists, every visual frame has `asset_candidates`, `SCRIPT.md` exists when narration is needed, and the user approved the frame-by-frame plan (autonomous: the summary was posted as a heads-up).
---
## Step 3.1: Audio
Goal: Generate narration, word timings, music, and audio metadata from the approved script.
Start audio after Step 3 approval. Run it in the background, then continue to Step 4.
**Choose the narration provider and voice from the user's ask before invoking.** Pass the provider selected in Step 0 with `--provider <provider>` (or set `HF_TTS_PROVIDER`). If the request named a voice, gender, or tone, pick a matching voice id and pass it with `--voice <id>`. The pipeline default is otherwise **Marcia (female)** on HeyGen / `am_michael` on Kokoro — so a request like "a male voice" is silently ignored unless you pass the flag. Voice ids are provider-specific; resolve against whichever provider Step 0's sign-in status selected: **HeyGen** (signed in) via `node <MEDIA_DIR>/audio/scripts/heygen-tts.mjs --list` (or `GET /v3/voices?engine=starfish`); **Kokoro** (offline) via the voice table in `<MEDIA_DIR>/audio/references/tts.md` (prefixes `am_`/`bm_` male, `af_`/`bf_` female). When the user expressed no preference, fall back to the remembered voice (brief contract § 2) before the pipeline default, and say which one you used; omit `--voice` only when neither names one. When the user explicitly picked a voice this run, record it (`prefs.mjs record --key voice`).
`node <SKILL_DIR>/scripts/audio.mjs --script ./SCRIPT.md --storyboard ./STORYBOARD.md --hyperframes . --out ./audio_meta.json --provider <provider> --voice <voice-id> &`
The audio script handles narration, word timings, BGM lookup from HeyGen's music library, and timing metadata. BGM mood comes from the storyboard's `music:` field; **`music: none` turns BGM off**. This uses the HeyGen Audio API for retrieval, not generation, and uses the same `~/.heygen` credential as TTS. For provider details, read `../media-use/audio/references/tts.md`.
If there is no narration and no `SCRIPT.md`, skip voice generation. BGM may still run if the storyboard has a music mood.
**The canonical fully-silent marker:** `music: none` in the STORYBOARD.md top YAML block **and** no `SCRIPT.md`. That combination marks the project silent — no narration, no BGM, no SFX. `audio.mjs` recognizes it and generates nothing (it removes any stale `audio_meta.json`; an absent `audio_meta.json` is what assemble treats as silent), so Step 3.1 is a clean skip. Use it when the user asks for a silent / music-free video — don't improvise other spellings.
**Gate:** audio job has started, or the project is marked silent (`music: none` + no `SCRIPT.md`).
---
## Step 4: Frame Visual Design
Goal: Add the visual direction, layout intent, and motion choices to each storyboard frame.
**Sketch the board first (collaborative only).** The moment the plan is approved, run the sketch pass — `../hyperframes-core/references/review-loop.md` § 2 (don't wait on Step 3.1; sketches don't use timings): wireframe every frame yourself, mark each `built`, pause for the one layout question when the board is full, and revise only the sketches named until the board is confirmed. Stand-ins: plain labeled blocks for the captured assets — the real files arrive with Step 5's workers. Only then write the visual design below onto the confirmed layouts. In autonomous mode, or when the user chose to skip sketches at Step 3, skip this pass — frames go straight from `outline` to `animated` at Step 5.
Edit `STORYBOARD.md` in place. Do not create another storyboard. Use `frame.md` as source of truth for color, type, layout feel, and style.
Read `references/visual-design.md`, `../hyperframes-animation/blueprints-index.md`, `references/motion-language.md`, and `../hyperframes-animation/rules-index.md`. Use `visual-design.md` for the method (the time-coded shot sequence, the inline Layout vocabulary, and the required `## Video direction` block). Use `../hyperframes-animation/blueprints-index.md` to pick each frame's shot shape. Use `motion-language.md` (the motion vocabulary + the motion doctrine) and `../hyperframes-animation/rules-index.md` (valid rule names) for motion — do not invent motion names.
For every visual frame, write a **time-coded shot sequence** into `STORYBOARD.md` per `visual-design.md`'s method: pick the frame's blueprint (or compose), instantiate it with THIS product's content, and pace each Scene's reveal to the voiceover so the frame develops across its full duration instead of front-loading then freezing. State layout and motion **inline** per Scene (vocabularies in `visual-design.md` and `motion-language.md`). Add one video-wide `## Video direction` block.
When an element visibly continues across a frame boundary, give both workers the same numerical handoff in `STORYBOARD.md`: add `handoff_out:` to the outgoing frame and a matching `handoff_in:` to the incoming frame. Name the element and its exact x/y position, scale, opacity, and motion direction/speed at the cut — state every field even when it does not change, because a constant is `opacity: 1`, not an omission. Omit the whole block only for a deliberate clean cut. The goal is simple: parallel workers must not invent two different versions of the same seam.
Do not change story, script, asset choices, `asset_candidates`, `transition_in`, or captured source material. Do not write HTML in this step.
Stage named assets after visual design is locked:
`node <SKILL_DIR>/scripts/stage-assets.mjs --storyboard ./STORYBOARD.md --hyperframes .`
**Gate:** every visual frame has a time-coded shot sequence whose reveals are paced to the voiceover (no front-loading); `## Video direction` exists; `assets/` contains the named assets. Collaborative: the sketch board was confirmed.
---
## Step 5: Build Frames
Goal: Build every storyboard frame as an HTML composition and assemble the playable video.
Wait for Step 3.1 audio to finish if audio was started. Then sync durations and fetch SFX; skip both if silent.
`node <SKILL_DIR>/scripts/audio.mjs sync-durations --audio-meta ./audio_meta.json --storyboard ./STORYBOARD.md`
`node <SKILL_DIR>/scripts/audio.mjs fetch-sfx --storyboard ./STORYBOARD.md --hyperframes .`
Duration sync is mechanical: real voice duration wins; silent frames keep estimates; never hand-edit synced durations.
Check the music against the final cut before assembly. A library track can match the requested mood but open on a quiet build that drains the first seconds of a short launch video. Compare the opening with later five-second sections; when a later section has a stronger, musically clean start, trim from there and keep a short fade-in plus a longer fade-out. If frame or narration timing changes, redo this check against the new final duration so the music never ends early or leaves silence at the tail.
Before dispatch, read `../hyperframes-core/references/subagent-dispatch.md`. Build the per-frame packets and the worker role payload:
`node <SKILL_DIR>/scripts/frame-packets.mjs --project "$PROJECT_DIR" --storyboard "$PROJECT_DIR/STORYBOARD.md"`
The builder writes one bounded packet per frame under `.hyperframes/frame-packets/` (the frame's exact storyboard block + the blueprint body + every cited rule recipe, inlined) and `_role.md` (`../hyperframes-core/references/frame-worker-core.md` + this skill's `sub-agents/frame-worker.md`, concatenated verbatim — the complete worker role). Dispatch one sub-agent per frame, in parallel if possible; otherwise run workers in waves. Each worker gets exactly one frame: its prompt carries `_role.md` and that frame's packet — paste both in full, or hand the two file paths for the worker to read first (equivalent; the worker starts from exactly those two documents either way) — plus a dispatch context with `PROJECT_DIR`, `frame_id`, whether the frame has a **confirmed sketch** on disk (the worker dresses that layout rather than redrawing it — frame-worker core § When a confirmed sketch exists), canvas size, and caption status + keep-out band if captions are enabled.
Workers read only their packet and `frame.md`; they never open `STORYBOARD.md` or the skill documents (the packet inlines what was selected upstream). Each worker writes only `compositions/frames/NN-*.html`. Workers must never edit `STORYBOARD.md`.
**Full-bleed backgrounds ride on a `class="clip"` layer, never the `#root`.** A frame's ground (color field / gradient / grid) is its own full-duration background clip — a `background` set on the `#root` / `data-composition-id` element is clip-gated to the frame's window and is not a dependable ground, so dark content can land on the black host `body` and render invisible. The video's base ground is painted by the assembler from `frame.md`'s `canvas` color onto the index `#root`. (Full rule + self-check: `../hyperframes-core/references/frame-worker-core.md`.)
As each worker returns, the orchestrator marks that frame as `animated` in `STORYBOARD.md`.
After audio timings exist, build captions in the background and assemble the index:
`node <SKILL_DIR>/scripts/captions.mjs build --storyboard ./STORYBOARD.md --audio-meta ./audio_meta.json --hyperframes . --out ./caption_groups.json &`
`node <SKILL_DIR>/scripts/assemble-index.mjs --storyboard ./STORYBOARD.md --hyperframes .`
`captions.mjs` uses the project's `.hyperframes/caption-skin.html` (copied in Step 2) as the caption look, injecting brand tokens from `frame.md`; with no skin present it renders the built-in default pill. `captions: skipped (<reason>)` is valid. Continue without captions when explicitly skipped.
**Gate:** every frame is marked `animated` (collaborative: the sketch board was confirmed at Step 4), `index.html` exists, and captions are built or explicitly skipped.
---
## Step 6: Finalize
Goal: Verify the assembled video, get user approval, and render the final MP4.
Inject transitions, run checks, pause for review, then render.
`node <SKILL_DIR>/scripts/transitions.mjs inject --storyboard ./STORYBOARD.md --hyperframes .`
`node <SKILL_DIR>/scripts/transitions.mjs verify --storyboard ./STORYBOARD.md --index ./index.html`
`npx hyperframes lint`
`npx hyperframes check`
`npx hyperframes snapshot --at <frame-midpoints-and-each-cut-minus-0.1s-and-plus-0.2s>`
`snapshot` stitches the captured frames into one contact sheet (`snapshots/contact-sheet.jpg`). Inspect the midpoint frames for layout failures, then compare the two images around every cut. A continuing element must keep the promised position, scale, opacity, and direction; fix any visible pop before rendering.
If a command fails, surface stderr and stop — don't pile on recovery commands. Fix it yourself: the cheapest safe edit to `compositions/frames/NN-*.html`, then rerun the failed check.
After checks pass, pause for user review — the review loop's final look (`../hyperframes-core/references/review-loop.md` § 4): one question, on the Studio that has been open since Step 3 — render now, or what changes? (Autonomous: the one kept question, preview first or render.) Then deliver the MP4 with the contact sheet and the frame ids so revisions can target a single frame.
Preview: `npx hyperframes preview --background`
Render only after user approval (autonomous mode: after the preview-or-render question):
`npx hyperframes render --skill=product-launch-video --quality high --output renders/video.mp4`
Do not rerun `lint`, `check`, or `snapshot` after rendering unless the user asks.
**Gate:** `lint` and `check` passed and the snapshots were inspected before render; user approved at the review pause (autonomous: checks passed and the delivery includes the contact sheet); `renders/video.mp4` exists. Final reply states MP4 path and final duration.
---
## Quick Reference
**Formats:** landscape `1920x1080`; portrait `1080x1920`; square `1080x1080` — derived from the destination (brief contract § 2). Set the format once in the storyboard frontmatter.
**Background scripts:** the workflow ships only these scripts under `scripts/`: `build-frame` for adopting + brand-remixing a frame preset into `frame.md` (+ caption skin); `audio` for TTS, transcription, BGM, SFX, and duration syncing; `captions`; `transitions` for inject and verify; `stage-assets` for copying frame-named assets into `assets/`; and `assemble-index`. Everything else is handled by the `hyperframes` CLI.
The reusable, product-agnostic shot shapes live in `../hyperframes-animation/blueprints/` (indexed by `../hyperframes-animation/blueprints-index.md`).
| Read | When |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| `[../hyperframes-core/references/brief-contract.md](../hyperframes-core/references/brief-contract.md)` | Gate types, mode derivation from `BRIEF.md`, field semantics. |
| `[../hyperframes-creative/references/story-spine.md](../hyperframes-creative/references/story-spine.md)` | Step 3: story doctrine — hook language, value-before-evidence, proposal shape, source-traceable visuals. |
| `[../hyperframes-creative/frame-presets/](../hyperframes-creative/frame-presets/)` | Step 2: choose and adopt a frame preset. |
| `[../hyperframes-creative/references/design-spec.md](../hyperframes-creative/references/design-spec.md)` | Step 2: apply brand tokens correctly. |
| `[references/story-design.md](references/story-design.md)` | Step 3: plan the product-launch story. |
| `[../hyperframes-animation/blueprints-index.md](../hyperframes-animation/blueprints-index.md)` | Step 3: role→blueprint menu. Step 4: pick the shot shape. |
| `[../hyperframes-core/references/storyboard-format.md](../hyperframes-core/references/storyboard-format.md)` | Step 3: write `STORYBOARD.md`. |
| `[../hyperframes-core/references/script-format.md](../hyperframes-core/references/script-format.md)` | Step 3: write `SCRIPT.md`. |
| `[../media-use/audio/references/tts.md](../media-use/audio/references/tts.md)` | Step 3.1: choose or understand TTS providers and voices. |
| `[references/visual-design.md](references/visual-design.md)` | Step 4: write the frame's shot sequence (+ Layout vocabulary). |
| `[references/motion-language.md](references/motion-language.md)` | Step 4: the motion vocabulary + the motion doctrine. |
| `[references/cut-catalog.md](references/cut-catalog.md)` | Step 4-5: the cut catalog (worker builds within-frame seams). |
| `[../hyperframes-animation/rules-index.md](../hyperframes-animation/rules-index.md)` + `[../hyperframes-animation/rules/](../hyperframes-animation/rules/)` | Step 5: local rule recipe bodies for the cited motions. |
| `[../hyperframes-core/references/frame-worker-core.md](../hyperframes-core/references/frame-worker-core.md)` | Step 5: the shared worker contract (packet builder prepends it to the delta). |
| `[sub-agents/frame-worker.md](sub-agents/frame-worker.md)` | Step 5: the workflow's frame-worker delta. |
| `[../hyperframes-core/references/subagent-dispatch.md](../hyperframes-core/references/subagent-dispatch.md)` | Step 5: dispatch sub-agents safely. |
sub-agents/frame-worker.md
# Frame worker — product-launch delta
> The shared law is the core contract above (the packet builder prepends `../hyperframes-core/references/frame-worker-core.md` to this file as `_role.md`) — read the two as one role. This file carries only what's specific to a product-launch frame; you run N-up, **one frame each** — your dispatch carries exactly one packet. Tempted to add a generic GSAP / timeline rule here? Wrong home — it belongs in the core contract or `hyperframes-core`.
## Your `focal:` / `roles:` — real captured media
- `focal:` — which candidate is the hero.
- `roles:` — each candidate's role: `cutout` foreground / `background` full-bleed / supporting — plus the real media available (each `public/<basename> — description`; a **`[video]`** tag marks a `.mp4` motion source that cannot be mounted by this sub-composition worker).
- **Asset paths are project-root relative everywhere.** This includes HTML attributes and CSS `url(...)` values such as `@font-face src`: use `assets/...` (or the supplied `capture/assets/...` path), never `../` or `../../`. Shared fonts are staged before dispatch so every parallel frame resolves the same local files; never fall back to a network `@import`.
## Placing candidates (product-launch constraint)
**Place each candidate by its `roles`** (the `focal` is the hero): a `cutout` is a foreground subject — respect the 83% keep-out, lay text around it, not over its face; a `background` is full-bleed and dimmed ~30–50% so foreground content stays legible. Frame files are sub-compositions. Audio remains orchestrator-owned: never author `<audio>` in a frame. An approved `[video]` candidate may be declared as a frame-local `<video data-frame-video="approved" data-start="..." data-duration="..." data-track-index="...">`; `assemble-index.mjs` hoists it to the host root and translates its timing. Give every approved video explicit host geometry with numeric `data-frame-video-x`, `data-frame-video-y`, `data-frame-video-width`, and `data-frame-video-height`; optionally set `data-frame-video-fit="cover|contain|fill|none|scale-down"` (default `cover`). The assembler converts only those values to host CSS: frame-local classes and inline styles do not cross the hoist boundary. Do not use this declaration for audio, unapproved URLs, or videos without explicit timing and geometry. If no approved video is supplied, use an explicitly supplied static still/key art (`[video-still]` or another image candidate) as `<img>`; do not extract a frame, fabricate a URL, or silently embed an unapproved clip.
When a frame showcases a captured website, use the supplied screenshot as the page. Do not rebuild the full site in HTML: even a close recreation can change the real layout, spacing, or branding. If the shot needs motion inside the page, overlay the supplied real assets at measured positions or rebuild only the moving component.
Brand text comes from your frame's `scene` / narrative — never from `frame.md` (a style spec, not the product's content). Place the named assets, and never invent new ones.
## Cross-frame handoffs
If the packet includes `handoff_in:` or `handoff_out:`, treat those values as a hard boundary contract. Start or end the named element at the exact x/y position, scale, opacity, and motion direction/speed provided; a field the packet states as unchanged is still binding, not optional. Do not restyle or reinterpret that boundary state. The neighboring frame is being built by another worker and will use the matching values.