One word at a time, big white keys, small lime filler

Tall condensed uppercase sans, a single word per beat: key words large white, filler words tiny in acid lime, centred low in frame. Fast business, coaching and fitness explainers that should read as one-word-per-beat.

Add word-timed captions to my video in this style: "One word at a time, big white keys, small lime filler" from the 24fps video editor library.
Preset JSON: https://24fps.dev/video-editor/captions/single-word-condensed-lime-filler/preset.json (pass its "preset" to the engine; its "script" is demo text).
Engine (MIT, no dependencies, ES module): https://24fps.dev/editor/engine.js. Download the engine next to your page (curl -O) rather than importing cross-origin, then import { mount } from "./engine.js" in a <script type="module"> served over http.
Wiring: root = a position:relative box at the video's size. media = { main }: main is a wrapper holding my video, filling the root. Then const fx = mount(root, { kind: "captions", preset, script: words, media }) and call fx.render(seconds) every frame. It is seek-safe: drive it from the GSAP timeline in HyperFrames or call render(frame / fps) in Remotion. Before capturing frames, await (fx.ready ?? document.fonts.ready). Check it on my footage: if text is hidden behind the subject or doesn't read against my colours, nudge the layer positions or colours.
Transcribe my audio to word timestamps first (e.g. Whisper with word timestamps) and pass words as [{ w, s, e, emph }] with s/e in seconds; set emph: true on words to stress. A plain string also works (evenly timed, *word* = emphasis).
Field reference: https://24fps.dev/video-editor/llms.txt. Keep sizes as fractions of the frame height so it works at any resolution. Credit: 24fps.dev/video-editor
Same idea
Also called: Impact II (AI Edit)
Technique
Single-word condensed captions with size contrast: stressed words big, connective words small and coloured.
Footage
Man talking and gesturing, Pexels

Similar caption styles