Two-line Montserrat pop, green hit

Black-weight Montserrat in two stacked lines of three words, current word flashes green. Dense explanations that need more words on screen at once.

Add word-timed captions to my video in this style: "Two-line Montserrat pop, green hit" from the 24fps video editor library.
Preset JSON: https://24fps.dev/video-editor/captions/pop-montserrat-green-two-line/preset.json (pass its "preset" to the engine; its "script" is demo text).
Engine (MIT, no dependencies, ES module): https://24fps.dev/editor/engine.js. Download the engine next to your page (curl -O) rather than importing cross-origin, then import { mount } from "./engine.js" in a <script type="module"> served over http.
Wiring: root = a position:relative box at the video's size. media = { main }: main is a wrapper holding my video, filling the root. Then const fx = mount(root, { kind: "captions", preset, script: words, media }) and call fx.render(seconds) every frame. It is seek-safe: drive it from the GSAP timeline in HyperFrames or call render(frame / fps) in Remotion. Before capturing frames, await (fx.ready ?? document.fonts.ready). Check it on my footage: if text is hidden behind the subject or doesn't read against my colours, nudge the layer positions or colours.
Transcribe my audio to word timestamps first (e.g. Whisper with word timestamps) and pass words as [{ w, s, e, emph }] with s/e in seconds; set emph: true on words to stress. A plain string also works (evenly timed, *word* = emphasis).
Field reference: https://24fps.dev/video-editor/llms.txt. Keep sizes as fractions of the frame height so it works at any resolution. Credit: 24fps.dev/video-editor
Same idea
Also called: Devin
Technique
Extra-bold Montserrat is the most cited caption default in tutorials.
Footage
Man talking and gesturing, Pexels

Similar caption styles