Word-by-word captions where the spoken word fills with a sliding two-colour gradient inside a contained glow, words already said stay white and the next ones wait dimmed
This is the format it was composed for. Switch above to see it lay itself out for another: same animation, re-arranged for the frame.
Everything below can be changed without touching the animation. Hand the list to an agent, or edit the values at the bottom of the component file.
Paste an SRT or WebVTT file's text. When it has cues it replaces `words`, and each cue's time is shared across its words by length. A transcript longer than the clip plays only as far as the composition runs, so lengthen the composition to show all of it.
Exact per-word timings in seconds, straight from any transcription tool that exports word timestamps. Used when `srt` is empty.
now: 21 items
Most words on screen at once. Three reads fastest on a phone, two feels punchier, one gives every word its own beat. A screen also ends early at a full stop or a pause.
now: 3
Words that keep the accent colour after they are spoken instead of turning white. Matched ignoring case and punctuation, so 38412 finds 38,412. Save it for the figures and names the video is about.
now: 2 items
Where the captions sit. At 9:16, lower stays above the platform's own caption and button area (the lowest 18%), and every position keeps out of the right-hand 12% where the like and share buttons are.
now: lower
Opacity of the words not yet spoken. 0.45 lets viewers read ahead without competing with the lit word; below about 0.35 they start to vanish over bright footage.
now: 0.45
Strength of the glow behind the spoken word. It is sized to its own word and stops short of the next one at every setting, so turning it up makes it brighter, not wider. 0 turns it off.
now: 1
Dark halo under every word, which is what keeps white and dimmed words readable over a bright screen or a sky. Lower it only over footage that stays dark.
now: 0.7
Type size relative to the frame. One size is shared by every screen of the transcript, so a long word shrinks the whole set a little rather than jumping on its own.
now: 1
Draw the stand-in shot behind the captions: a desk at night lit by a monitor, with a keyboard edge and a mug, all out of focus. It is a demo backing and not part of the captions. Switch it off to put them over your own video.
now: true
Paint a plain brand ground behind the captions when the stand-in is off. Off by default: with the stand-in off as well, the captions render on a transparent background and drop straight onto your footage.
now: false
Brand token id, one of the keys of BRANDS. The gradient runs from the brand's accent to its second accent, so brands with two bright accents on a dark ground suit it best.
now: dusk
No Tailwind, no CSS, no asset files. Inline styles only.
I make videos about building apps with AI and I want captions that look like the tech I'm talking about. Two or three words at a time, the word I'm saying filled with a purple-to-blue gradient that actually moves while I say it, with a soft glow behind it that doesn't smear into the other words. Words I've already said should stay white and the next ones faded so people can read ahead. It has to take my real subtitle file with word timings, keep clear of the like and share buttons on a vertical video, and come out transparent so I can put it over my screen recordings.