Make the words easy to follow.

Timed captions let a viewer follow your narration with the sound off. Remixure can create a draft from an English or Spanish WAV soundtrack or use an SRT or VTT file you already have. You can correct each cue before it becomes part of your exported video.

See the finished exports.

Narration and timed captions

Watch narration with timed captions on its own page.

Six seconds with synthetic English narration and manually reviewed timing. The captions are part of the picture.

Transcript: Give your next story room. Start with one clear idea. Then shape it into a video worth sharing.

The original example uses the studio renderer at 720×1280, H.264 and 30 frames per second.

Start with the recording you will use.

  1. Add your own PCM WAV soundtrack or an approved AI narration. The transcription trial supports English and Spanish speech, up to two minutes and 6 MB.
  2. Choose the audio start time and scene lengths, then save the video. A source trim of two seconds means captions are aligned to the recording from that point onward.
  3. Open Captions, choose the recording language for uploaded audio (saved AI narration supplies its language), confirm your recording rights and speaker permissions, and choose Create captions. The recording is sent to Cloudflare for transcription.
  4. Play the soundtrack and review the words and timings. Confirm your review, then choose Use captions. The draft does not overwrite your track automatically.

The trial allows five transcription attempts per workspace per UTC day. Failed and cancelled attempts count. A failed or uncertain request is not retried automatically. No voice cloning or speaker identification is performed.

Correct the cue, then watch it in context.

Select a cue to edit its text, start, or end. For example, a phrase spoken from 1.20 to 2.80 seconds should have those same times in the editor. Keep cues in order and prevent their times from overlapping. Use Preview caption to watch the correction with your current soundtrack.

Each cue can contain up to 100 characters; the track can contain up to 120 cues within the video’s two-minute limit.

Keep Original default or choose a dark box, white outline, or yellow highlight, then put the track at the top or bottom. Small, standard, and large text sizes apply to the whole track; long cues shrink to fit with a preset. Scene text leaves room for the selected caption area.

Restore default captions returns to white text on a dark box near the bottom. Undo restores an earlier edit or imported track.

Changing the audio trim or scene timing can change what the viewer hears at a particular moment. Review the captions again after those edits. Replacing a recording clears the current track; Undo can restore the recording and its captions together.

Shift the whole caption track without changing the recording.

If every cue is early or late by the same amount, open Align the whole caption track. Choose Earlier in the video or Later in the video, enter seconds with up to three decimal places, and select Review caption alignment. For example, a 0.250-second later shift changes 1.200–2.800s to 1.450–3.050s. Every existing word start and end moves by the same amount; wording, cue durations, gaps, appearance, scene timing and source audio stay fixed.

Compare every original and proposed cue, confirm the review, then Apply caption alignment. Changing the amount, direction or video requires another review. Discard caption alignment leaves the track unchanged; one Undo restores it after Apply. Preview each boundary and save a new version. Autosave stays off during review.

The available earlier/later range comes from the first and last cue. A shift outside the video is rejected as a whole; no cue is cropped or clamped. A zero shift makes no edit. Restore removed video timing before using this control. A constant offset does not stretch speech, infer new word timings or edit the immutable transcript candidate; adjust individual cues if the alignment varies across the recording.

Choose a font for the caption track.

Keep Sans or choose Serif or Mono in Caption text font. The whole track uses that family in the preview and exported picture, including spoken-word highlighting. Scene headlines and supporting text keep their own font.

Choosing a font on Original default starts with the dark-box preset. Font changes preserve cue wording, times, word timings, colors and placement. Reset caption font to Sans changes only the family; Undo restores the previous edit. SRT and VTT contain text and cue times without fonts.

Serif and Mono use the studio’s fixed licensed DejaVu files and support the checked English and Spanish characters. If a character is unsupported, choose Sans or correct it before saving. If a font file cannot load, the preview pauses until Retry text fonts succeeds or you choose Sans. Private uploads use a separate permission and preview review, described below.

Bring a permitted caption font.

Watch and rebuild the four-second private-font example with accented text, authored word highlights and a separate licensed font step.

Where private fonts are available, open Use a private text font after adding timed captions. Choose Timed captions in Private font text target. Choose one static regular or bold TTF file up to 1 MiB. Confirm that its license permits upload, browser preview and finished videos, then choose Upload and review font. Uploaded originals stay private in your workspace and count toward its media allowance.

Inspect the preview of your first cue, confirm readability and permission, then choose Apply private text font. Preview every cue before Save video. Apply makes one undoable edit; Discard leaves the video unchanged and keeps any stored original in the library. You can also choose an existing workspace font, remove the private font, or select a catalog family to recover.

The current boundary is one upright static TrueType face, regular or bold, for Latin text and common punctuation. Variable, collection, color, bitmap, compressed, italic and CFF fonts are not supported. Unsupported characters stop preview/export instead of silently substituting a face. Check actual coverage and keep your license with the original.

The caption target changes the whole timed track, including word highlighting and reveals. In the same target selector, you can separately review a scene heading and supporting line, a callout, or a bar, line or pie chart. The scene uses the original font weight for both lines. Each target needs its own review and Apply. Watch the separate bar, line and pie font example with editable scenes and source notes.

Label review shows all words without timing or movement so you can inspect the letters. Apply preserves actual timing, movement, opacity and reveals; recheck the complete video. One exact private original can be shared across these targets in a video. A chart applies the supplied regular or bold face to its title, labels, values, scales and source. Scene, caption and callout brand defaults can reuse a reviewed private original; chart brand defaults keep their catalog faces and need a separate private-font review.

Private fonts export as MP4, with up to 12 scenes and 60 seconds at 720p or 30 seconds at 1080p. SRT/VTT retain words and cue times without fonts. If the capability is unavailable, remove the private setting and save a new version. This does not delete its original.

Match captions to your colors.

Choose Customize caption colors to set the track’s text, box, outline and active-word colors. Use the picker or type a six-digit hex value such as #163729. An unfinished or invalid value leaves the saved color unchanged. The selected palette overrides preset colors and stays when you switch presets.

Outline captions use their outline color instead of a box; active-word colors apply when word highlighting is on. Keep words easy to distinguish from their background, then inspect the finished MP4 over your scene. Reset caption colors returns to the selected preset’s colors while keeping placement, size, timings and animation. If custom colors are temporarily unavailable, reset them and save a new version, or try again later.

Highlight the word being spoken.

Watch word highlighting with Spanish narration and see the transcript correction and timing gaps in the finished eight-second video.

Choose Caption animation → Highlight spoken words to follow individual words while keeping the full cue readable. Automatic captions retain candidate word times; open Word timings to correct each start and end. The highlight disappears during gaps between words.

Changing text while keeping the same word count retains the existing times. Adding or removing words, or changing a cue’s boundaries, clears that cue’s word timings. For an imported or manual track, choose Prepare evenly spaced word timings on each cue, then listen and adjust. Even spacing is an approximation, not speech alignment.

Word highlighting supports up to 240 timed words across the track and 20 per cue. It works with all three styles, positions and sizes. Choose None to keep static captions; Restore default captions also removes animation. SRT and VTT export cue text and times without word timing or highlighting.

Reveal words while keeping their place.

Choose Caption animation → Reveal words in place to show each word at its existing start time. Revealed words remain visible through pauses until the cue ends. Before the first word starts, the caption area is empty. The full cue sets the line positions, so the words stay in place as the sentence appears.

Open Word timings to review starts and ends against your recording. Both word animations require timing for every cue and support up to 240 words across the track. Reveals keep your wording, cue times, font, placement and colors. Active-word colors apply only to highlighting. Undo restores the previous animation; choose None for static captions or Highlight spoken words to keep the whole cue visible.

Save a new version, export MP4 and inspect the finished picture and sound. SRT and VTT retain the original complete cue text and cue times without animation. If reveals are temporarily unavailable, the saved track stays readable; choose another animation and save a new version, or try again later.

Keep a passage from your captions.

Open Trim to a caption passage and choose the first and last cues to keep. Read the selected text, check the displayed interval and preview its start, then choose Trim video to passage. The video before and after that interval is removed; retained footage, audio, captions, word times and scene graphics move together.

Scene boundaries use 0.1-second steps and each retained scene lasts at least one second, so the displayed interval may include extra footage. If that would overlap a neighboring cue, include that cue in your selection. Use cuts between scenes, set audio fades to zero and turn off background-music repeat before trimming. These restrictions keep existing transition and audio timing from silently changing.

This is one continuous passage from the current short video. It does not delete individual words, regenerate speech or find highlights automatically. Undo restores the earlier edit; saved versions remain available. Save, then listen at both ends of the finished MP4 before sharing.

Remove a reviewed word from the video.

Open Captions → Remove a word from the video. Choose a word with start and end times you have checked by listening. Review the proposed source interval and finished length, choose Listen before this word, then Remove this word. The edit removes that moment from picture and sound together and removes its caption token. Original media and earlier saved versions stay intact. Scene text stays as written; review it separately.

Caption and word times keep referring to the original recording. The preview clock and SRT/VTT downloads use the shorter finished timeline. Existing scene graphics, transitions, audio fades, music repeat and ducking follow the retained source moments. Listen at every join, save a new version, and watch the actual MP4 before sharing.

Word removal supports up to 12 separate sections in original videos of 30 seconds at 720p or 15 seconds at 1080p, with at least one second retained. Cuts align to 30 fps video frames. If rounding would affect a neighboring word, correct the timing or leave that word in. Even spacing and automatic timings are candidates that need review; removing a word does not regenerate speech.

Restore removed video timing restores picture and sound while retaining your caption wording edits. Undo or an earlier saved version restores both. Restore timing before changing scene order, scene lengths, transitions, or the primary recording, and before creating or applying another transcript. If word-removal export is temporarily unavailable, restore timing and save a new version to export, or return when available.

Review pauses and filler words together.

Open Clean up pauses and filler words in the studio. Find filler words suggests timed “um”, “uh” and “erm” tokens from your reviewed captions. Analyze quiet sections also checks the decoded primary recording in your browser, before volume or mixing. No new model request is made.

Quiet candidates require a low peak level in both channels, keep 0.1 seconds at each edge and protect timed words or untimed cues. Start with the conservative −42 dBFS threshold and 0.4-second minimum; adjust the threshold or minimum when needed. Quiet speech, breaths and meaningful pauses can still be misclassified. Background music and silent footage are not analyzed. Use an MP3 or WAV primary recording within 60 seconds and 6 MiB.

Listen around each candidate and check its context. Select up to twelve sections, confirm your review, then choose Remove selected sections. This makes one reversible edit to picture, all audio and caption wording. Scene text stays as written; check it separately. A changed video or recording makes the review stale. If nothing fits, your video remains unchanged.

Original video-cut limits apply: 30 seconds at 720p or 15 seconds at 1080p, with at least one second retained. Play the joins, save a new version and inspect its finished export. One Undo restores the edit. Restore timing in this video restores removed picture and sound while keeping caption wording edits, even when new cut exports are unavailable; Undo or an earlier saved version restores both.

Translate the captions while keeping the recording.

Save your caption track, then open Translate subtitles. Choose the original subtitle language and a different available target: English, Spanish, French or Portuguese. Changing either choice clears its permission confirmation. Confirm permission to send the caption wording to Cloudflare, then create the candidate. Compare each original and translated cue before choosing Use translated subtitles. Correct wording in the editor and save a new video version.

For example, our original “Give your next story room.” produced the candidate “Dale espacio a tu próxima historia.” Its existing cue still runs from 0.00 to 2.80 seconds. The English wording and any English recording remain unchanged until you deliberately edit them.

The translation trial accepts up to 30 cues and 1,800 source characters, with five accepted attempts per workspace per UTC day. Cue times stay unchanged; the audio and scene text remain in their original language. Translated captions are unverified wording, not fact checking or dubbed speech. Check names, qualifications and disclosures.

Applying a candidate clears word animation because translated words have no source speech alignment. Undo restores the prior edit; the original track remains in its saved version and candidate history. If captions, audio or video length change before applying, restore the original saved version or create another translation. Failed, cancelled and uncertain attempts count and are not retried automatically.

Bring a transcript you already have.

Choose Import SRT or VTT for a subtitle file under 100 KB. The file is read in your browser, and the reviewed track is saved with your private video project. Imported styling is removed. You can use the free caption editor to prepare a longer transcript separately or convert formats with the subtitle converter.

MP3 soundtracks still work for video playback and export; use an imported subtitle track for them. Automatic MP3 transcription, wider-language translation and phoneme effects are not part of this release.

Check the space before the timing.

Use the caption-space and format tool to compare portrait, landscape, square and 4:5 layouts. It shows the studio’s reserved caption strip and provides a four-second example document and VTT. Then check the actual picture and sound in your video.

Correct repeated wording in one reviewed edit.

Open Find and replace caption text in the studio. Enter the literal text and replacement, then choose Match case and Whole words as needed. Review every proposed before/after cue and confirm the changes before choosing Apply caption replacements. Changing the search or video requires another review. A no-match search, Discard or review alone leaves your video unchanged.

Cue times, recordings, scene text, music and effects stay fixed. When a cue keeps the same word count, its existing word times stay; listen and check them against the corrected wording. Adding or removing words clears that cue's word times. If it used word animation, the review explains that animation will change to None for the whole track; other appearance and remaining word times stay. Prepare and review new timings before enabling an animation again.

Each resulting cue must contain 1–100 characters, across up to 120 cues. A blank replacement removes matched text but cannot empty a cue. Restore removed video timing before a bulk replacement. Review pauses autosave, Apply makes one undoable edit, and Save video keeps it as a new version. Correcting captions does not change or regenerate recorded speech.

Split crowded captions or join two cues.

Open Split or merge caption cues, choose an existing caption, and select Split into two cues or Merge with next cue. A split happens after a complete word at the time you choose. Checked word times stay exact, so the split must fit between that word’s end and the next word’s start. A cue without word times uses a duration-based suggestion; listen and correct it yourself.

Choose Review caption boundaries to compare all original and proposed words and cue bounds. A merge with a gap explicitly warns that the joined wording can remain visible across the formerly empty interval. Two timed cues retain all word times; a mixed timed and untimed pair becomes one untimed cue. The complete track must still fit 120 cues, 100 characters and 20 timed words per cue. Nothing divides or regenerates recorded speech.

Confirm the review, then choose Apply caption boundaries. Scene, soundtrack, music and effect clocks stay unchanged; one Undo restores the earlier cues. A changed video or selection requires another review. Restore removed video timing first, preview the join, then save a new version with autosave off.

Review shorter groups across a timed track.

Open Regroup timed captions, then choose the maximum words and characters per cue. Review caption groups shows every changed original cue and its proposed shorter groups. Complete words stay together. Each original cue's outer bounds and gaps stay, and all existing word timestamps are retained; boundaries fall between the adjacent words.

Untimed cues stay unchanged, even if they exceed your chosen limits. A word longer than the character limit or a result exceeding 120 cues is refused without editing the track. This does not identify speech, create alignment, join separate original cues or alter the recording.

Confirm the complete review and choose Apply caption groups for one undoable edit. Scene, soundtrack, music, effects and caption appearance stay fixed. Shorter cues change what remains visible during a pause or reveal, so preview and listen at their boundaries. A changed video or limit requires another review. Restore removed video timing first; autosave stays off until you deliberately enable it. Save the reviewed result as a new version.

Check both exports.

Save the video and choose Export MP4 to burn the timed text into the picture. Download SRT or VTT when a destination accepts a separate caption track. Watch the finished MP4 and check names, punctuation, timing gaps, and readable placement before sharing.

Add captions in the studio →