The Caption Slicer takes a human-edited transcript that already has timecode at speaker changes and a granular, machine-generated Timecode Reference file (WebVTT or SRT). It then slices the human transcript into smaller caption cues that follow the reference file’s timing structure, while using only the text from your primary transcript.
All processing occurs locally in your web browser—no files are uploaded or stored—and the output stays on your device.
.txt with [HH:MM:SS] timestamps at speaker changes (e.g., Synchronifier output).
How to Use the Caption Slicer (v1.2.1)
1) Load your Primary Transcript (UTF-8 .txt).
• This transcript should already have [HH:MM:SS] timecodes at each speaker change.
• Example:
[00:00:05]
INTERVIEWER: Can you tell me about your childhood?
[00:00:12]
AL YOUNG: Sure, I grew up in...
2) Load your Timecode Reference (WebVTT or SRT).
• Typically a machine-generated captions file.
• The Slicer uses only its timecodes and cue structure (start/end times).
• The reference should be fairly granular (short cues) for best results.
3) Configuration (optional).
• Minimum characters per cue:
- Default 0: the Slicer follows the reference cue structure as-is.
- >0: tries to give each non-empty cue at least this many characters
(by allocating more words to that cue from the segment).
• Skip empty cues in VTT export:
- Off (default): all reference cues are kept in the VTT, even if they end up empty.
- On: empty reference cues are omitted from the final VTT export.
• Speaker names:
- Default (unchecked): keep speaker names and bias them to the first cue in each segment.
- Checked: remove speaker names altogether from the final captions.
4) Click “Slice Captions”.
• The tool:
• Parses your speaker-level transcript into segments based on [HH:MM:SS].
• Associates reference cues with those segments by time.
• Distributes your human transcript words across those cues, guided by
the reference cue word-lengths and the optional min-character rule.
• Keeps the reference start/end times for each cue.
5) Export TXT and/or VTT outputs.
• TXT shows [HH:MM:SS] per cue plus text (non-empty cues only).
• VTT is a standard WebVTT captions file using the reference timings
(with optional skipping of empty cues).
6) (Optional) Load media and click any [HH:MM:SS] in the TXT preview to seek & verify.
Notes
• All processing occurs locally—no files are uploaded or stored.
• The Primary Transcript is text-authority; the reference file is time-authority.
• When "Minimum characters per cue" is > 0, the Slicer still honors the overall
cue structure but may give some cues more text than the reference suggested.
Disclaimer
This tool is provided for non-commercial and archival purposes and is used at your own risk.
No warranty is expressed or implied; results should be reviewed for accuracy before publication or deposit.