The VTT Cue Combiner merges consecutive WebVTT cues (also called segments) that share the same speaker label into unified paragraphs, updating each combined section’s timecode to reflect the full duration. The resulting WebVTT file is optimized for use with OHMS (the Oral History Metadata Synchronizer), which prefers transcript timecodes aligned to speaker changes. In addition to VTT files, this tool can also process SRT transcripts — provided the SRT includes identifiable speaker labels.
Your data stays private: This tool processes transcript files entirely within your web browser — nothing is uploaded to any server and nothing is saved.
No files selected.
ConfigureView Configure
Cue Combining & Export
Unlabeled VTT cues inherit the last labeled speaker until a new label appears.
Unlabeled SRT cues inherit the last labeled speaker until a new label appears.
Recommended to avoid overwriting originals.
Blather Buster
Enabled by default. Splits long single-speaker sections into more manageable continued sections.
Default: split speaker sections longer than 3 minutes near the 2-minute mark.
seconds
seconds
seconds
seconds
Batch ResultsHide Batch Results
No files loaded.
Preview Processed File
Choose one processed file to preview both the original input and grouped output.
Original / Input Preview
No file loaded.
Grouped Speaker Output VTT Preview
The VTT Cue Combiner was created by Douglas A. Boyd. This tool is provided for non-commercial and archival purposes and is used at your own risk. No warranty is expressed or implied; results should be reviewed for accuracy before publication or deposit. For questions or feedback, please use the contact form.|Help / About
Help / About
How to Use the VTT Cue Combiner
1) Choose one or more WebVTT or SRT transcripts exported from your speech-to-text platform. 2) If SRT is provided, it must include speaker labels (e.g., “Speaker: …” or <v Speaker>). 3) Click Group Speakers to batch merge consecutive cues by the same speaker. 4) Preview any processed file using the single preview menu. 5) Export all processed VTT and TXT files as a ZIP.
Batch Processing
Version 3.0 supports multiple VTT/SRT files at once. The batch table reports the original cue count, final section count, and how many long speaker sections were split for each file. You can choose any file to preview the original input and grouped output before export.
Toggles: Sticky-speaker options for both VTT and SRT inherit the last labeled speaker across subsequent unlabeled cues until a new label appears.
Long Speaker Sections
The Blather Buster option is enabled by default in its own Configure section. It splits speaker sections longer than 3 minutes near a sentence boundary around the 2-minute mark, unless the remaining continued section would be shorter than 1 minute. Continued sections receive CONT. after the speaker label.
WebVTT Speaker Tags
Exported VTT files use standard WebVTT voice spans for speaker identification, for example <v Speaker 1>Transcript text</v>. Plain-text exports continue to use readable Speaker 1: labels.
Notes
All processing occurs locally in your web browser; no files are uploaded or stored. The combined VTT associates timecodes at speaker changes for OHMS optimization.
Disclaimer
This tool is provided for non-commercial and archival purposes and is used at your own risk. No warranty is expressed or implied; results should be reviewed for accuracy before publication or deposit.