Overview
Timecode Eraser removes embedded timecode references from oral history transcripts so the text can be read, analyzed, and deposited without timing marks.
Targets include:
- Bracketed/parenthesized:
[HH:MM:SS], (HH:MM:SS.mmm), [1:23]
- SMPTE:
HH:MM:SS:FF and HH:MM:SS;FF
- Standalone and line‑initial codes (optionally stripping a following dash)
Enclosing brackets/parentheses are removed along with the timecode.
Workflow
1) Load
- Click Load Transcript Files and select one or more
.txt, .docx, or legacy .doc files.
- Use the left-hand file list to select a file and preview the Transcript (Original).
- Confirm extraction looks correct (especially for .docx).
2) Process
- Click Process All Files to generate Transcript (Without Timecode) for every loaded file.
- In the Original preview, click any highlighted timecode to jump to the corresponding paragraph in the cleaned preview (highlight persists until another is clicked).
3) Export
- Export All as
.txt or .docx, or create ZIP bundles for each format.
- Exported filenames are suffixed with
_no_tc.
Preferred Formats & Tips
- .docx (preferred for Word): parsed with Mammoth; JSZip fallback extracts text from
word/document.xml.
- .txt: use UTF‑8 encoding when possible.
- .doc (legacy): treated as plain text; content may include binary artifacts—consider converting to .docx first.
- Use a blank line between paragraphs for best paragraph mapping between the two previews.
Regex Configuration (Optional)
Open the configuration panel to adapt the patterns/toggles to house styles:
- Core pattern: HH:MM or HH:MM:SS, optional milliseconds (HH:MM:SS.mmm).
- SMPTE pattern: HH:MM:SS:FF or HH:MM:SS;FF.
- Toggles: remove bracketed/parenthesized, line‑initial (and optional dash), standalone core, standalone SMPTE.
After changes, click Apply & Re‑process to re-run processing.
Troubleshooting
- Original shows “PK … [Content_Types].xml”: a .docx was read as plain text. Re‑save as .docx and reload.
- No cleaned preview: click Process All Files after loading.
- Unusual timecode style: adjust patterns in Regex Configuration and re‑process.
- Paragraph mapping seems off: ensure blank lines separate paragraphs in the source.
- Large files: very large .docx can take longer to parse in‑browser.
Attribution & Privacy
Timecode Eraser was created by Douglas A. Boyd. This tool is provided for non-commercial and archival purposes and is used at your own risk. No warranty is expressed or implied; results should be reviewed for accuracy before publication or deposit.