Transcription Pro
Whisper transcribes, GPT-5.4 tags speakers and chapters. Every export format your workflow needs — SRT, VTT, TXT, DOCX, PDF, JSON.
How Transcription Pro works
.mp3 / .wav / .m4a / .mp4 / .mov / .mkv / .webm up to 25 MB (Whisper's hard limit).
Set the source language (or auto-detect) and — for Pro — an expected speaker count (e.g. 2 for a two-person interview).
Free tier gets the raw Whisper transcript + SRT/VTT. Pro tier chains through GPT-5.4 for speaker labels and 3–8 topical chapter markers, then exports as DOCX/PDF/JSON too.
Frequently asked
How accurate are the speaker labels?+
GPT-5.4 approximates speakers from turn-taking cues (questions vs answers, self-references). It's fine for interviews and podcasts. For studio-quality diarization use pyannote — we can plug that in if you need it.
What's the length limit?+
25 MB (Whisper's cap) — that's roughly 25 minutes of MP3 audio at 128 kbps or 12 minutes at 256 kbps. Trim first if longer.
Are my recordings kept?+
No. Audio flows through our backend to Whisper and is discarded once transcription finishes. We log a metering row only — byte count, elapsed ms, user_id.
Can I edit before exporting?+
Yes — every segment is inline-editable. Change the text or the speaker letter, then hit Export. Changes only live in your browser tab; nothing is synced back to the server.
Related utilities
AI subtitles + editable transcript + burn-in export. 100% browser.
Multi-track timeline, split/drag/resize clips, effects, transitions, animated titles, keyboard shortcuts, autosave — pro-grade in your browser.
All-in-one video editor: trim, speed, rotate, flip, crop, resize, add text, watermark remover, split audio — in your browser.