100% Browser · Zero Uploads· No Tracking
Go Pro — batch 20 files at once, ad-free.
Upgrade
ConverterPro logo
ConverterPro
Back to tools
Module · Audio · Pro

Transcription Pro

Whisper transcribes, GPT-5.4 tags speakers and chapters. Every export format your workflow needs — SRT, VTT, TXT, DOCX, PDF, JSON.

Language (optional)
Speaker hint (optional)
Pro-only
Pro unlocks
Speaker labels (A/B/C…), chapter markers, DOCX & PDF exports. Free tier gets the raw transcript + SRT/VTT.
Unlock Pro
Advertisement
Guide

How Transcription Pro works

Step 01
Upload audio or video

.mp3 / .wav / .m4a / .mp4 / .mov / .mkv / .webm up to 25 MB (Whisper's hard limit).

Step 02
Optional hints

Set the source language (or auto-detect) and — for Pro — an expected speaker count (e.g. 2 for a two-person interview).

Step 03
Whisper → GPT → Export

Free tier gets the raw Whisper transcript + SRT/VTT. Pro tier chains through GPT-5.4 for speaker labels and 3–8 topical chapter markers, then exports as DOCX/PDF/JSON too.

FAQ

Frequently asked

How accurate are the speaker labels?+

GPT-5.4 approximates speakers from turn-taking cues (questions vs answers, self-references). It's fine for interviews and podcasts. For studio-quality diarization use pyannote — we can plug that in if you need it.

What's the length limit?+

25 MB (Whisper's cap) — that's roughly 25 minutes of MP3 audio at 128 kbps or 12 minutes at 256 kbps. Trim first if longer.

Are my recordings kept?+

No. Audio flows through our backend to Whisper and is discarded once transcription finishes. We log a metering row only — byte count, elapsed ms, user_id.

Can I edit before exporting?+

Yes — every segment is inline-editable. Change the text or the speaker letter, then hit Export. Changes only live in your browser tab; nothing is synced back to the server.

Advertisement