Audio to Text

Turn a podcast, interview, lecture or voice memo into text you can read, search and quote — free, with no signup and no upload. The speech-recognition model is a one-time download to your browser; the recording itself never leaves your device.

Why there's no per-minute limit

Transcription services meter free plans by the minute because every minute costs them server time. This runs on your own machine instead, so the only thing spending anything is your laptop — which is also the honest catch: a long recording takes real time, and the tool tells you roughly how long before you commit to it.

Working with the transcript afterwards

The transcript is a real subtitle track underneath, not just a wall of text, so every line keeps its timing. If the recording came off a video call, you can export it as SRT and drop it straight back onto the picture. If you only have the video, pull the sound out with the audio extractor first — or just feed the video in here, which works too.

Frequently asked questions

Is my recording uploaded anywhere?

No. The speech-recognition model (Whisper) is downloaded to your browser and runs there, so the recording never leaves your device. That's also why there's no signup: there's no server bill to cover.

What file types does it take?

MP3, WAV, M4A, FLAC, OGG and Opus, plus any video file if what you have is a recording of a call. If your recorder produced something unusual, convert it first and come back.

How long can the recording be?

The cap is shown before you start, and it's a cap on running time rather than file size — the work scales with minutes of speech, not megabytes on disk. A long recording is not refused for being large, only for being long.

Does it separate speakers?

No. This transcribes what was said, not who said it — speaker diarization is a second model and isn't built here. For an interview you'll get the full text in order, without labels.

Can I fix mistakes before downloading?

Yes. The transcript is editable in place: correct a word, retime a line, delete a false start. Every download and every copy takes the corrections with it.

Can I get timestamps?

Yes — switch the output to SRT or VTT for timed lines, or tick timestamps on the plain-text output to get them inline. The timing is word-level underneath, so lines land where the words actually fall.