Transcribe English Audio
Get editable text and timed subtitles from a short recording, entirely in your browser.
Drop an audio file to transcribe
Turn spoken English into editable text and timed subtitles in your browser.
up to 100 MiB / 5 min · first use downloads a speech modelHow audio transcription works
Download the model once, then review each result.
- 01
Choose an audio recording up to 100 MiB and 5 minutes.
- 02
Start English transcription. The first run downloads a speech model; inference then runs on your device.
- 03
Review each timestamped segment and download TXT or SRT.
What affects transcription accuracy?
Speech stays on your device
Only the speech model is downloaded. Your recording and transcript are processed locally in your browser.
Audio transcription questions
Up to 100 MiB and 5 minutes per file. Short, clear English speech gives the best results.
Does the recording leave my device?
No. The model files are downloaded, but your audio is decoded and transcribed in the browser. The recording is not sent to an API.
Can I use it for another language?
The first version uses an English-only model. Non-English speech is not supported.
Will it recognize every word?
No. The lightweight model can miss words, names, accents, or speech behind noise. Review and correct the text before using it.
What is the difference between TXT and SRT?
TXT contains the spoken text. SRT includes timestamps so the segments can be used as subtitles.