Audio to text converter: free, private, right in your browser
Drop an audio or video file and get a transcript. The Whisper speech model runs inside your browser, so your recording never leaves your device.
- No upload
- No sign-up
- 90+ languages
Drop your audio or video file here
or click to choose a file
- MP3
- WAV
- M4A
- AAC
- OGG
- FLAC
- MP4
- MOV
- WEBM
Everything runs locally: your file is processed in your browser and never uploaded. Only the speech model (about 77 MB) is downloaded once from Hugging Face.
Want to skip the recording and dictate straight into any app?
Ownvox types what you say at your cursor on Mac and Windows: one hotkey, any app, and just as private as this tool.
How it works
- 01
Pick a file
Drop an MP3, WAV, M4A, OGG or a video file onto the page. Nothing is uploaded; the file stays in your browser's memory.
- 02
Whisper loads once
The OpenAI Whisper model (about 77 MB) is downloaded a single time and cached. From then on it starts instantly, even offline.
- 03
Copy or download
Watch the text appear, then copy it or download it as .txt or as .srt subtitles with timestamps.
Why transcribe locally?
Most free online converters upload your recording to a server, keep it for a while and often train on it. This one does not, because it cannot: the model runs on your machine.
Nothing to upload
Interviews, meetings, voice memos and lectures often contain other people's words. Keeping them on your device is the simplest way to stay GDPR-compliant.
No sign-up, no limit
No account, no minutes to buy, no watermark. Long files work too; a modern GPU transcribes many times faster than real time.
Works offline
Once the model is cached, transcription needs no connection at all. You can switch off Wi-Fi and check for yourself.
Formats, languages and limits
Formats
- MP3
- WAV
- M4A
- AAC
- OGG
- FLAC
- MP4
- MOV
- WEBM
The tool accepts any file your browser can decode: MP3, WAV, M4A/AAC, OGG and FLAC audio, plus the audio track of MP4, MOV and WEBM videos. There is no hard length limit, but very long recordings need a lot of browser memory and, without a GPU, patience: on the CPU path, expect roughly real time or a bit faster.
Languages
Whisper understands more than 90 languages. Choose the spoken language before you start; it tells the model which language to expect, which matters most for short clips and strong accents. The model transcribes, it does not translate.
Accuracy
This page uses Whisper "base", a good fit for clear speech in a quiet room. Strong accents, background noise, several people talking at once and specialist vocabulary bring down the quality. It does not label speakers; timestamps are available in the .srt export.
Frequently asked questions
You're not slow.Your keyboard is.
One hotkey, one voice, one cursor. Text lands where you're working, instantly.
7 days free, then from €9.99/mo · cancel anytime, effective at period end