Free tool

Audio to text converter: free, private, right in your browser

Drop an audio or video file and get a transcript. The Whisper speech model runs inside your browser, so your recording never leaves your device.

  • No upload
  • No sign-up
  • 90+ languages

Drop your audio or video file here

or click to choose a file

  • MP3
  • WAV
  • M4A
  • AAC
  • OGG
  • FLAC
  • MP4
  • MOV
  • WEBM

Everything runs locally: your file is processed in your browser and never uploaded. Only the speech model (about 77 MB) is downloaded once from Hugging Face.

Want to skip the recording and dictate straight into any app?

Ownvox types what you say at your cursor on Mac and Windows: one hotkey, any app, and just as private as this tool.

Discover Ownvox
3 steps

How it works

  1. 01

    Pick a file

    Drop an MP3, WAV, M4A, OGG or a video file onto the page. Nothing is uploaded; the file stays in your browser's memory.

  2. 02

    Whisper loads once

    The OpenAI Whisper model (about 77 MB) is downloaded a single time and cached. From then on it starts instantly, even offline.

  3. 03

    Copy or download

    Watch the text appear, then copy it or download it as .txt or as .srt subtitles with timestamps.

Privacy

Why transcribe locally?

Most free online converters upload your recording to a server, keep it for a while and often train on it. This one does not, because it cannot: the model runs on your machine.

Nothing to upload

Interviews, meetings, voice memos and lectures often contain other people's words. Keeping them on your device is the simplest way to stay GDPR-compliant.

No sign-up, no limit

No account, no minutes to buy, no watermark. Long files work too; a modern GPU transcribes many times faster than real time.

Works offline

Once the model is cached, transcription needs no connection at all. You can switch off Wi-Fi and check for yourself.

Good to know

Formats, languages and limits

Formats

  • MP3
  • WAV
  • M4A
  • AAC
  • OGG
  • FLAC
  • MP4
  • MOV
  • WEBM

The tool accepts any file your browser can decode: MP3, WAV, M4A/AAC, OGG and FLAC audio, plus the audio track of MP4, MOV and WEBM videos. There is no hard length limit, but very long recordings need a lot of browser memory and, without a GPU, patience: on the CPU path, expect roughly real time or a bit faster.

Languages

Whisper understands more than 90 languages. Choose the spoken language before you start; it tells the model which language to expect, which matters most for short clips and strong accents. The model transcribes, it does not translate.

Accuracy

This page uses Whisper "base", a good fit for clear speech in a quiet room. Strong accents, background noise, several people talking at once and specialist vocabulary bring down the quality. It does not label speakers; timestamps are available in the .srt export.

FAQ

Frequently asked questions

Last step

You're not slow.Your keyboard is.

One hotkey, one voice, one cursor. Text lands where you're working, instantly.

Start free trial

7 days free, then from €9.99/mo · cancel anytime, effective at period end