Transcribe Audio to Text

Transcribe audio to text free, right in your browser with no upload. 99 languages or auto-detect, timestamps, and TXT, SRT, VTT or Word files to download.

  • Free, no sign-up
  • Files stay on your device
  • No watermark

By using this tool you agree to our Terms and Privacy Policy.

Drag & drop your files or Processed on your device · no size limit
How to

Audio to Text in 3 steps

01

Upload one or more recordings: MP3, WAV, M4A, OGG, OPUS, FLAC, WEBM or the audio of an MP4 video.

02

Pick Fast or Accurate, and the spoken language or let it be detected. Watch the text appear as it is transcribed.

03

Listen, correct any line, then download TXT, SRT, VTT or Word. Your audio never leaves your device.

Features

Private speech to text that you can check and fix

Whisper speech recognition, on your device

The open Whisper model runs inside your browser, so interviews, meetings and voice notes are never uploaded. The model downloads once (about 41 MB for Fast or 77 MB for Accurate) and is then reused from the browser cache.

99 languages, auto-detect or translate

Choose the spoken language or let the tool detect it and show how sure it is. You can also get an English translation of speech in another language instead of a transcript.

Timestamps, subtitles and Word in one go

Every line keeps its start time. Download plain text, text with timestamps, SRT or WebVTT subtitles with readable two-line cues, or a Word document, for one file or a whole batch as a ZIP.

Listen and edit before you download

Click any time to play the audio from that moment, while the current line is highlighted. Fix words, clear lines you do not need and use Find and replace for names that were misheard.

How to transcribe audio to text in your browser

BestConverter turns speech into text with Whisper, an open speech recognition model, running directly in your web browser. Upload a recording, choose how accurate it should be and download the transcript as text, subtitles or a Word document. Nothing is sent to a server, which makes it a good fit for confidential interviews, medical or legal dictation, meetings and personal voice notes.

What happens to your recording

  • The audio is decoded and converted to 16 kHz mono, the format Whisper expects.
  • Long recordings are cut into pieces of up to 29 seconds at natural pauses, so words are rarely split in half. Silent parts are skipped, which saves time and stops the model from inventing text in silence.
  • Each piece is transcribed with timestamps, normally in a background thread so the page stays responsive, and you see the text as it is produced.

Tips for the best transcript

  • Pick the spoken language yourself when you know it. It is faster and avoids a wrong guess on short or noisy clips.
  • Use Accurate for accents, several speakers, phone calls and languages other than English.
  • If the result says no speech was found in a quiet recording, open Change settings and turn off Skip silent parts.
  • Record close to the microphone and avoid music under the voice.
  • Keep the tab open during long recordings. The tool asks the browser to keep the screen on while it works.

Good to know

Transcription on your own device is slower than on a large server, and very long recordings need a lot of memory, so split multi-hour files on phones. Speakers are not labelled, and the English translation option gives a rough translation only.

FAQ

Frequently asked questions

Is it safe to transcribe confidential recordings here?

Yes. Decoding and speech recognition both run in your browser, so the recording and the transcript stay on your device. The only download is the speech model itself, fetched once from Hugging Face and then cached by your browser.

How long does it take to transcribe an hour of audio?

It depends on your device and the model. On a typical laptop the Fast model often needs only a fraction of the recording length, Accurate takes roughly twice as long, and phones are slower. The real speed and the time left are shown while it runs, and you can stop at any moment and keep the text that is done.

Which model should I use, Fast or Accurate?

Fast (Whisper tiny) is fine for clear speech in widely spoken languages. Accurate (Whisper base) makes noticeably fewer mistakes, especially with accents, background noise and less common languages. Both are small models, so check names, numbers and technical terms before you share the text.

Can I transcribe MP3, WAV or M4A files?

Yes, plus AAC, OGG, OPUS, FLAC and WEBM, and the sound track of MP4 and MOV videos. Formats the browser cannot decode itself, such as WMA or AMR, are converted on your device by a built-in converter that downloads once (about 31 MB).

How do I turn audio into SRT or VTT subtitles?

Transcribe the file, then tick SRT or VTT under Download as. Long sentences are split into cues of at most two lines and about seven seconds, and you can choose 32, 42 or 50 characters per line.

Does the transcript label different speakers?

Not yet. The transcript is a single stream of timed lines without speaker names. Overlapping voices, music and very noisy recordings also reduce accuracy.

Need a tool we do not have yet?

Tell us what you are converting and we will build it. Write to support@bestconverter.online.

Contact us