Transcribe audio to text with AI, in your browser

Open your recording and choose Analyze ▸ Transcribe…. Auditius runs Whisper, an AI speech recognition model, on your own device: you download it once, when you click Download, and from then on your audio never leaves the browser. Every word keeps its time, so you can click it to hear it and export subtitles.

  • Free
  • No install, no account
  • Files never leave your computer
  • English and Spanish

How to transcribe a recording, step by step

Transcription works on the file open in the editor, or on the part of it you select.

  1. Click Open a recording to transcribe, or use File ▸ Open… (Ctrl+O), and choose your audio file.
  2. To transcribe only part of it, drag across that part. Otherwise the whole file is used.
  3. Choose Analyze ▸ Transcribe…
  4. Pick a Model: Accurate (small, 252 MB) or Fast (base, 80 MB). Set Language to Auto-detect, Spanish or English, and Range to Selection or Whole file.
  5. The first time, click Download. The model is stored in this browser, so you only download it once.
  6. Click Transcribe. When it finishes, the text appears in the Transcript panel next to the waveform.
  7. Click any word to play from it, or Shift+click another word to select the audio up to it.

What you can do with the transcript

The transcript is more than text: every word keeps its start and end time in the recording.

  • While the file plays, the current word is highlighted in the Transcript panel.
  • Export SRT and Export VTT save subtitles. Lines break at the end of sentences and at long pauses, stay within 42 characters and 6 seconds, and follow the file’s timeline even when you transcribed only a selection.
  • Add cues per sentence creates one range cue for each sentence, in a single step you can undo. You find them in Edit ▸ Cue List…, ready to jump from sentence to sentence.
  • If you close the panel, Window ▸ Transcript shows it again. File ▸ Save Session… keeps the transcript with the file.

Choosing a model and getting better text

Accurate (Whisper small) is the one to use for Spanish. Fast (Whisper base) is about three times faster but rough in Spanish, so keep it for quick drafts.

  • Auto-detect decides the language from the first 30 seconds. If the recording starts with music or another language, set the language yourself.
  • Very short clips are transcribed poorly. Give it at least a few seconds of speech.
  • Long recordings are processed in 30-second windows while a progress bar advances. The browser uses the graphics card through WebGPU when it can and the processor otherwise, which is slower.
  • If you edit the audio after transcribing, the panel warns that word times may be off. Click Transcribe again to refresh them.

Where your recording goes

Nowhere. The model files are downloaded the first time you click Download and kept in this browser. Transcription then runs on your device, and both the audio and the text stay there, which makes it a good fit for interviews, meetings and other private recordings.

Frequently asked questions

Which languages can Auditius transcribe?

You can choose Spanish or English, or let Auto-detect decide from the first 30 seconds. For Spanish, use the Accurate model.

Do I have to download the transcription model every time?

No. After the first download it stays in this browser and is reused. The first model you download also brings a 27 MB runtime that the other models share. You can delete models later from the models window in the Options menu.

How do I make SRT or VTT subtitles from a recording?

Transcribe it with Analyze ▸ Transcribe…, then click Export SRT or Export VTT at the top of the Transcript panel. The subtitle lines are built from the word timings.

Is my recording uploaded to transcribe it?

No. The speech recognition model runs in your browser on your own device. Only the model itself is downloaded, once; your audio and the resulting text are never sent anywhere.

Can I transcribe the audio of a video?

Auditius does not open video files. Export the audio track from your video software, open it here and transcribe it; subtitles exported as SRT or VTT line up with the audio.

Audio tools