AiPal

Speech to Text

Transcribe MP3, WAV and M4A audio to text with Whisper, fully locally.

Runs LocallyNo Upload100% Private
How it works
  1. 1Select an audio file or record from your microphone — nothing is uploaded.
  2. 2Audio is decoded and resampled to 16 kHz mono inside your browser.
  3. 3The Whisper tiny model (~40 MB) runs in a Web Worker on your device.
  4. 4Read, copy or download the transcript with optional timestamps.
FAQ

Is my audio uploaded anywhere?

No. Decoding and transcription run entirely inside your browser. Your audio never leaves your device.

Which languages are supported?

Whisper is multilingual — it auto-detects the spoken language, including English and Chinese. English works best on the tiny model; other languages may come out rough because the model itself is small.

Why download a model the first time?

The Whisper tiny model (~40 MB) is fetched once from a CDN and cached in your browser. After that it loads instantly and works offline.

How long does transcription take?

It runs on WASM in your browser — typically roughly real-time length (a 2-minute clip takes 1–3 minutes depending on your device).