Video & Audio to Text

Transcribe speech to text and subtitles.

Loading tool…

About Video & Audio to Text

Video & Audio to Text transcribes the speech in any audio or video file into text — right in your browser. Upload a file and get a full transcript you can copy or download, plus ready-made .srt and .vtt subtitle files with timestamps.

It runs an in-browser speech-recognition model, so your file is never uploaded to a server. It's free, with no sign-up. The first run downloads the model, which is then cached for next time.

Frequently asked questions

Is the transcriber free?
Yes — completely free with no sign-up and no limits.
Is my file uploaded to a server?
No. The audio is transcribed entirely on your own device. Your file never leaves your browser — only the speech model itself is downloaded (once) from a content delivery network.
What files can I transcribe?
Audio files like MP3, WAV, M4A, AAC, OGG and FLAC, and video files like MP4, MOV, WEBM and MKV. For videos, the audio track is extracted automatically.
Can I get subtitles?
Yes. Along with the plain-text transcript, you can download .srt and .vtt subtitle files with timestamps, ready to drop into a video editor or player.
How accurate is it?
It's very good on clear English speech and weaker on noisy, quiet or heavily-accented audio. There are two quality levels — a faster one and a more accurate, slower one.
Why is it slow, and why is the first run longer?
The first run downloads the speech model (about 40–80 MB), then it's cached. Transcription is heavy work: a computer that supports WebGPU is much faster than a phone, and long files take several minutes.
Which languages does it support?
It's optimised for English. Other languages may partly work but accuracy will be lower.