Video & Audio to Text
Transcribe speech to text and subtitles.
Loading tool…
About Video & Audio to Text
Video & Audio to Text transcribes the speech in any audio or video file into text — right in your browser. Upload a file and get a full transcript you can copy or download, plus ready-made .srt and .vtt subtitle files with timestamps.
It runs an in-browser speech-recognition model, so your file is never uploaded to a server. It's free, with no sign-up. The first run downloads the model, which is then cached for next time.
Frequently asked questions
- Is the transcriber free?
- Yes — completely free with no sign-up and no limits.
- Is my file uploaded to a server?
- No. The audio is transcribed entirely on your own device. Your file never leaves your browser — only the speech model itself is downloaded (once) from a content delivery network.
- What files can I transcribe?
- Audio files like MP3, WAV, M4A, AAC, OGG and FLAC, and video files like MP4, MOV, WEBM and MKV. For videos, the audio track is extracted automatically.
- Can I get subtitles?
- Yes. Along with the plain-text transcript, you can download .srt and .vtt subtitle files with timestamps, ready to drop into a video editor or player.
- How accurate is it?
- It's very good on clear English speech and weaker on noisy, quiet or heavily-accented audio. There are two quality levels — a faster one and a more accurate, slower one.
- Why is it slow, and why is the first run longer?
- The first run downloads the speech model (about 40–80 MB), then it's cached. Transcription is heavy work: a computer that supports WebGPU is much faster than a phone, and long files take several minutes.
- Which languages does it support?
- It's optimised for English. Other languages may partly work but accuracy will be lower.