Transcribe
Audio or video → text, optionally with timestamps.
Drop an audio or video file, or click to choose
MP3, M4A, WAV, MP4, WebM. Up to 25 MB.
About Transcribe
Convert any audio or video file into accurate text with optional timestamps and speaker labels. Powered by Groq's Whisper-Large-V3 — among the fastest and most accurate transcription models available. Upload an MP3, MP4, WAV, M4A, or paste a YouTube link.
Frequently asked questions
Which languages does it support?+
100+ languages, with English, Spanish, Hindi, Chinese, French, and German being the strongest.
How long can my file be?+
Up to 2 hours per file. Longer recordings should be split first.
Does it identify speakers?+
Basic speaker turn detection, yes — full speaker diarization (Speaker 1, Speaker 2) is roadmapped.
How accurate is it?+
Whisper-Large-V3 hits 95%+ accuracy on clear English audio. Noisy or heavily accented audio degrades gracefully.
Can I get timestamps?+
Yes — toggle word-level or segment-level timestamps in the output.