You convert an MP3 to text by uploading it without an account and reading the free preview of the transcript. Size is driven by bitrate, so the 5 GB limit holds roughly 37 hours of audio at 320 kbps and over 370 hours at 32 kbps. The result comes with timestamps and speaker labels.
Upload a recording. Create an account only after you see the result.
Drop a file here or choose one from your device. We will show you the beginning of the transcript first.
Which transcription model gets the fewest words wrong?
skryba.ai transcribes on Scribe v2, which gets 2.2% of words wrong in Artificial Analysis' independent index — the lowest of every model measured. Gemini 3.5 Transcribe gets 2.6% wrong, Whisper large-v3 4.1%. It is the same model that writes the transcript preview you see before paying.
Word Error Rate — words transcribed incorrectly — lower is better
Model
WER
skryba.ai (Scribe v2)the model we transcribe on
2.2%
MAI-Transcribe-1.5Microsoft Azure
2.4%
Gemini 3.5 TranscribeGoogle
2.6%
Universal-3 ProAssemblyAI
3.1%
GPT TranscribeOpenAI
3.3%
Whisper large-v3OpenAI
4.1%
Nova-3Deepgram
5.2%
Source: Artificial Analysis, AA-WER v2 index, non-streaming mode — read 27 August 2026. Seven of the 40-plus models measured: the top of the table and the names you already know. The index is measured mostly on English audio; for other languages, Polish included, no comparable public ranking exists. That is why we show you the start of every transcript before we ask for anything.
How many hours of MP3 fit in the limit
The upload limit is 5 GB and it is the same for everyone — anonymous, trial and paid alike. With MP3 that is almost never the constraint, because size is decided by bitrate rather than length.
The figures below are estimates. MP3 is often encoded with a variable bitrate, where the nominal setting is an average rather than a constant, and tags and cover art add bytes of their own.
Bitrate
Typical source
Per hour
Fits in 5 GB
32 kbps
dictaphone, voice mode
14 MB
~370 h
64 kbps
mono recordings, speech podcast
29 MB
~186 h
128 kbps
default export, stereo
58 MB
~93 h
192 kbps
music podcast
86 MB
~62 h
320 kbps
maximum MP3 quality
144 MB
~37 h
Scroll the table sideways to see the remaining columns.
Not sure of your file's bitrate? Divide its size in megabytes by its length in hours and compare with the “per hour” column. An hour of dictaphone audio is usually 30–60 MB.
Does MP3 compression hurt the transcript
At typical settings — effectively not. The model recognises phonemes, not instrument fidelity, and what MP3 discards first sits outside the speech band. The difference between 128 and 320 kbps does not show up in accuracy. Very low bitrates and heavy compression are another matter: the artefacts can eat word endings, especially in noise.
So there is no point re-encoding a file “upwards” before uploading: you cannot recover from a 128 kbps MP3 information that is not in it, and the file only grows. The reverse is worth doing — if you have a WAV, converting to 128 kbps MP3 shortens the upload with no loss of accuracy.
What actually degrades the result: a distant microphone, overlapping voices, reverb in an empty room.
What does not: sensible compression, the container format, the absence of video.
Where MP3 files usually come from
Digital dictaphones — usually mono, 32–128 kbps, ready to upload without conversion.
Exports from editing software and podcast platforms.
Call recordings from phone systems and call-centre platforms.
A conversion from another format, when the original was too large or unsupported.
We also accept M4A, WAV, AAC, OGG, Opus, FLAC and WMA, and MP4, MOV, MKV and AVI on the video side, so converting to MP3 “just in case” is usually unnecessary — except when you want a smaller file to upload.
The limit is size, not length: 5 GB. At a typical 128 kbps that is around 93 hours of audio, and around 37 hours at 320 kbps. In practice no single recording comes close.
Does low MP3 quality reduce accuracy?
At typical settings the difference is negligible, and recording conditions — microphone distance, overlapping voices and reverb — matter far more. At very low bitrates, though, compression artefacts can reduce accuracy.
Do I have to convert my file to MP3?
No. Besides MP3 we accept M4A, WAV, AAC, OGG, Opus, FLAC, AIFF, WMA and AMR, plus video files — MP4, MOV, MKV, AVI, WMV and WebM. Converting only makes sense to shrink a large WAV before uploading — 128 kbps is more than enough.
Does the MP3 transcript have timestamps?
Yes. Every segment carries a timestamp and a speaker label, so you can jump back to a specific point in the audio. SRT export — useful for subtitles — is available on paid plans.