You convert an AMR file to text by uploading it without conversion. Honestly: it is the hardest material we accept, because AMR-NB writes an 8 kHz band at a few kilobits per second — enough for a phone call and no more. Recognition will work, but expect more mistakes than from a dictaphone or a phone held to the mouth.
Upload a recording. Create an account only after you see the result.
Drop a file here or choose one from your device. We will show you the beginning of the transcript first.
Which transcription model gets the fewest words wrong?
skryba.ai transcribes on Scribe v2, which gets 2.2% of words wrong in Artificial Analysis' independent index — the lowest of every model measured. Gemini 3.5 Transcribe gets 2.6% wrong, Whisper large-v3 4.1%. It is the same model that writes the transcript preview you see before paying.
Word Error Rate — words transcribed incorrectly — lower is better
Model
WER
skryba.ai (Scribe v2)the model we transcribe on
2.2%
MAI-Transcribe-1.5Microsoft Azure
2.4%
Gemini 3.5 TranscribeGoogle
2.6%
Universal-3 ProAssemblyAI
3.1%
GPT TranscribeOpenAI
3.3%
Whisper large-v3OpenAI
4.1%
Nova-3Deepgram
5.2%
Source: Artificial Analysis, AA-WER v2 index, non-streaming mode — read 27 August 2026. Seven of the 40-plus models measured: the top of the table and the names you already know. The index is measured mostly on English audio; for other languages, Polish included, no comparable public ranking exists. That is why we show you the start of every transcript before we ask for anything.
Why AMR is hard
AMR was built to fit human speech into a GSM channel, not to record it faithfully. In its narrowband form it samples at 8 kHz with a bitrate between 4.75 and 12.2 kbps — for comparison, an ordinary dictaphone MP3 has several times the bitrate at twice the bandwidth.
The practical effect is always the same: the high frequencies that separate sibilants disappear. In Polish “s”, “sz” and “ś” suffer most, and after them the inflectional endings — precisely what carries a sentence's grammar. Surnames and proper nouns come out worse than the rest, because context cannot help guess them.
We say this plainly because the preview is free, and it is better that you judge the quality on your own recording before starting a trial rather than after. If the result does not convince you, you pay nothing.
What can still be done
Pick the language yourself. At telephone bandwidth, automatic language identification is often less certain than word recognition itself.
Do not convert to WAV or to a high-bitrate MP3. The frequencies that were cut cannot be added back, and the file only grows.
Do not pass the file through a messenger. A second re-encode on already heavily compressed audio piles artefacts on artefacts.
If the same event was captured by anything else — a recorder in the room, a second phone — upload that file instead. It will almost always be better.
For next time: most phones today can record with the system voice recorder, which writes full-bandwidth M4A. If a conversation is going to be transcribed later, that is the difference between hard material and easy material — not a question of better transcription software.
A call recording is usually somebody else's words too. It is worth being clear on what basis it was made and who you share the finished transcript with — we store recordings privately, and your account settings can delete the audio file automatically after 1, 7 or 30 days.
It is, just with lower accuracy than an ordinary recording. The preview is free and needs no account, so the cheapest way to find out is on your own file — a glance at the first few sentences is enough to judge whether the result is usable.
Will converting AMR to WAV improve the result?
No. Converting enlarges the file without adding information that is not in it — an 8 kHz band stays an 8 kHz band. Upload the original.
I have a .3gp file, not .amr
We take that too. 3GP is a container from older phones, and the audio inside it is often AMR — sometimes with a picture, sometimes alone. Upload it without converting.