You convert an OGG file to text by uploading it without conversion and without an account. OGG is a container and Opus one of the codecs it can carry — and a WhatsApp or Telegram voice note happens to be the one inside the other, which is why the same recording turns up labelled .ogg or .opus. We take either extension identically.
Upload a recording. Create an account only after you see the result.
Drop a file here or choose one from your device. We will show you the beginning of the transcript first.
Which transcription model gets the fewest words wrong?
skryba.ai transcribes on Scribe v2, which gets 2.2% of words wrong in Artificial Analysis' independent index — the lowest of every model measured. Gemini 3.5 Transcribe gets 2.6% wrong, Whisper large-v3 4.1%. It is the same model that writes the transcript preview you see before paying.
Word Error Rate — words transcribed incorrectly — lower is better
Model
WER
skryba.ai (Scribe v2)the model we transcribe on
2.2%
MAI-Transcribe-1.5Microsoft Azure
2.4%
Gemini 3.5 TranscribeGoogle
2.6%
Universal-3 ProAssemblyAI
3.1%
GPT TranscribeOpenAI
3.3%
Whisper large-v3OpenAI
4.1%
Nova-3Deepgram
5.2%
Source: Artificial Analysis, AA-WER v2 index, non-streaming mode — read 27 August 2026. Seven of the 40-plus models measured: the top of the table and the names you already know. The index is measured mostly on English audio; for other languages, Polish included, no comparable public ranking exists. That is why we show you the start of every transcript before we ask for anything.
OGG or Opus — the container and what sits in it
OGG is a container — a box for audio tracks. Opus is a codec — the way the sound inside is encoded. Neither is the other: an OGG container can just as well hold Vorbis or FLAC, and Opus is packaged outside OGG too — in WebM, or as a bare .opus file. A messenger voice note happens to be the most common pairing, Opus in OGG, which is why the same recording turns up with one ending or the other.
In practice you do not have to settle it. We take .ogg, .oga and .opus, and OGG with Vorbis or FLAC inside it too — older files from games or audiobooks, say — and read the audio track in every case. Converting before upload improves nothing, for two different reasons: if what is inside is lossy, transcoding to MP3 piles artefacts on artefacts, and if it is lossless, MP3 introduces them for the first time.
Telegram saves voice notes as OGG with Opus audio.
WhatsApp uses the same codec; the exported file may be labelled .opus or .ogg depending on the phone.
Messenger and Signal more often hand over M4A or AAC — which we also take without conversion.
Why voice notes come out well
A voice note is usually the easiest material you can hand to speech recognition, and not because of the format — because of how it was recorded. One person speaks, the phone is held close to the mouth, nobody talks over anybody, and the whole thing lasts under a minute. That is the exact opposite of a meeting room with one microphone in the middle of the table.
Opus handles speech at low bitrates markedly better than older codecs, and was designed for exactly that, so a voice note's small size is not the problem here. The real trouble comes from elsewhere: background — a street, a car, wind — and from talking while moving.
When there are dozens of them
Voice notes rarely arrive one at a time. If you have a whole conversation to transcribe, chopped into a dozen recordings, two things are worth knowing.
Without an account you can upload 5 recordings a day from one address. That limit protects the free preview rather than pushing you towards more accounts — once you have one and the trial starts, only recording time counts.
A very short recording can be refused as too small: a file has to be at least 512 bytes, which in Opus is a fraction of a second. In practice that only ever catches an accidental tap on the record button.
If your messenger can export a whole conversation, the voice notes come out as separate files — joining them into one before uploading saves both the daily limit and a lot of clicking.
It is also worth remembering whose words these are. Somebody else recorded that voice note, and the transcript is a record of it — sharing it is governed by the same considerations as forwarding the recording itself.
Yes — both, and .oga as well. OGG is a container and Opus one of the codecs it can carry; with voice notes it is usually that pairing, which is where the two names for one recording come from. Do not convert the file before uploading.
How do I upload a WhatsApp voice note?
Export it from the conversation to disk — usually via “share” or “save” — and upload the file unchanged. It makes no difference whether your phone wrote it as .opus or .ogg.
Does a voice note's low bitrate hurt the transcript?
Effectively not. Opus was designed for speech at low bitrates and handles it markedly better than older codecs. Background noise — a street, a car, wind — matters far more.
I have a dozen short recordings. Can I upload them all?
Without an account the limit is 5 recordings a day from one address. Once you have an account and the trial starts, only total recording time counts rather than the number of files — or join the notes into one file before uploading.