Speech to text splits into two cases. Live dictation, where you speak and words appear immediately — the free tools built into your operating system handle that best. Transcription, where you already have a recording and want it written out — that is what skryba.ai does, with timestamps and speaker labels.
Upload a recording. Create an account only after you see the result.
Drop a file here or choose one from your device. We will show you the beginning of the transcript first.
Which transcription model gets the fewest words wrong?
skryba.ai transcribes on Scribe v2, which gets 2.2% of words wrong in Artificial Analysis' independent index — the lowest of every model measured. Gemini 3.5 Transcribe gets 2.6% wrong, Whisper large-v3 4.1%. It is the same model that writes the transcript preview you see before paying.
Word Error Rate — words transcribed incorrectly — lower is better
Model
WER
skryba.ai (Scribe v2)the model we transcribe on
2.2%
MAI-Transcribe-1.5Microsoft Azure
2.4%
Gemini 3.5 TranscribeGoogle
2.6%
Universal-3 ProAssemblyAI
3.1%
GPT TranscribeOpenAI
3.3%
Whisper large-v3OpenAI
4.1%
Nova-3Deepgram
5.2%
Source: Artificial Analysis, AA-WER v2 index, non-streaming mode — read 27 August 2026. Seven of the 40-plus models measured: the top of the table and the names you already know. The index is measured mostly on English audio; for other languages, Polish included, no comparable public ranking exists. That is why we show you the start of every transcript before we ask for anything.
Which of the two jobs do you have?
Signal
Live dictation
Transcribing a recording
When the text appears
While you speak
After you upload the file
How many people speak
One — you
Usually several
Typical use
A note, an email, a message
An interview, a meeting, a lecture, a hearing
What you need
A built-in OS tool, free
Transcription with timestamps
Scroll the table sideways to see the remaining columns.
If you want to dictate live
skryba.ai does not do this, and there is no point creating an account here for it. You already paid for dictation with the device in your hand.
iPhone and iPad: the microphone icon on the system keyboard.
Android: the microphone in Gboard, the same route.
macOS: Dictation in System Settings, by default double-tap Control.
Windows 11: Windows + H opens dictation in any text field.
Google Docs: Tools → Voice typing, if you are writing in a document anyway.
Where your voice goes depends on the tool, and it is not uniform: dictation in Windows 11 and in Google Docs is processed server-side, while Apple and Google can recognise speech on the device itself — but that depends on the model and the language. If the material is confidential, check the documentation for your specific system before you start dictating.
Come back here if it turns out you do have a recording to convert — a dictaphone file from a meeting, say, or a file somebody sent you.
If you have a recording to convert
This is the case skryba.ai was built for. The difference from dictation runs deeper than timing: a recording has several speakers, background noise, interruptions and room acoustics, and the output has to be a document rather than a stream of words.
Timestamps on every segment, so you can jump back to the point in the audio.
Speaker labels, so it is clear who said what.
Polish and English, with automatic language detection.
Browser editing — you fix the names and terms no model could have known.
The choice of tool matters less to the result than the quality of the recording. The same factors degrade transcription at every provider, so they are worth knowing before you record rather than after.
How far the microphone sits from the speakers — the single biggest factor.
Overlapping voices, which no model separates cleanly.
Reverb in an empty room, worse for accuracy than steady background noise.
Jargon, names and acronyms — here a manual fix is faster than hunting for a better model.
Frequently asked questions
Can skryba.ai do live dictation?
No. skryba.ai converts recordings you already have. For live dictation use your operating system's built-in tool — the keyboard microphone on iPhone and Android, Windows + H on Windows, Dictation in settings on macOS.
How is “speech to text” different from transcription?
“Speech to text” is the umbrella term covering both. Transcription specifically means converting an existing recording, usually with several speakers and timestamps. Dictation means turning your voice into text in real time.
I have a voice recording from my phone — does that work?
Yes. iPhone voice memos are M4A files and messenger voice notes are usually OGG or WebM — all of those are supported. Upload the file without an account and see a preview of the transcript.
Does speech to text work for Polish?
Yes. Polish is this tool's primary language alongside English, and the language is detected automatically. The best test is your own recording — the free preview exists for exactly that.