An hour-long interview becomes text in a few minutes. Every passage carries a timestamp and a speaker label, and you download the finished transcript as TXT or SRT, with or without speaker labels. One hour-long recording costs $6.00 as a one-off. Ten such interviews fit into a 10-hour pack for $29.99, which is $3.00 an hour. Recordings up to five minutes are free once you sign in.
Upload a recording. Create an account only after you see the result.
Drop a file here or choose one from your device. We will show you the beginning of the transcript first.
Which transcription model gets the fewest words wrong?
skryba.ai transcribes on Scribe v2, which gets 2.2% of words wrong in Artificial Analysis' independent index — the lowest of every model measured. Gemini 3.5 Transcribe gets 2.6% wrong, Whisper large-v3 4.1%. It is the same model that writes the transcript preview you see before paying.
Word Error Rate — words transcribed incorrectly — lower is better
Model
WER
skryba.ai (Scribe v2)the model we transcribe on
2.2%
MAI-Transcribe-1.5Microsoft Azure
2.4%
Gemini 3.5 TranscribeGoogle
2.6%
Universal-3 ProAssemblyAI
3.1%
GPT TranscribeOpenAI
3.3%
Whisper large-v3OpenAI
4.1%
Nova-3Deepgram
5.2%
Source: Artificial Analysis, AA-WER v2 index, non-streaming mode — read 27 August 2026. Seven of the 40-plus models measured: the top of the table and the names you already know. The index is measured mostly on English audio; for other languages, Polish included, no comparable public ranking exists. That is why we show you the start of every transcript before we ask for anything.
Three situations, one series of recordings
An interview is rarely a single interview. Usually there are a dozen or more, and that changes the sum: what matters is not the price of one file but the price of the whole series, and how much work is left once the transcript arrives. Research interviews are the longest recordings we receive — an hour and up — and the most common reason anyone buys an hour pack.
Who uploads
A typical series
What is actually needed
A master's or doctoral student, IDI interviews
8–15 conversations of 45–90 minutes
Text to code, and the researcher told apart from the participant
An agency researcher, FGI sessions
4–8 sessions of 90–120 minutes
Timestamps for going back to the audio, and quotes for the report
A journalist
One conversation, today
A verbatim quote and certainty about who said it
Scroll the table sideways to see the remaining columns.
All three take the same route: you upload the file, a few minutes later you have text split by speaker, you correct it in the browser and export it. The only difference is what you pick from the price list.
Who is speaking — and when that stops working
Speaker recognition is on by default. The model splits the recording into people and assigns every passage to one of them; you give them names in the editor: “Interviewer”, “P1”, a surname.
An IDI interview, two people at one microphone — the easiest and most common case.
Three or four people speaking in turn — still good, as long as they do not talk over each other.
An FGI session, six voices and parallel conversation — here the split degrades and needs fixing by hand.
One microphone in the middle of the table — the weakest input there is, whatever the number of people.
Stating it plainly, because the difference between two voices and six is real, and better checked before paying than after. The free preview shows the beginning of the transcript together with the speaker split — on your recording, not on our sample.
What you get for the work that follows
A browser editor — you fix names, terms and speaker labels while listening to the audio in the same window.
TXT to load into MAXQDA, NVivo or ATLAS.ti, in two variants: with speaker labels and without.
SRT with timestamps, if you want to go back to a specific minute of the recording.
A share link for a supervisor or an editor, revocable at any time.
Search across the text of every recording — you type a sentence and get the recording it was said in.
If you cut recordings into pieces — by question, by participant — every file is its own recording: its own preview, its own speaker labels, its own export. The file name stays, so “P2_q3.m4a” is findable by name and by content. One caveat for the sum: every file is billed on its own. Bought one-off, each carries its own $1.99 minimum; paid with pack hours, each rounds up to a started minute separately — thirty two-minute fragments come to $59.70 as one-offs, one hour-long file to $6.00. If you can, upload whole interviews and cut the text afterwards.
Export and the cleaned-up version are part of the paid offer. The transcript itself and the editor also open after a free unlock of a recording up to five minutes.
Verbatim quotes, the cleaned-up version, and anonymisation
The default transcript is verbatim: the “erm”, the repetition and the unfinished sentence all stay in. For qualitative analysis that is usually a feature — a participant's hesitation is data, not a transcription error.
The cleaned-up version removes fillers and stammers, and nothing else. It changes no word order, no wording and no meaning: the model may only delete, and every answer it gives is checked mechanically before it is stored. If the check fails, the passage stays exactly as it was said.
What that version does not do: it does not remove names, addresses or any other personal data. You pseudonymise the transcript yourself in the editor, before exporting — we do not swap names for codes. Worth knowing before you describe the procedure in an ethics committee application.
The transcript is produced automatically, so names, proper nouns and figures are always worth checking against the audio. Downloaded files carry a note that the text was generated by AI — when you publish quotes, that information is sometimes required.
What writing out the whole series costs
A single transcript costs $0.10 per started minute, with a $1.99 minimum. An hour-long interview therefore comes to $6.00. From the fourth hour of audio a pack works out cheaper, and pack hours never expire.
How much audio you have
What to choose
Cost
One interview, an hour
A single transcript
$6.00
Three hour-long interviews
The 5-hour pack
$19.99 ($4.00 an hour)
8–12 hour-long interviews
The 10-hour pack
$29.99 ($3.00 an hour)
Recordings every month, through a whole project
The Pro subscription, 10 hours a month
$19.99 a month ($2.00 an hour)
Scroll the table sideways to see the remaining columns.
A pack and a subscription differ in their deadline: subscription hours expire at the end of the billing period, pack hours stay on the account until you use them. For a series spread across a semester a pack usually fits better. There is also Starter at $9.99 a month with four hours.
Method
How long it takes
Per audio hour
skryba.ai
A few minutes
$2.00–6.00
A freelancer or a transcription agency
1–3 days
PLN 80–400 (~$20–110)
Typing it out yourself
4–8 hours of work
Your time
Scroll the table sideways to see the remaining columns.
The market rates are an order of magnitude, not a quote — they depend on the deadline, the audio quality and whether you order a verbatim record. For ten hour-long interviews they mean anything from several hundred to a few thousand złoty from a supplier, or at least forty hours of your own typing.
Why not Whisper on your own laptop
OpenAI's Whisper is free, open source, and transcribes Polish well. If you can install it and have the machine time, it is a real alternative — we say so plainly, because for a series of interviews it is Whisper, not another service, that we are actually competing with.
Setup: Python, ffmpeg and a model download. Usually an evening on a Mac or Linux, more often two on Windows.
Time: without a graphics card, ten hours of audio means several to a dozen-plus hours of computing, with the machine left on.
Speakers: Whisper alone does not tell voices apart. That takes a separate library and stitching its output into the text.
Silence and music: where nobody is speaking, the model can write sentences nobody said — they have to be caught against the audio.
Corrections: the output is a text file. An editor where you click a sentence and hear the audio is something you find separately.
skryba.ai is the same order of quality without that work: you upload the file, a few minutes later you have text split by speaker in an editor with a player, and ten hours of interviews cost $29.99. If you price your own time at zero and like a terminal, Whisper wins. If your defence is in a month — probably not.
Confidentiality, GDPR and what deletes itself
A research interview is usually personal data, sometimes special-category. Below are facts you can copy into a data-processing description, without our interpretation.
Recordings sit on Cloudflare R2 within the European Economic Area, in a private bucket.
Speech recognition is done by ElevenLabs, Inc. in the USA: it receives the audio through an expiring link and returns the text; the transfer rests on standard contractual clauses — details in the privacy policy.
You can have the audio file auto-deleted after 1, 7 or 30 days; the transcript stays.
You delete a recording together with its transcript and every share yourself, at any time. The same goes for the whole account.
A share link expires after 30 days, and you can revoke it sooner.
Preparing the cleaned-up version sends the transcript text alone to Anthropic in the USA — we ask about that separately, before the first use. The audio file is not sent there.
What we will not do for you: participants' consent, telling them they are being recorded, and anonymising the transcript. That stays on your side, and that is how to describe it in your study documentation.
You pay by card or BLIK, and a one-off purchase keeps no card on file. For 14 days after buying you can withdraw and get the whole amount back, even if you have already used some of the hours — nothing is deducted for minutes spent.
Four ways to transcribe an interview — by hand, by re-dictating, through an agency and with AI — are covered separately, with times, costs and a step-by-step walkthrough.
Will it tell who is speaking with two or four people?
Yes, and that range is where it works best. A two-person IDI interview, and three or four people speaking in turn, split cleanly; passages land under separate speakers, which you name in the editor. With six voices and the parallel conversation typical of an FGI, the split will need corrections. Check it on the free preview of your own recording before paying.
Do you accept a recording from a dictaphone, a phone or Teams?
Yes, all three, and with no conversion. Dictaphones usually write MP3, WAV or WMA, iPhone and Android write M4A, and Teams and Zoom hand back the recording as MP4 or M4A. You upload a video file whole and we pull the audio track out on our side. The limit is 5 GB, and an hour of dictaphone audio is usually 30–60 MB.
What happens to the recording after transcription?
It stays in a private bucket on Cloudflare R2 in the EEA until you delete it — or until the auto-delete setting does it for you after 1, 7 or 30 days. The transcript survives that deletion, because the transcript is the output. You can delete a recording together with its transcript and shares by hand at any time, and the same goes for the whole account.
Is the transcript verbatim, and do you strip names from it?
The transcript is verbatim: hesitations, repetitions and unfinished sentences stay in. The cleaned-up version, available on the paid offer, removes fillers and stammers only — it changes no word order and no wording, and it does not remove names, addresses or any other personal data. You anonymise it yourself in the editor, before exporting.
What formats can I download for MAXQDA or NVivo?
TXT and SRT, each in a variant with speaker labels and without. TXT loads into MAXQDA, NVivo or ATLAS.ti like any text document, and SRT is useful when you want to jump back to a specific minute of audio. There is no DOCX export — if you need a Word document, open the TXT file and save it in that format.
Will a two-hour focus group go through as one file?
Yes. There is no limit on the length of a recording, only on the size of the file: 5 GB, the same for everyone. Audio never comes close to that line; it only starts to bite on long, high-quality video, and then exporting the audio track alone is enough. You pay per started minute, so two hours is $12.00 as a one-off.