Independent benchmark · September 2026

Podcast transcription: what we measured

Twenty-one transcription systems compared on one real, hour-long podcast episode, named and published in full so the measurement can be checked — every figure below comes from that episode. Three more episodes were run to see whether the findings hold. A person was paid to transcribe the same hour, so that every score can be read against something other than another machine. On ordinary speech the field separates far less than on names — but it is not one flat plateau: two systems pull clear, a broad middle clusters within two and a half points, and a tail falls away below it. On proper names the spread is sixfold. And they differ on everything that comes after both: speakers, omissions, cleanup, timestamps, silent failures. The single number usually called accuracy is measured where the differences are smallest.

Exhibit

“That’s Pogačar for everybody else. They’re like, ‘Pogačar, I heard of that guy.’”

Warren Cycling, “2026 Tour de France Preview!” · 34:12 · spoken clearly, at a calm pace, twice in a row

ElevenLabsPogačar
Gemini 3.5 TranscribePogacar
GladiaPogacar
RSS.comPogacar
AushaPogacar
Apple PodcastsPogacar
TurboScribePogacar
BuzzsproutPagatcha
RiversidePagatcha
MacWhisper · ParakeetPagatcha
gpt-4o-transcribePagaccia
MacWhisper · Base 150 MBPogaccia
DescriptGotcha
Castmagicgotcha
Podbeangotcha
Speechmaticsgotcha
Sonixgotcha
Rev AIgotcha
Captivategotcha
MacWhisper · Whisper Large v3Gotcha
Spreaker / iHeartGach

1 of 21 spelled the name exactly right6 lost only the diacritic14 wrote a different word

Scoring folds diacritics on both sides, so those seven all count as correct in the tables. The split is shown because it is what a reader sees in the file.

What we found

Four lines, before the method.

  • Ordinary words: two systems clear of the field, then nine inside two and a half points that nothing here can order, then a tail — and two failures that are not about recognition at all.
  • Proper names: 12% to 72%, a sixfold spread. Five systems form a top group; which of the five is first, this study does not claim to know.
  • Speakers: three systems match a paid human pass at telling two voices apart — and none of the twenty-one writes down who those voices are.
  • Price: almost no relationship to any of it. The cheapest API is mid-table; two APIs two cents apart on price are fourfold apart on names.

The axes disagree. Which one counts is decided by your show, not by us.

What we did

One podcast is named; three more were run as a check.

Almost every published accuracy figure for speech recognition is a vendor claim, usually produced on read speech or short clips. We asked a different question: how do these services behave on a real, hour-long podcast episode — with interruptions, an ad insert, accents and dense specialist vocabulary?

Every benchmark score and ranking on this page comes from a single episode: Warren Cycling, “2026 Tour de France Preview!”, 60 minutes, two hosts, studio recording. Published with the show owner’s written permission, and anyone can listen to the exact hour we measured. That matters: a benchmark that cannot be checked is no better than a vendor claim.

Three further episodes from other podcasts were run to see whether the findings reproduce: a remote four-person American conversation, an Irish accent in a field recording, and an episode with British hosts. Their owners have not granted permission to publish, so from those episodes we report only whether a finding reproduced — no scores, no titles, no speakers, nothing that identifies a show. The one exception is three de-identified fragments of a few words each, used once to demonstrate that a platform removes words from what was said; a behaviour like that cannot be shown any other way.

Selection

Why these services, and not the hundreds of others.

There are hundreds of products calling themselves transcription tools, with new ones every month. Testing all of them is impossible and unnecessary: as this benchmark shows, several of them return transcripts so alike that a shared engine, or an equivalent pipeline, is the only plausible explanation.

We selected on one principle — what a podcaster actually gets. First, transcription built into hosting platforms: in our own catalogue of 4.7 million podcast feeds — built from public RSS directories, not from this episode — the platforms measured here are used by one in four active podcasts. Second, the engines those platforms run on. Third, the tools most often recommended to podcasters, plus local options for people who would rather their recording went nowhere.

We left out tools built for meetings, calls and dictation — they solve a different problem. The second constraint was the price a podcaster actually pays: an entry plan in the fifteen-to-twenty-dollar range, pay-per-hour pricing, or a free local option.

Trint is the instructive case. A well-known newsroom tool, no permanent free plan, and its trial gave us exactly five minutes of an hour-long episode — twice, on two different podcasts. Nothing can be judged from that slice: we checked, and the ordering of services over five minutes does not match the ordering over a full episode at all. Trint is absent from the tables not because it is bad, but because a podcaster cannot try it on their own episode without paying.

Twenty-one systems are compared on the Warren episode, and every table below is the full list. Two well-known names are absent by their own rules: the terms of use of Deepgram and AssemblyAI restrict benchmarking and competitive analysis. We asked both, in writing and well before publication, for permission to include them and publish the result. Neither granted it, so neither appears here in any form. MacWhisper is listed as three entries rather than one: it is a local application holding several models, and we measured its two strongest plus the free-tier model most people actually reach for. Every number here comes from this one episode — figures only mean something when compared within the same audio.

The reference

Settled by ear, then audited against every machine.

The reference is an editorial transcript, and it is worth being exact about what that means, because the distinction decides how much weight it can carry. Its layout — which positions exist at all — was built from what the systems produced. Its content is human: the full hour was then listened to against the recording and settled word by word, and every turn attributed by ear, with the passages the machines dropped restored by the person doing it. So it is not a machine consensus, which would carry only what the systems agree on; and it is not a transcript typed from silence either. That matters most where the two hosts talk over each other — passages a vote between systems simply drops, because the systems do not agree that anything was said there.

It was then audited against the machines rather than by them. The set of scored name mentions has moved five times as the reference was corrected — 336, 275, 282, 278, 280 and now 283 across 111 names — and the ranking of systems did not change at any of those steps; the methodology carries each move and its effect. The last of them was this audit: every one of the 283 name occurrences was re-checked against what all twenty-one systems produced in that second, by their own clocks. That found twelve positions where the proper-noun canon had written a rider’s name over an ordinary word — “the first ascent of Alpe d’Huez” had become “the first ascent Èze Alpe d’Huez” — and two where a name was left unrecognised. A name nobody spoke is unreachable for every system at once, so each of those positions was quietly deflating the whole axis.

The same axes were then recomputed on two further references built without using any single system as the anchor — one from agreement between systems, one from the audio timeline. The substantive conclusions agree. The scores do not: both drop the overlapping speech, and on ordinary words they compress the field into a plateau that the editorial reference separates. Where the constructions disagree, this page reports the editorial figure and the methodology carries all three.

That last claim used to rest on our own reference, which is a weak place for a claim to rest. So we bought a second opinion — and it holds: of the words the consensus constructions drop, three quarters are words an independent human transcriber, working from the audio alone, also heard. Half of them he wrote exactly as we did. The two full transcripts agree on volume to within one percent, 12 047 words against 11 933, while the consensus grids sit 1 100 to 1 500 words below both.

The line

We paid a person to do the same hour.

A benchmark of machines against each other tells you who is ahead. It does not tell you whether being ahead is any good. So we commissioned a transcript of the same episode from a transcription house that has nothing to do with this study — full verbatim, two speakers, and one instruction that costs the transcriber dearly: write proper names as you hear them, and never look one up.

He is not a competitor and he is not ranked. He is the line the machines are read against, and he appears in the tables as one — highlighted, out of the running. Where a system stands relative to him is the question this page could not otherwise answer.

The obvious objection first: how do you know it is not a machine transcript someone tidied up? We asked the provider, who said in writing that their transcribers work from audio directly, and then we checked it rather than believing it. Two tests, both fixed before the file arrived.

One. If a file were a cleaned-up copy of some system’s output, it would sit closer to that system than our own reference does. We measured that excess for every system in the study. It is negative for all twenty-one — between 2.8 and 7.4 points — meaning the independent human sits further from every machine than our own reference does, which is what you would expect of a reference whose layout came from machine output before a person settled it.

Two. How often does he reproduce a mistake that exactly one system made, and no other? At most 5.7% of such cases — eleven of the 193 mistakes that belong to Descript alone — averaging 1.9% across the rest.

Neither test proves independence on its own — we did not build a corrected machine draft to calibrate what one would score. Both are consistent with the provider’s written statement that the transcript was typed from the audio, and neither shows the pattern a cleaned-up copy of one system’s output would leave.

The order, the provider’s answer, both tests and the transcript itself are published with the data. What he is not is a better reference: we constrained his names on purpose, and one commercial pass does not replace an hour settled by ear.

Axis 1 · Ordinary words

Two clear of the field, then one tight middle.

On ordinary speech — the 98% of the episode that is not proper names — the field splits into two. ElevenLabs at 96.9% and Descript at 94.6% sit clear of everyone; then a gap of nearly five points, and then nine systems inside two and a half points of each other, between 87.2% and 89.6%.

Inside those nine, no system separates from the one next to it: Rev AI at the top and MacWhisper’s Parakeet at the bottom are two and a half points apart, and every step between them is smaller than the measurement can resolve. The order there means nothing.

Whether those nine are one band or two is not something this measurement settles. The only cut inside them falls between Castmagic and Riverside, and the bottom of its interval sits at +0.01 points — a hundredth of a point. It clears zero, so the statistics call it a separation; we still do not print it as one. Our rule is stated rather than improvised: a boundary is published only if the bottom of its interval reaches 0.05 points, and this is one of exactly two gaps in the whole study that fall between those two thresholds. For scale, two quantisations of one and the same model differ by 0.8 points on this axis — sixteen times the threshold. Both verdicts, statistical and editorial, are published for every adjacent pair, so a reader who disagrees with where we drew the line can see the number and draw their own. Read the nine as one broad middle group. What is not thin is the distance across the whole stretch: Rev AI at the top of it clears MacWhisper’s Parakeet at the bottom by 2.4 points, and that survives resampling whole passages of the episode rather than single words. Nine systems, a slope of two and a half points, one hairline cut in the middle of it — against a sixfold spread on names.

An independent human transcript scores 90.5% here — third of everything measured, ahead of nineteen of the twenty-one systems. Two machines are above a commercial human pass on ordinary speech and the rest are below it. Read that as a floor rather than a verdict: our reference sits closer to the machines than to this independent human, because its layout came from machine output before a person settled it by ear, so the ground here tilts towards the machines and he clears most of them anyway.

Below it the decline is gradual — MacWhisper Large v3 85.9%, Gladia 85.4%, Ausha 84.7%, Gemini 84.1%, TurboScribe 83.8%, RSS.com 82.4%, Apple 81.9%, gpt-4o 81.5% and the free MacWhisper base at 77.7%. One sits far below and is not about recognition quality at all: the Spreaker transcript collapsed into a repeated phrase. The market is two leaders, a tight middle and a handful of dropouts, not a smooth ranking — and not one flat plateau either.

Ordinary-word accuracy · 11 504 positions · proper names and fillers excluded · Warren Cycling. The bars start at zero. Most systems occupy a much narrower range here than they do on proper names, but the field is not flat. The WER column — word error rate — is the standard measure computed the standard way — substitutions, deletions and insertions over the whole hour, names and fillers included, true minimum edit distance — so that these numbers can be laid beside any other published benchmark. Lower is better there
ServiceOrdinary wordsWER
ElevenLabs96.9%5.0%
Descript94.6%9.1%
Unaided human transcript (not a participant)90.5%17.1%
Rev AI89.6%14.9%
Castmagic88.9%13.8%
Buzzsprout (hosting)88.9%14.5%
Riverside (recording)88.1%17.3%
Podbean (hosting)88.0%14.4%
Speechmatics87.6%15.3%
Sonix87.6%16.2%
MacWhisper · Parakeet 1.24 GB (local)87.2%16.9%
Captivate (hosting)87.2%17.0%
MacWhisper · Whisper Large v3 (local)85.9%17.5%
Gladia85.4%18.0%
Ausha (hosting)84.7%18.7%
Gemini 3.5 Transcribe (preview)84.1%18.0%
TurboScribe83.8%19.7%
RSS.com (hosting)82.4%21.3%
Apple Podcasts81.9%21.9%
gpt-4o-transcribe81.5%21.5%
MacWhisper · Whisper Base 150 MB (free)77.7%27.3%
Spreaker / iHeart (free)53.3% *58.6%

The two columns disagree, and that is the point of printing both. They agree on the leaders and on the last place, and they reorder the middle: Rev AI is third on words heard and sixth on WER, Riverside sixth and eleventh, Podbean seventh and fourth. The reason is what each measure charges. Ordinary-word accuracy walks the reference and asks whether each spoken word came back; it never charges a system for words it added. WER charges those too, and charges every substitution at full price — so heavy paraphrasing and invented words are cheap in one column and expensive in the other.
Read the human’s WER with its cause. At 17.1% he sits below every machine here, and the reason is the brief he was given, not his ear. He was told to write full verbatim and to look nothing up, so his transcript runs 12 282 words against the reference’s 11 966 — every stutter, repetition and false start the reference settles into clean speech — and he leaves the hardest passages marked inaudible rather than guessing. The decomposition says the same: 652 insertions against 336 deletions, so more of his distance from the reference is words he wrote down and it does not carry than words he missed. WER measures agreement with a reference’s transcription conventions as much as it measures hearing, which is the single most useful thing to know about the number before comparing anyone by it.
The highlighted row is not a competitor. It is an independent human transcript of the same hour, commissioned from a house outside this study — the line the machines are read against, not a ranked entry.
gpt-4o-transcribe was re-run, and the earlier figure was our fault. It has no timestamps and no speaker labels, and its endpoint takes 25 MB at a time, so the hour has to be sent in pieces. We first sent it in twelve-minute pieces — and the binding limit turned out to be not the file size but the model’s output cap, which silently returned each piece with its tail missing. That cost it 18.6 points. Re-run in three-minute pieces, with every response and its token usage published, it scores 81.5% here and 50.2% on names. What remains is not truncation but tidying: it drops stutters, repetitions and false starts, and only 149 of the 1 989 reference words it omits are fillers.
* Spreaker: on this episode its transcript collapsed into a repeated phrase (see below), so 53.3% is less a recognition score than a trace of that failure. Read it as failure, not as a percentage.

This is the number people usually call accuracy. It is measured where the differences are smallest — not where they are absent.

Axis 2 · Proper names

Here the spread is sixfold.

We scored every mention, not “did the service ever spell this name right at least once”. A name said ten times is ten trials.

The objection that a couple of misspelled surnames hardly matter runs into two measurements from our own data. First, this episode is dense with them: 283 mentions in 11 933 words — 24 proper names per thousand, one every forty-two. It was chosen by its owner as a name-heavy hour, so treat that as the demanding end of the range rather than a typical one — but a cycling, football or politics show lives there permanently. Second, people search through names: of 575 real questions readers asked of the podcast pages we run — our own logs, not this episode — 66% contained at least one proper name and 42% contained the host’s name. Names appear in two out of three real questions — this is not an edge case, and it is where a podcast transcript earns its visibility in search: a name spelled wrong is a query the page will never answer.

A person transcribing the same hour by ear gets 14.8% of them. We commissioned an independent transcript from a house outside this study — full verbatim, and the transcriber was told to write names as heard and never to look one up. That instruction is published with the file. It puts him below twenty of the twenty-one machines — only the collapsed free transcript scores lower — and the comparison is a fair one in the way that matters: nothing here looks anything up. A transcription model does not browse the web mid-file either; both sides produce a spelling from the audio in front of them.

One asymmetry is worth stating rather than burying: a model has read the internet already. “Pogačar” sits in its training data with that spelling, which is not a lookup but is not an unaided ear either. So read the line as what it is — proper names are close to unrecoverable from audio alone. The result suggests that prior exposure to a name is a large part of the advantage; it cannot tell us how much of the rest comes from hearing the sound better, and we have not separated the two.

Of the twenty joints in the table below, three are real. Take any service and the one directly beneath it, and ask whether the gap survives resampling: it does in three places out of twenty. The rest is a slope you cannot cut — which is why this page publishes groups and refuses to publish an order.

At the top, five services are a group and not an order. ElevenLabs, Descript, Podbean, Castmagic and Gemini run from 72.4% down to 65.0%, and those gaps do not survive the test: 283 mentions is a small number of trials, and resampling whole names leaves an interval of roughly ±9 points around each figure. First against second is three points — well inside it. What the axis does support is the distance from that group to everything below: 14.8 points to the sixth service, which holds under every scheme we ran.

Proper-name accuracy · 283 mentions · 111 names · Warren Cycling · all 21 systems compared on this episode. MacWhisper is listed as three entries because three of its local models were measured separately
ServiceNames
ElevenLabs72.4%
Descript69.3%
Podbean (hosting)66.4%
Castmagic65.7%
Gemini 3.5 Transcribe (preview)65.0%
gpt-4o-transcribe50.2%
TurboScribe48.1%
Gladia41.7%
Ausha (hosting)41.7%
Speechmatics38.9%
Sonix38.9%
RSS.com (hosting)36.4%
Apple Podcasts35.7%
MacWhisper · Whisper Large v3 (local)28.3%
Captivate (hosting)25.4%
Buzzsprout (hosting)25.4%
Riverside (recording)18.4%
MacWhisper · Parakeet 1.24 GB (local)18.4%
Rev AI15.5%
MacWhisper · Whisper Base 150 MB (free)15.2%
Unaided human transcript (no name lookup · not a participant)14.8%
Spreaker / iHeart (free)12.0%

The highlighted row is not a competitor. It is an independent human transcript of the same hour, commissioned from a house outside this study, and it is placed by its proper-name figure like every other row — on ordinary words it stands third. The order it was given, and the checks that established it is not a corrected machine draft, are published with the data.
Two notable services are missing from this comparison — Deepgram and AssemblyAI. Why is set out below.
gpt-4o-transcribe carries no timestamps and no speaker labels, and its published figure here is a re-run: the first one was cut short by our own chunking, which is set out where its number appears.
* Spreaker: on this episode its transcript collapsed into a repeated phrase (see below), so 53.3% is less a recognition score than a trace of that failure. Read it as failure, not as a percentage.

The newest arrival

A model announced on Tuesday, measured on Wednesday.

Google announced Gemini 3.5 Transcribe on 26 August, while this study was being finished. It went through the same pipeline as everything else the next day: the same file, byte for byte, the same reference, the same scorers. It is in public preview, at $0.30 an hour.

On proper names it joins the leading group — 65.0%, inside the band with ElevenLabs, Descript, Podbean and Castmagic, which no measurement here can order. On ordinary speech it is near the bottom third, 84.1%, below all but two of the hosting platforms in the table. That combination is unlike anything else we measured, and it comes with a caveat worth more than the ranking: weigh every name equally instead of every mention and Gemini drops to 56.3% while the other four stay near 68%. Its strength is concentrated in the names this episode repeats, not spread across the names it contains.

It can label speakers and time every word — but not on an hour. Google caps both features at thirty minutes. Ask for them on a sixty-minute file and the API returns HTTP 200, a complete transcript, and not one annotation: the feature is dropped in silence, and the response says nothing about a limit. Ask for both at once and you get a 400 reading “Invalid input received”, which also says nothing about a limit. This is why the speaker column has no figure for it — and why seven of twenty-one is really “six returned nothing, and one was never allowed to answer”.

Measured inside its own boundary, it separates voices well and then forgets whose they are. That result belongs to the speaker axis and is set out there in full, with the numbers.

A shared blind spot

Six names nobody got right — and three that fell this week.

Of the 111 names in this episode, six are never spelled correctly by any service — not once, by any of the twenty-one systems. Three of them occur more than once:

zero correct spellings across all attemptsCian, Lejarreta, Auvergne-Rhône-Alpes — 0 of 42 each · Télégraphe, Asgreen, Bol — 0 of 21 each

All six are non-English names or place names. This is not a spread in quality between services but a shared blind spot in the technology: choosing a different engine will not fix it. A vocabulary list or human review will.

Until last week the list was nine. Gemini 3.5 Transcribe, announced on 26 August and measured here the next day, spells three names that twenty systems and a human transcriber had never once produced: Girmay, Narváez, and Seixas — ten mentions out of ten, against 0 of 200 attempts by everyone else. That is not luck at those odds; something in that model knows the name. It is also the sharpest demonstration we have that this axis measures something real and moving: the blind spot is a property of the technology at a date, not a law of nature.

A vocabulary list helps less than you would hope. Feeding the leading service a realistic list — built from earlier episodes of the same show, the best a podcaster could assemble without knowing the answers — moves it from 75.3% to 76.3%. (That experiment ran on the vendor’s newer model, whose no-vocabulary baseline is 75.3%; the 72.4% in the table above is the older one, which is what the default endpoint returns.) That is one point.

We also ran the unfair version: the model was handed the correct spellings in advance, taken from the reference itself. That is not something anyone can do in practice, and it is the ceiling of the method. It reaches 86.2% — so even when told the answers, close to one mention in five is still wrong.

Who is missing

We offered them an open comparison. We did not receive permission.

Deepgram and AssemblyAI are major players in speech recognition, and a reader is entitled to ask why they are not shown here. Both companies’ terms of service contain clauses restricting benchmarking and competitive analysis:

Deepgram · terms of service · version dated 6 August 2026“use or access our Services or any Output for competitive purposes, including model training, benchmarking and other competitive analysis, or developing competing models, products or services”
AssemblyAI · terms of service“…or engage in competitive analysis or benchmarking

We are not going to interpret those clauses on the companies’ behalf. We did the only appropriate thing: back in July, well before publication, we asked both in writing for explicit permission to include them and publish the result — with the full methodology attached, before we had a single final score for anyone, and with a right of reply we would have printed verbatim beside our own conclusions.

AssemblyAI sent an automated acknowledgement promising a reply soon. No substantive reply followed.
Deepgram had a substantive exchange with us about the methodology — and did not answer the question about permission.

In August we followed up with both, to a named contact where we had one, saying plainly that silence would be treated as no permission and giving a date. That date has passed.

Neither company granted permission, and the composition is now final. Neither appears in the comparison at all: no scores, no ranking position, no suggestion that they are better or worse than anyone else. We hold measurements for both and will not publish them — not a figure, not a rank, not a hint. We will not invent motives on their behalf either: a company is entitled to decline, and the same standard was applied to everyone. One vendor whose terms carry a similar clause was asked the same question and said yes; it is in the study in full.

Every service in this study is compared in the open. Two are absent: we asked for permission and did not receive it.

We do not know why companies that sell speech-recognition accuracy did not permit an independent comparison of that accuracy. We will not speculate — but we do think the fact is worth publishing. Readers can draw their own conclusions.

Copies of both documents, with retrieval dates and checksums, are archived alongside the terms of every other service named in this study. None of the others restricts the publication of comparisons.

Axis 3 · The gap between axes

An excellent engine can be nearly empty for search.

The gap between the two axes is a metric no vendor publishes. The clearest case is Rev AI: a solid mid-table performer on ordinary speech, and lowest on names of every paid API in every episode where we built a reference.

Ordinary words89.6%third of the twenty-one
Names15.5%lowest of the paid APIs
Speaker labels98.4%near-perfect

By the classic single number it is an ordinary result and no warning at all: 14.9% WER, sixth of the twenty-one, inside the pack. But if a podcast transcript is meant to work as a search layer, 15.5% on names is a serious limitation — and neither overall figure will ever show it. That is the whole argument for splitting the axes: Rev AI is third on the words a listener said, sixth on WER, and nineteenth on the words a listener would search for.

Axis 4 · Verbatimness

Some services silently remove fillers — and you cannot put them back.

“Uh”, “um”, “you know” — roughly what separates a live conversation from edited prose. Services treat them in opposite ways and almost none says so in advance.

Fillers per 1,000 words · same episode, same file
ServiceFillersWhat that means
Buzzsprout (hosting)15verbatim
ElevenLabs13verbatim; switchable off by parameter
Descript13verbatim
Rev AI13verbatim
MacWhisper · Parakeet (local)13verbatim
Speechmatics11verbatim
Gemini 3.5 Transcribe (preview)10verbatim by default; a `smart` mode strips fillers, and cannot be combined with speakers or timestamps
Podbean (hosting)7partial
MacWhisper · Whisper Large v3 (local)6partial
Castmagic5partial
Gladia4suppresses
Ausha (hosting)4suppresses
MacWhisper · Whisper Base 150 MB (free)2suppresses
RSS.com (hosting)1suppresses
Spreaker / iHeart (free)1suppresses — but this transcript collapsed, so the count means little
Apple Podcasts0cleans
Riverside (hosting)0cleans
TurboScribe0cleans
Captivate (hosting)0cleans
Sonix0cleans
gpt-4o-transcribe0cleans — but it also drops about a sixth of the episode

The meaningful question is not who is more verbatim, but who gives you the choice.

Of everything measured, only ElevenLabs offers cleanup as a switch you can throw without losing anything else: with cleanup on it removes 97% of fillers and every “you know”, without touching names. Elsewhere the mode is hard-wired — you get either every “uh” or polished prose, and the second cannot be undone.

Which is better depends on why you need a podcast transcript. Captions and quotes want verbatim text, or the subtitle stops matching the audio. A page read by people and indexed by search engines is easier to use clean. The problem is that the only way to discover your speech was tidied is to compare the transcript with the recording.

Axis 5 · Who is speaking

They can separate voices. Nobody names them.

Two different abilities need separating. Separating voices — working out how many people are speaking and where each turn begins. And naming them — linking “Speaker 2” to an actual person.

Two thirds do the first. None of the twenty-one systems tested does the second. For bare APIs that is fair: we handed them an audio file and nothing else, so they have no way to know a guest’s name. But hosting platforms and finished tools already hold the episode description in the same system — and still return “Speaker 1” and “Speaker 2”.

Before any of that, the plain split: on this episode — two hosts, nothing exotic — fourteen of the twenty-one separated the voices, six did not, and one could not be asked — Gemini caps diarization at thirty minutes and this episode runs sixty.

Separated the two voices — 14 of 21, plus the human line, ranked by the share of words given the right speaker. The proof-read reference records what was said, not who said it, so we built a second one: a person listened to the full hour against the recording and settled every turn by ear. That matters more than it sounds. A speaker reference elected by the engines themselves would grade each of them partly on their own vote, and would hand every contested moment to the majority — which is exactly where majorities fail, because a transcript cannot show who said a one-word “yeah”. Every figure here is measured against the ear, over 11 901 words.

Riverside99.4%
ElevenLabs99.4%
An independent human · the line99.3%
Apple Podcasts99.1%
Sonix98.9%
Speechmatics98.8%
Rev AI98.4%
TurboScribe96.5%
MacWhisper · Parakeet95.9%
Buzzsprout95.7%
MacWhisper · Base 150 MB93.9%
Gladia93.1%
Ausha92.6%
Descript92.4%
MacWhisper · Large v392.2%

Returned no usable speaker information — 7 of 21: no speaker field at all. Gemini is a seventh and a different case, marked differently below: it has the feature, refuses it above thirty minutes, and this episode runs sixty — measured inside its own limit it is in the section above.

The list is not kept by hand: a system appears on this axis when its output carries two distinct speaker labels, and in the list below when it does not.

Castmagicno speaker field
Podbeanno speaker field
Captivateno speaker field
RSS.comno speaker field
Spreaker / iHeartno speaker field
gpt-4o-transcribeno speaker field
Gemini 3.5 Transcribecapped at 30 min · not asked

Measured inside its own limit, Gemini separates the voices as well as anything here — and then loses track of whose they are. Cut the episode into two halves that fit its thirty-minute cap, each cut at a speaker turn, and the labels score 51.1% on the first and 91.9% on the second — forty points apart on one conversation, the same two hosts, the same settings.

The reason is visible in how you judge it. Over thirty-second windows the labels are 93–98% correct; over a minute, 87–96%; over five minutes, 74–92%; over a whole half they come apart. The turn boundaries themselves are excellent — it finds 88–95% of the real speaker changes and 93% of the changes it marks are real. What does not survive is the identity attached to a label: spk:0 is one host in the third minute and the other host in the twelfth, and across those 58 minutes the two swap places eighteen times, roughly once every three minutes. One five-minute stretch scores 0.3% — which is not a failure to hear anything, it is perfect separation with the labels perfectly inverted.

That is worse than returning nothing. The seven cards above are honest about being empty; a transcript with a speaker on every turn looks finished, breaks its paragraphs in the right places, and hands minutes at a time to the wrong host without a mark anywhere to say where it started. Full account, with the method and the caveats, in the segment note published with the data.

The top is a photo finish, and a person is standing in it. Diarization does not err word by word — it hands a whole turn to the wrong person, so the honest test resamples turns. Under it Riverside and ElevenLabs are tied at 99.4%, and Apple Podcasts at 99.1% cannot be told apart from either — the one comparison that came close put the bottom of its interval at +0.00, a margin that rounds to nothing. Three systems, one photo finish.

Below them the picture is not a ladder either, and the reason is worth stating. Sonix at 98.9% is clear of Riverside and ElevenLabs by about a quarter of a point — but it is not clear of Apple Podcasts, the third member of that photo finish, where the bottom of the interval sits at −0.23. So Sonix opens a new group without being separated from everything above it: that is what non-transitivity looks like in real data, and our banding rule records which comparison forced each break rather than hiding it. Speechmatics sits with Sonix — the gap between them is +0.02, a fiftieth of a point, which we do not print as a boundary. The first break we are willing to call a break is Rev AI, below both.

And the independent human transcript sits at 99.3%, inside that photo finish. Three machines have reached a commercial human pass at telling two voices apart — and eleven have not, seven of them by three points or more. But the human does one thing none of the twenty-one does: he writes down Dean Warren and Randy Warren. The machines separate the voices and leave them as “Speaker 1” and “Speaker 2”.

The spread below them is real and it is wide. Six services sit at 98.4% and above, where the remaining errors are a word here and there. Below them a second group loses one word in twenty-five, and at the bottom Descript and MacWhisper’s Large v3 misattribute roughly one word in thirteen — in an hour of conversation that is several hundred words put in the wrong person’s mouth.

The proportion does not predict the accuracy. Descript splits the episode 61/39, almost exactly the true ratio, and still misplaces those words — a ratio says how much each person talked, never which words were theirs. Parakeet looks worse on the split at 63/37 and places speakers better.

Separation itself varies widely and, unexpectedly, the ranking flips. Being best at names does not make a service best at voices, and the clearest case is the name leader itself.

Being first on names is no guarantee here. ElevenLabs leads the name axis and shares the top of this one — but on a different recording we ran, one with several people in the room, the same service invented speakers who were never there and split a single person across labels. Re-running the identical file produced a different number of them each time. That episode is not published here, so take it as a caution rather than a measurement: two hosts in a studio, as below, is the easy case, and a good score on it does not carry over to a crowded room.

The reverse holds too: services near the bottom of the name axis separated the same voices cleanly. None of the three axes predicts the others, which is why a single “accuracy” number for a transcription service is not a meaningful thing to quote.

One more trap, visible in the table above: the raw count of labels means nothing without their distribution. Speechmatics and Sonix report three speakers on this two-host episode and Buzzsprout four — but those surplus labels hold three words and ten. Count alone would tell you this episode had three or four people in it.

Axis 6 · Can you find the quote

A timestamp on the phrase instead of on the word.

Only six of the twenty-one timestamp every word. The rest mark a chunk of eight to twelve: you can jump to the segment but not to the word, so a clip cannot be cut exactly around a quote without trimming it by hand.

The extreme case is Riverside: 39.7 words per timestamp, more than three times coarser than the next service. A quote is located to within a paragraph, not a sentence.

Gemini is a case of its own on this axis too. It times every word — but only on audio of thirty minutes or less. Ask for word timestamps on an hour and the request is refused outright; ask for speaker labels alone and they are dropped in silence. On this episode it returns no timestamps at all, so a quote cannot be located in it by any means.

One caution about this axis, since it caught us: the same product can differ sevenfold between its own export formats. Buzzsprout gives a mark every 8.5 words in VTT and every 58 in its plain-text export — same transcript, same speakers, different precision. Every figure here is taken from the finest export the product offers, or the axis would measure which button we clicked.

Price

What you are actually buying.

Dozens of brands in this market rest on a small handful of engines. We inferred likely lineage from outputs: near-identical text on the same audio, the same rare errors, and matching timestamp boundaries were treated as strong evidence of a shared engine or an equivalent pipeline. Of the seven platforms here that bundle transcription with the product — six hosts and one recording tool — four produce output matching a known third-party engine; for three we could not identify the source. We found no evidence of proprietary recognition in any of them. That is not a problem in itself — the question is what the markup costs.

Prices cannot be compared head-on: some services sell a bare API, where you pay per hour and receive a file, others a finished product with an editor, storage and collaboration. Hence two tables.

API · price per hour of recognition · public rates, August 2026
ServicePer hourNamesNote
Speechmatics$0.1338.9%cheapest of all
Rev AI$0.2015.5%lowest of the paid APIs
ElevenLabs$0.2272.4%best result measured
Gemini 3.5 Transcribe$0.3065.0%public preview; caps diarization and timestamps at 30 min
gpt-4o-transcribe$0.3650.2%no timestamps, no speakers
Gladia$0.6141.7%$0.20 on an annual commitment; 50 € of free credits
Finished products · effective price per hour of audio, from the plan
ServicePer hourNamesNote
MacWhisper (local)$028.3%€64 once, free thereafter
Castmagic$0.35–0.6365.7%$19–139/month for 30–400 hours
TurboScribe$0.6748.1%$20/month unlimited — rate shown at 30 h/month, cheaper above that
Descript$0.60–1.6069.3%$24–65/month for 10–40 hours
Sonix$2.00–10.0038.9%$2/hour on annual Pro, $10 without
Hosting platforms

You pay for a limit, not for an hour.

With hosting, transcription comes inside the plan you already pay for. So the useful question is not what an hour costs but whether your volume fits the allowance. A weekly hour-long episode needs four hours a month.

Platforms that include transcription with the plan · rates, August 2026 · Riverside is a recording platform and Apple a directory, not hosts — they are here because the transcript comes with the product
PlatformPer monthIncludedEnough forNames
Buzzsproutfrom $15unlimited, on every paid planany volume25.4%
RSS.comfrom $15no limit statedany volume36.4%
Podbean$12600 credits = 2 hours2 episodes66.4%
Podbean Plus$292,400 credits = 8 hours8 episodes66.4%
Ausha Launch$1560 minutes1 episode41.7%
Ausha Boost$29120 minutes, then $4/hour2 episodes41.7%
Captivate$5 onceAssistant add-on, 5 hours, no expiry5 episodes25.4%
Riverside$24unlimited, from the Pro plan upany volume18.4%
Apple Podcasts$0automatic on every published episodeany volume35.7%
Spreaker / iHeart$0no in-house ASR, sends audio outany volume12.0%

The limits differ, and the difference matters. Podbean’s base plan includes two hours a month — a weekly hour-long podcast does not fit, so you need the plan that costs more than twice as much. Ausha’s entry plan includes sixty minutes: exactly one episode, then four dollars an hour. Buzzsprout transcribes without a limit on every paid plan — and sits in the bottom third on names.

Captivate is a special case: five dollars buys five hours that never expire, cheaper per hour than any other platform — and Captivate is also where we found words silently removed from the transcript.

Price predicts accuracy neither within a group nor between groups.

The markup

The same money, a fourfold difference.

Rev AI costs 20 cents an hour, ElevenLabs 22. Two cents apart in price; 15.5% against 72.4% on names — fourfold. There is no trade-off being made here: these are simply different results for the same money.

The Sonix and Speechmatics transcripts are practically indistinguishable. Word agreement sits above our control threshold, and 10 849 of the 10 858 words they share carry the same timestamp to the millisecond — 99.9%, not similar but identical, and both score exactly 38.9% on names. Not merely the same total: they agree on 281 of the 283 individual name mentions, and on the single name where they differ the disagreement splits two ways and cancels. On speaker attribution they are not quite twins — 98.9% against 98.8%, a hair that holds up when whole turns are resampled — which is what a shared engine with different handling around it would look like. That strongly indicates a shared engine or an equivalent pipeline — we have no confirmation from either company.

Speechmatics charges 13 cents an hour directly. Sonix charges from two dollars an hour on an annual plan to ten without one. If those matching outputs do come from a shared upstream engine, the same recognition is being resold at fifteen to seventy-seven times the price. It is a legitimate product: you are paying for an editor, storage and collaboration that an API does not have. The point is only this — the markup does not buy accuracy. It buys the interface.

And the reverse: the transcript included with Podbean hosting scored 66.4% on names — above paid Speechmatics and Sonix, and nearly level with the leader. The podcaster pays nothing extra for it.

Silent failures

Two failures nobody notices.

A class of problem that appears in no accuracy ranking: the service raises no error, the file looks valid, and the content is spoiled. The only way to catch it is to read the transcript.

A free transcript can be half repetition. On this episode the transcript Spreaker and iHeart give away free repeated one phrase 834 times — 41% of all words, nineteen minutes solid. The file looked perfect: valid SRT, evenly spaced timestamps. On the other three episodes the failure did not occur at all.

We then reproduced the same failure in a second, unrelated tool. The free tier of MacWhisper, running the small 150 MB Whisper model on the same audio, locked onto its own phrase and repeated it 587 times — 36% of the 1 631 subtitle blocks; count every repeated line and it is 697 blocks, 42.7%. The two are unrelated as products, and we did not take the match on faith: we downloaded a small quantised Whisper build from its public source, ran it over the same hour on our own machine, and its output landed unusually close to the free transcript — closer than anything else in the study. That is evidence about behaviour, not an identification: we do not know what model this service runs. Two unrelated products, the same distinctive failure on the same hour of speech — this is not one vendor being careless.

What makes this one worth dwelling on is that the collapse is perfectly reproducible. We ran the same episode through the same free model three times, twenty minutes of computation each, once on automatic language detection and twice with English declared explicitly. All three runs returned a byte-identical file — the same checksum, the same phrase, the same 587 repetitions. The failure was deterministic in our three runs, and the language setting changes nothing.

Yet the same application, running the same 150 MB model on the same audio, produced a clean transcript at 77.7% with no repetition — on a different computer. The difference between the two machines is the processor: the Intel Mac collapsed on every attempt, the Apple Silicon one succeeded. We did not open the application to confirm which compute path it took, so we report the correlation and not the mechanism — but it is the same app, the same model, the same file, opposite outcomes, and each one repeatable.

That is the honest cost of the free local route. It is genuinely free and it can be genuinely good — the table above lists its clean result, 77.7%, because that is what the tool does when it works. But the result you personally get depends on the machine you run it on, it is silent when it goes wrong, and retrying will reproduce the same failure exactly. If you are transcribing on an older Intel Mac, read the transcript before you trust it. We cannot see what Spreaker and iHeart run, and we did not try to find out: we make no claim about the engine behind that transcript, only about the transcript itself.

A hosting platform can silently remove words. This episode is clean-spoken, so the finding is not visible here — but on other episodes we ran, ones with conversational swearing, Captivate returned a transcript holding a fraction of the swear words every other service heard in the same audio, and the effect repeated on a second show. Nothing is starred out and nothing is bleeped: the word simply disappears and the sentence closes over the gap.

heard in the recording → delivered to the author · every other service transcribed the word; “[gone]” is our marker, Captivate leaves nothing at all in its place“I don’t give a fuck if it’s humid”“I don’t give a [gone] if it’s humid”
“where the fuck do I start here?”“where the [gone] do I start here?”
“a bit of a shit show”“a bit of a [gone] show”

These three fragments are the only thing quoted from an episode this study does not publish, and they carry nothing that identifies it: no title, no show, no speaker, no timestamp, no date, and no score for any service on that material. They are here because a claim that a platform edits speech has to be shown to be believed. The episode whose owner declined to take part is excluded from the study in full and nothing from it appears anywhere, including here.

A transcript is a record of what was said. There it is altered without warning and without a marker, leaving broken sentences that then get published and indexed.

And a limit you only find by hitting it: Sonix enforces an upload ceiling, and an hour recorded at high bitrate can exceed it — on another episode we hold, the file came back “too large” and had to be recompressed before the service would take it at all. That did not happen here: on this episode Sonix accepted the same bytes as every other system, and its score is computed from them. We state the limit without a figure because the figure belongs to an episode this study does not quote. Why it matters is measurable here: recompression is not free. The same engine on this episode scores 85.5% at 96 kbps and 82.3% at 32 kbps.

Limits

What these numbers do not mean.

  • Ordinary-word accuracy clusters much more tightly than proper-name accuracy — on clean studio audio. On the field-recorded accented episode, even that clustering broke apart and the systems separated sharply.
  • There is a stable group at the top, not a stable winner. Five services pull clearly ahead of the market on names, but their order changes from episode to episode — and on this episode the five do not separate from one another at all. Choosing between them by points, let alone tenths of a point, is meaningless. Which of them suits your show is a different question.
  • A gap smaller than the interval is not a result. Every comparison here was resampled twenty thousand times, in the unit the errors actually arrive in — a whole name, a whole passage, a whole turn — because a name is missed every time it is said and a turn is misattributed as a turn. Anything that does not survive that is printed as a band rather than as an order. The intervals are published with the data.
  • The reference is ours, and that is the one place our judgement enters. Every figure is measured against an editorial transcript: the hour listened to end to end against the recording, disputed passages replayed on the show’s own video, names fixed from public records. We also hold a second human transcript, commissioned from a house outside the study and written under an instruction to look nothing up — which is what makes it a good independence check and a poor reference, since it leaves thirty passages marked inaudible and a reference that declines to resolve a passage marks everyone wrong there. Scored against that transcript instead, the top two swap by a fraction of a point, because it is fully verbatim and rewards writing more stutters down. Both transcripts are published; the methodology sets out why the table uses the one it uses.
  • Running the same command twice does not give the same number. We repeated one system’s run with identical settings: the ordinary-word figure moved 0.8 points and proper names 1.0. The published figure is the first clean run, not the better one. It is a second reason to read the bands rather than the order inside them.
  • One system was not run at its default, and we say which. Speechmatics was run at its higher accuracy tier (operating_point: enhanced) rather than the default, because a publisher ordering this transcript would use it. Every other system ran as it comes. The exact request sent to each service is published beside its output.
  • Difficulty is not set by vocabulary alone. The hardest episode in our set held the fewest rare names — accent and recording conditions decided it.
  • Four episodes are four episodes. A reference was built for each, and group-level conclusions were checked across three or four of them depending on the axis. The specific percentages belong to the specific material and do not generalise.
Disclosure

We have an interest, so we say so first.

Donato builds semantic episode pages for podcasts, so we buy transcription ourselves and have a direct interest in its quality.

This study began as our own procurement question: we needed the most accurate transcription we could get for what we build, and the published claims did not answer it. Having run the comparison, we are moving our own pipeline to one of the systems at the top of it — the same result we are publishing here, applied to ourselves.

That is a conflict of interest and we would rather name it than have it noticed: we are a customer of this market, our choice follows our own measurement, and a study in which the buyer is always proved right is not worth much. Everything needed to check ours is published — the reference, every system’s output, the scoring code and the intervals: github.com/Donato-Digital/podcast-transcription-benchmark. Every file carries a checksum, and reproduce.py recomputes the table from the published outputs.

The methodology is published as a separate document, written to be audited: every decision capable of moving a result is stated, and the episode was selected against criteria fixed before the runs — it is METHODOLOGY.md in the same package.

On rights we drew the line this way. For the episode published here we asked the owner and have written permission. From the other episodes we publish no scores, no titles, no speakers and nothing that identifies a show — only whether a finding reproduced. Three de-identified fragments of a few words appear once, in the section on silently removed words, because that finding cannot be demonstrated without showing a sentence; they carry no title, speaker, timestamp or date, and no service is scored on that material. One owner we approached declined, and that episode was removed from the study in full rather than anonymised: nothing from it appears anywhere, including there.