We Need to Stop Waiting for AI to Be Perfect

Ten years ago, experts said automated live subtitling would take 20 years. It arrived in five. So why are we still sitting on our hands?
By Ben Anchor — Tuesday, 4 November 2025 · Listen to the podcast episode
I’ve spent the better part of a decade watching people find reasons not to use AI. The accuracy isn’t quite there. The tools don’t plug into our existing workflows. YouTube’s automated captions are rubbish, therefore all speech-to-text must be rubbish. The excuses are endless, and most of them are grounded in a fundamental misunderstanding: that technology has to be flawless before it can be useful.
Paul Markham’s journey through 25 years of broadcast innovation—from commercial radio’s shoestring budgets to running Red Bee’s subtitling platform, and now shepherding the industry through Cloud Native Media—offers a counterpoint I find more compelling. Not because he’s a cheerleader for AI (he isn’t), but because he’s watched enough hype cycles to know the difference between what’s real and what’s theatre.
The Speech-to-Text Wake-Up Call
In 2015, when Paul and his colleagues at Red Bee started seriously discussing automated speech-to-text for live subtitling, the consensus was clear: maybe in 15 to 20 years. The technology was there in principle—speech recognition engines had existed since the turn of the century—but the accuracy was so poor that professional subtitlers dismissed it outright. They’d spend as much time correcting the AI’s output as they would re-speaking the audio themselves.
Five years later, the conversation had shifted entirely.
Now, I’m not suggesting the technology is perfect. Panel shows remain a nightmare—jokes that can’t be misrepresented, people talking over each other, laughter drowning out dialogue. Political debates like Question Time are a minefield where a single transcription error can land a broadcaster in accusations of bias. But for the bulk of broadcast content? Speechmatics and similar engines handle accents, complex vocabulary, and sporting events with pre-seeded names without breaking a sweat.

The YouTube canard drives me mad. Yes, YouTube’s automated captions are terrible. But that’s not because automated speech-to-text is terrible—it’s because Google isn’t using state-of-the-art tools for it. Anyone who’s fed audio into OpenAI’s Whisper or ChatGPT’s voice mode knows the gap between what YouTube offers and what’s actually possible. Yet this one poor implementation has become the stick with which people beat the entire category.
What strikes me about Paul’s retelling isn’t just that the timeline collapsed from 20 years to five. It’s that even now, with demonstrably capable tools in the market, adoption remains patchy. The technology moved faster than our ability—or willingness—to integrate it.
The Hype Versus the Reality
Paul draws a distinction I think is crucial: the technology itself isn’t overhyped. The financial promises around it are.
Big tech CEOs are raising billions on the premise that AI will slash headcount and mint billionaires. That’s the bubble. The actual capabilities of the technology—semantic metadata generation, compliance tagging, real-time sign language avatars—are progressing rapidly and proving useful in ways that don’t make for splashy investor decks.
Take Synapse (the pun on ‘sign’ and ‘synapse’ is chef’s kiss). They’ve trained avatars using real British Sign Language signers, and now those avatars can overlay live BSL interpretation on any content. This isn’t replacing human signers out of cost-cutting zeal—it’s enabling access at a scale that was previously impossible. There simply aren’t enough native BSL interpreters to cover every live sports event, every news bulletin, every niche broadcast. The Synapse approach doesn’t just make signing more affordable; it makes it feasible.
That’s the glass-half-full case for AI, and it’s the one I find myself returning to. Yes, hallucinations are real. Yes, you can’t blindly trust outputs. Yes, integration is hard. But if your benchmark is perfection, you’ll wait forever. If your benchmark is ‘does this solve a problem that’s currently unsolved or prohibitively expensive,’ the conversation changes.

The Interoperability Problem
Here’s where Paul’s Red Bee and Cloud Native Media background really shines through. The reason we’re not seeing more effective AI adoption isn’t just reluctance or scepticism—it’s that most broadcast workflows weren’t designed to accommodate it.
Twenty years ago, you had SDI cables. Any vendor’s kit could plug in and, broadly speaking, it worked. Today’s software-based workflows—cloud or otherwise—lack that plug-and-play interoperability. Projects like the EBU’s DMF (Distributed Media Fabric) are starting to address this, but we’re still in the early days. Without open standards that let different tools talk to each other, you’re locked into single-vendor ecosystems. And if that vendor hasn’t prioritised AI features, or if their AI plugin hasn’t been approved by IT, you’re stuck.
This is the unsexy part of the AI conversation, but it’s the bit that determines whether any of this actually matters. Sprinkling AI on top of legacy systems rarely works. Rebuilding workflows to be cloud-native, modular, and interoperable is the hard graft that makes everything else possible.
Paul mentioned MovieLabs’ 2030 vision in passing, and I think that’s the north star here. It’s not about waiting for AI to be good enough. It’s about building the infrastructure so that when AI is good enough—and in many cases, it already is—you can actually deploy it.
So What Are We Waiting For?
The doom-mongers will tell you AI is going to steal your job, flood the world with deepfakes, and turn us all into passive consumers of algorithmically generated slop. The hype merchants will tell you AI is two years away from superintelligence and infinite wealth.
Both are missing the point.
AI is a tool. Like steel, as someone astutely put it (I wish Paul could remember who—I’d credit them in a heartbeat). You can use it to make life-saving equipment or you can make bombs. The choice isn’t whether AI happens to us; it’s whether we engage with it thoughtfully, critically, and creatively, or whether we abdicate that responsibility to bad actors and rent-seeking platforms.
Paul’s career arc—from DAB radio innovation on a shoestring to architecting access services at scale to now convening the Cloud Native Media community—demonstrates something I think the industry needs to internalise: constraints breed innovation, and the willingness to experiment when the stakes are lower (or the budgets tighter) is what separates the leaders from the laggards.
We didn’t wait 20 years for speech-to-text. It arrived in five because people rolled up their sleeves and made it work. The next wave—semantic search, compliance automation, real-time signing, the lot—will follow the same pattern, but only if we stop demanding perfection and start building systems that can evolve.
The glass is half full. But only if we’re willing to pick it up.
Ancast Intelligence — AI in broadcast consulting by Ben Anchor.
Featured Blog Posts
The real threat isn’t convincing fakes fooling us—it’s real footage being dismissed as fake. Provenance, not detection, is the only way out of this mess.
At MPTS 2026, I heard the same safety blanket wheeled out over and over: ‘We’ll keep humans in the loop.’ But are we clinging to control because we genuinely need to—or because we’re scared to trust the machines?
ChAIse & AIva walked me through one of the gnarliest delivery challenges in modern broadcast. Turns out the real battle isn’t code — it’s making humans cooperate at scale.