Voice AI already beats human agents on some calls. PolyAI’s Nikola Mrkšić on managing expectations, the handoff, and deploying where wins compound.
Enterprise voice AI lives or dies at the handoff. A voice agent can pick up every call, day or night, and still lose the customer the moment it passes a request to a human who makes them start over. That gap — the lost context between the bot and the person, the requests routed wrong — is where trust breaks and where the coordination tax quietly runs up. It’s also the throughline of this episode of Untangled Ops, Front’s series where operators, founders, and investors talk through the work between the work while playing a not-so-complex game of Snake — an honest picture of what happens when coordination goes unmanaged: it just keeps growing, until it runs into a wall.
Nikola Mrkšić is CEO and founder of PolyAI, the enterprise voice AI company he spun out of his research at Cambridge. His agents now handle millions of customer calls a year for brands like Marriott, Caesars, and FedEx. He joins Dan O’Connell, CEO of Front, who built AI infrastructure at TalkIQ and DialPad before Front. The two trade notes on why expectations are the first thing to manage, what a handoff looks like when it’s done right, the calls AI already handles better than people, and why the real limits on AI now are legal and cultural rather than technical.
Key takeaways
Manage expectations around AI first. People expect AI at the ninetieth percentile before they’ll credit it at the fiftieth — set narrow, provable goals first.
The handoff is where trust is won or lost. Make someone repeat everything to a human after the bot fails, and they’ll never try your automation again.
Start where the wins compound. Deploy on high-volume, low-glory work first, so early proof gives everyone permission to go further.
AI can outperform humans in certain call types. A voice agent scored higher CSAT on bereavement calls because callers wanted the transaction, not the sympathy.
The technical limits with voice AI are mostly gone; permission isn’t. What holds AI back now is legal liability, regulation, and buy-in — not raw capability.
Full transcript
Transcript has been edited for brevity and clarity.
Dan O’Connell (host): Hey everyone, I’m Dan O’Connell, CEO of Front and thrilled to be your host for another episode of Untangled Ops — the show where we invite operators, founders, CEOs, and investors to talk through the work between the work. I’m excited to have Nikola Mrkšić, the CEO and founder of PolyAI here today. They’re the enterprise voice AI company that handles millions of customer calls a year for brands like Marriott, Caesars, and FedEx and more. Nikola, welcome to the show. Are you ready to play some Snake?
Nikola Mrkšić (guest): I haven’t done this in a long time, so thank you for having me. The cognitive load of this is a very interesting thing to bring to an interview.
Dan: It’ll put you to the test. So, first question. You and I have both spent our careers building the AI infrastructure that customer service now runs on — you from the dialogue systems and voice agent side, me from my time at TalkIQ and DialPad, and now here at Front. You’ve deployed at large enterprises like FedEx, Marriott, and Caesars. What’s one thing about running AI at scale that nobody understands until they’ve actually done it?
Nikola: It’s really like running a large workforce where people have the expectations of technology. They think it’s software that, once implemented, just works — but it’s a living, breathing thing, much like your workforce. People will expect AI to be at the ninetieth percentile before they’ll acknowledge it’s at the fiftieth, so for the longest time you’re just blocking and tackling those expectations. Its ability to scale and do everything reliably can be a huge boon — people even give you credit for things that aren’t really within your merit, just stuff that happens because you pick up the phone every time, day or night. But equally, someone will come to you convinced the system’s behavior has changed, when really they’ve just found an error for the first time after years of running the system. That’s where you have a lot of explaining to do, and it’s not always easy.
Dan: Do you have any stories or data points you or the team lean on to help convince the AI skeptics out there?
Nikola: We’ve got a bunch. We’ve been at it eight years now, and the first five were heavy evangelical selling — you had to convince people this would work one day. At one point a large regulated financial services client in the UK worked with us on a system that picked up all sorts of call types and dispositions. About one in eight calls had to do with inheritance and bereavement — mostly people who’d lost a loved one and had to take care of an administrative task. We told them initially, look, you shouldn’t do this with AI. It was nowhere near as good as it is today, and it was hard to imagine it doing better than humans. But their call queues were so swamped they had no choice — they said, do your best, let’s see how it does. Very surprisingly, we got the highest incremental CSAT in exactly those call types. At first I thought, are we measuring this right? Then you think about it: if you’re a human agent, especially a new one, and one in eight of your calls is with someone who’s lost a loved one, they don’t really want to talk to you about it. They’re calling because the company’s process requires it. So it makes sense they just want a transactional conversation. AI is capable of non-transactional conversations too, but that’s when I realized the sky’s the limit to what’s possible. Especially back then, and this was 3-4 years ago. The limit is more about what people will allow you to do with AI — in some cases they police it way too hard, in others not carefully enough. That’s the newer phenomenon: a group of buyers now believe AI should be able to do anything, and we’re not there yet either. It’s a fun time to be doing this.
Dan: That’s a super interesting point, and I totally agree. Getting into our next question — our research found that 71% of companies using AI hit a significant issue in the last three months. And the failures weren’t always automation failures; they were coordination failures. Lost context at handoffs, requests routed wrong. What does a successful handoff look like when a voice agent needs to pass something to a human agent?
Nikola: You could say that’s the most important thing to get right if you’re a company that really cares about CSAT. Most companies will tell you they do, but it’s very clear from the first days of deployment which kind you’ve got. A lot of this is change management — my joke is that I started as an AI researcher and I’ve since graduated into an essential partner, telling people how to think about all this. On the handoff itself: the moment you lose context, you’re making the customer repeat themselves — even though they already did you the favor of trying your AI, which not many people want to do. About a quarter of people immediately look for a way around the AI, especially in the US, because Americans have historically been the most abused by bad automation, so they’ve got very little patience for it. If, at the end of it, you disrespect them by making them repeat everything to a human, they’re never going to try your automated experience again — and why would they? You’ve clearly not integrated things. The opposite is: "Hello Dan, I see you were trying to do that — did the bot fail here? No problem, let me do that for you right now, and we’re done." At that point you still have a shot at delighting them, if you ever had one. They might be calling about something so serious there’s no way to win them back, but these are things you learn. The people doing CX — you know this better than I do — have been through a lot. The work in a contact center is genuinely tough. So that handoff really matters, and for the most part it’s done quite poorly.
Dan: I wholeheartedly agree. We saw in our own research that all this coordination — the lack of context, the frustration that comes from it — is the biggest problem to be solved. That anecdote about different expectations across cultures and geos is fascinating; we can dig into it another time. Where do you think AI now outperforms humans, and what’s your go-to story to convince the skeptic?
Nikola: We’ve got a number of cases that are highly non-trivial — like lead follow-up — where AI is just able to do something people made of flesh and blood can’t, because it comes down to the speed of reaction, or scouring an API and figuring out the results. And then there’s the hard, high-volume business of just picking up phone calls, where sometimes you do better with AI simply because of the volume. Those are the lowest-hanging fruit. It’s really all about finding the right place to deploy so you get the first wins and the whole thing compounds.
Take the bereavement piece — what was crazy is we got higher CSAT there, and we’ve had that deployment for years now, so it’s very much not a pilot. Compared to a lot of what people talk about now, where they think AI is in its infancy — sure it is, relative to where it’s going to end up, but where we are today has already driven tens of millions in extra revenue, and helped companies dig themselves out of really difficult retail periods over Christmas and Black Friday. For a lot of it, it’s all about setting the right expectation, implementing within the wider CX estate, and knowing what to measure. For the most part it can do most things you commit to doing — you just really have to commit. And that’s the hard part for a lot of people trying it for the first time.
Dan: You said some really well-pointed things there — you’ve got to have the right expectations, start small, show momentum. People get excited about the potential of AI and then want to start on the most complex, hardest thing to solve. It’s counterintuitive, but it makes sense when you step back: start small, keep expectations reasonable, know it’s going to be a journey.
Okay, last question before we wrap. Five years out, where does AI handle something end to end, stop, and a human has to own the start? It seems like that line is blurring and getting closer together faster — or do you think it’s actually slower than the market thinks?
Nikola: It depends which market. Wall Street and San Francisco don’t have the same expectations. For the most part, the technical barriers to doing most things with AI are already gone. There are still real technical blockers to some things working great straight out of the box — especially in voice, that’s where some things will probably keep irking people. I used to think voice might have a final reservoir of problems that would plague us for a decade. I’m no longer sure that’s right, because our latest model can write down my name perfectly even over a very noisy line — and it was in no way prepared for the monstrosity of a Serbian last name that I have. So at this point the sky’s the limit on capability. Where we don’t know yet is legal liability. You’ll have a few industries that hold out — whether it’s regulatory capture, or just places where it’s really hard to get buy-in for AI, or where workers’ councils get in the way, or mission-critical work. For example, we handle about a quarter of all non-urgent medical transport in the US — and by non-urgent I still mean serious situations: oncology treatments, physiotherapy after major surgery, pre-op things. It’s funny how it moves slowly and then moves fast. Once there’s a precedent for something being done with AI, it goes really quickly. The first one is often really hard, but if a vendor does a good job advertising what they’ve pulled off, that gives everyone else permission. Frankly, I’m almost surprised we haven’t had more serious blowback from some of the high-profile incidents — a competitor deployed a public chatbot for a very well-known retailer that went full racist, and it blew over in about three days. The appetite for AI is just very high right now.
Dan: Looking back at my time in speech recognition, it blows my own mind how far we’ve come in ten years — pronunciation, machines understanding accents. I don’t think it’s completely solved, but the progress is incredible.
Nikola: Despite the huge progress in deep learning, speech recognition moved a lot slower, and frankly it’s still not solved. A lot of the word error rates advertised by vendors singularly focused on speech recognition are complete hogwash and bad evaluation — we’re not at a two percent word error rate. That doesn’t mean you can’t design a system that does really, really well. It’s an interesting problem — still one of the most interesting ones to work on in all of AI.
Dan: Wholeheartedly agree with that. Let’s take a look at where you landed on the leaderboard. And that’s a wrap on this round of Untangled Ops. Thanks so much for giving us a behind-the-scenes look at enterprise voice agents — all while being thrown into the snake pit.
For more conversations on wrangling the coordination tax, subscribe to Front’s YouTube channel @FrontHQ. Thanks again for watching.

