HockeyStack’s Buğra Gündüz on getting AI to do high-judgment work: encode taste from the systems you already have, capture feedback, close the loop.
Sometimes the coordination tax shows up as work that only one person is qualified to sign off on. The CTO who reviews every pull request. The agent that re-reads your entire Salesforce instance before it can answer a single question, then answers it wrong. Both are the same problem wearing different clothes: judgment and context live in one head or one system, and everything else queues up behind them. That’s the through-line of this episode of Untangled Ops, Front’s series where operators, founders, and investors talk through the work between the work while playing a retro game of Snake — a fair picture of what unmanaged coordination does to an operation. It keeps growing until it runs into a wall.
Buğra Gündüz is co-founder and CEO of HockeyStack, the revenue AI platform that shows B2B teams what’s actually driving pipeline rather than what their CRM claims. The company has raised over $50 million from Bessemer and Y Combinator to build an AI revenue agent that runs prospecting, new business, and expansion around the clock. He joins Front CEO Dan O’Connell to walk through the "Arda brain": the model HockeyStack built from its CTO’s own review history, which lifted code review recall from 40% to over 80%. They also get into why connecting your tools is a memory problem rather than an integration project, why AI will never reach 100% acceptance rate, and what happens to an organization that stops reviewing what its machines produce.
Key takeaways
Train AI to have taste from the systems you already have. Modeling one person’s judgment from past reviews and merged code got HockeyStack to 40% recall; the feedback loop took it past 80%.
Every high-judgment loop needs three parts. A base model, a learning loop, and humans who engage with it — the third part is most often skipped.
Simply connecting your AI tools to your systems won’t cut it. Without the institutional context for why it was built that way, an agent keeps missing the mark.
Approval rates and feedback themes show whether judgment is landing. Full autonomy is rare, and you accept an error rate the way you do with people.
More output is not the win. Prioritize what machines do with the least effort and keep reviewing what ships, or you start solving the wrong problems.
Full transcript
Transcript has been edited for brevity and clarity.
Dan O’Connell (host): Hey everyone, and welcome to Untangled Ops, the show where operators, founders, CEOs, and investors talk through the work between the work — the clunky handoffs and lost context that pile up between people, and now between people and AI.
That’s the coordination tax: the overhead that nearly every business pays and almost nobody tracks. I’m Dan O’Connell, CEO of Front, and your host for this episode. I’m thrilled to have Buğra Gündüz, the co-founder and CEO of HockeyStack, here. We use HockeyStack, and we love it. It’s the revenue AI platform helping B2B teams see what’s actually driving pipeline, not just what their CRM shows.
They’ve raised over $50 million from Bessemer and Y Combinator to build the very first AI revenue agent that runs prospecting, new business, and expansion around the clock. These interviews are usually done remotely, so I’m thrilled to have Buğra here, sitting awkwardly close to me, as we dive in. Are you ready to play some Snake?
Buğra Gündüz (guest): Yeah, let’s do it.
Dan: All right, on to the questions. First up, walk me through the Arda brain. Your CTO reviews every PR, and you’ve turned that judgment into a model engineers hit 80% accuracy before it ever reaches a human. How did you build that, and how has it improved productivity?
Buğra: I’ll start by saying I think the AI tools we have on our hands are really good at low-judgment tasks. For example, you want to do prospecting. It’s a very simple thing — you have an ICP, you have a set of people you want to reach within those companies, you have specific messaging, and you can turn that into a very good prospecting flow that runs autonomously end to end. When you reach high-judgment tasks such as code review, it becomes a little bit harder than that, because now you have much more of the human taste in play, and much more of the human tribal context that the machine actually doesn’t know.
So we were running into this problem where my CTO, Arda, was spending all of his time doing code review, and that was taking him away from his job. I think lots of early-stage CTOs run this way, and lots of engineering managers at large companies run this way. The way we approached the problem was, number one: How can we encode the taste that Arda has, based on all of his past interactions within our digital systems? You have GitHub code reviews, you have things that Arda himself pushed, you have your Notion architecture write-ups, you have your Linear roadmap, you have your entire change history in Linear. That can be turned into a brain model of sorts, which you then feed into your code review model.
When we first did this, we had very low recall on our code reviews. We were basically comparing the code review produced by the agent to what Arda actually did, and getting a percentage. Our scores were hovering around 40, 45%, which is pretty high for a code review agent, but not high enough to make actual impact. So then we introduced another parameter, which is that you start giving feedback to the model itself. You run it in production for a bit, and you build a learning loop that takes that feedback and pushes it back into the brain model. Without you ever prompting the agent anymore, or changing anything about its architecture, the model just shoots up from 40, 45% recall to 80% plus.
Dan: That’s amazing. You see me smiling and laughing because I’ve been working on an agent to basically be able to write for myself, and no matter what I do in ChatGPT or Claude, the agents still get stuck trying to sound like me. So what I’ve built is a tool along similar lines — provide it with a bunch of information, it learns how I write, and then it drafts a response for me. It allows me to edit that same response side by side, captures all of the edits I make, and then lets me review and accept suggestions: make this a golden rule, make this a hypothesis to test. And then it builds my model. As I go back and do that constant loop, the next time I generate something it’s already got my voice, and the same thing happens again continually. I’ve only done that over the past week, and I’m pretty surprised — initially it’s bad, and then progressively it gets faster and better.
Buğra: I just think it generalizes very well. And not just in writing or in coding. You can think of any high-judgment process that involves a lot of human taste, and you can turn that into a similar loop. You need a base model, a learning loop, and you need actual humans to engage with the model. I think that’s the biggest primitive we’re missing in our AI tools today — what does the human-agent interaction look like? The labs definitely have not tracked this.
Dan: I think a lot of times, as new technologies come out, we get really excited: this is going to be autonomous, this is great, right out of the box. There was probably some sentiment in the market a couple of years ago where you’re asking, what do you need the humans for? And then you suddenly realize you need the humans to go and do a lot of this work — the training, the fine-tuning, the judgment. All of the things that make us unique and ultimately valuable.
Second question for you. Most teams treat connecting their tools as an integration project. You’re treating it as a memory problem. What’s the difference in how you built it, and what outcomes have you seen since implementing it just a couple of months ago?
Buğra: That’s a very good call-out. Think about a company adopting Claude Cowork. They go and integrate a bunch of tools into Claude Cowork. Every time you want to do a task, it goes and retrieves the same data. It doesn’t know how the connected tools actually work. For example, you connect your Salesforce into it — it doesn’t know what your Salesforce looks like, or the retrieval patterns, because that’s unique to your company. And then every single task just rediscovers that from the beginning. This is massive waste, number one. Number two, they do it wrong, so the task output isn’t good.
So I do think there’s benefit to normalizing and transforming the data you get from these connected tools into a unified system. If you’re not using HockeyStack for your revenue connections, you should at least use a data warehouse, and you should spend some data engineering cycles just to build that unified —
Dan: But use HockeyStack.
Buğra: Use HockeyStack — that is the best way to do it. But I’m just saying you need some way of doing this. And it’s not just in revenue tech. Whatever task you’re making AI do, the data retrieval patterns of these agents are not good enough today. Maybe that changes tomorrow. We don’t know. But today it just doesn’t work.
Dan: What are your beliefs on whether the agents get drastically better over time?
Buğra: There are some solvable problems, and there are some unsolvable problems. For example, the fact that your Salesforce has a bunch of custom knowledge and setup in it — that is just inherent to the Salesforce data access problem, and that will never get solved. But you have things like, AI will definitely get good at writing. I believe that goes away as a problem. Or AI will write better, more bug-free code in the future. It will just never get good at understanding organizations deeply, because that is not a model intelligence problem. That’s a different set of problems that need to be solved.
Dan: Do you have a second brain set up at HockeyStack?
Buğra: Of course. I mean, I would say this is not as unified today in our company as it should be. We use HockeyStack for revenue, we use a bunch of internal build tools for coding, and then we use some tools for marketing. The project going on now is to sort of unify.
Dan: Unify them. I think like most businesses. You know, on LinkedIn we read that every business is perfect and has all of this success. You and I are both building businesses, and much like what we talk about here — we’re about to play Snake, and we talk about complexity and running into walls — those are very real challenges I think every business has. I run a second brain personally, and we’re working on how we build one within Front to better understand the context, as you think about agents plugging into different systems. These are really complex challenges, especially at scale. And you’re at a reasonable scale. I think it’s generally really hard. Pretty interesting, but hard.
Buğra: I also think as the company gets larger, there’s more risk to ingesting more data and making it more accessible to all employees. Ideally, in a company brain you want very wide access, but I just think larger companies won’t be able to get there.
Dan: I agree. We were talking about Claude Tag just the other day. Claude Tag can go do some pretty amazing things, and then you get into permissions and access and understanding everything. Interesting times.
All right, hopping into question three. You’ve described an automation spectrum in three stages. Tell AI what to do. AI acts and asks for approval. Then judgment gets fully encoded, and review disappears. What’s the line between stage two and stage three? And how do you know judgment’s actually being captured, not just approximated?
Buğra: By definition of high-judgment tasks, there’s no clear scorecard other than the human review. Therefore, when you’re running the agent in production, you need to have human approval or denial, and you need to capture feedback. Then you just measure the approval rates and the feedback themes over time. There’s a very interesting thing, which is that you can build a self-learning system such that this very quickly moves from stage two to stage three. And sometimes — actually in most cases — you don’t get to full stage three, full autonomy. You don’t get to a full hundred percent acceptance rate, but you accept some error rate, just like you accept some error rate with humans. When you have a new hire, it takes time.
Dan: Not everything is going to be perfect out the gate.
Buğra: I’ve hired salespeople and put them in front of customers on week two, and they mess up, but it’s fine. That’s what has to happen.
Dan: Last question. More AI-generated output isn’t automatically a win if the person still has to sign off on it all. How do you decide what’s worth reviewing?
Buğra: Well, the age-old prioritization problem. People’s immediate instinct is, we have the infinite work machine, let’s make it do more work. As it turns out, if more work produced more results, then Google would take over the world and win at every single sector, because they have infinite resources to do infinite work.
Basically, you have to know your goals, and then you have to break down your goals into what is doable using machines, what is doable with the least amount of effort using machines. Then prioritize those things, and deeply push back on people creating a culture of AI slop just being thrown around the company. Look at generated code, for example. We generate so much code, but we look at all of the code using human eyes, and we review all of the code. As soon as you let that go, you devolve into an organization that solves the wrong problems, and that just produces low-quality work.
Dan: Or just creates work for the sake of creating work. All right — you ready to play some Snake?
Buğra: Let’s do it.
Dan: Let’s go into the rapid-fire round. First question: what’s one thing you think AI should have been able to solve long ago, but still can’t pull off?
Buğra: I think it’s a lot of these open problems in mathematics. Damn, this is definitely very hard. That, I think, should have been solvable in previous versions of models, but it’s only just starting now. I’m deeply excited to see actual scientific progress made using AI.
Dan: Where do most teams go wrong with AI and customer experience?
Buğra: Trusting it too much right off the bat, without review. That’s definitely the biggest challenge.
Dan: Claude or ChatGPT? And why?
Buğra: ChatGPT, just because of Codex now.
Dan: Self-driving cars on all freeways, or we get to Mars first?
Buğra: Mars first.
Dan: Data centers in space?
Buğra: I’m not educated enough, to be honest, but I trust Elon.
Dan: Less than five years?
Buğra: Oof. No — he always gets the timelines wrong.
Dan: What’s one AI tool you can’t live without?
Buğra: Codex.
Dan: Do you speak to your computer?
Buğra: Not as much as I think I should.
Dan: Why not?
Buğra: I just find it unnatural to do in the middle of the office.
Dan: Name one AI feature everyone’s excited about that you think is overrated.
Buğra: The Slack bots. I don’t think Slack is going to be the medium for AI.
Dan: What’s the hardest operational decision you’ve had to make in your tenure as founder and CEO?
Buğra: Layoffs.
Dan: Anything you would do differently?
Buğra: I would ignore the short term completely and focus on the long term, always.
Dan: Favorite book or podcast?
Buğra: Favorite book is The Hard Thing About Hard Things. I don’t really do podcasts.
Dan: Except for this one.
Buğra: Well, I like to speak on podcasts.
Dan: Okay, that’s a wrap.
Buğra: Damn, I didn’t even get close.
Dan: And that’s a wrap on this round of Untangled Ops. Buğra, thanks for being here today and sharing all of your wealth and wisdom in terms of what you’re doing out there building HockeyStack. Let’s see where you stack up on the leaderboard so far. Roll the leaderboard!
Dan: For more conversations on where AI is actually cutting coordination tax and where it’s quietly adding to it, subscribe to Front’s YouTube channel @FrontHQ. We look forward to sharing more.
