There's a number making the rounds right now with a convenient message attached: most companies already use AI agents, and a majority are scaling them. Success, apparently. In our leadership rounds, though, we talk less about experimenting and more about cleaning up, where an agent actually creates value, and where it just defers work that a human later has to pick up again at higher cost.
My position is a calm one: we use AI, and we get real productivity out of it. It's no cure-all, least of all on the harder projects. We'd never hand it the wheel.
The test we ran
Rather than settle for an opinion, we tried it. We took one of the more complex projects, which one, I'll leave open on purpose , and built parts of it using AI agents pushed to their limit. Not as a gimmick, but to find where the line runs between a genuine performance gain and "an agent can't do this."
Internally I call the result Don Quixote mode. You build a new feature, an old one breaks. The agent fights windmills it has mistaken for requirements. It doesn't grasp the real requirements with any precision, because too much nuance drops out along the way, the same nuance you pick up in client conversations and carry into the spec. In full-stack development, that can almost never be fully pinned down in numbers.
Function is quantifiable. Nuance isn't.
That distinction is the whole thing. Functionality is quantifiable, and it usually gets specified that way too. An agent gets a long way there. Where it tips over is front-end design, UI, interaction design. As soon as specific requirements enter the picture, the results break, and the user-experience expectation simply goes unmet.
Here I'll be blunt. Even the latest frontier models don't yet reach, and probably never will reach , the point where you can say engineers have been replaced by real AI agents. "Probably never" is a strong phrase, especially from the CEO of a digital product engineering firm. I'll say it anyway, because our test points that way, and we don't serve clients wishful thinking.
«It's about the nuance, front-end design, UI, interaction design. That's simply not what an agent delivers once the requirements get specific. . Jakob Kaya»
The real trap is called sunk cost
The most interesting part isn't that AI can't do something. The real risk is the sunk-cost trap. You appear to be so far along that putting a real engineer on it now feels wrong. So you keep prompting. And prompting. Instead of stopping, you dig the hole deeper.
A cultural problem reinforces this, and it's worth naming plainly: a good front-end developer has no appetite for fixing AI slop. Nobody enjoys that. So you stick with the agent, because the handover to a human feels too slow and too involved, which is precisely what feeds the spiral.
Honestly, we have no clean rule of thumb for when to stop. No single trigger exists. It's the overall scope that eventually shows its hand: too much still lies ahead, the problems are too big, the ratio no longer adds up. Then the call is made, this goes to humans. Not an elegant metric, but an honest one.
What we take from this
AI stays in use with us because it measurably helps. We don't give it the wheel, and we don't sell it as a stand-in for craft. Whether a product meets the user-experience expectation, that responsibility can't be delegated to an agent.
For product leaders under pressure right now to scale everywhere, my advice is pragmatic. Use agents where function is cleanly quantifiable. Stay skeptical the moment nuance enters. And decide up front when you'll stop, before sunk-cost logic decides for you.






