Gartner's projection that over 40% of agentic AI projects will be canceled by the end of 2027 got a lot of coverage and surprisingly little introspection. From the build side, the number doesn't shock us at all — because the failures are not random. The same handful of decisions kill these projects over and over, usually before the first line of code. Here are the failure modes we see, and the specific antidote to each.
01Agent-washing: buying the word, not the capability
A meaningful share of "agentic AI" purchases in 2025–26 were rebrands — a chatbot with a new label, an RPA script with an LLM stapled on. Gartner itself flagged that only a fraction of vendors claiming agentic capability actually deliver it. The project then fails not because agents don't work, but because nothing agentic was ever bought.
Antidote: define the capability in operational terms before evaluating anything. Can it plan multi-step work? Use tools? Recover from failure without a human rewriting the prompt? If a vendor demo can't show state, tool calls, and error recovery, it's a chatbot with better marketing.
02No baseline, no business case — value that can't be proven wasn't there
Gartner's stated cancellation drivers are escalating costs, unclear business value, and inadequate risk controls. "Unclear value" is almost always a measurement failure that predates the project: nobody recorded what the process cost before the agent, so nobody can prove what it saved after. The project dies in a budget review it was never equipped to survive.
Antidote: a baseline metric — minutes per case, cost per ticket, error rate — captured before the build starts, and a kill criterion agreed in writing: "if we can't beat X by date Y, we stop." Projects with explicit kill criteria paradoxically survive more often, because they earn trust.
03Autonomy maximalism: automating the hard 30% first
The most expensive instinct in the field: full end-to-end autonomy as the day-one target. The rewarding cases — Klarna-scale support, Hostinger's 75% autonomous resolution — all did the opposite: they automated the repetitive majority and built honest escalation for the rest. Teams that chase the hard 30% first spend their entire budget on edge cases while the easy 70% sits unautomated.
Antidote: tiered autonomy. Ship the boring majority, escalate the rest with context, expand the boundary only as evals prove it's safe to.
04No evaluation harness — the demo that never becomes a product
This is the single most common technical failure we're asked to rescue. The agent looked brilliant in a demo, shipped, and then quality drifted with every model update and prompt tweak — with nobody able to say whether it got better or worse. An agent without a scored, versioned test suite is unfalsifiable, and unfalsifiable systems lose arguments with CFOs.
Antidote: evals before agents. The test harness is the first deliverable, not the last; every change runs against it; the score is the shared language between engineering and the business.
05Guardrails as an afterthought
The numbers here are genuinely alarming: 82% of organizations now use AI agents but only 44% have security policies governing them, and 80% report agents taking unintended actions — including unauthorized system access and data exposure. This is the "inadequate risk controls" leg of Gartner's cancellation triad, and it's the one that kills projects after launch, when an incident lands the whole program in front of legal.
Antidote: permission boundaries, audit trails, and human-in-the-loop checkpoints designed in from day one. We build with healthcare-grade habits regardless of industry — not because every workflow is regulated, but because the discipline is what keeps agents in production.
06Bolting agents onto processes designed for humans
IBM's research captures the deepest failure: 78% of executives agree agentic AI needs new operating models, yet roughly the same share of investment goes to optimizing existing processes. An agent dropped into a workflow designed around human handoffs inherits every one of its inefficiencies — and adds latency. The technology works; the org chart doesn't.
Antidote: redesign the workflow around what agents are good at (parallelism, tirelessness, structured data) and what humans are good at (judgment, exceptions, accountability) — then automate. This is scoping work, and it's why we insist strategy and build live in the same hands.
—The uncomfortable summary
The 40% will mostly be projects that were never set up to succeed: vague capability, no baseline, maximal autonomy, no evals, no guardrails, old workflows. The inverse list is not exotic — it's just discipline applied before enthusiasm. That's the entire reason our engagements start with a readiness sprint that's allowed to conclude "don't build this." A canceled project costs six figures. A week-one "no" costs a conversation.
Mid-flight project wobbling? Or trying to avoid the 40% before you start?
A 30-minute fit call gets you a straight answer and a free written point-of-view memo.
Talk to Medellis