The agentic AI market will roughly double from ~$7.3B in 2025 to a projected $139B+ by 2034, and 88% of organizations now report regular AI use in at least one business function. Yet only about 23% have scaled an agentic system into production. We build these systems — our own products and our clients' — and the gap between those two numbers is where all the interesting lessons live. Here's what we're seeing actually ship and pay for itself in mid-2026.
01Document-heavy workflows are the runaway winner
The single most reliable category we've shipped: agents that read, extract, cross-reference, and draft against piles of documents. Clinical notes, lab reports, contracts, claims, compliance filings. Our own medical scribe generates evidence-linked clinical notes where every statement traces back to its source in the encounter; the document-analysis platform we built for a hospital-software client answers clinician questions with citations to the underlying record.
Why this works: the output is verifiable. A citation either points to a real passage or it doesn't. When an agent's work can be checked cheaply, trust compounds instead of eroding — and the economics are obvious, because you're replacing hours of skilled reading, not a judgment call.
02Tiered customer service, with honest escalation
The public proof points are real: Klarna's assistant absorbed the workload of roughly 850 full-time agents, and Hostinger's support agent autonomously resolves about 75% of ~750,000 monthly conversations. The pattern that works is tiered autonomy — the agent fully owns the repetitive 70%, and escalates the rest to humans with full context attached. The projects that fail are the ones that chase 100% automation and burn trust on the hard 30%.
03Intake and triage — deciding what deserves human attention
Inbound leads, support tickets, referrals, claims, patient messages: agents are extremely good at classifying, enriching, prioritizing, and routing. Nobody's job is "read the queue," so nobody mourns automating it — adoption friction is near zero, and the failure mode (an occasional mis-route) is cheap and correctable. This is the pattern we most often recommend as a first build: high volume, low blast radius, measurable from day one.
04Research-and-draft loops with a human sign-off
Agents that gather, synthesize, and produce a draft — a prior-auth request, a market brief, a compliance summary, a proposal — for a human to approve. The human stays the author of record; the agent removes the blank page and the legwork. In regulated industries this isn't a compromise, it's the design: the approval step is where liability wants to live anyway.
05The eval harness — the unglamorous thing that separates shippers from demos
Every production system above has one thing in common that no demo does: a battery of scored test cases that runs on every change. When a client asks why our first two weeks are spent building evaluation before building the agent, the honest answer is that an agent without an eval harness is a demo, whatever it costs. PwC finds 66% of adopters report productivity gains — but the adopters who can prove the gain are the ones who measured a baseline first.
—What this means if you're buying
Start where the work is high-volume, document-shaped or queue-shaped, and checkable. Demand a baseline metric before any build starts, and a scored eval suite as a deliverable — not a promise. Treat "the agent hands off to a human" as a feature, not an admission of failure. And be suspicious of anyone selling full autonomy on day one; the compounding wins in 2026 are all coming from teams that automated the boring 70% brilliantly.
Have a workflow that fits one of these patterns?
A 30-minute fit call gets you a straight answer and a free written point-of-view memo.
Talk to Medellis