
According to Gartner, 30% of generative AI projects are abandoned after proof of concept, and 40% of agentic AI projects will be cancelled by 2027. These figures don't describe technology that doesn't work. They describe organisations that didn't have the conditions to make what they deployed actually work. The distinction matters, because it completely changes what needs to be done.
The pilot went well. The deflection rate climbed. The technology vendor congratulated the teams. Then, a few months after going live, the gains levelled off. CSAT didn't improve as expected. The operational savings are there, but more modest than in the initial presentation. And no one on the team can quite explain why.
This scenario is more common than conference talks suggest. According to Gartner, at least 30% of generative AI projects are abandoned after proof of concept, and more than 40% of agentic AI projects will be cancelled by 2027, for the same reasons: underestimated costs, unclear business value, absent governance. It's not a failing technology. These are deployments that lacked the right conditions.
Five causes come up systematically when you look closely at AI customer service projects that plateau. They aren't independent of each other, and it's often their combination that produces the feeling of a dead end.
AI was deployed to reduce the volume of human interactions, and that's what it did. The deflection rate climbs, HR teams breathe easier, leadership is satisfied. But deflection isn't a satisfaction indicator: it's a volume indicator. A "deflected" customer who didn't get their answer isn't an avoided contact, they're a frustrated customer who will call back, send an email, or quietly give up.
The problem is that the KPI that justified the project at board level (deflection rate) doesn't measure what the customer actually experienced. And in many organisations, no one measures CSAT specific to bot interactions, separately from overall CSAT. The result: operational ROI seems to hold, but relational value erodes without showing up in the dashboards. As David Gillaux, Chairman of Armatis, puts it: "The question isn't 'does it cost less?' but 'is the ROI good?'" The distinction isn't semantic: it determines whether you're managing the right thing.
The fix is known: measure bot CSAT and human CSAT separately, cross-reference with the re-contact rate (customers who get back in touch within 48 hours of a bot interaction), and make the autonomous resolution rate the primary KPI rather than deflection alone. A well-performing chatbot resolves between 60 and 80% of the cases it handles. Below 60%, the scope is too broad or the scenarios are poorly calibrated.
At the time of deployment, the knowledge base was up to date. Products matched the reference sheets. Processes were properly documented. AI answered well. Then time passed: an offer changed, a process was modified, a price shifted. The knowledge base, meanwhile, stayed frozen at go-live day.
It's the most underestimated blind spot in AI customer service projects. AI is only as accurate as what it's fed. As knowledge degrades, AI starts giving outdated or contradictory answers. The customer loses trust, the transfer-to-human rate climbs, bot CSAT falls. And the problem is hard to diagnose because the symptoms look like a technology problem when it's actually an editorial one.
Armatis's CX Horizon 2030 study, conducted among CX directors at major French brands (ENGIE, Volkswagen, SFR, LVMH, MACIF, Matmut, La Banque Postale, Carrefour), puts it directly: the knowledge base is no longer a static repository, it's the engine of AI performance. Agentic AI doesn't compensate for knowledge that's vague, contradictory, or not kept current.
By removing simple cases from the human queue, you didn't lighten it: you hardened it. Advisors now handle a concentrated flow of complex, emotionally-charged, hard-to-resolve situations, often after the customer has already been frustrated by a bot interaction that went nowhere. That's precisely where relational quality and resolution ability matter most.
Yet in many contact centres, advisor training hasn't evolved alongside the AI deployment. According to Zendesk's CX Trends 2024 data, more than one agent in two has received no training on new tools and technologies, and among those trained, only 21% consider it satisfactory. An advisor who only receives complex cases without being trained to handle them, with tools that don't give them the context of what happened earlier in the interaction, produces degraded CSAT precisely on the most visible interactions.
The combined effect is perverse: AI improves efficiency on simple cases, but degrades perceived quality on complex ones, which are exactly the ones customers remember. The net balance can be negative on retention even when it's positive on cost.
At launch, the scope was defined conservatively: a few simple, well-documented request types with accessible data. That was reasonable. Then no one reassessed that scope. Neither to expand it once early successes were confirmed, nor to reduce it once certain cases proved too complex for the bot.
The result is a frozen scope that no longer matches either the initial ambitions or operational reality. The cases that create value aren't automated. The cases that create friction still are. And leadership notices ROI has stopped progressing, without understanding the problem is the absence of active scope governance, not the technology itself.
According to MIT CISR, 95% of generative AI pilot projects have no measurable impact on the P&L. The distinction between the 5% that succeed and the 95% that stagnate largely comes down to this continuous governance: regular scope reviews, documented decisions on cases to expand or withdraw, and impact metrics that go beyond deflection rate.
When a new advisor joins a contact centre, they're coached, evaluated, and continuously trained. Their CSAT score is tracked. Their mistakes are analysed. A development plan is built. When a bot is deployed, it produces its interactions in a relative black box, with no one systematically reviewing poorly-handled conversations, no continuous enrichment of scenarios, no clearly identified owner to decide on evolutions.
This governance gap is the common denominator across every project that plateaus. The project was delivered, the vendor moved on, and internal teams don't have the bandwidth or skills to keep the system in optimal operating condition. The bot gradually degrades with no visible warning sign, until CSAT starts falling and no one connects the dots.
The solution isn't technical: it's organisational. It requires appointing a bot performance owner with a clear mandate, weekly reviews of poorly-handled conversations (exactly like quality listening for human advisors), and a continuous process for enriching the knowledge base and scenarios. The 15 essential contact centre KPIs apply to AI setups too: CSAT, FCR, re-contact rate, quality score.
Armatis's CX Horizon 2030 study, conducted among CX directors at major French brands, offers valuable insight because it moves beyond vendor talking points to gather what practitioners actually say. None of the decision-makers interviewed envisions a 100% automated customer service by 2030. All of them position AI as a lever to augment the advisor, not replace them. Two quotes sum up the field's mindset better than any market projection.
Thierry Suquet, Head of Customer Experience at Volkswagen France: "AI has no room for error. A human agent can apologise and recover a situation. A bot that gets it wrong destroys trust." Dominique Russo, Head of Customer Experience at MACIF: "The sincerity of a relationship is incompatible with a bot pretending to be human."
These positions aren't hostile to AI. They simply describe the conditions under which AI creates durable value: on simple, well-defined cases, with reliable data, in a hybrid architecture where humans remain visible and accessible. And with, according to the study, 85% of decision-makers now viewing customer service as a profit centre rather than a cost centre, the real question is no longer "how do we cut costs with AI" but "how does AI help generate more value in every interaction."
The five identified causes call for five concrete actions, in an order that matters.
The first is to redo the metrics audit. Replace deflection rate as the primary KPI with autonomous resolution rate, and add bot CSAT, post-bot re-contact rate, and transfer rate. These four indicators give an honest picture of what AI actually produces.
The second is to audit the knowledge base. Identify outdated reference sheets, undocumented processes, products whose description hasn't been updated. It's often this editorial work that unlocks the most AI value in the short term.
The third is to train advisors on the new cases now landing with them. These aren't the same interactions as before the bot deployment. They're more complex, more emotional, more demanding. Training needs to keep up.
The fourth is to reassess the bot's scope with a simple grid: on which cases does the autonomous resolution rate exceed 70%? Those are the cases to keep. On which is it below 50%? Those are the cases to remove from scope or rework before keeping them in it.
The fifth is to install formal governance: an identified owner, regular reviews, a continuous enrichment process. AI isn't a product you install. It's a living system that degrades without active maintenance.
Because the metrics used to justify the project (deflection rate, cost per interaction) don't capture all the effects of deployment. A deflected customer who didn't get their answer generates invisible costs: a callback, an email, or silent churn. A robust ROI incorporates both operational efficiency and relational value, measured through bot CSAT, re-contact rate, and FCR.
Deflection gains are visible within the first few weeks. Durable gains, which incorporate customer satisfaction and loyalty, generally consolidate between 6 and 12 months, provided governance is active from the start. Projects that don't set up this governance in the first month are the ones that plateau after 6 months.
No, and it's one of the most consistent lessons from real deployments. AI creates value on simple, repetitive, well-documented, low-emotion cases. On complex, emotional, or high-commercial-stakes cases, humans remain irreplaceable. As Armatis's CX Horizon 2030 study notes, none of the CX directors interviewed envisions a 100% automated service by 2030.
By treating the knowledge base as a living asset, not a project deliverable. That requires an identified owner, a systematic update process for every product or process change, and regular consistency audits. Organisations that invested in this editorial discipline before deploying AI get significantly better results than those that did it afterward.
The pilot almost always succeeds: it's run on favourable cases, with clean data, mobilised teams, and close monitoring. Deployment has to hold up under ordinary conditions: imperfect data, teams under pressure, a broader scope. It's this scaling-up that reveals the real conditions for success. According to Gartner, 30% of generative AI projects are abandoned after POC precisely because these conditions weren't anticipated.
Sources
Armatis is a European specialist in customer relations and business process outsourcing (BPO), operating across multiple continents with thousands of employees serving companies of all sizes and sectors. The company designs and manages end-to-end customer service operations: multichannel contact centres, complaints handling, technical support, back-office and digitised processes. Backed by integrated technology infrastructure and the ability to adapt to any sectoral and regulatory context, Armatis helps its clients combine operational performance, quality of experience and cost control, wherever they need it.
Contact our teams to discuss your challenges and find out how we can support you
Join the leaders who trust our multilingual and technological expertise.