A board asked me last month what success looks like for the GenAI pilot their marketing team was running, and I did not have a reassuring answer. The honest one is that most pilots like it will quietly end, their findings will get written up as learnings, and the budget will roll into next year’s slide. That is not a forecast. It is the base rate.
Six people with better data than mine have said the same thing publicly this year, and the figures line up suspiciously well. The question for anyone commissioning this work is not whether to invest in GenAI. It is how to recognise, early, which pilot is sitting in the 5% that scales and which is in the 95% that will not.
So I went and read what six named voices have put on the record. Four of them, who sell AI advisory for a living, say the same uncomfortable thing. The fifth adds the piece everyone else leaves out. The sixth is a Nobel laureate who thinks the whole market is sized wrong. All of it is useful if you are the one signing the statement of work.
The contributors
- Rita Sallam, Distinguished VP Analyst, Gartner
- Aditya Challapally, lead researcher, Project NANDA, MIT
- Alex Singla, Senior Partner, QuantumBlack / McKinsey & Company
- Jim Rowan, Applied AI leader, Deloitte Consulting
- Nick South, Managing Director and Senior Partner, Boston Consulting Group
- Daron Acemoglu, Institute Professor, MIT, and 2024 Nobel laureate in economics
The number nobody puts in the pilot pitch
The clearest figure in the room comes from Gartner, who in July 2024 predicted at least 30% of GenAI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value. Rita Sallam, their Distinguished VP Analyst, framed the mood plainly: “After last year’s hype, executives are impatient to see returns on GenAI investments, yet organisations are struggling to prove and realise value.”
Gartner put the typical build-and-deploy cost of a GenAI project at $5 million to $20 million. That number matters because it changes the shape of the question. A failed $50,000 experiment is a learning. A failed $8 million one is a board topic.
A year on, the real-world number came in worse than the forecast. MIT’s Project NANDA, in the July 2025 report The GenAI Divide: State of AI in Business 2025, based on 150 leader interviews, a survey of 350 employees and analysis of 300 public deployments, found that only 5% of custom enterprise AI tools reached production. Aditya Challapally, who led the research, attributes the gap not to the quality of the models but to a learning gap: generic tools do not learn from or adapt to a company’s workflows, so a model that works in a sandbox does not survive contact with production.
The industry number tracks the research. S&P Global Market Intelligence’s 2025 survey of over 1,000 enterprises in North America and Europe found that 42% of companies abandoned most of their AI initiatives this year, up from 17% in 2024, and that the average organisation scrapped 46% of its AI proofs of concept before they reached production. The direction of travel is the thing to look at. Pilots are being killed earlier and in larger numbers than they were a year ago.
For a buyer, those three figures taken together are the right starting point for any pilot conversation. The expected outcome is cancellation. Any proposal that does not account for that is a brochure.
Adoption is not the problem. Scale is.
The two largest consulting shops have reached the same diagnosis from different angles.
McKinsey’s annual The State of AI in 2025: Agents, Innovation, and Transformation, published in November 2025, reports that 88% of organisations now use AI in at least one business function. The headline looks like victory. The second number undoes it: only about a third are scaling AI enterprise-wide, 62% of respondents are exploring AI agents, and just 39% report any noticeable profit improvement from them. Alex Singla, a Senior Partner at McKinsey and a co-author of the survey, summarised it in a line worth memorising: “AI is now used almost everywhere, but very few organizations are capturing real value from it.”
Deloitte reaches the same shape of answer from the other direction. The Q4 edition of its State of Generative AI in the Enterprise, published 21 January 2025, surveyed 2,773 director-to-C-suite respondents across 14 countries. More than two-thirds reported that 30% or fewer of their experiments would be fully scaled in the next three to six months. Jim Rowan, Deloitte’s Applied AI leader, writes that “future-thinking organizations are as bullish as ever in building bridges to ROI, all while understanding the need for nuance, and patience, as we embrace this next wave of GenAI.” Rowan’s “patience” is the diplomatic word. Deloitte is telling its clients that scaling will not happen on this budget cycle.
Taken together, Singla and Rowan name the same thing the Gartner and MIT figures imply. The industry does not have an adoption problem. It has a value-capture problem, and value capture is a problem of operations, not models.
“AI is now used almost everywhere, but very few organizations are capturing real value from it.”
Alex Singla, Senior Partner, QuantumBlack / McKinsey & Company
What moves the needle has nothing to do with the model
The piece the vendor analysts tend to leave implicit, BCG makes explicit. In the June 2025 AI at Work report, surveying 10,635 employees across 11 countries, BCG found that 72% of leaders, managers and frontline workers are now regular GenAI users, while frontline usage has stalled at 51%, down one percentage point from 2023. The regression is the finding worth reading twice.
Nick South, Managing Director and Senior Partner at BCG, is explicit about where the lever sits: when leaders demonstrate strong support for AI, the share of employees who feel positive about it rises from 15% to 55%. Only a quarter of frontline employees currently receive that level of support. The report’s own conclusion is that merely introducing AI tools into existing ways of working is not enough. Real value is generated when businesses redesign the workflow end to end and the leadership visibly stands behind it.
Read this through the buyer’s lens. The gap between the 5% that scales and the 95% that stalls is not primarily technical. It is a leadership and operations gap, and vendors do not sell that line because it does not have a licence fee attached.
A measured sceptic worth hearing
A survey of a market where every participating analyst is pricing in a floor is not a complete survey. The useful counterweight comes from Daron Acemoglu, the MIT Institute Professor whose work on technology, inequality and productivity won him a share of the 2024 Nobel in economics.
In The Simple Macroeconomics of AI, published May 2024, Acemoglu built a task-based model and worked through what AI would plausibly do to productivity in the next ten years. His headline: no more than a 0.66% increase in total factor productivity over the decade, and quite possibly less than 0.53%. He estimates AI will reach roughly 5% of tasks, not 50%, and argues the way out is for the industry to prioritise reliable information that increases the marginal productivity of specific workers rather than general human-like conversational tools.
Acemoglu is not telling buyers to pass. He is telling them to size the ambition correctly, invest in use cases with legible per-worker productivity gains, and discount the forecasts coming from parties who need the number to be large. For anyone commissioning a GenAI pilot in 2026 that is the sentence to put at the top of the brief.
What a WordPress or WooCommerce estate needs to actually ship a GenAI pilot
Enough strategy. Here is what the work looks like inside the systems most of the organisations reading this already run. Four buckets of effort decide whether a pilot survives contact with your own operations.
The data layer, which is where pilots quietly die. On WordPress, a model needs a clean, machine-readable view of your content and your WooCommerce catalogue: structured product data on every variant, consistent taxonomies, published schema, and a feed that reflects live stock rather than yesterday’s export. Expect two to three weeks of catalogue remediation on a 2,000 to 10,000 SKU store before any model gives a usable answer. If the brief does not budget this, the pilot’s output is bad because the input was bad.
The access layer. A real pilot needs a per-use-case authenticated endpoint, usually a WordPress REST route or a custom WooCommerce Store API extension, rate-limited and audited. Shipping it against an authenticated service costs roughly one engineering sprint. Shipping it against an authenticated service your legal team is willing to sign off on costs two.
The measurement layer. This is the one every vendor glosses over. You need a locked KPI defined in advance, a baseline taken before the pilot, and an attribution path that survives the AI output being invisible to GA4 (crawlers and assistants do not execute JavaScript, so most dashboards do not see them). On a WordPress estate this means a server-side log plus a reconciliation job against orders. Plan on a week for the instrumentation and a four-week minimum observation window. Shorter and the result is not a result.
The governance layer. Someone owns the content the model is reading from your site, someone owns the prompts, someone owns the escalation path when a model says something wrong about you, and someone owns the kill-switch when a cost line surprises finance. On WordPress VIP the editorial team plus platform engineering carries it naturally. On a managed WordPress stack it is a role that has to be invented. If it is not named in week one, it will be nobody’s job by month three.
Rough aggregate on a mid-market estate: a disciplined pilot that answers one clean operational question, with the four layers above shipped properly, is six to eight weeks of combined platform and marketing effort, not six months. That is the honest size of the smallest useful experiment. Anything smaller is theatre. Anything larger belongs to the 95%.
The question that stays unresolved
The six voices agree that most pilots will not scale. They disagree on why. Deloitte and McKinsey name it as a problem of time and change management. MIT NANDA names it as model fit to workflow. BCG names it as leadership. Acemoglu names it as the underlying technology being smaller than advertised. We do not know which of those is the dominant cause, and the sampling in each of these studies has its own slant. Any buyer commissioning GenAI work in 2026 should hold all four as live hypotheses on their own pilot and plan the measurement so that the result can tell them which one applied to them.
If you are weighing an AI pilot on top of a WordPress or WooCommerce estate right now, that is the conversation we have most often with marketing and digital leads. Our side of it is the implementation plan in the middle of this piece, done against the four layers rather than against the vendor’s demo. If that is useful, the WordPress VIP agency page is the easiest route in, and the wp-vip archive carries the related argument.
Last modified: October 6, 2026
United States / English
Slovensko / Slovenčina
Canada / Français
Türkiye / Türkçe