One of the harder conversations I have with marketing leads goes like this. We pull a traffic report, point at a page that still earns thousands of visits a month, and I say the quiet part. That page is costing you money, and the fix is to delete it. The usual reaction is a long pause, because you do not often get told that a page performing on last decade’s dashboard is now the thing teaching an AI model to recommend you to the wrong buyer.
Twenty years ago we treated a website the way a print house treated an archive. More pages, more coverage, more surface for Google to crawl. The economics of that era rewarded hoarding. The economics of this one punish it. If a buyer’s first touch is a conversation with an assistant, every page you leave up is a sentence in the brief that assistant is being trained on, and some of those sentences are now actively working against you.
Last week I wrote about the gap between what AI says about you and what you would say. The piece below is the operational half. One expert, one specific problem, and the method he uses to find the pages a buyer should never see again.
The pages making you money can be the pages hurting you
Paxton Gray is the CEO of 97th Floor, the agency that ran the Enterprise AI Discoverability Series with WordPress VIP earlier this autumn. In session two, at the thirteen-minute mark, he says the thing most content leads do not want to hear:
There are cases where you will have pages that have good traffic, they’ve got good engagement rate, but it is not aligned with your business. That page should not exist, and it is steering AI in the wrong direction.
Paxton Gray, CEO, 97th Floor, session two, 13:26
That is the catch. Delete a page with real traffic and you lose the sessions. Keep it and you keep the recommendation engine’s reason to think of you as last year’s product. For twenty years the right answer was clearly to keep it. The question is whether the right answer is still the right answer.
Half of your buyers asked a model about you before they called
The reason the question is live now, and was not two years ago, is a shift in buyer behaviour that is already measured. According to G2’s 2026 Buyer Behavior Report, based on 1,076 B2B software buyers surveyed in March 2026, 51% now start research with an AI chatbot more often than with Google, up from 29% a year earlier. On that same guidance, 69% chose a different vendor than they initially planned, and eight in ten say the chatbot accelerated the decision.
Pause on that for a moment. The majority of buyers in your pipeline are forming an opinion of you before any form fill. The pages a model reads to form that opinion are the ones you have published, the ones you forgot you published, and the ones your business unit shipped in 2019 and nobody shut down. Paxton’s framing is the right one: a brand surfing a twenty-year wake of content is handing an assistant hundreds of inputs it never agreed to be assessed on.
Lever 1: Find the pages where a model is telling buyers the wrong thing
Paxton describes the triage method twice in session two, most clearly at the twenty-eight-minute mark. The first pass is anomaly hunting. Pull visits and engagement for every page. Look for the shape that is interesting rather than the one that is big: a page that ranks and does not convert, an old release note from a product line you no longer sell, a glossary entry that defines your category in the vocabulary of a competitor.
The second pass is sharper and new. Run an AI sentiment analysis across the inventory to find the places where a model is stating something wrong about your brand and citing your own site back at itself. That is the mechanical form of the problem: the model is not inventing the misdescription, you are feeding it. Paxton gave an unnamed example on the record. A client that wanted to move into smaller venues found a model calling them the IBM of their category, well established but clunky and not for small venues. The description was not malicious. It was accurate to the pages they had left up.
HubSpot made a version of this move public in 2020, when they recounted deleting large tranches of old blog content that no longer matched their product positioning and reported organic traffic rebounding on the pages that remained. The AI search version of that move is the same instinct with a different metric: you are not pruning to concentrate link equity, you are pruning to narrow the surface a model is averaging over. Jodi Cerretani, CMO of WordPress VIP, puts it bluntly in session one: an average brand is a diluted brand.
Lever 2: Rank pages by what still fits the business, not by what still ranks
The slow pass is the honest work. Page by page, filtered by traffic and engagement, score each one on relevance to the business you run now, not the one you ran when the page was written. Three questions carry most of the weight. Does this page describe a product line we still sell to a buyer we still want. Does the vocabulary on it match the vocabulary our current buyer uses. Does the page’s view of our category match the view we are putting into every new asset this quarter.
Three outcomes come out of that pass. The first is deletion with a 410 Gone, used sparingly, for pages that misrepresent you and have nothing worth keeping. The second is a redirect, where the intent is still relevant and a current page serves it. The third, which is the one most teams undercount, is a rewrite that keeps the URL and replaces the content wholesale. That option matters because an AI crawler, like a human reader, trusts an older URL more than a new one. The authority is in the address, the message is in the body, and you can change the body without discarding the address.
Paxton describes a cruise line in session two that wanted to own Alaska. The agency ran this pass, then rebuilt a cluster covering times of year, animals, ports and excursions mapped to the buyer journey. His description of the outcome is that the brand went from trading the top position back and forth with a competitor to being the most cited cruise line for anything Alaska-related. The brand is not named, no metric is attached, and I would not run that as a case study. As the shape of what a decent prune enables, though, it is instructive.
Lever 3: The one nobody has cracked, measurement
The lever that does not work yet is attribution. Paxton is unusually honest about this on the record. In session one, at twenty-four minutes, he says plainly that showing up in an AI answer is one thing and showing it moved the bottom line is a different thing, and the honest answer is a lot of correlation and not causation. His analogy, from five minutes later, is a billboard. You do not measure a billboard with a conversion pixel. You look at whether brand search went up, whether direct traffic to the site rose, whether the people who did arrive converted at a better rate.
Jodi undercuts her own category of product in the same session, which is worth quoting because very few vendors do. She notes that current measurement tools, hers included at WordPress VIP, really measure humans, because they run on JavaScript in the browser, and AI crawlers do not execute JavaScript. The implication is uncomfortable. The number your dashboard is showing you about AI referral is almost certainly undercounting the AI crawl, and the number is drawn from the one surface those crawlers do not touch.
Session three of the series, on measurement, airs on 20 October. Whatever is on offer there, I would be surprised if it closes this gap in the next two quarters. Plan accordingly: do the pruning on the strength of the diagnosis, not on the strength of the attribution report.
What a prune looks like inside a WordPress and WooCommerce estate
Here is where this stops being an argument and starts being a budget line. Pruning content on an enterprise WordPress or WooCommerce estate is not a one-afternoon piece of work, and the reason it is not is that the data you need to decide what to prune is scattered across four systems that do not naturally talk to each other.
The data model, in order. The post and page inventory lives in the WordPress database, keyed on post ID and post type. On a WooCommerce catalogue the product pages are a second post type and the category and tag archives are a third surface worth auditing. Analytics lives in GA4, keyed on page path, which does not match the WordPress permalink reliably once trailing slashes and parameters enter the picture. Engagement signals, if you have Search Console connected, come keyed by canonical URL. And the AI sentiment pass, if you run one, produces findings keyed by whatever URL the model happened to cite, which may or may not be the canonical one.
The practical build is a reconciliation layer. Export the WordPress post inventory with WP-CLI, normalise the URL column against GA4 and Search Console exports, join on canonical URL rather than on path, flag everything in the bottom quartile of engagement and the top quartile of volume as the first anomaly bucket. That reconciliation alone is where most content prune projects quietly die on an enterprise estate, because nobody budgeted for it and the alternative, picking URLs by eye out of a list of forty thousand, is nobody’s job description.
The deletion step is the easy half. On WordPress a prune is a status change, not a disappearance: a 410 Gone served through a redirection rule, a note in the migration log, and a change to the sitemap. For a WooCommerce store, there is one more wrinkle. Deleting a product page breaks the order history for everyone who ever bought it, because the WooCommerce order record points at the post ID. The pattern we run is to unpublish the product, remove it from the catalogue and from the sitemap, and leave the record addressable to the customer who ordered it. The buyer-facing surface narrows, the historical record stays intact.
Rough effort, on an estate of ten to fifty thousand URLs with a product catalogue of two to five thousand active SKUs. The reconciliation and anomaly bucket, two to three weeks of senior content and analytics time. The AI sentiment pass, one week if you are willing to automate the prompt pass over a reproducible model and have the budget for the token spend. The slow page-by-page relevance pass, four to eight weeks depending on how many people are allowed to decide. The implementation of deletions, redirects and rewrites, one sprint. The cadence Jodi recommends, quarterly, is realistic once the pipeline exists and gets expensive if you keep rebuilding it from scratch.
What the brief for your next content review should actually say
The old brief was length and freshness. The new brief has three lines. First, identify the pages a model is citing against you and remove them. Second, keep the URLs that have earned authority and replace the content under them where the content no longer fits. Third, build the pipeline that lets you do the first two every quarter without starting over. There is a version of this your current agency will tell you is a bigger piece of work than it needs to be, and there is a version your in-house team will tell you is smaller than it is. The honest number for a mid-market enterprise is a quarter of senior content time, and it pays for itself the first time a buyer stops being told you are the clunky, well-established option that is not for them.
If that is the kind of work your business needs doing before buyers finish forming an opinion of you through someone else’s model, that is the kind of work we take on, and the WordPress VIP archive is where the related reading sits.
Last modified: October 2, 2026
United States / English
Slovensko / Slovenčina
Canada / Français
Türkiye / Türkçe