🇹🇷 Türkçe: Bu yazının Türkçesini oku →

A six-month engineering plan that ships in a month is the kind of number a board hears twice. Once when the vendor presents it, and once when procurement writes it down as either aspirational marketing or evidence you can hold them to. Which of the two it turns out to be depends less on the first month than on the second.

The case I want to put in front of you today is one of those compressions, published by WordPress VIP in May. Two reasons I am picking it out. The name on it is one your board already trusts as a research authority rather than a technology vendor, and the argument transfers even though the organisation itself does not sell anything.

The angle I want to hold, before anyone gets to the ChatGPT-as-second-referrer figure at the end, is this. What the case actually shows is what happens when you put the right specialist next to your existing team, on a platform that already carries the surface area the work needs. It is a resourcing story on top of a platform story, and both halves matter. Neither one on its own reproduces the outcome.

AI Turned Pew’s Factual Accuracy Into a Distribution Problem

The Pew Research Center is the American nonpartisan fact tank funded by the Pew Charitable Trusts, best known for the surveys and reports that news organisations, academics and policymakers cite when they need a neutral number. Its whole product is factual accuracy on politics, religion, science, technology adoption and demographics. That reputation is the asset, and everything downstream, subscriptions to its newsletters, invitations to brief legislators, citations in the press, rests on it.

Which is why the pressure Pew’s engineering team was under, in Seth Rubenstein’s own words in the WordPress VIP case study published in May 2026, is worth reading closely rather than skimming.

AI has put a lot of pressure on our business as far as capturing audiences, but also pressure from up on high on our team of what they expect us to accomplish. We’re concerned mostly about how LLMs are going to factually ingest our data and get the topic correct.

Seth Rubenstein, Head of Engineering, Pew Research Center

Read that twice. The concern is not visibility, which is where the AI-search conversation usually stops. It is fidelity. A language model that paraphrases a Pew finding on immigration attitudes, or on generational technology use, and gets the topic slightly wrong is not just costing Pew a citation. It is putting a wrong number under Pew’s name in front of a policymaker who now believes it. That is a bigger business problem than losing search traffic.

Every enterprise that trades on authority, and that includes law firms, banks, healthcare providers, industry analysts and any commerce brand whose product story is about safety or provenance, has this same problem now. Pew is the clean case because the accuracy stakes are unambiguous.

The Six-Month Plan and the One-Month Plan Were the Same Plan, Differently Staffed

Rubenstein’s scope, before VIP got involved, was a six-month internal build. Not a small pilot: a proper rebuild of how Pew’s research content was structured so that a model could ingest it accurately. His own line from the case:

To get our site LLM ready was probably gonna be a six-month project just on our own. Having this additional engineering support directly from WordPress VIP, we were able to plan it and execute it basically within two weeks, and we launched a fully featured product within a month.

Seth Rubenstein, Head of Engineering, Pew Research Center

The change here is not a change of scope. The change is a change of resource. VIP put a Forward Deployed Engineer alongside the Pew team under an engagement they call Fully Distributed Engineering: an embedded senior specialist, working on the client’s own repository, on the client’s own deploy cadence, for the length of the project. This is not a managed hosting relationship and it is not a support ticket queue. It is one experienced person, sitting inside your team for a defined window.

Why the compression is realistic rather than marketing. A six-month DIY estimate on this kind of work is honest, from an engineering team that has never shipped it before. Most of that six months is spent finding out which of your assumptions are wrong. A specialist who has built the same layer on ten other estates already knows which of your six weeks were actually two-day tasks in disguise, and which two weeks were secretly the whole project. Removing that discovery time is where the 5x comes from. It is not that anyone typed faster.

1.5 Million Lines in a Month, and ChatGPT Arriving as the Number Two Referrer

The result, as the WordPress VIP case study records it. Pew shipped a release of about 1.5 million lines of changed code, the largest single release in the organisation’s history, in that same one-month window. Within thirty days of the launch, ChatGPT was Pew’s number two traffic referrer, coming from effectively zero. Google Discover visibility rose alongside it. Rubenstein’s summary, in the same case:

Look, the money spent already on the FDE is paying off because we are three, four times ahead of where we should be and 10X ahead of where we would’ve been with just us developing by code. As an engineering leader, it still breaks my brain to think about.

Seth Rubenstein, Head of Engineering, Pew Research Center

The ChatGPT-as-number-two figure is the one that has been travelling in the trade press since May, and it is the one worth being careful about, because it is easily misread. It does not mean ChatGPT is now Pew’s second-biggest revenue source. Pew does not sell subscriptions or ads in a way that makes referral traffic a direct revenue line. What it means is that a large language model with tens of millions of weekly users is now regularly sending people to Pew’s site because it is citing Pew in its answers, and the surface Pew rebuilt is what made that possible. Before the rebuild, the same model was there and Pew was invisible to it.

That is what the number is: evidence that a specific system, that was already talking about your subject to millions of people, is now willing to cite you as its source. For a research institution whose product is accuracy under its own name, that is a bigger operational win than an equivalent bump in Google organic traffic would have been.

What Pew’s Case Does Not Say, and the Questions a Buyer Should Put to a Vendor

Every good case study leaves things out. Three that the Pew write-up leaves out, and that a mid-market or enterprise buyer should put on the table before they sign a Statement of Work with anyone, VIP or otherwise.

First, no revenue attribution. ChatGPT as the number two referrer is a visibility metric. Whether the model referrals convert to newsletter signups, event attendance, report downloads or any other Pew-relevant outcome is not in the case. For a research organisation, this is genuinely fine, because their success measure is authority rather than commerce. For a commerce brand copying the pattern, the same rebuild has to be paired with an attribution model before the board reads the same story and starts asking why your revenue chart did not move.

Second, reproducibility on a non-research organisation is untested in this case. Pew’s content, long-form research reports with clear methodology sections, is unusually well-suited to structured ingestion. An e-commerce catalogue, a services page tree, or a healthcare brand’s therapy-area content is a harder shape. The 5x compression comes partly from Pew having a clean content model to start from. Ask what shape your own content is in before you assume the same timeline.

Third, VIP’s Fully Distributed Engineering is a real engagement with a real cost that scales with the specialist’s seniority and the length of the embed. That figure is not in the case study. Get the cost in writing, alongside a scoped work plan that ties the specialist’s hours to specific deliverables, before you compare it against what your existing agency or in-house team would charge to do the same work over the longer window. The comparison is not FDE against zero. It is FDE against your six-month plan.

What the Same Rebuild Costs Inside a WordPress Install You Already Own

Here is the claim I want to leave with you, because it is the one that is easy to miss in a vendor case study. Everything Pew shipped is possible on a WordPress installation you already own, on hosting you already pay for. VIP compressed the timeline. It did not create the capability. If your organisation is not on VIP and does not plan to be, the work is still doable, on your existing stack, over your six-month window rather than one, at roughly a quarter of the specialist cost. What follows is what the same job actually consists of.

Structured content models, not free-text blobs. This is the biggest single lever and the one most likely to be underestimated. Convert your custom post types so the important facts live in fields (author, date, methodology, sample size, geography, topic taxonomy) rather than in the body copy. A model reads fields cleanly. It reads inference from prose imperfectly. Effort: three to six weeks depending on how much legacy free-text content has to be migrated. This is where a real content editor, not just an engineer, has to sit in the working group.

Schema.org markup on every content template. Article, Organization, Person, FAQPage, HowTo where relevant. Ship it at the template level so it renders on every new page automatically, rather than as a plugin an editor has to remember to fill in. Effort: two to three weeks of front-end and template work for a mid-market publisher. Longer if your theme is deeply customised.

Sitemap alignment, and text fallbacks behind JS-rendered content. The sitemap has to list what you are actually promoting, and every JavaScript-rendered component that carries factual content needs a plain HTML fallback that a non-JS crawler can read. Most current AI crawlers do not execute JavaScript. Effort: one to two weeks, depending on how much of your site was built as a single-page application.

Crawler policy at the edge. Your CDN or web server needs an explicit allowlist for the AI crawlers you want to be cited by, GPTBot, ClaudeBot, PerplexityBot, GoogleOther, Applebot-Extended and a handful of others. And you need the logs surfaced so someone on your team can actually see what is being fetched. If your platform makes this easy, this is days of work. If your platform makes it hard, and quite a lot of managed WordPress hosts do, this is the argument for either changing platform or negotiating a support ticket with your existing one.

A page inventory job. Not glamorous, and not optional. What pages exist, which of them still reflect your current positioning, which are orphaned from any navigation, which are contradicting each other. Effort: weeks of editorial time, not engineering time, and usually the longest single line item on the plan. This is the piece that Pew got a head start on because their content is inherently disciplined. Your organisation may not be so lucky.

What WordPress VIP actually buys, in that shopping list, is not the ability to do any of it. Every WordPress install can. What it buys is the platform already carrying the edge, the object cache, the multisite plumbing and the deploy discipline, plus a specialist who knows which two weeks of the plan were secretly the whole project. If you are running WordPress on managed hosting today and evaluating VIP, this is the honest way to size the difference: the platform work above is one bucket of cost, the specialist time is another, and VIP prices them together rather than separately.

If VIP is not on the table for you, either because the volume does not justify it or because the procurement route is closed for the year, this is the shape of engagement my team runs on standard managed WordPress, on the customer’s existing stack, in the six-month window rather than the one. The output is the same content-model-and-schema layer. What differs is who is doing the work, and how much of your calendar it takes.

What the Pew case actually proves, once you strip the compression story off the top, is that being cited by a large language model in thirty days is a resourcing decision made against a technically feasible plan, and that the plan does not require moving platforms. It requires a serious content model, real schema, honest sitemap hygiene, and a crawler policy someone on your team is willing to own. On VIP that is one month with the right specialist embedded. Off VIP that is six months with the right agency alongside. Both are buyable. The one that is not buyable is the version where nobody does the work and hopes the model figures your brand out on its own, because it will not. If your organisation is at that decision, thirty minutes with my team is worth booking.

Leave a Reply

Your email address will not be published. Required fields are marked *

Close Search Window
↑