In March 2025,
Airbnb migrated around 3,500 React test files from one testing library to another. They had estimated a year and a half of engineering time. It took six developer-weeks.
The interesting part is not the model. It is the loop they built around it. Each file moved through a fixed sequence of checks: refactor it, then run Jest, then ESLint, then TypeScript. A file only advanced when the current check passed. When one failed, the model was handed the exact error and the current version of the file and asked again, usually about ten times, occasionally fifty. What lifted the success rate from 75% to 97% was not better prompting. It was feeding each retry more of the surrounding code.
Meta published the correction to any optimism you might be forming. Their test-generation system puts every candidate through three gates: it has to build, it has to pass five times running without flaking, and it has to actually improve coverage. Three quarters of what the model writes never survives. Of the quarter that does and reaches an engineer, engineers still reject about a quarter of those.
So the model writes four tests to get one worth reading.
Airbnbtest migration
- Jest
- ESLint
- TypeScript
97%
after tuning, from 75%
Who decided
Engineers reviewed the output
Metaunit tests
- builds
- passes 5x
- improves coverage
25%
cleared all three gates
Who decided
Engineers still rejected 27% of those
Googlecode completion
- compile
- reference
- type check
80%
of Go suggestions would not compile
Who decided
Developer accepts or ignores
pull requests
- structured errors
- agent-tagged keys
- audit trail
1,000-1,300/wk
reach review
Who decided
Human merge approval, same as for people
security fixes
- sandboxed execution
- read-only validation
75.8%
catch rate on a bounty corpus
Who decided
Security engineers triage
Shopifyworkflow steps
- deterministic gates
- AI steps
Not published
Who decided
Human PR review, mandatory
production PRs
- OS-level sandbox
- video proof
35%
of merged internal PRs
Who decided
Human review
unit tests
- runs tests
- checks coverage
~11%
of all new tests in the codebase
Who decided
Engineers review the MR
“The filter does not replace the reviewer. It decides what the reviewer sees.”
What they built is the boring part
Look at what is actually in those loops. A build step. A test runner. A linter. A type checker. A coverage report. Airbnb's pipeline is Jest, ESLint and TypeScript in a particular order with a retry between them. None of it is exotic, and none of it was built for AI. These are the checks those teams already had, wired so that a machine has to satisfy them before a human spends attention.
This is what a harness is. It decides what the agent can see, what it can touch, what it can run, and what it has to prove before a change is accepted. The model supplies judgement in the middle. Everything around it is code you own and version, and at the end of it a person still decides.
WordPress has the boring checks too. PHPCS with the WordPress standards, PHPStan, Plugin Check, PHPUnit, wp doctor. You can run all of them from a command line and get JSON back today.
The difference is that WordPress's checks have holes in them, and an agent walks straight through.
Five things that pass
Here is an agent's pull request. PHPCS is green. PHPStan is green. Plugin Check is green. The tests pass.
MODEL ALONE
TICKET
Add a related-posts block to article pages
MODEL
interchangeable
PULL REQUEST
- PHPCS
- PHPStan
- Plugin Check
- Tests pass
route answers without login
settings written separately on every site
every page now heavier
block markup the editor can't open
MODEL IN A HARNESS
PHASE 1 · BEFORE IT WRITES
Change what the model is shown.
Read the site
WesperWhich blocks, tokens, plugins and settings exist, and in what order they load. The model starts from the site, not the repo.
ours, v1 · includes load order
Bound the ticket
contract“Add related posts” becomes: these files, article pages only, works on multisite, done means X. Out of scope is written down.
Hand it the slice
context packetNot the codebase. The hooks, options and templates this change touches, plus one similar past change.
Airbnb went from 75% to 97% by changing what the model was shown, not how it was asked.
MODEL
interchangeable
PHASE 2 · WHILE IT WRITES
Make the wrong write impossible.
Build the blocks
Block RunnerModel says which blocks go where. The tool writes the markup.
ours, benchmarked above
block markup the editor can't open · never written
Register the route
typed opA form, not free text. Permission callback is a required field. __return_true on a mutating route won’t submit.
the pattern, not a product · nothing off the shelf validates permission callbacks
route answers without login · field required
Add the setting
typed opSame. Scope (site or network) and autoload are required fields, with a size budget.
the pattern, not a product
settings forked per site · scope required
every page heavier · budget on the form
The mistake isn’t caught. It can’t be expressed.
PHASE 3 · BACKSTOP
Check the code
PHPStan · PHPCS
Run it for real
wp-env · multisite on
Look at it
Playwright
A person approves
It opened a route to the internet. The agent wrote a REST endpoint and set its permission_callback to __return_true, which means anyone who finds the URL can call it. Nothing checks this. Not PHPCS, not PHPStan, not Plugin Check. There is no linter in the WordPress ecosystem that validates permission callbacks, so a mutating endpoint with no authentication passes every gate you have.
It forked your network's settings. On multisite, get_option and get_site_option look almost identical and mean something entirely different. Use the wrong one and configuration meant to be network-wide gets written separately on every site, which nobody notices until the sites disagree.WordPress.org concedes this gap in its own tracker. Plugin Check, the official tool, "does not have any specific checks to validate a plugin's claim of Multisite compatibility." That issue has been open since July 2025.
It made every page heavier. The agent stored a value one page needs as an autoloaded option, so now it loads on every request across the site. wp doctor will tell you your autoload size if you ask it. Nothing fails a build over it. Reporting is not gating, and an agent shipping fifty changes a week will never be stopped by a number in a report nobody reads.
Those three are holes in checks you already run. The fourth is different, because no static check could ever catch it.
It depends on what else is installed. Two plugins hook the same action at the same priority. Which one wins depends on which loaded first, and load order comes from the active_plugins option, which WordPress sorts alphabetically when a plugin is activated. Your hook order is a function of folder names. That fact exists nowhere in your repository, and a coding agent reading your files cannot see it.
That is the real shape of the problem. The repository is not the site. Blocks, tokens, hooks, capabilities and option scope are assembled at runtime, by whatever is installed, in an order set by a sort function. An agent reading files is guessing about all of it.
“The repo is not the site. A harness reads the site.”
The fifth thing, and the only one we fixed
The four above are gaps in checking. The last one takes a different route: stop checking, and make the mistake impossible to express.
Converting HTML into Gutenberg blocks is a job models are bad at. Given a section of markup, five frontier models scored between 17 and 48 on our benchmark, and produced markup the Gutenberg validator rejected on between 25 and 41 of 63 fixtures. In the editor that is the recovery prompt your clients send you screenshots of.
So we stopped asking them to write markup. The model emits a typed description of which blocks go where, and a deterministic assembler builds them with createBlock(). The markup is valid because it cannot be anything else. The same five models score 97 to 99.
The number that says the most is the one with no model in it at all. Run the assembler on its own, with the intent step removed, and it scores 31, ahead of three of the five frontier models writing markup by hand.
One task, two mechanisms
63 fixtures · 24 layouts · 3 independent producers
every run at low reasoning effort
Making the wrong answer unrepresentable
That is the general move, and it is not ours.
Meta's team working on design systems tested five ways of stopping models generating off-system UI and concluded that typed component APIs beat linting the output afterwards. Same principle: a check tells you after the fact that something is wrong, a type stops it being written.
Prefer the type where you can. Where you cannot, gate it.
What to build first
Start with the environment, not the agent. In July 2025 an AI agent deleted a production database while explicitly instructed not to, because development and production were the same database and nothing prevented it. The fix was separation and permissions, not a better model.
Then wire the checks you already have so a machine has to pass them.
PHPCS / WPCS
What it catches: Standards, escaping, sanitisation
Machine-readable: --report=jsonCovers the holes?No
Plugin Check
What it catches: Plugin-directory requirements
Machine-readable: CLI, JSON claimedCovers the holes?No multisite
PHPUnitWhat it catches: Behaviour you wrote a test for
Machine-readable: JUnit XMLCovers the holes?Only what you wrote
wp doctor
What it catches: Site health, autoload size
Machine-readable: --format=jsonCovers the holes?Reports, never gates
wp profile
What it catches: Hook, DB and stage timing
Machine-readable: --format=jsonCovers the holes?Reports, never gates
PlaywrightWhat it catches: What the page actually renders
Machine-readable: JSON/JUnit reportersCovers the holes?Not WordPress-aware
Put them in CI rather than only in a git hook or an agent hook: hooks are bypassable, and Anthropic's own documentation notes that a hook which times out fails open rather than closed.
If the work is visual, let the agent look at it. The evidence here is real but narrower than the enthusiasm suggests: one controlled study moved success from 22% to 28% by giving a model screenshots and interaction, while another found the same kind of tool helped one model and made another worse. Measure it on your own work rather than assuming.
Then add what WordPress does not give you: something that reads the running install rather than the repository, and typed operations for whatever your agents get wrong most.
Ours are Wesper, which reads an install and reports what it actually accepts, and Block Runner, which is the tool in the chart above. Block Runner is the only one with a published number. Wesper is a v1 scaffold and we would rather tell you that than imply otherwise.
The queue moved, the decision did not
Every team above built machinery that throws out bad answers automatically. Not one of them let it approve anything.
Meta's tests survive three gates and engineers still reject a quarter. Stripe applies the same merge approval to agent and human work.
Google's 2025 developer survey found 90% adoption and 24% of developers trusting the output a great deal.
What the checks buy is not autonomy. It is that the wrong answers stop arriving at a person's desk, and the ones that do arrive are worth the attention. That is the whole trade, and in WordPress most of the machinery to make it has not been built.
Further reading
How we measured: 63 HTML fixtures across 24 layouts from three independent producers. Each model ran the corpus twice at low reasoning effort, writing markup directly and then emitting intent through Block Runner. Scores combine structural and content fidelity, halved when the Gutenberg validator rejects the markup.

