Your WordPress agent passes every check. That's the problem.

The teams shipping AI-written code at scale do not trust the model. They surround it with checks that throw out wrong answers before a person sees what is left. WordPress has the model and almost none of the checks.

Noel Tock

Noel Tock

3 September 2026 · 8 min read

In March 2025, Airbnb migrated around 3,500 React test files from one testing library to another. They had estimated a year and a half of engineering time. It took six developer-weeks.

The interesting part is not the model. It is the loop they built around it. Each file moved through a fixed sequence of checks: refactor it, then run Jest, then ESLint, then TypeScript. A file only advanced when the current check passed. When one failed, the model was handed the exact error and the current version of the file and asked again, usually about ten times, occasionally fifty. What lifted the success rate from 75% to 97% was not better prompting. It was feeding each retry more of the surrounding code.

Meta published the correction to any optimism you might be forming. Their test-generation system puts every candidate through three gates: it has to build, it has to pass five times running without flaking, and it has to actually improve coverage. Three quarters of what the model writes never survives. Of the quarter that does and reaches an engineer, engineers still reject about a quarter of those.

So the model writes four tests to get one worth reading.

Airbnb

test migration

  • Jest
  • ESLint
  • TypeScript

97%

after tuning, from 75%

Who decided

Engineers reviewed the output

Meta

unit tests

  • builds
  • passes 5x
  • improves coverage

25%

cleared all three gates

Who decided

Engineers still rejected 27% of those

Google

code completion

  • compile
  • reference
  • type check

80%

of Go suggestions would not compile

Who decided

Developer accepts or ignores

Stripe

pull requests

  • structured errors
  • agent-tagged keys
  • audit trail

1,000-1,300/wk

reach review

Who decided

Human merge approval, same as for people

Figma

security fixes

  • sandboxed execution
  • read-only validation

75.8%

catch rate on a bounty corpus

Who decided

Security engineers triage

Shopify

workflow steps

  • deterministic gates
  • AI steps

Not published

Who decided

Human PR review, mandatory

Cursor

production PRs

  • OS-level sandbox
  • video proof

35%

of merged internal PRs

Who decided

Human review

Uber

unit tests

  • runs tests
  • checks coverage

~11%

of all new tests in the codebase

Who decided

Engineers review the MR

every team runs a filter. No team lets the filter approve anything.
“The filter does not replace the reviewer. It decides what the reviewer sees.”

What they built is the boring part

Look at what is actually in those loops. A build step. A test runner. A linter. A type checker. A coverage report. Airbnb's pipeline is Jest, ESLint and TypeScript in a particular order with a retry between them. None of it is exotic, and none of it was built for AI. These are the checks those teams already had, wired so that a machine has to satisfy them before a human spends attention.

This is what a harness is. It decides what the agent can see, what it can touch, what it can run, and what it has to prove before a change is accepted. The model supplies judgement in the middle. Everything around it is code you own and version, and at the end of it a person still decides.

WordPress has the boring checks too. PHPCS with the WordPress standards, PHPStan, Plugin Check, PHPUnit, wp doctor. You can run all of them from a command line and get JSON back today.

The difference is that WordPress's checks have holes in them, and an agent walks straight through.

Five things that pass

Here is an agent's pull request. PHPCS is green. PHPStan is green. Plugin Check is green. The tests pass.

Same model. The difference is what it knew before it started and what it was allowed to write.

It opened a route to the internet. The agent wrote a REST endpoint and set its permission_callback to __return_true, which means anyone who finds the URL can call it. Nothing checks this. Not PHPCS, not PHPStan, not Plugin Check. There is no linter in the WordPress ecosystem that validates permission callbacks, so a mutating endpoint with no authentication passes every gate you have.

It forked your network's settings. On multisite, get_option and get_site_option look almost identical and mean something entirely different. Use the wrong one and configuration meant to be network-wide gets written separately on every site, which nobody notices until the sites disagree.WordPress.org concedes this gap in its own tracker. Plugin Check, the official tool, "does not have any specific checks to validate a plugin's claim of Multisite compatibility." That issue has been open since July 2025.

It made every page heavier. The agent stored a value one page needs as an autoloaded option, so now it loads on every request across the site. wp doctor will tell you your autoload size if you ask it. Nothing fails a build over it. Reporting is not gating, and an agent shipping fifty changes a week will never be stopped by a number in a report nobody reads.

Those three are holes in checks you already run. The fourth is different, because no static check could ever catch it.

It depends on what else is installed. Two plugins hook the same action at the same priority. Which one wins depends on which loaded first, and load order comes from the active_plugins option, which WordPress sorts alphabetically when a plugin is activated. Your hook order is a function of folder names. That fact exists nowhere in your repository, and a coding agent reading your files cannot see it.

That is the real shape of the problem. The repository is not the site. Blocks, tokens, hooks, capabilities and option scope are assembled at runtime, by whatever is installed, in an order set by a sort function. An agent reading files is guessing about all of it.

“The repo is not the site. A harness reads the site.”

The fifth thing, and the only one we fixed

The four above are gaps in checking. The last one takes a different route: stop checking, and make the mistake impossible to express.

Converting HTML into Gutenberg blocks is a job models are bad at. Given a section of markup, five frontier models scored between 17 and 48 on our benchmark, and produced markup the Gutenberg validator rejected on between 25 and 41 of 63 fixtures. In the editor that is the recovery prompt your clients send you screenshots of.

So we stopped asking them to write markup. The model emits a typed description of which blocks go where, and a deterministic assembler builds them with createBlock(). The markup is valid because it cannot be anything else. The same five models score 97 to 99.

The number that says the most is the one with no model in it at all. Run the assembler on its own, with the intent step removed, and it scores 31, ahead of three of the five frontier models writing markup by hand.

One task, two mechanisms

63 fixtures · 24 layouts · 3 independent producers

every run at low reasoning effort

Our tool, our benchmark, one narrow task: 63 HTML fixtures, five models, every run at low reasoning effort.

Making the wrong answer unrepresentable

That is the general move, and it is not ours. Meta's team working on design systems tested five ways of stopping models generating off-system UI and concluded that typed component APIs beat linting the output afterwards. Same principle: a check tells you after the fact that something is wrong, a type stops it being written.

Prefer the type where you can. Where you cannot, gate it.

What to build first

Start with the environment, not the agent. In July 2025 an AI agent deleted a production database while explicitly instructed not to, because development and production were the same database and nothing prevented it. The fix was separation and permissions, not a better model.

Then wire the checks you already have so a machine has to pass them.

  • PHPStan

    What it catches: Type and API errors

    Machine-readable: --error-format=json

    Covers the holes?No

  • PHPCS / WPCS

    What it catches: Standards, escaping, sanitisation

    Machine-readable: --report=json

    Covers the holes?No

  • Plugin Check

    What it catches: Plugin-directory requirements

    Machine-readable: CLI, JSON claimed

    Covers the holes?No multisite

  • PHPUnit

    What it catches: Behaviour you wrote a test for

    Machine-readable: JUnit XML

    Covers the holes?Only what you wrote

  • wp doctor

    What it catches: Site health, autoload size

    Machine-readable: --format=json

    Covers the holes?Reports, never gates

  • wp profile

    What it catches: Hook, DB and stage timing

    Machine-readable: --format=json

    Covers the holes?Reports, never gates

  • Playwright

    What it catches: What the page actually renders

    Machine-readable: JSON/JUnit reporters

    Covers the holes?Not WordPress-aware

All seven emit structured output today. None of them knows what your site assembles at runtime.

Put them in CI rather than only in a git hook or an agent hook: hooks are bypassable, and Anthropic's own documentation notes that a hook which times out fails open rather than closed.

If the work is visual, let the agent look at it. The evidence here is real but narrower than the enthusiasm suggests: one controlled study moved success from 22% to 28% by giving a model screenshots and interaction, while another found the same kind of tool helped one model and made another worse. Measure it on your own work rather than assuming.

Then add what WordPress does not give you: something that reads the running install rather than the repository, and typed operations for whatever your agents get wrong most.

Ours are Wesper, which reads an install and reports what it actually accepts, and Block Runner, which is the tool in the chart above. Block Runner is the only one with a published number. Wesper is a v1 scaffold and we would rather tell you that than imply otherwise.

The queue moved, the decision did not

Every team above built machinery that throws out bad answers automatically. Not one of them let it approve anything. Meta's tests survive three gates and engineers still reject a quarter. Stripe applies the same merge approval to agent and human work. Google's 2025 developer survey found 90% adoption and 24% of developers trusting the output a great deal.

What the checks buy is not autonomy. It is that the wrong answers stop arriving at a person's desk, and the ones that do arrive are worth the attention. That is the whole trade, and in WordPress most of the machinery to make it has not been built.

Further reading

How we measured: 63 HTML fixtures across 24 layouts from three independent producers. Each model ran the corpus twice at low reasoning effort, writing markup directly and then emitting intent through Block Runner. Scores combine structural and content fidelity, halved when the Gutenberg validator rejects the markup.