A pedestrian crossing a sunlit city street, framed by the Empiric Earth Delta

Perspective · Verifiable AI

The Study of Things That Did Not Happen

In the past few weeks, Evan Hubinger, Anthropic's Alignment Science Lead, put his personal odds that AI kills every human in the next decade at better than one in ten. Then more than 75,000 people, including hundreds of scientists, faith leaders and public figures signed a statement calling for a prohibition on developing superintelligence until there is broad scientific consensus that it can be done safely and controllably, and strong public buy-in. Days later, Dario Amodei made the case for slowing frontier development on purpose, so safety work can catch up.

I run a company building AI that operates where mistakes are physical. These debates keep me up, and I think they are legitimate. Physical AI data sets, and the implications for how these applications move safely around and interact with humans, is far more consequential than LLMs ever were. This is not a lab argument. Whatever standard gets set at the frontier is the standard the rest of us will inherit.

The question is legitimate. The answer needs proof.

That argument, once you set the timelines aside, is about proof. Not a company's account of its own work. Proof that somebody outside the company can check.

There is nothing philosophical about that. We expect it from a drug trial, a food inspection and a set of audited accounts. AI is the exception, and nobody has explained why.

Apply the same standard to a different kind of AI, the kind that already left the lab and is around us daily. There is no pause button on this one. It is already helping vehicles brake, warning drivers, making calls in places where a mistake has a physical cost. That deployment happened quietly, over a decade, while the rest of the argument was about models that do not exist yet. The question is not whether those systems are out there. It is how anyone knows they work.

We measured the road.

To know that, you need something to measure them against. Start with what we can actually count. We did not measure an autonomous vehicle. We measured the road.

From September 2025 through August 2026 we observed 1.105 billion miles of driving across 74.3 million rides in the United States. In that record: 95,805 collisions and 1,815,827 events that crossed our near-miss threshold. Almost nineteen near misses for every collision.

That is a particular set of drivers. Mostly professional, mostly in dense cities, with more than half the events in three states. It is not the whole country and I am not going to pretend it is.

Nothing in those numbers says an automated system is safe. They say what the road does, which is the part that has been missing. We still cannot tell you why most of those near misses happened. That is the next problem.

A busy city crossing at golden hour, motorcycles and cars moving through, with stopping-distance formulas overlaid as an illustrative model

A safety claim needs a baseline.

You cannot claim a system made a road safer unless you know what that road does without it. Every safety claim in this industry, ours included, is a comparison against a baseline, and for a hundred years that baseline has been missing most of what happens.

The collision is one event in twenty, and it is the one we have always been best equipped to record. It gets a police report, a claim, a file. The other nineteen disappear. Those are the ones that tell you where the risk sits before anybody gets hurt.

That is the evidence base my industry makes safety claims on, and it is why so many of them sound like faith.

Bar chart of share of reported events: near-miss events 94.99%, collisions 5.01%

The study of things that did not happen.

Safety is the study of things that did not happen. You cannot prove a system saved a life by pointing at the life. You have to count the events that used to end in a collision and no longer do.

That cuts at our claims too, so here is ours with the scope attached: safety models built on this record helped one national fleet cut its most serious collisions by 67%.

Bar chart of most serious collisions as an index: 100 before, 33 after

Build for verification. Protect the person.

Privacy cannot be a promise layered onto a dataset after the dataset exists. It has to decide how the record gets built in the first place. We designed ours to describe what happened on a road without needing to know who was there.

A record built to be verified and a record built to be turned on a person are two different designs. We picked one, and no part of our business depends on the other being possible.

And we should not be the ones grading it. Nearly every company building AI for the physical world tests its own work on its own data and publishes the result. The entire point of an audit is third-party validation. We kept the word and dropped the part that made it mean anything.

We build none of the systems we measure. We make no vehicles and we compete with none of our buyers. That is why a carmaker, a fleet, a city and an insurer can work off the same record without asking whose side it is on. That is not something we should get credit for saying. It should be visible in how the company is built and how the results get validated. If it is not, do not take my word for it.

The answer is verifiable AI.

Here’s what I’d suggest (and it applies to us too): If you sell a system that acts in the physical world, publish the baseline you measured it against, and let somebody outside your company check it.

The people worried about superintelligence and the people worried about the intersection two blocks from here are asking the same question in different clothes.

How would we know? The answer is not necessarily slower AI. It is verifiable AI.

I do not know what the frontier does in ten years. I know what happened across 1.105 billion miles of road last year, because we counted it.

That is what I have to offer the argument. Not reassurance. A method.