WhichBot / Blog / INTELLIGENCE
INTELLIGENCE · 2026-08-08

Inside the Delivery Scores: How We Grade Robot Companies, and What the First Scoreboard Reveals

Yesterday we launched the Robot Database — and with it, the feature we expect to get yelled at about: the Delivery Scores, a 0-100 grade for every robot company on one blunt question: do you ship what you promise?

Within the first day, the obvious questions arrived. Who are you to score Tesla? Why does a lawnmower company outrank the most famous robotics lab on Earth? Isn’t this unfair to startups doing hard things?

Fair questions, all. So this is the long answer: exactly how the scores work, why the scoreboard looks the way it does, and the five lessons buried in the first thirty-seven grades — some of which surprised even us.

The one thing we’re measuring

The Delivery Score does not measure how good a robot is. It doesn’t measure engineering brilliance, ambition, or how impressive the demo video looked. Our regular review scores handle machine quality; the methodology page covers that system.

The Delivery Score measures something narrower and, we’d argue, more useful to anyone about to spend money: the relationship between a company’s public promises and observable reality.

That’s it. When a company says “shipping in 2025,” did robots reach customers in 2025? When a price is announced, is it the price? When a preorder window is given, do the boxes arrive inside it? Every input to the score is a public statement the company chose to make, held against a public outcome anyone can check. We don’t grade secret roadmaps. We grade what was said out loud.

How the grades actually work

Each company on the Watch starts from its record of announcements — launch events, CES reveals, executive statements, preorder pages — and we track four things against them:

Did it ship at all? The heaviest weight. A machine that reached paying customers, even late, sits in a different universe from one that exists only on stages. Shipping late is a hardware sin; not shipping is a hardware identity.

How late, and how often? One slipped quarter is normal hardware turbulence. Timelines that move every time they’re restated — where the ship date stays perpetually 12-18 months away across years of restatements — are a pattern, and the score treats patterns as information.

Did the terms hold? Price at announcement versus price at checkout. Specs at the keynote versus specs in the box. Availability “for everyone this year” versus invite-only lists that never open. Quiet term-shifting erodes a score even when hardware eventually arrives.

What happens after shipping? Delivery isn’t the finish line. Machines need parts, updates, and a company that still exists. The scoreboard’s grimmest entry — Moxie, the child-companion robot that shipped, sold to families, and then bricked when its maker shut down — earns a 15 despite having delivered, because the promise of a robot includes the promise it keeps working.

The result is an editorial judgment, expressed as a number, receipts attached. We say that plainly: these are not algorithm outputs pretending to be objective. They’re graded the way a good analyst grades — from the public record, with the reasoning shown on every card so you can disagree with specifics instead of vibes. When companies ship, scores rise. When receipts correct us, we correct the board. That’s the whole governance model.

Lesson one: the boring companies are the honest ones

The top of the scoreboard is a parade of the unglamorous. Roborock at 95 — CES promise, spring shelf, every single year. Husqvarna at 94, which has been shipping robot lawnmowers since before most humanoid founders had driver’s licenses. Sony’s aibo, Whisker’s Litter-Robot, iRobot’s Roombas — categories the press stopped finding exciting a decade ago, and delivery records approaching perfection.

The pattern isn’t coincidence. Companies with real manufacturing and retail muscles make promises sized to those muscles. Their announcements are logistics statements, not fundraising instruments. It turns out the strongest predictor of a kept promise is a company that has kept a thousand before it — boring, in the best possible sense.

Lesson two: shipping is a choice some startups actually make

The cynical read on the humanoid category — all demos, no deliveries — dies against two entries. Unitree promised a $16,000 humanoid and then did the strangest thing in robotics: sold it. Order, pay, receive. That’s a 92, earned the only way. Agility Robotics’ Digit scores 90 for the least cinematic reason imaginable: its robots go to work in actual warehouses, and the company built a factory to make more of them.

These two matter because they delete the excuse. Humanoids are brutally hard — and shipping one is still possible, this year, at a real price. Every future “the technology just isn’t ready” claim on the scoreboard now has two counterexamples standing next to it.

Lesson three: a trillion dollars doesn’t ship a robot

The bottom of the board is where the household names live, and that’s the scoreboard’s most uncomfortable finding. Amazon’s Astro: announced in 2021, still effectively unbuyable five years later, its business version quietly killed — a 25, wearing the deepest pockets in retail. Samsung’s Ballie: announced at CES 2020, re-announced at CES 2024, release windows still walking — a 12, our current vaporware mascot. Tesla’s Optimus sits at 38: the demos are real, the internal pilots are real, and the customer deliveries remain forever next year, across multiple restatements.

The lesson isn’t that these companies can’t build robots. It’s that corporate resources don’t convert to delivery discipline — if anything, giants can afford to announce forever without shipping, because no single product’s failure threatens them. A startup that misses ship dates dies; a giant that misses them holds another keynote. The score reflects the behavior, not the balance sheet.

Lesson four: the middle of the board is where buyers get hurt

Vaporware at the bottom rarely takes anyone’s money — you can’t preorder what never opens orders. The genuinely dangerous zone is the 50-70 band: companies that do ship, eventually, after windows stretch and terms drift. Matic’s vacuum arrived — after extended preorder windows. Dyson’s robot arrived — years behind its own saga. Mammotion’s mowers graduated from crowdfunding chaos to real retail cadence.

This band is where “preorder now” buttons meet real credit cards, which makes it exactly where a Delivery Score earns its keep. Our working rule for readers: above 85, preorder with normal confidence; 60-85, preorder money you can wait on; below 60, treat every preorder as a donation with a possible robot attached.

Lesson five: the record only works if someone keeps it

The strangest part of building the scoreboard was discovering how hard the receipts were to assemble. Announcement coverage is everywhere; outcome coverage barely exists. Tech media writes ten thousand words when a robot is revealed and almost zero when its ship date quietly slides — the incentives all point at the keynote, none at the follow-through.

Which is the entire reason the Watch exists. Promises with no scoreboard aren’t accountability; they’re content marketing. From here, the record deepens weekly: score changes when reality reports in, new companies as they enter the arena, and every movement documented in the Physical AI Brief — the free weekly digest built from this database.

Argue with us — that’s the point

Every score on the board shows its reasoning, and every card is an invitation: if we’ve got a receipt wrong, show us the shipment and the score moves. What we won’t do is grade the sizzle. The robot era is arriving one kept-or-broken promise at a time, and somebody finally wrote them all down.

The scoreboard is live. The record only deepens from here: the Robot Database →

Keep reading

Explore: The Robot Database Robot Waitlist Tracker Best Robots 2026 Robot Match Quiz