Price per million tokens does not predict your bill. We built the same seven jobs on seven models and the cheapest tier won four of them outright — scraping and spreadsheet work cost the same on every model tested, a 36× price difference for an identical result. What actually decides the bill is turns to done, not the rate card. Pick per job, not per account.
Everyone compares models on price per million tokens. We built the same seven jobs on seven different models and found that number predicts the bill badly — because what you actually pay for is turns, and the models differ enormously in how many they need.
How we measured it
Each job has a set of acceptance checks that run in a real browser: does it fit a phone screen, are there script errors, is the arithmetic right, is the technical SEO complete. The agent works, the checks run, and whatever failed goes back to it as the next instruction — phrased the way a person would complain, not the way a test runner would. That repeats until it passes or gives up.
Every model gets the identical prompt and the identical form of feedback. The only thing that differs between two runs is what the models produced. And the price under each result is what our own account was charged, our markup included, because it went through the same code path a customer uses.
Scraping and data: everything ties
Scraping a few hundred rows into a CSV, and cleaning a messy spreadsheet, cost the same in effort on every model tested — all of them finished on the first turn. The cheapest did it for $0.0029. The most expensive did it for $0.27.
That is a 36× price difference for an identical outcome. If your job is fetching, parsing and looping, paying frontier rates buys nothing you can measure.
Websites: the cheap model won anyway
We expected the opposite here, and did not get it. The cheapest tier passed every page build on the first turn for $0.0072 to $0.01 each, and when we put the screenshots side by side the work stood up next to a model costing eighty times more.
Two models did fail several page builds — but not by producing something ugly. They wrote perfectly good HTML into a subdirectory, so the preview served a directory listing and a person would reasonably conclude nothing had been built. That is a file-placement problem, not a quality one, and we have since made the instructions explicit about it.
Where paying more finally earns it
One job type broke the pattern: an app that has to save something. A booking system that actually stores bookings was built correctly by exactly one model, in three turns, for $1.24. Both cheap models produced something that looked right and lost the data — one returned a server error on every booking, the other forgot them on reload.
That is fifty times the price of a page build, and the first job type in the set where it is worth paying.
The number that actually matters
One model costs a third as much per turn as our recommendation and still loses, because it needs two turns where the other needs one. Per-turn price is what everybody compares. Turns-to-done is what decides the bill.
Which is why the app suggests a model per kind of job rather than picking one for everything — and tells you why, with the measurement behind it.
See all 53 builds What it costs
Get the next one.
Working notes on building with AI, with the measurements and the bills attached. Sent when there is something worth sending, which is not weekly.
No spam, and one click to leave. We do not sell or share the list.
Pick a model, or let it pick one.
Every model shows its price before you spend anything, and Amelia suggests one per job.
Free to start · no card · every answer shows its price