Suggested per job
Type what you want and it names a model before you start, with the reason and the measured cost. Nobody should have to know a model catalogue to build a website.
Models
The answer to "which model is best" is not one model. We ran seven of them across seven kinds of job — 53 builds — and the useful finding was where paying more buys nothing at all.
Prices here run from $0.18 to $50 per million tokens — a 278× spread. Scraping and data work came out the same on all of them. Pages did not.
The catalogue
Per million tokens, at the rate our supplier bills us. You pay that plus 20–30% depending on your plan, and the figure for your turn is on your answer.
In full
| Model | Tier | In / 1M | Out / 1M | What it is for |
|---|---|---|---|---|
| DeepSeek V4 Flash | Budget | $0.09 | $0.18 | The cheapest here by a distance. Strong on shell work and extraction. |
| MiMo V2.5 | Budget | $0.14 | $0.28 | Very cheap, million-token context. Good for chewing through a lot of input. |
| GPT-5.6 Luna | Budget | $0.20 | $1.20 | Cheapest per finished build. Good at straightforward pages and small fixes. |
| MiniMax M3 | Budget | $0.30 | $1.20 | Fast and cheap. Fine for a landing page, thin on harder logic. |
| Gemini 3.7 Flash | Budget | $0.38 | $1.88 | Google's fast tier. Reliable at summarising and reshaping text. |
| GLM 5.3 | Mid | $1.40 | $4.40 | Built for coding agents. Best value once a build has real logic in it. |
| Grok 4.6 | Mid | $2.00 | $6.00 | Strong general coder, quick turns. |
| GPT-5.6 Sol | Mid | $2.00 | $10.00 | OpenAI's flagship. Reliable across multi-step work. |
| Claude Opus 5 | Frontier | $5.00 | $25.00 | The best output here, and roughly twenty times the cheap tier. |
| Claude Fable 5 | Frontier | $10.00 | $50.00 | For long, ambiguous jobs that would otherwise take a person days. |
What we found
That is the finding worth having. Data work and extraction passed first time on everything we tested, so frontier rates there buy nothing but a bigger bill. Design-led pages are the opposite — that is where a better model earns its keep, and even then the gap was far smaller than the price gap.
Each one taken to a working state in a single turn.
In full
| Job | Suggested tier | Reason |
|---|---|---|
| A website | Budget tier | Passed all four page builds first time, for $0.0072 to $0.01 each. |
| A tool or app | Budget tier | Built a working calculator with correct arithmetic in one turn. |
| Scraping | Budget tier | One turn, for $0.0075 — every model tested passed this first time. |
| Spreadsheets and data | Budget tier | Deduped, reformatted and totalled correctly for $0.0035. |
| Fixing something broken | Not measured yet | No build behind it, so no recommendation. |
| Reviewing what you have | Not measured yet | No build behind it, so no recommendation. |
Choosing
The ten above are the ones we ran 53 builds through, so they are the ones we will make a claim about. They are not the limit: the picker in the app lists the whole catalogue — over 300 models from around 60 providers, a couple of dozen of which cost nothing per token — and the agent will run any of them. What it will not do is pretend we have measured one when we have not.
Type what you want and it names a model before you start, with the reason and the measured cost. Nobody should have to know a model catalogue to build a website.
Pick a different model per turn. The price updates before you send, not after you are billed.
A small allowance pointed at a $10/$50 model buys about three turns, and running out mid-build is the worst place to stop. Cheap models make the trial feel generous.
Every model, every job, every screenshot, and what each one really cost.
Questions
For a website, a tool, scraping or spreadsheet work the measured answer is one of the cheap models — all four of those jobs finished in one turn for $0.0035 to $0.01. Paying frontier rates bought nothing measurable on data work.
No. It suggests one for the job before you start and tells you why. You can change it per turn, and the price updates before you send.
Sometimes, and by far less than the price gap. Across the same seven jobs five of seven models produced work that stands comparison, at prices ranging over 270x.