Sign in and build · Free to start, no card needed · add a card when you want more credit · see pricing →

Models

Over 300 models. Seven of them measured.

The answer to "which model is best" is not one model. We ran seven of them across seven kinds of job — 53 builds — and the useful finding was where paying more buys nothing at all.

Prices here run from $0.18 to $50 per million tokens — a 278× spread. Scraping and data work came out the same on all of them. Pages did not.

ameliasagent.com/app
The model picker, showing over 300 models each with its published price.
300+models in the catalogue
10measured across 53 builds
278×price spread, cheapest to priciest

The catalogue

What each one costs to run.

Per million tokens, at the rate our supplier bills us. You pay that plus 20–30% depending on your plan, and the figure for your turn is on your answer.

Budget Budget Budget · Budget Budget Mid Mid Mid Frontier Frontier
$0.18 per 1M out$50.00 per 1M out
Output price per million tokens for the ten we measured, on a logarithmic rule — drawn linearly, the first eight would sit on top of each other. The large point is the one suggested for most jobs, and it is at the cheap end, which is the whole finding.

In full

Every model we measured, side by side.

The ten we measured, cheapest first
ModelTierIn / 1MOut / 1MWhat it is for
DeepSeek V4 FlashBudget$0.09$0.18The cheapest here by a distance. Strong on shell work and extraction.
MiMo V2.5Budget$0.14$0.28Very cheap, million-token context. Good for chewing through a lot of input.
GPT-5.6 LunaBudget$0.20$1.20Cheapest per finished build. Good at straightforward pages and small fixes.
MiniMax M3Budget$0.30$1.20Fast and cheap. Fine for a landing page, thin on harder logic.
Gemini 3.7 FlashBudget$0.38$1.88Google's fast tier. Reliable at summarising and reshaping text.
GLM 5.3Mid$1.40$4.40Built for coding agents. Best value once a build has real logic in it.
Grok 4.6Mid$2.00$6.00Strong general coder, quick turns.
GPT-5.6 SolMid$2.00$10.00OpenAI's flagship. Reliable across multi-step work.
Claude Opus 5Frontier$5.00$25.00The best output here, and roughly twenty times the cheap tier.
Claude Fable 5Frontier$10.00$50.00For long, ambiguous jobs that would otherwise take a person days.
The two frontier models need a paid plan. Everything else is available on the free allowance.

What we found

Scraping costs the same on every model.

That is the finding worth having. Data work and extraction passed first time on everything we tested, so frontier rates there buy nothing but a bigger bill. Design-led pages are the opposite — that is where a better model earns its keep, and even then the gap was far smaller than the price gap.

Cleaning data $0.0035
A calculator $0.0072
Scraping $0.0075
A portfolio $0.0082
A service site $0.0092
A SaaS page $0.0096
An agency site $0.0100

Each one taken to a working state in a single turn.

What a finished job cost on the recommended model, in dollars. Every figure is what our own account was charged, markup included — nothing here is estimated.

In full

The exact model, before you spend.

What the app suggests, and why
JobSuggested tierReason
A websiteBudget tierPassed all four page builds first time, for $0.0072 to $0.01 each.
A tool or appBudget tierBuilt a working calculator with correct arithmetic in one turn.
ScrapingBudget tierOne turn, for $0.0075 — every model tested passed this first time.
Spreadsheets and dataBudget tierDeduped, reformatted and totalled correctly for $0.0035.
Fixing something brokenNot measured yetNo build behind it, so no recommendation.
Reviewing what you haveNot measured yetNo build behind it, so no recommendation.
The app names the exact model, and why, before you spend anything — this table stops at the tier. Two rows say "not measured yet" on purpose: a recommendation with no build behind it is an opinion with a price tag attached.

Choosing

It picks. You can override.

The ten above are the ones we ran 53 builds through, so they are the ones we will make a claim about. They are not the limit: the picker in the app lists the whole catalogue — over 300 models from around 60 providers, a couple of dozen of which cost nothing per token — and the agent will run any of them. What it will not do is pretend we have measured one when we have not.

Suggested per job

Type what you want and it names a model before you start, with the reason and the measured cost. Nobody should have to know a model catalogue to build a website.

Change it any time

Pick a different model per turn. The price updates before you send, not after you are billed.

Why cheap models on a free trial

A small allowance pointed at a $10/$50 model buys about three turns, and running out mid-build is the worst place to stop. Cheap models make the trial feel generous.

See all 53 builds.

Every model, every job, every screenshot, and what each one really cost.

Questions

Asked, answered

Which AI model should I use?

For a website, a tool, scraping or spreadsheet work the measured answer is one of the cheap models — all four of those jobs finished in one turn for $0.0035 to $0.01. Paying frontier rates bought nothing measurable on data work.

Do I have to choose a model?

No. It suggests one for the job before you start and tells you why. You can change it per turn, and the price updates before you send.

Is a more expensive model better?

Sometimes, and by far less than the price gap. Across the same seven jobs five of seven models produced work that stands comparison, at prices ranging over 270x.