October 10, 202611 min readMultimodel Chat TeamUpdated October 10, 2026

The Effective Price Index: $1.52–$215 per 1,000 Chats

Short answer: On our fixed one-message profile, GPT-6 Luna and Claude Haiku 5.5 tie for cheapest at $1.52 per 1,000 messages, DeepSeek V4.1-Flash follows at $1.87, and GPT-5.6 Cyber tops the index at $215.25. Cache reads cut every row by 13–22%, and output tokens are 54–78% of each bill, about two-thirds on most rows.

Thirteen models, one chat message: 8,000 input tokens, 2,000 out, 40% of the prompt cached. Two of them tie at $1.52 per 1,000 messages. The most expensive costs $215.25, which is 142 times the cheapest.

This is the first edition of a monthly index, not a model review. Same profile every month, every rate read from the vendor's own page, and the arithmetic printed so you can check it. Sticker prices per million tokens are the wrong unit for a chat product; nobody budgets in millions of tokens.

What one chat message costs across 13 models

The profile: 8,000 input tokens, 2,000 output tokens per message, 40% of input tokens served as cache reads, 5% paid as cache writes. That is the profile we use for every cost post, and it is close to how a normal chat session actually bills once a system prompt and a few turns are in context.

Model$ / 1,000 messages$ / 1M inputSource
GPT-6 Luna$1.52$0.10OpenAI pricing
Claude Haiku 5.5$1.52$0.10Anthropic pricing
DeepSeek V4.1-Flash (off-peak)$1.87$0.15DeepSeek pricing
Gemini 3.5 Flash-Lite$6.42$0.30Google pricing
Gemini 3.8 Flash$11.04$0.75Google pricing
Grok 4.7$22.40$2.00xAI models
GPT-6.1 Sol$30.12$2.00OpenAI pricing
Claude Sonnet 5.5$30.12$2.00Anthropic pricing
Claude Opus 5.5$60.24$4.00Anthropic pricing
GPT-5.6 Sol$60.88$4.00OpenAI pricing
Claude Fable 5.1$149.80$10.00Anthropic pricing
GPT-6 Astra$152.20$10.00OpenAI pricing
GPT-5.6 Cyber$215.25$12.50OpenAI pricing

Cost per 1,000 chat messages across 13 models at 8K input, 2K output and 40% cache reads, from $1.52 on GPT-6 Luna and Claude Haiku 5.5 to $215.25 on GPT-5.6 Cyber

One caveat before the analysis: every row is the vendor's published rate card, and Anthropic's current tokenizer emits about 30% more tokens for the same text. On identical content that moves the Haiku 5.5 row to roughly $1.98, off the $1.52 tie. We keep the table on published rates and flag the adjustment under What you give up.

Three things fall out of the table.

First, the mid-tier is a plateau, not a slope. GPT-6.1 Sol and Claude Sonnet 5.5 tie to the cent at $30.12, and Claude Opus 5.5 ($60.24) sits within about 1% of GPT-5.6 Sol ($60.88). Two years of price cuts have converged the cost of a "good enough" message to about three cents.

Second, the top of the index is a different product. GPT-5.6 Cyber is $215.25 per 1,000 messages because it was built for security work where a wrong answer costs more than the tokens. Almost nobody reading a BYOK cost index should be in that column, and the people who are already know.

Third, the cheap tier is not uniform. Gemini 3.5 Flash-Lite costs 4.2x what Luna costs per message. Google's Flash models were never the cheapest option; they were the cheapest option that also does video, grounding and live audio in the same API call. Different purchase.

What nobody tells you: cache reads move the index 13–22%

Every model here charges less for a cached read than for a fresh input token, and the discount differs by vendor. Anthropic's cache reads run at 10% of input on Haiku 5.5, 5% on Sonnet 5.5 and Opus 5.5, and 2.5% on Fable 5.1. OpenAI charges 10% on Astra and Luna but 5% on GPT-6.1 Sol. Google sits at 10%, xAI at 25%, and DeepSeek's automatic cache hit at 2% of its input rate.

Run the same profile with caching switched off and the index moves more than the model choice does at the cheap end:

ModelNo cache40% cached80% cached
Claude Haiku 5.5$1.80$1.52$1.23
GPT-6 Luna$1.80$1.52$1.23
DeepSeek V4.1-Flash$2.40$1.87$1.40
Claude Sonnet 5.5$36.00$30.12$24.04
GPT-6 Astra$180.00$152.20$123.40

That is a 13–22% swing on identical usage, purely from how you structure the prompt. On the cheap tier it is the whole difference between Haiku 5.5 and DeepSeek V4.1-Flash. If you take one operational thing from this index: structure prompts for cache hits before you shop for a cheaper model.

The second structural fact is where the money goes. Output tokens are 66% of the bill on most rows in the index, and 78% on Gemini 3.5 Flash-Lite, because every model charges input and output on different scales and your reply is the expensive half.

Share of each model's bill that comes from output tokens: 54% on Grok 4.7, 78% on Gemini 3.5 Flash-Lite, and 64-70% on the rest of the index

Practically: if your assistant writes 2,000-token answers to questions that need 200, you are paying ten times for the part nobody reads. Trim the reply length and the index collapses toward the input column, where every model is cheap.

The cliff map: where each rate changes

A dollar per million is only true inside a band. Here is where the bands end, verified today:

  • Claude Haiku 5.5 — over 100,000 prompt tokens, input goes $0.10 → $0.50 and output $0.50 → $2.50, a 5x cliff measured per request, and a cache hit inside that prompt does not exempt it.
  • OpenAI's flagship trio — short context up to 272,000 tokens, then 2x input and 1.5x output across the line.
  • Gemini 3.8 Flash and the 3.x Flash line — $0.75/$3.75 through December 31, 2026, then $1.50/$7.50 from January 1, 2027. That moves this index's Gemini 3.8 Flash row from $11.04 to $22.08 per 1,000 messages, and it is the largest scheduled change on the calendar.
  • DeepSeek — the $1.87 row in this index is the off-peak rate. Weekdays 01:00–04:00 and 06:00–10:00 UTC double every number, which puts DeepSeek at about $3.74 per 1,000 messages and above every Google row.
  • Gemini 4 Argon — excluded from this index. Its published $2/$10 is introductory, the standard rate is $4/$20 with no date attached, and it has not appeared on Google's public pricing page in four checks.

The pattern: cliffs cluster at 100K–272K input tokens and on January 1. If your workload sits still and your prompts stay short, this index is your bill. If your prompts grow, re-run the number before you re-architect.

Which model for which job

If you…Best row in the index
Chat, summarize, classify, routeLuna or Haiku 5.5 at $1.52
Run the same prompt against five modelsCheap tier for drafting, Sonnet 5.5 / 6.1 Sol for the final answer
Summarize 150K-token documentsLuna — 5x cheaper than Haiku 5.5 above 100K tokens
Do agentic coding with long tool loopsOpus 5.5 or Sonnet 5.5; the cheap tier pays for itself in retries
Need video, grounding and audio in one callGemini 3.8 Flash, and budget for January
Batch over weekendsDeepSeek V4.1-Flash off-peak, or any OpenAI Batch job at half price

The honest verdict: for ordinary chat, the cheap tier is now so close in price that the model choice barely matters, and the two decisions that actually move your bill are prompt caching and answer length. That is an unusual conclusion for a pricing index, and it is the one we keep arriving at.

What you give up

  • The profile is ours. Change the output length to 4,000 tokens and every row roughly doubles; change the input to 30,000 and the cheap tier separates from itself. The full arithmetic is reproducible: input × 4,400 uncached + 3,200 cache reads + 400 cache writes + 2,000 output, per million tokens.
  • Cache-write accounting is not fair across vendors. Google bills cache storage per hour rather than per write token, xAI charges nothing for writes, and DeepSeek's cache is automatic with no write charge. We scored those at zero, which nudges their rows down by a few percent.
  • Tokenizers differ. Anthropic's current tokenizer emits about 30% more tokens for the same text than its older one, and we have no cross-vendor count for arbitrary content. If that average holds for your text, add it to Haiku 5.5, Sonnet 5.5, Opus 5.5 and Fable 5.1.
  • We did not test output quality, latency or reliability. This is a bill, not a benchmark.
  • Catalog churn is real: OpenRouter listed 458 models on October 10, down from 465 three days earlier. Any "best model" table older than a week is stale.

FAQ

What is the cheapest AI model per chat message right now? GPT-6 Luna and Claude Haiku 5.5 tie at $1.52 per 1,000 messages on our profile. DeepSeek V4.1-Flash is $1.87 off-peak, and Gemini 3.5 Flash-Lite, the cheapest Google model, is $6.42.

How much do cache reads actually save? 13–22% on the same usage, depending on the model's cache-read discount. Most vendors charge 10% of the input rate for a cached read, Anthropic charges 2.5% on Fable 5.1, and DeepSeek's automatic cache hits bill at 2%.

Why is GPT-5.6 Cyber 142 times the cheapest model? Because it is a different product at $12.50 input and $75 output per million tokens, built for security work. Its price reflects a customer who would rather pay $215 per 1,000 messages than get an answer wrong.

Does this index hold if my prompts are long? Only until the first cliff. Above 100,000 prompt tokens Claude Haiku 5.5 multiplies by five, OpenAI's models double at 272,000, and Google's Flash line doubles for everyone on January 1, 2027.

Why is Gemini 4 Argon missing? Google has not published Argon's rates on its public pricing page in four checks since the September 30 announcement, and the $2/$10 figure is introductory against an undated $4/$20 standard rate. An index with unpublished rates in it is not an index.

If you want to run your own numbers instead of ours, the method above is reproducible with any vendor's rate card: 8,000 input tokens a message (4,400 uncached, 3,200 cache reads, 400 cache writes), 2,000 output, and the provider's published rates and cache discounts. The plan math behind the $2/month workspace fee, and when a subscription beats per-token billing, is in is BYOK cheaper than subscriptions. For the quality side of the same question, the October coding reprice prices 200 tasks instead of one message, and Gemini's pricing page is the source for Google's two clocks.


Start your free trial → — 7 days, all providers, no credit card required.

The workspace is $2/month or $39 once — AI providers always bill your key directly at their own rates.

See pricing
Share this story
Stay in the loop

The Multimodel Journal

Get the latest AI insights, model comparisons, and product updates delivered to your inbox.

Subscribe