Reference

Models

Everything below is servable today. Call https://api.dawnstackai.com/v1/models for the same list at runtime.

Maker
Price
Context
Can do
Licence
Model Maker Context Input Output Speed Answers Licence
llama-3.1-8b Meta 131,072 $0.03 $0.06 652 ms · 16.8 tok/s 13 Sept directly Llama 3.1 Community License
llama-3.3-70b Meta 131,072 $0.15 $0.48 916 ms · 11.6 tok/s 13 Sept directly Llama 3.3 Community License
mistral-small-3.1-24b Mistral AI 128,000 $0.112 $0.3 564 ms · 33.4 tok/s 13 Sept directly Apache 2.0
gemma-4-26b Google 262,144 $0.105 $0.51 12,416 ms · 61.3 tok/s 13 Sept after reasoning Gemma 4 Terms of Use
gpt-oss-20b OpenAI 131,072 $0.045 $0.21 481 ms · 136.5 tok/s 13 Sept after reasoning Apache 2.0
gpt-oss-120b OpenAI 131,072 $0.055 $0.255 552 ms · 64.2 tok/s 13 Sept after reasoning Apache 2.0
kimi-k2.6 Moonshot AI 262,144 $1.125 $5.25 8,205 ms · 39.2 tok/s 13 Sept after reasoning Modified MIT
deepseek-v3.2 DeepSeek 163,840 $0.39 $0.57 347 ms · 41.7 tok/s 13 Sept directly MIT
deepseek-v4-flash DeepSeek 1,048,576 $0.135 $0.27 6,883 ms · 38.8 tok/s 13 Sept after reasoning MIT
qwen3-235b Qwen (Alibaba) 262,144 $0.135 $0.825 632 ms · 15.2 tok/s 13 Sept directly Apache 2.0
qwen3-vl-30b Qwen (Alibaba) 262,144 $0.225 $0.9 176 ms · 7 tok/s 13 Sept directly Apache 2.0
mistral-nemo Mistral AI 131,072 $0.028 $0.045 687 ms · 27.5 tok/s 13 Sept directly Apache 2.0
gemma-3-27b Google 131,072 $0.12 $0.24 236 ms · 29.1 tok/s 13 Sept directly Gemma Terms of Use
mistral-small-24b-2501 Mistral AI 32,768 $0.075 $0.12 1,075 ms · 73.1 tok/s 13 Sept directly Apache 2.0
gemma-3-4b Google 131,072 $0.075 $0.15 327 ms · 40.7 tok/s 13 Sept directly Gemma Terms of Use
gemma-4-e4b Google 131,072 $0.03 $0.15 11,267 ms · 45.6 tok/s 13 Sept after reasoning Apache 2.0
phi-4 Microsoft 16,384 $0.105 $0.21 436 ms · 73.7 tok/s 13 Sept directly MIT
gemma-3-12b Google 131,072 $0.075 $0.225 116 ms · 98.9 tok/s 13 Sept directly Gemma Terms of Use
qwen3-14b Qwen (Alibaba) 40,960 $0.18 $0.36 175 ms · 76.8 tok/s 13 Sept after reasoning Apache 2.0
qwen3-32b Qwen (Alibaba) 40,960 $0.12 $0.42 552 ms · 30.3 tok/s 13 Sept after reasoning Apache 2.0
gemma-4-31b Google 262,144 $0.195 $0.57 10,742 ms · 32.4 tok/s 13 Sept after reasoning Apache 2.0
hy3 Tencent 262,144 $0.21 $0.87 63,996 ms · 15.8 tok/s 13 Sept after reasoning Apache 2.0
hermes-3-llama-3.1-70b Nous Research 131,072 $1.05 $1.05 250 ms · 34 tok/s 13 Sept directly Llama 3 Community License
qwen3-vl-235b-a22b Qwen (Alibaba) 262,144 $0.3 $1.32 277 ms · 18.5 tok/s 13 Sept directly Apache 2.0
qwen3.6-35b-a3b Qwen (Alibaba) 262,144 $0.15 $1.425 6,612 ms · 181.4 tok/s 13 Sept after reasoning Apache 2.0
hermes-3-llama-3.1-405b Nous Research 131,072 $1.5 $1.5 315 ms · 28 tok/s 13 Sept directly Llama 3 Community License
qwen3-next-80b-a3b Qwen (Alibaba) 262,144 $0.135 $1.65 312 ms · 145.8 tok/s 13 Sept directly Apache 2.0
step-3.7-flash StepFun 262,144 $0.3 $1.725 6,088 ms · 130.4 tok/s 13 Sept after reasoning Apache 2.0
glm-4.7 Z.ai 202,752 $0.6 $2.625 41,780 ms · 17.2 tok/s 13 Sept after reasoning MIT
mimo-v2.5 Xiaomi 262,144 $0.6 $3 1,794 ms · 22.8 tok/s 12 Sept after reasoning MIT
deepseek-r1-0528 DeepSeek 163,840 $0.75 $3.225 572 ms · 17 tok/s 12 Sept after reasoning MIT
glm-5.2 Z.ai 1,048,576 $1.125 $3.6 11,604 ms · 62.1 tok/s 12 Sept after reasoning MIT
glm-5.3-flash Z.ai 1,048,576 $0.225 $0.75 3,748 ms · 13 tok/s 12 Sept directly MIT
glm-5.3 Z.ai 1,048,576 $1.8 $6 13,296 ms · 6.9 tok/s 12 Sept after reasoning GLM-5.3 License
deepseek-v4-pro DeepSeek 1,048,576 $1.95 $3.9 339 ms · 40.7 tok/s 12 Sept directly MIT
mimo-v2.5-pro Xiaomi 1,048,576 $1.5 $4.5 3,384 ms · 42.1 tok/s 12 Sept after reasoning MIT
qwen3.6-27b Qwen (Alibaba) 262,144 $0.48 $4.8 45,836 ms · 83.6 tok/s 12 Sept after reasoning Apache 2.0
glm-5.1 Z.ai 202,752 $1.575 $5.25 16,811 ms · 39.2 tok/s 12 Sept after reasoning MIT
inkling Thinking Machines 524,288 $1.425 $6.075 1,343 ms · 104.7 tok/s 13 Sept after reasoning Apache 2.0

Prices are per million tokens, in USD. You are charged in the currency you pick when you top up, converted at the day's rate.

Speed is two numbers: how long until the first token arrives, and how fast tokens come after that. We measure both once a night from a Cloudflare edge, over the public internet, on the same prompt for every model, so these are slower than a vendor's own lab figures and closer to what you will see. Each row carries the day it was taken.

Speech

Transcription and synthesis, on OpenAI's own endpoint paths. Priced in their own units rather than converted into tokens, because that is how they are metered: you are billed for the seconds of audio you send, or the characters of text you send.

Model Maker Does Price Endpoint Licence
whisper-large-v3-turbo OpenAI speech to text $5 / M seconds /v1/audio/transcriptions MIT
qwen3-asr-1.7b Qwen (Alibaba) speech to text $11.25 / M seconds /v1/audio/transcriptions Apache 2.0
kokoro-82m Hexgrad text to speech $0.93 / M characters /v1/audio/speech Apache 2.0

Transcription is rounded up to the whole second, so a three-minute voice note costs about $0.0009. Repeated synthesis of the same text in the same voice is served from cache at a quarter of the price.

How well speech recognition really works

Character error rate on the FLEURS test split, measured 31 July 2026. 20% means one character in five is wrong. For reference, the same pipeline scores 8% on French. These numbers are not good, and we would rather you saw them than discovered them.

Language whisper-large-v3-turboqwen3-asr-1.7b
Amharic 98% 120%
Xhosa no language code 92% not measured
Northern Sotho no language code 75% not measured
Igbo no language code 70% 48%
Fula no language code 67% not measured
Luo no language code 65% not measured
Wolof no language code 63% not measured
Somali 55% not measured
Zulu no language code 53% 38%
Chichewa no language code 52% not measured
Shona 51% not measured
Umbundu no language code 49% not measured
Yoruba 48% 58%
Kamba no language code 47% not measured
Hausa 43% 47%
Luganda no language code 33% not measured
Lingala 33% not measured
Swahili 20% 15%
Afrikaans 13% not measured

Eleven of these languages have no language code on any model we can serve. Whisper and Qwen3-ASR accept exactly the same eight African languages and reject exactly the same eleven, so you cannot tell either one it is listening to Igbo, and it will decide on its own that it is hearing English. Every African language in this table was auto-detected as something else.

Pick per language rather than by reputation: Qwen3-ASR is clearly better on Hausa, Igbo and Zulu, Whisper on Swahili. We do not route between them for you, because that would change your bill and your quality on a guess you never saw. Full method and the rest of the landscape in the write-up.

Cost by language

How many tokens the same sentence takes compared with English. 1.95x means a Swahili sentence costs almost twice as much as the identical English one.

Model SwahiliHausaWolofZuluSomaliIgboYorubaAmharic
llama-3.1-8b 1.95x 2.01x 1.91x 2.15x 2.2x 2.4x 2.77x 7.69x
llama-3.3-70b 1.95x 2.01x 1.91x 2.15x 2.2x 2.4x 2.77x 7.69x
mistral-small-3.1-24b 1.85x 1.95x 1.9x 2.08x 2.18x 2.41x 2.87x 7.6x
gemma-4-26b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
gpt-oss-20b 1.47x 1.58x 1.75x 1.67x 1.72x 1.65x 2.2x 5.85x
gpt-oss-120b 1.48x 1.58x 1.74x 1.67x 1.72x 1.66x 2.21x 5.82x
kimi-k2.6 1.91x 1.98x 1.95x 2.16x 2.2x 2.43x 2.66x 4.27x
deepseek-v3.2 1.96x 1.98x 1.91x 2.15x 2.24x 2.43x 2.89x 6.07x
deepseek-v4-flash 1.96x 1.98x 1.91x 2.15x 2.24x 2.43x 2.89x 6.07x
qwen3-235b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
qwen3-vl-30b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
mistral-nemo 1.85x 1.95x 1.9x 2.08x 2.18x 2.41x 2.88x 7.65x
gemma-3-27b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
mistral-small-24b-2501 1.85x 1.95x 1.9x 2.08x 2.18x 2.41x 2.87x 7.6x
gemma-3-4b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
gemma-4-e4b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
phi-4 1.97x 2.03x 1.95x 2.17x 2.22x 2.47x 3x 7.69x
gemma-3-12b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
qwen3-14b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
qwen3-32b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
gemma-4-31b 1.65x 1.76x 1.79x 1.95x 1.95x 2.17x 2.45x 1.94x
hy3 2.01x 2.02x 1.95x 2.18x 2.25x 2.48x 2.84x 6.01x
hermes-3-llama-3.1-70b 1.95x 2.01x 1.91x 2.15x 2.2x 2.4x 2.77x 7.69x
qwen3-vl-235b-a22b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
qwen3.6-35b-a3b not measured not measured not measured not measured not measured not measured not measured not measured
hermes-3-llama-3.1-405b 1.95x 2.01x 1.91x 2.15x 2.2x 2.4x 2.77x 7.69x
qwen3-next-80b-a3b 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
step-3.7-flash 1.96x 1.98x 1.91x 2.15x 2.24x 2.43x 2.89x 6.07x
glm-4.7 1.95x 2.01x 1.93x 2.15x 2.2x 2.41x 2.78x 7.69x
mimo-v2.5 not measured not measured not measured not measured not measured not measured not measured not measured
deepseek-r1-0528 1.96x 1.98x 1.91x 2.15x 2.24x 2.43x 2.89x 6.07x
glm-5.2 1.95x 1.98x 1.92x 2.15x 2.2x 2.41x 2.77x 7.69x
glm-5.3-flash not measured not measured not measured not measured not measured not measured not measured not measured
glm-5.3 not measured not measured not measured not measured not measured not measured not measured not measured
deepseek-v4-pro 1.96x 1.98x 1.91x 2.15x 2.24x 2.43x 2.89x 6.07x
mimo-v2.5-pro 1.96x 1.98x 1.93x 2.16x 2.21x 2.39x 2.76x 4.11x
qwen3.6-27b 1.84x 1.93x 1.86x 2.09x 2.1x 2.36x 2.66x 4.46x
glm-5.1 1.95x 1.98x 1.92x 2.15x 2.2x 2.41x 2.77x 7.69x
inkling 1.42x 1.52x 1.66x 1.6x 1.63x 1.58x 2.06x 5.3x

Measured on FLORES-200, 1012 professionally translated sentences aligned across languages, run against each model's own tokenizer. Reported as the mean per-sentence ratio. A model reading 'not measured' is one whose tokenizer we could not obtain. We would rather say so than publish a number we have not run.

A cheaper tokenizer is not always a cheaper request

Models marked "after reasoning" think before they answer, and that thinking is billed even though it never reaches you. The same short Swahili reply costs a fraction of a cent on llama-3.1-8b and several hundred on kimi-k2.6, because the reasoning dwarfs the answer.

So a better multiplier above pays for itself on long answers and not on short ones. For one-line replies, a direct model is usually cheaper despite the worse tokenizer. Both numbers are here so you can work out which applies to your workload.

Reasoning models also need room. Send a small max_tokens and the budget goes on thinking, leaving you an empty answer you paid for, so we raise the ceiling for them. You are still only billed for the tokens actually produced.

Why this column exists

A price per million tokens assumes a token is worth the same in every language. It is not. The same sentence in Swahili or Amharic splits into more tokens than in English, so identical meaning costs more. A single flat rate hides that difference rather than removing it, and the rate above hides it too.

Some open models with the strongest African-language performance carry non-commercial licences. No provider can serve them to you, ourselves included, so they are absent from this list rather than swapped silently for a weaker model.

Every model here is open-weight

The weights are published, so you can download any model in this list, run it on your own hardware, and check the prices and language multipliers above against your own measurement. That is the difference between a number you are given and a number you can reproduce.