Reference
Models
Everything below is servable today. Call https://api.dawnstackai.com/v1/models for the same list at runtime.
Maker
Price
Context
Can do
Licence
| Model | Maker | Context | Input | Output | Speed | Answers | Licence |
|---|---|---|---|---|---|---|---|
| llama-3.1-8b | | 131,072 | $0.03 | $0.06 | 652 ms · 16.8 tok/s 13 Sept | directly | Llama 3.1 Community License |
| llama-3.3-70b | | 131,072 | $0.15 | $0.48 | 916 ms · 11.6 tok/s 13 Sept | directly | Llama 3.3 Community License |
| mistral-small-3.1-24b | | 128,000 | $0.112 | $0.3 | 564 ms · 33.4 tok/s 13 Sept | directly | Apache 2.0 |
| gemma-4-26b | | 262,144 | $0.105 | $0.51 | 12,416 ms · 61.3 tok/s 13 Sept | after reasoning | Gemma 4 Terms of Use |
| gpt-oss-20b | | 131,072 | $0.045 | $0.21 | 481 ms · 136.5 tok/s 13 Sept | after reasoning | Apache 2.0 |
| gpt-oss-120b | | 131,072 | $0.055 | $0.255 | 552 ms · 64.2 tok/s 13 Sept | after reasoning | Apache 2.0 |
| kimi-k2.6 | | 262,144 | $1.125 | $5.25 | 8,205 ms · 39.2 tok/s 13 Sept | after reasoning | Modified MIT |
| deepseek-v3.2 | | 163,840 | $0.39 | $0.57 | 347 ms · 41.7 tok/s 13 Sept | directly | MIT |
| deepseek-v4-flash | | 1,048,576 | $0.135 | $0.27 | 6,883 ms · 38.8 tok/s 13 Sept | after reasoning | MIT |
| qwen3-235b | | 262,144 | $0.135 | $0.825 | 632 ms · 15.2 tok/s 13 Sept | directly | Apache 2.0 |
| qwen3-vl-30b | | 262,144 | $0.225 | $0.9 | 176 ms · 7 tok/s 13 Sept | directly | Apache 2.0 |
| mistral-nemo | | 131,072 | $0.028 | $0.045 | 687 ms · 27.5 tok/s 13 Sept | directly | Apache 2.0 |
| gemma-3-27b | | 131,072 | $0.12 | $0.24 | 236 ms · 29.1 tok/s 13 Sept | directly | Gemma Terms of Use |
| mistral-small-24b-2501 | | 32,768 | $0.075 | $0.12 | 1,075 ms · 73.1 tok/s 13 Sept | directly | Apache 2.0 |
| gemma-3-4b | | 131,072 | $0.075 | $0.15 | 327 ms · 40.7 tok/s 13 Sept | directly | Gemma Terms of Use |
| gemma-4-e4b | | 131,072 | $0.03 | $0.15 | 11,267 ms · 45.6 tok/s 13 Sept | after reasoning | Apache 2.0 |
| phi-4 | | 16,384 | $0.105 | $0.21 | 436 ms · 73.7 tok/s 13 Sept | directly | MIT |
| gemma-3-12b | | 131,072 | $0.075 | $0.225 | 116 ms · 98.9 tok/s 13 Sept | directly | Gemma Terms of Use |
| qwen3-14b | | 40,960 | $0.18 | $0.36 | 175 ms · 76.8 tok/s 13 Sept | after reasoning | Apache 2.0 |
| qwen3-32b | | 40,960 | $0.12 | $0.42 | 552 ms · 30.3 tok/s 13 Sept | after reasoning | Apache 2.0 |
| gemma-4-31b | | 262,144 | $0.195 | $0.57 | 10,742 ms · 32.4 tok/s 13 Sept | after reasoning | Apache 2.0 |
| hy3 | | 262,144 | $0.21 | $0.87 | 63,996 ms · 15.8 tok/s 13 Sept | after reasoning | Apache 2.0 |
| hermes-3-llama-3.1-70b | | 131,072 | $1.05 | $1.05 | 250 ms · 34 tok/s 13 Sept | directly | Llama 3 Community License |
| qwen3-vl-235b-a22b | | 262,144 | $0.3 | $1.32 | 277 ms · 18.5 tok/s 13 Sept | directly | Apache 2.0 |
| qwen3.6-35b-a3b | | 262,144 | $0.15 | $1.425 | 6,612 ms · 181.4 tok/s 13 Sept | after reasoning | Apache 2.0 |
| hermes-3-llama-3.1-405b | | 131,072 | $1.5 | $1.5 | 315 ms · 28 tok/s 13 Sept | directly | Llama 3 Community License |
| qwen3-next-80b-a3b | | 262,144 | $0.135 | $1.65 | 312 ms · 145.8 tok/s 13 Sept | directly | Apache 2.0 |
| step-3.7-flash | | 262,144 | $0.3 | $1.725 | 6,088 ms · 130.4 tok/s 13 Sept | after reasoning | Apache 2.0 |
| glm-4.7 | | 202,752 | $0.6 | $2.625 | 41,780 ms · 17.2 tok/s 13 Sept | after reasoning | MIT |
| mimo-v2.5 | | 262,144 | $0.6 | $3 | 1,794 ms · 22.8 tok/s 12 Sept | after reasoning | MIT |
| deepseek-r1-0528 | | 163,840 | $0.75 | $3.225 | 572 ms · 17 tok/s 12 Sept | after reasoning | MIT |
| glm-5.2 | | 1,048,576 | $1.125 | $3.6 | 11,604 ms · 62.1 tok/s 12 Sept | after reasoning | MIT |
| glm-5.3-flash | | 1,048,576 | $0.225 | $0.75 | 3,748 ms · 13 tok/s 12 Sept | directly | MIT |
| glm-5.3 | | 1,048,576 | $1.8 | $6 | 13,296 ms · 6.9 tok/s 12 Sept | after reasoning | GLM-5.3 License |
| deepseek-v4-pro | | 1,048,576 | $1.95 | $3.9 | 339 ms · 40.7 tok/s 12 Sept | directly | MIT |
| mimo-v2.5-pro | | 1,048,576 | $1.5 | $4.5 | 3,384 ms · 42.1 tok/s 12 Sept | after reasoning | MIT |
| qwen3.6-27b | | 262,144 | $0.48 | $4.8 | 45,836 ms · 83.6 tok/s 12 Sept | after reasoning | Apache 2.0 |
| glm-5.1 | | 202,752 | $1.575 | $5.25 | 16,811 ms · 39.2 tok/s 12 Sept | after reasoning | MIT |
| inkling | Thinking Machines | 524,288 | $1.425 | $6.075 | 1,343 ms · 104.7 tok/s 13 Sept | after reasoning | Apache 2.0 |
| No model matches every filter. Turn one off to widen the search. | |||||||
Prices are per million tokens, in USD. You are charged in the currency you pick when you top up, converted at the day's rate.
Speed is two numbers: how long until the first token arrives, and how fast tokens come after that. We measure both once a night from a Cloudflare edge, over the public internet, on the same prompt for every model, so these are slower than a vendor's own lab figures and closer to what you will see. Each row carries the day it was taken.
Speech
Transcription and synthesis, on OpenAI's own endpoint paths. Priced in their own units rather than converted into tokens, because that is how they are metered: you are billed for the seconds of audio you send, or the characters of text you send.
| Model | Maker | Does | Price | Endpoint | Licence |
|---|---|---|---|---|---|
| whisper-large-v3-turbo | | speech to text | $5 / M seconds | /v1/audio/transcriptions | MIT |
| qwen3-asr-1.7b | | speech to text | $11.25 / M seconds | /v1/audio/transcriptions | Apache 2.0 |
| kokoro-82m | Hexgrad | text to speech | $0.93 / M characters | /v1/audio/speech | Apache 2.0 |
Transcription is rounded up to the whole second, so a three-minute voice note costs about $0.0009. Repeated synthesis of the same text in the same voice is served from cache at a quarter of the price.
How well speech recognition really works
Character error rate on the FLEURS test split, measured 31 July 2026. 20% means one character in five is wrong. For reference, the same pipeline scores 8% on French. These numbers are not good, and we would rather you saw them than discovered them.
| Language | whisper-large-v3-turbo | qwen3-asr-1.7b |
|---|---|---|
| Amharic | 98% | 120% |
| Xhosa no language code | 92% | not measured |
| Northern Sotho no language code | 75% | not measured |
| Igbo no language code | 70% | 48% |
| Fula no language code | 67% | not measured |
| Luo no language code | 65% | not measured |
| Wolof no language code | 63% | not measured |
| Somali | 55% | not measured |
| Zulu no language code | 53% | 38% |
| Chichewa no language code | 52% | not measured |
| Shona | 51% | not measured |
| Umbundu no language code | 49% | not measured |
| Yoruba | 48% | 58% |
| Kamba no language code | 47% | not measured |
| Hausa | 43% | 47% |
| Luganda no language code | 33% | not measured |
| Lingala | 33% | not measured |
| Swahili | 20% | 15% |
| Afrikaans | 13% | not measured |
Eleven of these languages have no language code on any model we can serve. Whisper and Qwen3-ASR accept exactly the same eight African languages and reject exactly the same eleven, so you cannot tell either one it is listening to Igbo, and it will decide on its own that it is hearing English. Every African language in this table was auto-detected as something else.
Pick per language rather than by reputation: Qwen3-ASR is clearly better on Hausa, Igbo and Zulu, Whisper on Swahili. We do not route between them for you, because that would change your bill and your quality on a guess you never saw. Full method and the rest of the landscape in the write-up.
Cost by language
How many tokens the same sentence takes compared with English. 1.95x means a Swahili sentence costs almost twice as much as the identical English one.
| Model | Swahili | Hausa | Wolof | Zulu | Somali | Igbo | Yoruba | Amharic |
|---|---|---|---|---|---|---|---|---|
| llama-3.1-8b | 1.95x | 2.01x | 1.91x | 2.15x | 2.2x | 2.4x | 2.77x | 7.69x |
| llama-3.3-70b | 1.95x | 2.01x | 1.91x | 2.15x | 2.2x | 2.4x | 2.77x | 7.69x |
| mistral-small-3.1-24b | 1.85x | 1.95x | 1.9x | 2.08x | 2.18x | 2.41x | 2.87x | 7.6x |
| gemma-4-26b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| gpt-oss-20b | 1.47x | 1.58x | 1.75x | 1.67x | 1.72x | 1.65x | 2.2x | 5.85x |
| gpt-oss-120b | 1.48x | 1.58x | 1.74x | 1.67x | 1.72x | 1.66x | 2.21x | 5.82x |
| kimi-k2.6 | 1.91x | 1.98x | 1.95x | 2.16x | 2.2x | 2.43x | 2.66x | 4.27x |
| deepseek-v3.2 | 1.96x | 1.98x | 1.91x | 2.15x | 2.24x | 2.43x | 2.89x | 6.07x |
| deepseek-v4-flash | 1.96x | 1.98x | 1.91x | 2.15x | 2.24x | 2.43x | 2.89x | 6.07x |
| qwen3-235b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| qwen3-vl-30b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| mistral-nemo | 1.85x | 1.95x | 1.9x | 2.08x | 2.18x | 2.41x | 2.88x | 7.65x |
| gemma-3-27b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| mistral-small-24b-2501 | 1.85x | 1.95x | 1.9x | 2.08x | 2.18x | 2.41x | 2.87x | 7.6x |
| gemma-3-4b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| gemma-4-e4b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| phi-4 | 1.97x | 2.03x | 1.95x | 2.17x | 2.22x | 2.47x | 3x | 7.69x |
| gemma-3-12b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| qwen3-14b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| qwen3-32b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| gemma-4-31b | 1.65x | 1.76x | 1.79x | 1.95x | 1.95x | 2.17x | 2.45x | 1.94x |
| hy3 | 2.01x | 2.02x | 1.95x | 2.18x | 2.25x | 2.48x | 2.84x | 6.01x |
| hermes-3-llama-3.1-70b | 1.95x | 2.01x | 1.91x | 2.15x | 2.2x | 2.4x | 2.77x | 7.69x |
| qwen3-vl-235b-a22b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| qwen3.6-35b-a3b | not measured | not measured | not measured | not measured | not measured | not measured | not measured | not measured |
| hermes-3-llama-3.1-405b | 1.95x | 2.01x | 1.91x | 2.15x | 2.2x | 2.4x | 2.77x | 7.69x |
| qwen3-next-80b-a3b | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| step-3.7-flash | 1.96x | 1.98x | 1.91x | 2.15x | 2.24x | 2.43x | 2.89x | 6.07x |
| glm-4.7 | 1.95x | 2.01x | 1.93x | 2.15x | 2.2x | 2.41x | 2.78x | 7.69x |
| mimo-v2.5 | not measured | not measured | not measured | not measured | not measured | not measured | not measured | not measured |
| deepseek-r1-0528 | 1.96x | 1.98x | 1.91x | 2.15x | 2.24x | 2.43x | 2.89x | 6.07x |
| glm-5.2 | 1.95x | 1.98x | 1.92x | 2.15x | 2.2x | 2.41x | 2.77x | 7.69x |
| glm-5.3-flash | not measured | not measured | not measured | not measured | not measured | not measured | not measured | not measured |
| glm-5.3 | not measured | not measured | not measured | not measured | not measured | not measured | not measured | not measured |
| deepseek-v4-pro | 1.96x | 1.98x | 1.91x | 2.15x | 2.24x | 2.43x | 2.89x | 6.07x |
| mimo-v2.5-pro | 1.96x | 1.98x | 1.93x | 2.16x | 2.21x | 2.39x | 2.76x | 4.11x |
| qwen3.6-27b | 1.84x | 1.93x | 1.86x | 2.09x | 2.1x | 2.36x | 2.66x | 4.46x |
| glm-5.1 | 1.95x | 1.98x | 1.92x | 2.15x | 2.2x | 2.41x | 2.77x | 7.69x |
| inkling | 1.42x | 1.52x | 1.66x | 1.6x | 1.63x | 1.58x | 2.06x | 5.3x |
Measured on FLORES-200, 1012 professionally translated sentences aligned across languages, run against each model's own tokenizer. Reported as the mean per-sentence ratio. A model reading 'not measured' is one whose tokenizer we could not obtain. We would rather say so than publish a number we have not run.
A cheaper tokenizer is not always a cheaper request
Models marked "after reasoning" think before they answer, and that thinking is billed even though it never reaches you. The same short Swahili reply costs a fraction of a cent on llama-3.1-8b and several hundred on kimi-k2.6, because the reasoning dwarfs the answer.
So a better multiplier above pays for itself on long answers and not on short ones. For one-line replies, a direct model is usually cheaper despite the worse tokenizer. Both numbers are here so you can work out which applies to your workload.
Reasoning models also need room. Send a small max_tokens and the budget goes on thinking, leaving you an empty answer you paid for, so we raise the
ceiling for them. You are still only billed for the tokens actually produced.
Why this column exists
A price per million tokens assumes a token is worth the same in every language. It is not. The same sentence in Swahili or Amharic splits into more tokens than in English, so identical meaning costs more. A single flat rate hides that difference rather than removing it, and the rate above hides it too.
Some open models with the strongest African-language performance carry non-commercial licences. No provider can serve them to you, ourselves included, so they are absent from this list rather than swapped silently for a weaker model.
Every model here is open-weight
The weights are published, so you can download any model in this list, run it on your own hardware, and check the prices and language multipliers above against your own measurement. That is the difference between a number you are given and a number you can reproduce.