Cost calculator
Costs are computed only from prices that were actually fetched from a source. Where a component (batch rates, for example) is not published by any ingested source, it is excluded and named below rather than guessed.
Workload
Ranked candidates
| Endpoint | $/1M in | $/1M out | Cost / request | Monthly | Annual |
|---|---|---|---|---|---|
| DeepSeek: DeepSeek V4 Flash Latest | $0.0048 | $0.071 | $0.00005 | $45.64 | $547.70 |
| Mistral: Mistral Nemo | $0.029 | $0.030 | $0.00007 | $71.50 | $858.02 |
| inclusionAI: Ling 3.0 Flash VL | $0.021 | $0.062 | $0.00007 | $72.11 | $865.37 |
| inclusionAI: Ling 3.0 Flash | $0.021 | $0.063 | $0.00007 | $72.83 | $873.94 |
| OpenAI: gpt-oss-20b | $0.018 | $0.090 | $0.00008 | $80.78 | $969.41 |
| IBM: Granite 4.0 Micro | $0.017 | $0.112 | $0.00009 | $90.07 | $1,081 |
| Nex AGI: Nex-N2.5-Mini | $0.025 | $0.100 | $0.00010 | $99.45 | $1,193 |
| Sao10K: Llama 3 8B Lunaris | $0.040 | $0.050 | $0.00010 | $103.02 | $1,236 |
| OpenAI: gpt-oss-20b (batch) | $0.024 | $0.112 | $0.00010 | $103.63 | $1,244 |
| Qwen: Qwen3.7 Flash | $0.030 | $0.130 | $0.00012 | $124.44 | $1,493 |
| OpenAI: gpt-oss-120b (batch) | $0.030 | $0.136 | $0.00013 | $126.72 | $1,521 |
| Inference.net: Schematron V2 Turbo | $0.030 | $0.150 | $0.00013 | $134.64 | $1,616 |
Spend decomposition
Cheapest candidate: DeepSeek: DeepSeek V4 Flash Latest · $45.64/month
| Fresh input | $9.12 |
|---|---|
| Output | $35.63 |
| Retries | $0.89 |
Sensitivity
Percentage change in total monthly cost when one assumption moves.
| Reasoning ×2 | +79.6% |
|---|---|
| Output length ±50% | +39.8% |
| Retry rate +10pt | +9.8% |
| Input length ±50% | +8.0% |
| Cache hit +30pt | +0.0% |
Excluded from this calculation
- Batch pricing: the ingested catalogue does not publish batch rates, so no batch discount is applied. Enabling a batch share would require inventing a number.
- Cache write costs: applied only when the source publishes them; otherwise a cache hit is priced at the published cached-read rate and a miss at the input rate.
- Token counts are your estimates. The same text tokenises differently per model family — use the tokenizer tool before trusting a cross-model comparison.