Methodology

Everything here is designed so a sceptical engineer can check our work. If a statement below is not true of the running site, that is a bug — report it.

What "real time" means here

We do not run a ticker. Different classes of data change at genuinely different rates, and pretending otherwise would be dishonest:

  • API list prices change on vendor announcement — days to months apart. We show when a value was last verified, not a blinking light.
  • GPU marketplace rates really are volatile minute to minute. They are not ingested yet, so we publish none.
  • Measured throughput and TTFT vary continuously with load. We publish none until our own probe harness runs, because the ready-made source is licence-restricted for redisplay.
  • FX updates daily from ECB reference rates; the rate date is always shown.

An honest "verified 4 hours ago, source unchanged" beats a dishonest green dot.

Collection

One connector per source, each implementing fetch → parse → validate. Every value is stamped with source id, source URL, tier, method, confidence, and the fetch timestamp, and that stamp is what the provenance popover shows. Requests identify themselves with a real User-Agent and a contact URL, use timeouts, and are cached server-side for 30 minutes so we never hammer an upstream.

Normalisation

Sources publish per-token USD strings; we multiply by 1,000,000 and store USD per million tokens. Prices are canonical in USD and converted at render time with a labelled, dated FX rate. Non-token components (per request, per web-search call, per image) keep their own unit and are never silently mixed into a per-token figure.

Blended rate

Blended $/1M = input × r/(r+1) + output × 1/(r+1), where r is your input:output ratio, default 3:1. That default reflects a typical chat/RAG workload where prompts dominate; it is adjustable in the URL because no single ratio is right for everyone. A blended figure is a convenience for ranking, not a bill.

Conflicts

Sources are tiered. A higher tier never has its value overwritten by a lower tier. Where two sources disagree by more than 5%, both numbers are displayed with a warning marker in the cell and the disagreement is stated in the popover. We do not pick a winner silently.

Breakage handling

A connector that returns zero records, an unparsable payload, or a value above a sanity ceiling is quarantined: the source is marked failed on the index, and its values simply do not appear. Nothing is zero-filled, back-filled from memory, or estimated. Gaps render as "—".

A token is not a universal unit

The same paragraph is a different token count under o200k, Claude's tokenizer, Gemini, and the Llama/Qwen BPEs, and the gap widens for code, CJK, Cyrillic, Hindi and emoji. Any cross-family comparison on token price alone is approximate. The tokenizer page computes exact counts for the encodings we actually have and explicitly declines to estimate the ones we do not.

What this index does not know

  • Measured tokens/sec, TTFT and intelligence indices — the obvious source restricts redistribution, so we publish nothing rather than breach it.
  • Per-provider endpoint variance for the same open-weight model, which is often extreme. Currently one row per catalogue model.
  • Price history, change detection and a changelog — these need persistent storage, which is not enabled yet. Every value on the site is a point-in-time observation.
  • GPU/self-host economics and break-even curves — no compute SKU source is ingested yet.
  • Batch pricing, rate limits and free-tier terms — not published by ingested sources.

When we are uncertain whether a number is right, we leave it out and say so. That is the whole point.