Methodology
v1 · effective July 2026Every score on this site is a number we can show our work for. Every workspace gets its own rating methodology, developed by the engine from that workspace's scoring configuration — the weights, normalization strategies, and data-quality policies below are read live from AI Model Rankings, and this page lets you reproduce any published score from the same data and math the ranking job uses. Worked examples use AI Model Rankings, the platform's reference workspace. See it applied: open the AI Model Rankings comparison.
- · Every metric value carries a source, a confidence level, and a verified date.
- · Scores are cohort-relative — normalized against the models published today, not a fixed scale.
- · The breakdown under every score is recomputed live from the same math on this page — never a cached explanation.
- · Changing the methodology gets a new version number and an effective date, not a silent edit.
Normalization
Each metric's raw value maps to a 0–100 score against the min and max of published models in the same cohort — not a hardcoded target. The mapping depends on the metric's normalization strategy:
- Linear
- Straight min–max scaling. A value at the cohort minimum scores 0, the maximum scores 100. Flipped when lower is better.
- Log-scaled
- Same min–max scaling, applied after a log(1+x) curve. Used for metrics that span orders of magnitude (context window, max output tokens) so a jump from 4K→128K tokens doesn't dwarf 128K→256K.
- Inverse
- A 1/x curve, min–max scaled. Used for price metrics — it already rewards cheap values, so the metric's "lower is better" flag isn't needed to flip it.
- Pass-through
- The raw value is clamped into 0–1 as-is, for metrics already expressed on that scale. Boolean metrics (does it support tool calling?) skip normalization entirely: true/false maps straight to 100/0, flipped when lower is better.
Weights
Each metric's normalized score is multiplied by a weight and summed. This table is read live from the workspace's scoring config — if a weight changes, this table changes with it, on the next request. In AI Model Rankings, benchmark scores and pricing carry the most weight because they are what the ranking claims to be about; capability booleans carry the least, because they gate what a model can do rather than how well it does it.
| Metric | Weight | Direction | Normalization |
|---|---|---|---|
| SWE-Bench Pro Scoreswe-bench-pro | 25% | Higher is better | Linear |
| Input Priceinput-price | 20% | Lower is better | Inverse (rewards low values) |
| Output Priceoutput-price | 15% | Lower is better | Inverse (rewards low values) |
| Context Windowcontext-window | 15% | Higher is better | Log-scaled |
| Tool Callingtool-calling | 8% | Higher is better | Pass-through |
| Extended Reasoningreasoning | 8% | Higher is better | Pass-through |
| Cache Read Pricecache-read-price | 5% | Lower is better | Inverse (rewards low values) |
| Max Output Tokensmax-output-tokens | 5% | Higher is better | Log-scaled |
Weights sum to 100%. This is the exact config compiled into the site right now — not a snapshot.
Missing data
Not every model has a verified value for every metric. One of three policies decides what happens — this vertical currently runs penalize:
- penalizeactive
- Missing metrics count as 0 against the full weight total. There's no way to improve a score by leaving a weak metric undocumented — an unverified value costs more than a disclosed bad one.
- exclude
- Missing metrics are dropped entirely and the remaining weights are renormalized to 100%.
- insufficient-data
- Behaves like
exclude, but returns no score at all when less than half of the total weight has data — better unranked than misleading.
Confidence adjustment
Every metric value is tagged with a confidence level when it's recorded. The base score is multiplied by the average of the confidence factors across every metric that had data:
| Confidence | Multiplier | Meaning |
|---|---|---|
| high | ×1.00 | Verified directly from the primary source. |
| medium | ×0.97 | Verified, minor uncertainty. |
| low | ×0.90 | Best available, not independently confirmed. |
| stale | ×0.85 | Was verified, but is past its freshness window. |
| conflicting | ×0.85 | Sources disagree. |
| unknown | ×0.80 | No confidence recorded. |
Enabled for this vertical: true.
Freshness adjustment
Data goes stale. Once the oldest verified metric value on an item passes 45 days, the score decays linearly — down to a 15% penalty once that data reaches 90 days old, where it bottoms out. One old metric drags down the whole item's score even if everything else was just verified — that's deliberate: it's the stalest data point that determines whether the score can be trusted.
In AI Model Rankings, prices are re-verified daily against three independent sources and models are canary-called with a real request — so a frontier model's score only decays if its lab's own published numbers stop being checkable.
Enabled for this vertical: true.
What we don't do
- — Affiliate and sponsored relationships never influence a score. The scoring pipeline above has no input for "paid," "sponsored," or "featured." This vertical doesn't run affiliate links at all.
- — No manual override on a published score. The only way to change a ranking is to change the underlying metric data, or the methodology itself — which gets a new version number.
- — Agent- and import-suggested data never publishes directly. It queues for human review first. Every public fact requires a cited source.
Reproduce this score
Pick any published model below. The breakdown is recomputed right now, from the live database, using the exact pipeline documented on this page — the same one the item page's score breakdown and the ranking job both use.
Claude Haiku 4.5 (latest)Anthropic
View full page →OptiSift Score breakdown
How this score was calculated — weighted metrics, normalized against the full dataset.
Every comparison, its own methodology
The engine develops a rating methodology with each workspace — same math, that workspace's own metrics, weights, and policies. Open any of them:
| Workspace | Items | |
|---|---|---|
| AI Model Rankings | 17 | Methodology·Comparison |
| Golf Cart Comparisons | 65 | Methodology·Comparison |
| Crypto Markets | 20 | Methodology·Comparison |
| Largest US Banks | 1 | Methodology·Comparison |
| Orbital Rockets | 19 | Methodology·Comparison |
| Top Federal Contractors | 18 | Methodology·Comparison |
| World Economies | 20 | Methodology·Comparison |
See the full rankings, or read more about data provenance and human review on the transparency page.