Tokens per Wh (efficiency)
Measured inference efficiency: output tokens generated per watt-hour of GPU energy, from the ML.ENERGY Leaderboard (minimum-energy serving configuration; hardware and workload in each value's citation). Higher is better. Closed API providers do not publish per-model energy — a missing value means no credible public measurement exists, not zero.
Measured in tokens/Wh · Higher is better · Weight in overall score: 0
Top models by Tokens per Wh (efficiency)
| # | Name | Value | Confidence |
|---|---|---|---|
| 1 | Llama 4 Maverick 17B InstructMeta | 7,435 | Single source |
| 2 | Qwen3 Coder PlusAlibaba | 2,903 | Low confidence |