Model energy index
Comparative energy use per query across leading large language models, ranked from most to least efficient.
- Gemini FlashGoogle
Small, efficient. Runs on TPU v5e — among the lowest reported energy per query.
0.24Tier A - Claude HaikuAnthropic
Lightweight Claude tier, good for high-volume tasks.
0.40Tier A - GPT-4o miniOpenAI
Distilled model, suitable for most everyday tasks.
0.50Tier A - Llama 3 8B (self-hosted)Meta
Energy depends on your grid; great on renewable-powered infra.
0.60Tier B - Mistral SmallMistral
Efficient European-hosted option.
0.70Tier B - Gemini 2.5 ProGoogle
Larger reasoning model — use selectively.
1.80Tier B - Claude SonnetAnthropic
Mid-tier Claude, balanced quality/energy.
2.20Tier C - GPT-4oOpenAI
Multimodal flagship; meaningfully more energy than 4o-mini.
2.90Tier C - Llama 3 70B (self-hosted)Meta
Heavy self-hosted footprint unless run on green energy.
5.50Tier D - Claude OpusAnthropic
Top-tier reasoning. Reserve for high-value tasks only.
6.50Tier D - GPT-5 / o-series reasoningOpenAI
Deep-reasoning runs can use 10–30× a standard query.
8.00Tier D
Tools like GitHub Copilot, Cursor, ChatGPT and Perplexity are products, not individual LLMs. They route each prompt through one or more underlying models, and that routing changes over time and across plans. We rank LLMs above by per-query energy use, and map popular products to their likely backing models below so you can estimate their footprint.
AI products & their likely models
IDE assistants, chat apps and agents don’t publish a fixed Wh/query because the model serving your request can vary by plan, prompt and time. The ranges below come from the LLMs each product is believed to route to — use them as order-of-magnitude guidance.
- GitHub CopilotGitHub / OpenAI
Code completion typically routes to a fast/small tier; chat & agent modes use larger GPT-4o-class models.
GPT-4o miniGPT-4o~0.50–2.90 Wh / queryIDE assistant - CursorAnysphere
User-selectable model; defaults vary across Auto, Sonnet and GPT-4o-class options.
GPT-4oClaude SonnetGemini 2.5 Pro~1.80–2.90 Wh / queryIDE assistant - ChatGPTOpenAI
Free tier routes to 4o-mini-class models; Plus/Pro reach GPT-4o and o-series reasoning.
GPT-4o miniGPT-4oGPT-5 / o-series reasoning~0.50–8.00 Wh / queryChat app - Claude (app)Anthropic
Tier and plan determine whether Haiku, Sonnet or Opus answers.
Claude HaikuClaude SonnetClaude Opus~0.40–6.50 Wh / queryChat app - Gemini (app)Google
Free uses Flash-class models; Advanced uses Pro for harder prompts.
Gemini FlashGemini 2.5 Pro~0.24–1.80 Wh / queryChat app - PerplexityPerplexity
Routes across providers; Pro lets users pick the backing model.
GPT-4o miniClaude SonnetGemini 2.5 Pro~0.50–2.20 Wh / querySearch