Directory

Model energy index

Comparative energy use per query across leading large language models, ranked from most to least efficient.

Model
Wh / query
Tier
  • Gemini FlashGoogle

    Small, efficient. Runs on TPU v5e — among the lowest reported energy per query.

    0.24
    Tier A
  • Claude HaikuAnthropic

    Lightweight Claude tier, good for high-volume tasks.

    0.40
    Tier A
  • GPT-4o miniOpenAI

    Distilled model, suitable for most everyday tasks.

    0.50
    Tier A
  • Llama 3 8B (self-hosted)Meta

    Energy depends on your grid; great on renewable-powered infra.

    0.60
    Tier B
  • Mistral SmallMistral

    Efficient European-hosted option.

    0.70
    Tier B
  • Gemini 2.5 ProGoogle

    Larger reasoning model — use selectively.

    1.80
    Tier B
  • Claude SonnetAnthropic

    Mid-tier Claude, balanced quality/energy.

    2.20
    Tier C
  • GPT-4oOpenAI

    Multimodal flagship; meaningfully more energy than 4o-mini.

    2.90
    Tier C
  • Llama 3 70B (self-hosted)Meta

    Heavy self-hosted footprint unless run on green energy.

    5.50
    Tier D
  • Claude OpusAnthropic

    Top-tier reasoning. Reserve for high-value tasks only.

    6.50
    Tier D
  • GPT-5 / o-series reasoningOpenAI

    Deep-reasoning runs can use 10–30× a standard query.

    8.00
    Tier D
A note on products vs models

Tools like GitHub Copilot, Cursor, ChatGPT and Perplexity are products, not individual LLMs. They route each prompt through one or more underlying models, and that routing changes over time and across plans. We rank LLMs above by per-query energy use, and map popular products to their likely backing models below so you can estimate their footprint.

Products

AI products & their likely models

IDE assistants, chat apps and agents don’t publish a fixed Wh/query because the model serving your request can vary by plan, prompt and time. The ranges below come from the LLMs each product is believed to route to — use them as order-of-magnitude guidance.

  • GitHub CopilotGitHub / OpenAI

    Code completion typically routes to a fast/small tier; chat & agent modes use larger GPT-4o-class models.

    GPT-4o miniGPT-4o
    ~0.502.90 Wh / query
    IDE assistant
  • CursorAnysphere

    User-selectable model; defaults vary across Auto, Sonnet and GPT-4o-class options.

    GPT-4oClaude SonnetGemini 2.5 Pro
    ~1.802.90 Wh / query
    IDE assistant
  • ChatGPTOpenAI

    Free tier routes to 4o-mini-class models; Plus/Pro reach GPT-4o and o-series reasoning.

    GPT-4o miniGPT-4oGPT-5 / o-series reasoning
    ~0.508.00 Wh / query
    Chat app
  • Claude (app)Anthropic

    Tier and plan determine whether Haiku, Sonnet or Opus answers.

    Claude HaikuClaude SonnetClaude Opus
    ~0.406.50 Wh / query
    Chat app
  • Gemini (app)Google

    Free uses Flash-class models; Advanced uses Pro for harder prompts.

    Gemini FlashGemini 2.5 Pro
    ~0.241.80 Wh / query
    Chat app
  • PerplexityPerplexity

    Routes across providers; Pro lets users pick the backing model.

    GPT-4o miniClaude SonnetGemini 2.5 Pro
    ~0.502.20 Wh / query
    Search