← Back to SOMA

Why estimates of the carbon footprint of a ChatGPT query differ by about 100x

Six published figures, what each of them actually counts, and a checklist for comparing any two of them. SOMA's own figures are computed from factor version soma-ai-ef-2026.09-lifecycle-v2, dated .

Why do published figures for one AI query differ so much?

Because they are not estimating the same thing, and two honest estimates can be about 100 times apart before anyone has made a mistake. The measured energy per text-generation inference published by Luccioni and co-authors is 0.047 Wh; the figure Sam Altman published for an average ChatGPT query is 0.34 Wh. That is 7.2 times, from the definition of a query alone. Then the electricity has to come from a grid, and the grids in SOMA's own table span 13.3 times between EU Sweden and Japan. Multiply the two and one query is 0.0018 g CO₂e at one end and 0.174 g at the other.

Nothing in that range is dishonest. It is what happens when the unit ("a query"), the boundary (what counts as the energy of serving it) and the grid are all left unstated. The rest of this page is the list of things that have to be stated before two numbers can be compared at all.

What does each published figure actually measure?

Published estimates for AI inference, as stated by their authors. Different units, different boundaries.
SourceFigureUnitWhat it includes
Luccioni, Jernite and Strubell (2023)0.047 kWhper 1,000 text-generation inferencesEnergy measured by the authors on their own hardware, as a benchmark across model types rather than an estimate of any commercial service.
Epoch AI (February 2025)0.3 Whper typical GPT-4o queryA modelled estimate: 500 output tokens, H100 hardware at about 70% of peak power, around 10% cluster utilisation. Long inputs are put at 2.5 to 40 Wh.
Sam Altman (2025)0.34 Whper average ChatGPT queryStated without a method. The same sentence gives 0.000085 gallons of water, about 0.32 mL, described as roughly one fifteenth of a teaspoon.
Google (August 2025)0.24 Wh · 0.03 g · 0.26 mLper median Gemini Apps text promptThe full serving stack: accelerators, host CPU and RAM, idle capacity held for spikes, and data-centre overhead, with Google's 2024 fleetwide grid carbon intensity.
Washington Post with UC Riverside (September 2024)519 mLper 100-word email written by GPT-4Water for a whole task rather than one prompt, counting the water used to generate the electricity as well as the water evaporated on site.
Li, Yang, Islam and Ren (2023)700to train GPT-3, onceTraining, not inference: 700,000 litres of clean freshwater evaporated directly in US data centres.
SOMA (this site)0.020 g · 0.16 mLper 400-token exchange, class B, US EastElectricity from benchmarked energy per token with facility overhead, plus embodied hardware and amortised training, on a location-based grid factor. Water is on-site cooling plus the water embedded in generation.

What are the seven things that actually differ?

Every gap between two AI footprint numbers comes from one of these, and usually from several at once.

  • What a query is. Epoch AI assumes 500 output tokens; SOMA's message conversion uses 400 tokens for a question and its answer; the Washington Post figure is a whole 100-word email. A figure per query and a figure per token only agree once you know the token count.
  • Which hardware, and how busy it is. The same model on newer accelerators, or at higher batch sizes, uses less energy per token. Epoch's estimate is explicit about assuming around 10% utilisation; a benchmark run on an idle test machine implies something different again.
  • What counts as the energy of serving. GPU energy alone, or GPU plus host CPU and memory, or all of that plus cooling and power distribution, or all of that plus the idle machines kept ready for traffic spikes. Google's figure includes the idle capacity; most others do not.
  • Which grid the electricity comes from. This is the largest single lever and the one most often unstated: 13.3 times between the cleanest and the dirtiest region in SOMA's table.
  • Location-based or market-based. A location-based figure uses the average grid; a market-based one nets off the provider's renewable contracts. The second is usually far lower and answers a different question.
  • Whether hardware and training are in scope. Almost every per-query figure is electricity only. Adding the embodied hardware and the amortised training raises the number, and raises it most where the grid is cleanest.
  • How water is counted. On-site evaporative cooling only, or that plus the water consumed generating the electricity. The second is several times larger in most regions, which is most of the distance between 0.26 mL and 519 mL.

A reasoning model adds an eighth: it generates 10 to 50 times more tokens for the same question, so a per-query figure from a standard model does not transfer to it at all.

The unit is the trap that catches people first. A figure per query and a figure per million tokens look like different levels of precision, but they are the same statement with a token count hidden inside one of them. Convert both to tokens and half the apparent disagreement disappears — which is also why SOMA publishes the factor per million tokens and derives the per-query figure from it, rather than the other way round.

Where do SOMA's own per-query figures fall, and why?

In the middle for electricity, and higher than most for everything after it. One exchange of 400 tokens comes to 0.0050 g CO₂e for a small model, 0.020 g for a mid-size one and 0.031 g for a frontier model, all served from US East. On a clean grid the same mid-size exchange is 0.0052 g, and on the dirtiest grid in the table 0.035 g.

The electricity behind those figures is about 0.06 Wh per exchange for a mid-size model, which sits below Google's 0.24 Wh and Epoch's 0.3 Wh. Two reasons, both boundary rather than optimism: the exchange here is 400 tokens rather than a 500-token answer, and the benchmark energy covers the serving hardware and facility overhead but not the idle capacity a provider holds in reserve. In the other direction, SOMA adds two terms most published figures leave out — the share of the server hardware and the amortised training — which is why its numbers do not fall as far as the cleanest grid alone would suggest.

Whether a figure is high or low matters less than whether it is reproducible. Every number here resolves to a versioned table with a DOI, so a reader can rebuild it rather than trust it. The arithmetic and the sources are on the methodology page, and the per-region and per-class tables at /ai-factors.

Why is water the widest disagreement of all?

Because two published figures for what looks like the same thing are about 2,000 times apart: 0.26 mL for a median Gemini text prompt, and 519 mL for a 100-word email written by GPT-4. Most of that gap is not a dispute about cooling towers. It is three separate differences stacked: a prompt against a whole task, on-site water against on-site plus the water used to generate the electricity, and one company's own fleet against a generic US data centre.

SOMA reports both components explicitly, which is why its per-exchange water figure of 0.16 mL rises to 0.39 mL on Sweden's grid even as the carbon falls. Where a water number comes with no statement of which components it contains, it cannot be compared with anything, and the honest response is to ask rather than to average.

If one query is a fraction of a gram, does any of this matter?

The per-query number is small and the aggregate is not. The IEA puts data centres at 415 TWh in 2024, about 1.5% of world electricity, and projects that to more than double to around 945 TWh by 2030 (IEA, Energy and AI, 2025). A figure that is negligible per use is not negligible per company per year, and it is the company-year figure that a Scope 3 inventory asks for.

For a reporting team the practical consequence is narrower still: the number has to be defensible, not dramatic. A small figure with a stated method, a versioned factor and a disclosed boundary survives assurance. A large figure copied from a headline does not, and neither does a small one. The CSRD guide works through the line itself.

Which of these figures should I use in my own reporting?

The one whose boundary matches the question you are answering, which for a corporate inventory is rarely a per-prompt headline. A Scope 3 Category 1 line needs a figure that scales with your own usage, uses the grid your provider actually serves from, and states what it leaves out — so it has to be per token or per message, location-based, and versioned.

A provider's published per-prompt number is a different instrument. It describes that provider's median prompt on its own fleet, not your traffic, and it cannot be multiplied by your message count without inheriting every assumption in it. There is one case where a provider figure wins outright: when the provider issues a certified figure for your account. That is the top tier of the hierarchy and it replaces the calculation rather than being averaged with it.

What you must not do is mix them inside one inventory — a market-based provider figure for one service and a location-based estimate for another, added into a single total, is a number that cannot be defended in assurance. Pick a basis, apply it to every service, and disclose the exceptions.

How do I compare any two AI footprint numbers?

Ask the same seven questions of both. If either one cannot answer a question, the comparison is not available.

  1. What is the unit — a prompt, a task, a thousand inferences, a million tokens? Convert both to tokens before anything else.
  2. How many tokens does the author's "query" contain, and is that stated or assumed?
  3. Which model, and which class of model? A small model and a frontier model differ by several times for identical work.
  4. Which region, and is the grid factor location-based or market-based?
  5. What is inside the energy boundary: accelerators only, the whole server, the facility, the idle capacity?
  6. Are embodied hardware and amortised training included, or is it electricity only?
  7. For water, is it on-site cooling only or also the water used to generate the electricity?

Two numbers that survive all seven can be compared. Two that do not are not in conflict; they are answers to different questions that happen to share a name.

Get the audit-ready figure for your organisation

SOMA turns your providers' usage exports into the finished Scope 3 Category 1 entry, with the tier, factor, version, sources and uncertainty attached, ready for your auditor — and stated in a way that answers the seven questions above before anyone asks them. Write to lili@somaai.earth, or start with the methodology and the CSRD guide.