The method behind every "carbon data powered by SOMA" figure, with the full factor tables, the sources, the uncertainty and the boundary. Factor version soma-ai-ef-2026.09-lifecycle-v2. Last updated .
They are estimates of the carbon and water cost of running one million tokens through a commercial AI model, built from published benchmarks and published grid data. SOMA does not meter anyone's data centre: every input is a figure somebody else published, assembled into one number a sustainability team can put in an inventory. There is one carbon factor and one water factor for each combination of model class and cloud region, so 3 classes across 16 regions gives 48 rows of each.
A carbon factor has three parts, added together: the electricity used to serve the tokens, the share of the server hardware those tokens used up, and the amortised share of the emissions from training the model. Most published AI figures stop at the first part. Water is reported beside carbon rather than instead of it, because the two do not move together — the cleanest grid in this table has the highest water intensity.
The method is described in a preprint, arXiv 2606.10660, which is under review at Sustainable Production and Consumption (Elsevier), manuscript SPC-D-26-04019. The factor tables themselves are open on Zenodo under CC BY 4.0, so a third party can reproduce any number on this page without asking SOMA for anything.
About 0.020 g CO₂e for one exchange with a mid-size model served from a US East data centre — one question and its answer, counted as 400 tokens, and including the hardware and training terms as well as the electricity. The figure changes with the model class and with the region far more than with anything the user types.
Car comparison: 3.93 × 10⁻⁴ metric tons CO₂e per mile for the average petrol passenger vehicle, and phone comparison: 0.019 kWh per smartphone charge, both from the US EPA greenhouse gas equivalencies calculator.
Two things move this number a lot. Moving the same model from EU Sweden to Singapore multiplies the electricity term by more than 12. And a reasoning model, which thinks in tokens before it answers, generates 10 to 50 times more tokens for the same question, so the per-exchange figure scales with it. For reasoning-heavy tools, use the token counts from the provider's billing export rather than a per-message estimate.
| Model class | g CO₂e / exchange | mL water / exchange | kg CO₂e / million tokens | Uncertainty |
|---|---|---|---|---|
| Class A — Small | 0.0050 | 0.040 | 0.01259 | ±40% |
| Class B — Mid | 0.020 | 0.162 | 0.05111 | ±50% |
| Class C — Frontier | 0.031 | 0.206 | 0.07771 | ±50% |
About 0.16 mL for one mid-size exchange served from a US East data centre, which is 0.16 litres per thousand exchanges. The same exchange served from EU Sweden uses 0.39 mL, more than twice as much, even though its carbon is far lower.
Two terms are added. WUE is the water evaporated on site to cool the servers, per kWh of electricity (Li et al. 2025). EWIF is the water consumed upstream to generate that electricity, per kWh (Reig et al. 2020). In most regions the upstream term dominates, which is why the water map follows the generation mix rather than the local climate. Hydro-heavy and nuclear-heavy grids have very low carbon and very high water; gas-heavy grids are the other way round.
Water values for regions that Li et al. 2025 do not cover are estimates and are marked as such in the dataset. Water is reported as a separate figure and is never converted into carbon or added to it.
Water does not belong in a GHG inventory line, so why produce it at all? Because it is the question that follows the carbon one, and because the two answers point in opposite directions often enough that answering only the first one gives bad advice. A team that moves a workload to the greenest grid it can find, on carbon evidence alone, can multiply its AI water footprint on the way. Having both numbers in front of you, from the same token count, is the only way to see that trade-off before the decision is made rather than after.
The tier is decided per service by the data you can actually obtain for it, not by the company as a whole. A typical organisation sits at three tiers at once: an API with token billing, a seat product with message counts, and a SaaS subscription where only the invoice exists.
Tiers are meant to be climbed. A service sitting at Tier 1 because nobody has asked the vendor for usage data moves to Tier 2b the moment an admin console exports seat and message counts, and to Tier 2a when the vendor exposes tokens per model. Each step narrows the band around the figure by roughly an order of magnitude, and the tier recorded against the line is what tells next year's reviewer whether the number moved because the usage changed or because the data got better.
The output is one Scope 3 Category 1 entry per service, carrying the tier, the factor, the factor version and the sources, which is what the GHG Protocol data-quality hierarchy that ESRS E1 refers to asks for. The step-by-step version of this is in the CSRD guide.
Commercial providers do not disclose model sizes, so models are grouped into three classes by their published tier and their benchmark behaviour. Class decides the factor, and the gap between the classes is large: a frontier model costs about 6.2 times a small model for the same number of tokens in the same region.
Putting a model in the wrong class is the largest single source of error in this method, worth about four to five times (arXiv 2606.10660, section 7.3) — more than the uncertainty on any individual factor. That is why SOMA maintains the mapping, updates it as models are renamed and re-released, and states in the output when a class has been assumed rather than recognised. Names are misleading on their own: a model called "mini" can be a reasoning model, and a mixture-of-experts model is classed by its active parameters, not its total.
It falls back, and it says that it has. An unrecognised model name is treated as class B, the middle class, and the answer is flagged as an assumed class rather than a detected one. A missing or unrecognised region falls back to the provider's usual serving region, or to the global average when even the provider is unknown, and is flagged the same way. Both flags travel with the figure into the API response and into the inventory line, so a reviewer can see which numbers rest on a lookup and which rest on an assumption.
Some workloads are out of scope for this factor version rather than merely unrecognised. Embedding, speech, image, video, moderation and reranking models are not token-based chat or completion work, their energy per unit behaves differently, and the tables here would be the wrong instrument. SOMA reports those separately rather than quietly pricing them at a chat factor.
Electricity. Energy per token from the ML.ENERGY Leaderboard v3, multiplied by a facility overhead of 1.20 (arXiv 2606.10660, section 3), multiplied by the carbon intensity of the grid where the provider serves the model (EPA eGRID 2023, Ember 2023). Uncertainty ±40% for class A and ±50% for classes B and C. This is the only term most published figures include.
Embodied hardware. Dell PowerEdge XE9680 PCF (Apr 2025), NVIDIA HGX H100 PCF (ISO 14067), 5 yr / 60% utilisation. Class A 0.75 g, class B 3.03 g and class C 3.85 g per million tokens, uncertainty -52% / +725%, skewed high.
Training, amortised. Meta Llama 3.1 model card, location-based, over 500 T tokens served. Class A 0.84 g, class B 4.08 g and class C 17.86 g per million tokens, uncertainty ±1 order of magnitude.
The two add-ons are region-independent, so their share of the total depends entirely on how clean the grid is. For a class B model in US East, electricity is about 86% of the factor. In EU Sweden the same model's electricity is only about 46% of it, and the hardware and training terms carry the rest. Moving a workload to a clean grid therefore cannot take the number to zero: the 12-times regional spread on electricity alone compresses to about 6 times once the whole factor is counted. A page that quotes the larger figure is quoting the electricity term only.
EU Sweden, at 0.006 kg CO₂e per million tokens of electricity for a class B model. The highest row in the table is Japan at 0.080. Where a model runs matters more than which model it is: the same class B model costs more than 12 times as much carbon in electricity in Singapore as in EU Sweden, while the spread from the smallest to the largest model class within one region is about 6.2 times.
| Region | Class A | Class B | Class C | Water A | Water B | Water C |
|---|---|---|---|---|---|---|
| EU Belgium | 0.005 | 0.022 | 0.028 | 0.138 | 0.558 | 0.710 |
| EU Finland | 0.003 | 0.013 | 0.017 | 0.138 | 0.558 | 0.710 |
| EU France | 0.002 | 0.009 | 0.011 | 0.138 | 0.558 | 0.710 |
| EU Germany | 0.015 | 0.059 | 0.075 | 0.115 | 0.464 | 0.590 |
| EU Ireland | 0.011 | 0.046 | 0.058 | 0.060 | 0.241 | 0.307 |
| EU Netherlands | 0.011 | 0.043 | 0.055 | 0.140 | 0.565 | 0.719 |
| EU Sweden | 0.002 | 0.006 | 0.008 | 0.244 | 0.985 | 1.253 |
| Global average | 0.016 | 0.065 | 0.082 | 0.144 | 0.582 | 0.740 |
| Japan | 0.020 | 0.080 | 0.101 | 0.087 | 0.350 | 0.446 |
| Singapore | 0.019 | 0.076 | 0.097 | 0.129 | 0.520 | 0.661 |
| UK (London) | 0.009 | 0.038 | 0.049 | 0.090 | 0.362 | 0.461 |
| US Central (Iowa) | 0.017 | 0.067 | 0.086 | 0.113 | 0.458 | 0.583 |
| US East | 0.011 | 0.044 | 0.056 | 0.100 | 0.404 | 0.514 |
| US South | 0.015 | 0.062 | 0.079 | 0.144 | 0.582 | 0.740 |
| US Texas | 0.013 | 0.054 | 0.068 | 0.060 | 0.242 | 0.308 |
| US West (Oregon) | 0.011 | 0.046 | 0.059 | 0.144 | 0.582 | 0.740 |
Water values for regions that Li et al. 2025 do not cover are estimates and are marked as such in the dataset. Read the two halves of the table together rather than one at a time: EU Sweden has the lowest carbon and the highest water of any region here.
Factors are location-based: they use the average carbon intensity of the grid serving the region, not the renewable contracts the provider has signed. A market-based figure, which nets off power purchase agreements and certificates, will be lower for the same work, and the two must not be compared with each other. Per-region and per-class pages are at /ai-factors.
Not in these tables. The factors are location-based, so they use the average carbon intensity of the grid that physically serves the region, whatever contracts the provider has signed. A market-based figure — one that applies power purchase agreements and energy attribute certificates — is a different number, usually much lower, and it answers a different question.
Both are legitimate and the GHG Protocol Scope 2 guidance expects a reporter to be clear about which one is on the page. Mixing them inside a single inventory line, or comparing a provider's market-based claim with a location-based estimate and calling the difference a saving, is the common error. If your provider gives you a certified account-level figure, that is Tier 3 and it replaces the calculation rather than being averaged with it — record which basis it uses.
Stated plainly, because a boundary that is not written down is not a boundary. The factor covers inference: the GPU energy measured in the benchmark, which is about 60% of total server energy at production batch sizes (arXiv 2606.10660), grossed up by the facility overhead, plus the embodied hardware and amortised training terms.
Outside it: data-centre construction, provider offices, staff and research overhead, network transport between the data centre and the user, and the user's own device. Also outside this factor version: embedding, speech, image, video, moderation and reranking models, which are not token-based chat or completion workloads and are reported separately rather than forced into this table.
Nothing here is a market-based claim, an offset or a net figure. A factor is a gross, location-based estimate from published benchmarks, and the page that consumes it says so.
They are in the same range, and they are not measuring the same thing. SOMA's class B factor works out at about 0.06 Wh of electricity for a 400-token exchange. Google reports 0.24 Wh, 0.03 g CO₂e and 0.26 mL of water for the median Gemini Apps text prompt (Google, 21 August 2025). Epoch AI estimates about 0.3 Wh for a typical GPT-4o query (Epoch AI, 7 February 2025), and Sam Altman has written that the average ChatGPT query uses about 0.34 Wh (The Gentle Singularity, 2025).
SOMA's electricity per exchange is derived from the same benchmark energy that produces the carbon factors. It covers the serving hardware and the facility overhead; it does not cover the idle capacity a provider keeps ready for traffic spikes, which Google's own figure does. That is one reason the two sit apart, and it is a boundary difference rather than a contradiction.
The differences between those numbers are mostly boundary and assumption, not disagreement about physics: how many tokens a "query" is, whether idle capacity and host CPU are counted, which grid the electricity comes from, and whether hardware and training are included at all. The published figures for one query span roughly two orders of magnitude once the grid is allowed to vary, which is worked through in why AI footprint estimates differ.
Take the class B row for US East in the table above. The electricity term is 0.044 kg CO₂e per million tokens; add the embodied term, 0.00303 kg, and the training term, 0.00408 kg, and you get 0.05111 kg per million tokens. Multiply by your own token count in millions and you have the figure SOMA would produce for that service, before any uncertainty band is applied.
The same three columns are in the Zenodo CSVs, so the arithmetic can be checked against the dataset rather than against this page. If a figure in a SOMA export does not reproduce, the factor version printed beside it says which edition of the tables to check, and older versions stay resolvable.
Method paper, plain text:
Llopis, G. (2026). Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting. arXiv:2606.10660. Under review at Sustainable Production and Consumption, manuscript SPC-D-26-04019. arxiv.org/abs/2606.10660
Factor dataset, plain text:
Llopis, G. (2026). AI Inference Emission and Resource Factors for Corporate GHG Inventories (version 2.0.0) [Data set]. Zenodo. CC BY 4.0. doi.org/10.5281/zenodo.20443585
Cite the concept DOI, 10.5281/zenodo.20443585, when you mean the dataset, and the version DOI, 10.5281/zenodo.22767475, when a reader has to reproduce a specific figure. BibTeX:
@misc{llopis2026aiinference,
author = {Llopis, Guillermo},
title = {Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting},
year = {2026},
eprint = {2606.10660},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.10660},
note = {Preprint. Under review at Sustainable Production and Consumption, manuscript SPC-D-26-04019}
}
@dataset{llopis2026factors,
author = {Llopis, Guillermo},
title = {AI Inference Emission and Resource Factors for Corporate GHG Inventories},
year = {2026},
publisher = {Zenodo},
version = {2.0.0},
doi = {10.5281/zenodo.20443585},
url = {https://doi.org/10.5281/zenodo.20443585},
note = {Concept DOI, resolves to the latest version. Version 2.0.0 DOI: 10.5281/zenodo.22767475}
}The same seven CSV files are mirrored on GitHub (with the consistency checker) and Hugging Face (with a dataset viewer). Zenodo remains the citable record.
Factor version soma-ai-ef-2026.09-lifecycle-v2, dated . Version 2.0.0 of the dataset was published on and is the version behind every number on this page.
Every factor SOMA serves carries this version string, in the product, in the API response and in the exported inventory line. When a table changes, the version changes, the old version stays resolvable on Zenodo, and partners are told what moved and why. That is what makes last year's inventory reproducible after the factors have been updated — an auditor can ask which version produced a figure and get an answer that resolves to a citable record.
Yes. Send a model name and the region it is served from, and the endpoint returns carbon and water per million tokens, split into the three terms, with the class, the region, the uncertainty, the factor version and the sources attached. The full table is 48 rows, 3 classes across 16 regions. The response says when a class or a region has been assumed rather than recognised, so a product can show the caveat instead of hiding it.
Documentation is at /docs/ai-factors, the browsable tables at /ai-factors. Products that display the figures label them "carbon data powered by SOMA" and link back to this page.
The numbers above are factors. SOMA turns your providers' usage exports into the finished Scope 3 Category 1 entry, with the tier, factor, version, sources and uncertainty attached, ready for your auditor. Write to lili@somaai.earth or visit somaai.earth.