Factor version soma-ai-ef-2026.09-lifecycle-v2, dated . Location-based grid averages, not renewable contracts. Estimates from published benchmarks.
A query of 400 tokens to a mid-size AI model (class B, such as gpt-4o) served from UK (London) emits about 0.018 g CO₂e, counting electricity, the share of server hardware and amortised training; a million tokens emit 0.0451 kg CO₂e. A small model (class A) emits 0.0042 g per query and a frontier model (class C) 0.028 g.
UK (London) ranks 5 of 15 regions for class B, where 1 is the lowest carbon. Compared with EU Sweden, the same query here emits 3.4 times as much; compared with Japan, 0.5 times as much. The grid, not the model, sets most of the difference.
| Model class | Electricity kg | Hardware kg | Training kg | Total kg CO₂e | g per query | Water L | mL per query |
|---|---|---|---|---|---|---|---|
| Class A (small) | 0.009 | 0.0008 | 0.0008 | 0.011 | 0.0042 | 0.090 | 0.036 |
| Class B (mid-size) | 0.038 | 0.0030 | 0.0041 | 0.045 | 0.018 | 0.362 | 0.14 |
| Class C (frontier) | 0.049 | 0.0039 | 0.0179 | 0.071 | 0.028 | 0.461 | 0.18 |
Example: a company of 100 people where 80 use a class B assistant for 2,000 messages a month each sends about 768 million tokens a year. Served from UK (London), that is 35 kg CO₂e a year, against 10 kg in EU Sweden.
A 400-token query to a class B model served from UK (London) consumes about 0.14 mL of water, and a million tokens 0.362 litres. The figure combines the water evaporated on site for cooling with the water embedded in generating the grid electricity. Regions that Li et al. 2025 do not cover directly use an estimated on-site intensity, marked as such in the dataset.
Class B, total kg CO₂e per million tokens, lowest first: EU Sweden 0.0131, EU France 0.0161, EU Finland 0.0201, EU Belgium 0.0291, UK (London) 0.0451, EU Netherlands 0.0501, US East 0.0511, EU Ireland 0.0531, US West (Oregon) 0.0531, US Texas 0.0611, EU Germany 0.0661, US South 0.0691, US Central (Iowa) 0.0741, Singapore 0.0831, Japan 0.0871.
Electricity term: GPU energy per token from the ML.ENERGY Leaderboard v3, a facility overhead of 1.20, and grid carbon intensity from EPA eGRID 2023 and Ember 2023. Embodied hardware and amortised training terms are documented on the methodology page. Water combines on-site cooling water (Li et al. 2025) with the water embedded in grid electricity (Reig et al. 2020). Uncertainty on the electricity term is ±40% for class A and ±50% for classes B and C.
Dataset: AI Inference Emission and Resource Factors for Corporate GHG Inventories, Guillermo Llopis, SOMA AI, September 2026, CC BY 4.0, concept DOI 10.5281/zenodo.20443585 (this version 10.5281/zenodo.22767475). Method: Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting, arXiv:2606.10660, under peer review.
Partners can fetch any row from the AI factors API.
When citing a single value, name the region and class, for example "0.0451 kg CO₂e per million tokens, class B, UK (London), SOMA AI factors September 2026", and link to app.somaai.earth/ai-factors/region/uk-london.