Factor version soma-ai-ef-2026.09-lifecycle-v2, dated . Location-based grid averages, not renewable contracts. Estimates from published benchmarks.
One query of 400 tokens to a mid-size model (class B: gpt-4o, claude-sonnet, llama-70b) emits about 0.029 g CO₂e on the global-average grid, counting electricity, the share of server hardware and amortised training; a million tokens emit 0.0721 kg CO₂e. Served from EU Sweden the same query emits 0.0052 g; from Japan, 0.035 g.
For comparison on the same grid, a class A (small) query emits 0.007 g and a class C (frontier) query emits 0.041 g. Class B uses 0.156 kWh of GPU electricity per million tokens before the facility overhead of 1.20.
| Region | Electricity kg | Total kg CO₂e | g per query | Water L | mL per query |
|---|---|---|---|---|---|
| EU Belgium | 0.022 | 0.029 | 0.012 | 0.558 | 0.22 |
| EU Finland | 0.013 | 0.020 | 0.008 | 0.558 | 0.22 |
| EU France | 0.009 | 0.016 | 0.0064 | 0.558 | 0.22 |
| EU Germany | 0.059 | 0.066 | 0.026 | 0.464 | 0.19 |
| EU Ireland | 0.046 | 0.053 | 0.021 | 0.241 | 0.096 |
| EU Netherlands | 0.043 | 0.050 | 0.02 | 0.565 | 0.23 |
| EU Sweden | 0.006 | 0.013 | 0.0052 | 0.985 | 0.39 |
| Global average | 0.065 | 0.072 | 0.029 | 0.582 | 0.23 |
| Japan | 0.080 | 0.087 | 0.035 | 0.350 | 0.14 |
| Singapore | 0.076 | 0.083 | 0.033 | 0.520 | 0.21 |
| UK (London) | 0.038 | 0.045 | 0.018 | 0.362 | 0.14 |
| US Central (Iowa) | 0.067 | 0.074 | 0.03 | 0.458 | 0.18 |
| US East | 0.044 | 0.051 | 0.02 | 0.404 | 0.16 |
| US South | 0.062 | 0.069 | 0.028 | 0.582 | 0.23 |
| US Texas | 0.054 | 0.061 | 0.024 | 0.242 | 0.097 |
| US West (Oregon) | 0.046 | 0.053 | 0.021 | 0.582 | 0.23 |
EU Sweden gives the lowest total for class B, 0.0131 kg CO₂e per million tokens; Japan the highest, 0.0871 kg, a 6.6 times spread. In low-carbon regions the hardware and training terms become the larger share of the total, which is why the spread on the total is smaller than the spread on electricity alone (13 times).
A 400-token query to a class B model consumes about 0.23 mL of water on the global average, from 0.096 mL in the lowest-water region to 0.39 mL in the highest. On-site cooling water plus the water embedded in grid electricity.
Class B covers gpt-4o, claude-sonnet, llama-70b. Providers do not disclose parameter counts, so the class comes from the published tier and benchmark behaviour, and SOMA keeps the mapping current. Reasoning modes (o1, o3, DeepSeek-R1, extended thinking) generate 10 to 50 times more tokens per answer; when token counts come from billing that is already in the number, and when they come from message counts SOMA applies the multiplier.
Hardware term uncertainty -52% / +725%, skewed high; training term uncertainty ±1 order of magnitude.
Electricity term: GPU energy per token from the ML.ENERGY Leaderboard v3, a facility overhead of 1.20, and grid carbon intensity from EPA eGRID 2023 and Ember 2023. Embodied hardware and amortised training terms are documented on the methodology page. Water combines on-site cooling water (Li et al. 2025) with the water embedded in grid electricity (Reig et al. 2020). Uncertainty on the electricity term is ±40% for class A and ±50% for classes B and C.
Dataset: AI Inference Emission and Resource Factors for Corporate GHG Inventories, Guillermo Llopis, SOMA AI, September 2026, CC BY 4.0, concept DOI 10.5281/zenodo.20443585 (this version 10.5281/zenodo.22767475). Method: Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting, arXiv:2606.10660, under peer review.
Partners can fetch any row from the AI factors API.