Carbon footprint of frontier AI models per query and per token

Factor version soma-ai-ef-2026.09-lifecycle-v2, dated . Location-based grid averages, not renewable contracts. Estimates from published benchmarks.

How much CO₂ does one query to a frontier model emit?

One query of 400 tokens to a frontier model (class C: claude-opus, gpt-4, o1, large MoE models) emits about 0.041 g CO₂e on the global-average grid, counting electricity, the share of server hardware and amortised training; a million tokens emit 0.104 kg CO₂e. Served from EU Sweden the same query emits 0.012 g; from Japan, 0.049 g.

For comparison on the same grid, a class A (small) query emits 0.007 g and a class B (mid-size) query emits 0.029 g. Class C uses 0.199 kWh of GPU electricity per million tokens before the facility overhead of 1.20.

Class C factors by cloud region

Per million tokens unless stated. Hardware 0.0039 kg and training 0.0179 kg are the same in every region and are included in the total. Per query = 400 tokens.
RegionElectricity kgTotal kg CO₂eg per queryWater LmL per query
EU Belgium0.0280.0500.020.7100.28
EU Finland0.0170.0390.0150.7100.28
EU France0.0110.0330.0130.7100.28
EU Germany0.0750.0970.0390.5900.24
EU Ireland0.0580.0800.0320.3070.12
EU Netherlands0.0550.0770.0310.7190.29
EU Sweden0.0080.0300.0121.2530.5
Global average0.0820.1040.0410.7400.3
Japan0.1010.1230.0490.4460.18
Singapore0.0970.1190.0470.6610.26
UK (London)0.0490.0710.0280.4610.18
US Central (Iowa)0.0860.1080.0430.5830.23
US East0.0560.0780.0310.5140.21
US South0.0790.1010.040.7400.3
US Texas0.0680.0900.0360.3080.12
US West (Oregon)0.0590.0810.0320.7400.3

Which region is best for frontier models?

EU Sweden gives the lowest total for class C, 0.0297 kg CO₂e per million tokens; Japan the highest, 0.123 kg, a 4.1 times spread. In low-carbon regions the hardware and training terms become the larger share of the total, which is why the spread on the total is smaller than the spread on electricity alone (13 times).

How much water does a frontier model use per query?

A 400-token query to a class C model consumes about 0.3 mL of water on the global average, from 0.12 mL in the lowest-water region to 0.5 mL in the highest. On-site cooling water plus the water embedded in grid electricity.

Which models are class C, and what about reasoning models?

Class C covers claude-opus, gpt-4, o1, large MoE models. Providers do not disclose parameter counts, so the class comes from the published tier and benchmark behaviour, and SOMA keeps the mapping current. Reasoning modes (o1, o3, DeepSeek-R1, extended thinking) generate 10 to 50 times more tokens per answer; when token counts come from billing that is already in the number, and when they come from message counts SOMA applies the multiplier.

Hardware term uncertainty -52% / +725%, skewed high; training term uncertainty ±1 order of magnitude.

Where do these numbers come from and how do I cite them?

Electricity term: GPU energy per token from the ML.ENERGY Leaderboard v3, a facility overhead of 1.20, and grid carbon intensity from EPA eGRID 2023 and Ember 2023. Embodied hardware and amortised training terms are documented on the methodology page. Water combines on-site cooling water (Li et al. 2025) with the water embedded in grid electricity (Reig et al. 2020). Uncertainty on the electricity term is ±40% for class A and ±50% for classes B and C.

Dataset: AI Inference Emission and Resource Factors for Corporate GHG Inventories, Guillermo Llopis, SOMA AI, September 2026, CC BY 4.0, concept DOI 10.5281/zenodo.20443585 (this version 10.5281/zenodo.22767475). Method: Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting, arXiv:2606.10660, under peer review.

Partners can fetch any row from the AI factors API.