Carbon footprint of mid-size AI models per query and per token

Factor version soma-ai-ef-2026.09-lifecycle-v2, dated . Location-based grid averages, not renewable contracts. Estimates from published benchmarks.

How much CO₂ does one query to a mid-size model emit?

One query of 400 tokens to a mid-size model (class B: gpt-4o, claude-sonnet, llama-70b) emits about 0.029 g CO₂e on the global-average grid, counting electricity, the share of server hardware and amortised training; a million tokens emit 0.0721 kg CO₂e. Served from EU Sweden the same query emits 0.0052 g; from Japan, 0.035 g.

For comparison on the same grid, a class A (small) query emits 0.007 g and a class C (frontier) query emits 0.041 g. Class B uses 0.156 kWh of GPU electricity per million tokens before the facility overhead of 1.20.

Class B factors by cloud region

Per million tokens unless stated. Hardware 0.0030 kg and training 0.0041 kg are the same in every region and are included in the total. Per query = 400 tokens.
RegionElectricity kgTotal kg CO₂eg per queryWater LmL per query
EU Belgium0.0220.0290.0120.5580.22
EU Finland0.0130.0200.0080.5580.22
EU France0.0090.0160.00640.5580.22
EU Germany0.0590.0660.0260.4640.19
EU Ireland0.0460.0530.0210.2410.096
EU Netherlands0.0430.0500.020.5650.23
EU Sweden0.0060.0130.00520.9850.39
Global average0.0650.0720.0290.5820.23
Japan0.0800.0870.0350.3500.14
Singapore0.0760.0830.0330.5200.21
UK (London)0.0380.0450.0180.3620.14
US Central (Iowa)0.0670.0740.030.4580.18
US East0.0440.0510.020.4040.16
US South0.0620.0690.0280.5820.23
US Texas0.0540.0610.0240.2420.097
US West (Oregon)0.0460.0530.0210.5820.23

Which region is best for mid-size models?

EU Sweden gives the lowest total for class B, 0.0131 kg CO₂e per million tokens; Japan the highest, 0.0871 kg, a 6.6 times spread. In low-carbon regions the hardware and training terms become the larger share of the total, which is why the spread on the total is smaller than the spread on electricity alone (13 times).

How much water does a mid-size model use per query?

A 400-token query to a class B model consumes about 0.23 mL of water on the global average, from 0.096 mL in the lowest-water region to 0.39 mL in the highest. On-site cooling water plus the water embedded in grid electricity.

Which models are class B, and what about reasoning models?

Class B covers gpt-4o, claude-sonnet, llama-70b. Providers do not disclose parameter counts, so the class comes from the published tier and benchmark behaviour, and SOMA keeps the mapping current. Reasoning modes (o1, o3, DeepSeek-R1, extended thinking) generate 10 to 50 times more tokens per answer; when token counts come from billing that is already in the number, and when they come from message counts SOMA applies the multiplier.

Hardware term uncertainty -52% / +725%, skewed high; training term uncertainty ±1 order of magnitude.

Where do these numbers come from and how do I cite them?

Electricity term: GPU energy per token from the ML.ENERGY Leaderboard v3, a facility overhead of 1.20, and grid carbon intensity from EPA eGRID 2023 and Ember 2023. Embodied hardware and amortised training terms are documented on the methodology page. Water combines on-site cooling water (Li et al. 2025) with the water embedded in grid electricity (Reig et al. 2020). Uncertainty on the electricity term is ±40% for class A and ±50% for classes B and C.

Dataset: AI Inference Emission and Resource Factors for Corporate GHG Inventories, Guillermo Llopis, SOMA AI, September 2026, CC BY 4.0, concept DOI 10.5281/zenodo.20443585 (this version 10.5281/zenodo.22767475). Method: Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting, arXiv:2606.10660, under peer review.

Partners can fetch any row from the AI factors API.