Energy Consumption of LLMs
iAs of: September 2026
Emission factors and training figures reflect the state as of September 2026. The German grid figure comes from the latest German Federal Environment Agency survey (344 g CO2/kWh for 2025) and keeps falling year over year.
Wissen
AI systems consume energy. A lot of energy. And consumption grows exponentially with model size. While the industry focuses on capabilities and benchmarks, the ecological footprint is often ignored -- or deliberately downplayed.
As an expert, you must not only understand the energy consumption of your AI systems but actively optimize it. This isn't just an ethical question but also an economic one: energy costs money, and under the EU AI Act, GPAI providers have been required to transparently report energy consumption since 2025.
Training vs. Inference
Training: The One-Time Energy Shock
| Model | Estimated Training Energy | CO2 Equivalent | Comparison |
|---|---|---|---|
| GPT-3 (175B, 2020) | ~1,287 MWh | ~552 t CO2 | about 112 cars for 1 year |
| GPT-4 (estimated, 2023) | ~50,000-62,000 MWh | ~25,000-30,000 t CO2 | Small town for 1 month |
| Llama 3.1 405B (2024) | ~21,600 MWh | ~11,390 t CO2eq | Officially reported by Meta |
Trend: Each model generation consumes 3-10x more training energy than the previous one.
iContext Matters
These numbers sound dramatic, but context matters: training happens once and is then shared by millions of users. The per-capita consumption is low. The question is whether the societal benefit justifies the energy expenditure -- and how to minimize it.
Inference: The Silent Ongoing Consumer
- Single request: ~0.001-0.01 kWh (depending on model size and response length)
- ChatGPT (estimated): With over 900 million weekly users, the estimated daily consumption is in the tens of GWh
- Entire inference industry: Inference consumption now exceeds training consumption by a factor of 3-10
Why inference is more problematic long-term: training is one-time, user numbers grow exponentially, agentic systems multiply requests, multimodal models consume significantly more.
Carbon Footprint
The carbon footprint critically depends on the electricity mix of the data center:
| Data Center Location | CO2 per kWh | Rating |
|---|---|---|
| Sweden, Norway (hydropower) | ~20-30 g | Excellent |
| France (nuclear) | ~50-80 g | Good |
| Germany (energy mix) | ~344 g (UBA, 2025) | Medium |
| USA (average) | ~380-420 g | Medium |
| India, China (coal-heavy) | ~600-900 g | Poor |
The same training produces 10x less CO2 in Norway than in India.
Water Consumption
Besides electricity, data centers consume significant amounts of water for cooling:
- GPT-3 training: Estimated 700,000 liters of freshwater
- Single ChatGPT conversation (20-50 questions): ~500 ml of water (2023 estimate — more recent analyses suggest significantly lower values, approx. 0.3-5 ml per individual query)
Verstehen
Optimization Strategies
1. Model Distillation -- A large model (teacher) trains a smaller one (student) for specific tasks. Energy savings: 80-95% for inference.
2. Quantization -- Reducing numerical precision:
- FP32 to FP16: ~50% less memory
- FP16 to INT8: Another ~50% reduction
- INT8 to INT4: Another ~50%, measurable but often acceptable trade-offs
- Total savings: INT4 needs ~8x less memory and ~4x less compute than FP32.
3. Caching and Prompt Caching -- Semantic caching for repetitive queries, prompt caching for long system prompts. Savings for repetitive workloads: up to 90%.
4. Routing: Right Model Size for the Task
Simple questions (FAQ, formatting) -> Small model (Haiku)
Medium complexity (summarization) -> Medium model (Sonnet)
Complex reasoning (analysis, planning) -> Large model (Opus)
An intelligent router saves 60-80% of inference costs because 70-80% of all queries are simple to medium.
5. Batch Processing Instead of Real-Time -- Better hardware utilization, dynamic batching, ability to wait for cheaper off-peak hours.
Interactive Energy Calculator
Calculate for yourself how much energy your AI usage consumes -- depending on the model, usage intensity, and data center location:
Energy Footprint Calculator
Calculate the energy consumption of your AI usage
Monthly Consumption
MediumThat equals...
18 km of driving
900 smartphone charges
0,9 days household electricity
900 hours LED bulb
Estimated values based on published data and industry estimates. Actual consumption may vary.
*Practical Rule of Thumb
Ask yourself with every AI deployment: Do I really need the largest model? Do I need real-time? Can I cache? These three questions alone can save 70-90% of energy -- without quality loss.
Anwenden
A company runs an AI chatbot with 100,000 queries per day. 75% of queries are FAQ responses that barely differ. Which optimization strategy has the greatest savings potential?
Reflect
The energy consumption of LLMs is real and measurable -- but also optimizable. Strategies like semantic caching, model routing, and efficient infrastructure can significantly reduce the footprint. In the next section, we will bring everything together with Responsible AI frameworks.