「では、AIを使う炭素コストは?」限定的開示下で大規模言語モデル推論のトークン当たり炭素強度を算出する有界推定フレームワーク:OpenAI GPT-5モデルファミリーを事例として
"So, What's the Carbon Cost of Using AI?" A Bounded Estimation Framework for Calculating the Per-Token Carbon Intensity of Large Language Model Inference Under Limited Disclosure, Using the Worked Example of OpenAI's GPT-5 Model Family (原題)
Manktelow, James
🤖 gxceed AI 要約
日本語
大手AI事業者が推論のエネルギー・炭素を開示しない中、公開データのみからトークン当たり炭素強度を推定する有界推定フレームワークを提示。GPT-5系に適用し、1,000トークン当たり約1.07〜14.80 gCO₂e(ロケーション基準)と約14倍のモデル間格差を推定。GHGプロトコルScope 2ガイダンスに沿いロケーション基準とマーケット基準を並行報告し、残存不確実性を乗法エンベロープに集約。モデル選択と推論努力設定が最大の削減レバーだと示す。
English
With no provider disclosing per-query inference energy or carbon, this paper offers a bounded estimation framework that triangulates public data to derive per-token carbon intensity. Applied to OpenAI's GPT-5 family, it estimates ~1.07–14.80 gCO2e/1,000 tokens (location-based), a ~14x spread across models, reporting location- and market-based figures in parallel under GHG Protocol Scope 2. Model choice and reasoning-effort setting emerge as the largest user-side levers.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
AI利用の炭素会計は日本企業のScope 3(購入したサービス)算定やSSBJ開示で未整備の領域であり、本フレームワークは推論由来排出の推定手法として実務に示唆を与える。特に生成AIを業務利用する企業が、調達・報告上の炭素インパクトをどう扱うかという論点を先取りしている。
In the global GX context
As ISSB/CSRD and SEC climate rules push companies to account for purchased digital services, this paper tackles the disclosure gap for AI inference emissions. It offers a replicable method for estimating Scope 2/3 impacts of LLM use, relevant to global debates on AI's climate footprint and provider transparency.
👥 読者別の含意
🔬研究者:限定的開示下でAI推論の炭素強度を推定する方法論と不確実性の扱いを学べる。
🏢実務担当者:生成AI利用の炭素コストを把握し、モデル選択・努力設定による削減策を検討できる。
🏛政策担当者:AI事業者への推論エネルギー・炭素開示義務の必要性を検討する根拠となる。
📄 Abstract(原文)
So, what is the carbon cost of using AI? No leading provider tells us: OpenAI, Anthropic and others publish no per-model, per-query energy or carbon figures. This means that nothing is disclosed about the operational carbon footprint of inference at the one unit where a user decides – a particular prompt, model and effort setting. We present a bounded estimation framework that answers the question from public and collected data alone. It triangulates an anchor (a published energy figure for one model) across three classes of public source, scales between models by observed throughput, applies query-length-conditional reasoning multipliers measured by our token-level campaign, corrects for the rise in per-GPU power between hardware generations, reports location-based carbon (what the local grid emitted) and market-based carbon (net of the provider's clean-energy contracts) in parallel under the GHG Protocol Scope 2 Guidance, and compounds residual uncertainty into a single multiplicative envelope (a range around each figure). Applied to OpenAI's GPT-5.x family, the framework estimates a fleet intensity of 353 gCO₂e/kWh location-based (319 – 500 across routing scenarios) and 52 gCO₂e/kWh market-based (47 – 65, rising to 159 in a particular scenario.) Per-model central estimates – per 1,000 tokens, about 750 words – span roughly 14×, from 1.07 gCO₂e/1,000 tokens location-based (0.16 market-based) for GPT-5.4-nano with reasoning off to 14.80 (2.16) for GPT-5 at high effort – within envelopes of roughly 10 – 19× top-to-bottom, median ~11×. All headline figures carry Low confidence (IPCC AR5 – the evidence is limited): a single line of audited per-model disclosure would collapse most of the envelope, and we invite it. What should a user do? A general user need not worry – twenty short prompts a day for a year is roughly 4 kg CO₂e, about 0.04% of the UK's average per-person footprint – however framing the task clearly and asking for shorter answers costs little. A power user, at about 122 kg a year (a short-haul plane flight), should choose the smallest capable model and the lowest effort setting the task allows – on GPT-5 to 5.4 the effort setting is the larger lever, from GPT-5.5 onward model choice is. An application designer holds the largest lever: at a million queries a month, routing from a flagship model at high effort to a mini model at medium could save on the order of 38 tCO₂e a year, the equivalent of just under four business-class return flights from Dallas to Delhi.
🔗 Provenance — このレコードを発見したソース
- Zenodo https://zenodo.org/records/23084616first seen 2026-10-02 04:14:17
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。