Sign In or Create an Account.

By continuing, you agree to the Terms of Service and acknowledge our Privacy Policy

Energy

Which AI Models Use the Most Energy?

A new dashboard from the Sustainable AI Group, founded by artificial intelligence alums, estimates the relative energy intensity of proprietary tools.

•
A Claude Pac-Man.
Heatmap Illustration/Getty Images

The rise of artificial intelligence is driving an historic surge in electricity demand that’s boosting fossil fuel use and threatening climate progress. All this electricity doesn’t power AI in some generalized, always-on way, though. Data centers’ energy consumption is a function of the millions of individual queries users submit to AI programs such as Claude and ChatGPT.

When it comes to how efficiently models process those queries and generate responses, AI models are not interchangeable. Some are more like gas guzzlers, others more like Priuses. When a user engages an AI chatbot or AI agent, however, there’s essentially no way for them to know which kind of vehicle they are stepping into. They may know which company built it, and even the precise model name and number, but no AI company has published information about how much energy one model uses compared to another.

In the absence of corporate disclosure from the big three proprietary AI developers — Anthropic, OpenAI, and Google — researchers with the Sustainable AI Group, a research and advisory company, developed a backdoor method to estimate and compare the amount of energy these developers’ models consume. They published their findings on Tuesday in an interactive dashboard that ranks AI programs by energy intensity.

“We think this is an important next step to get some science-based information out there to help folks start making better decisions,” Boris Gamazaychikov, the CEO of the Sustainable AI Group, told me. “We also hope that if the model providers think that this is really wrong, that they can come out and prove it with some actual data.”

In general, the researchers found that larger, higher-capability models, such as Anthropic’s Opus and OpenAI’s Sol, used nearly four times as much energy on average as smaller, nimbler models from those companies, Haiku and Terra. Newer iterations of each model also weren’t necessarily more efficient than their predecessors.

While the group has yet to evaluate the latest models that hit the market during the research period, so far the researchers found that for the same task, the least efficient models can consume more than 30 times the energy of the most efficient models. They also found a significant difference between “chat” sessions, where a user asks an AI chatbot a question, and “agentic” sessions,” where a user asks the AI to perform a series of tasks. A typical agentic session used 27 times more energy, on average, than a typical chat session conducted using the same AI model.

The Sustainable AI Group was founded by Sasha Luccioni, the former AI and climate lead at the open source AI platform Hugging Face, and Gamazaychikov, who previously led AI sustainability at Salesforce. In their earlier roles, the two collaborated on a project called AI Energy Score, which is similar in spirit to the Environmental Protection Agency’s EnergyStar program for home appliances. They developed a method to directly measure the energy efficiency of “open-weight” AI models, or those that fully disclose their inner workings, and published the results in a public leaderboard.

Luccioni and Gamazaychikov founded the Sustainable AI Group because they wanted to give AI users, particularly large corporate users, the tools to understand the relative emissions impacts of proprietary AI models. Gamazaychikov told me that Salesforce had tried to get energy-use data from its AI providers for years to no avail.

Their first hire was Nidhal Jegham, a graduate student at the University of Rhode Island who published a landmark paper last year called “How Hungry is AI?” Jegham and his co-authors developed a method to estimate the energy, water, and carbon effects of proprietary models at the level of a single prompt or query. The paper was accepted by the journal Communications of the Association for Computing Machinery, and the peer-reviewed version will come out in January.

The approach the Sustainable AI Group developed builds on both Jegham’s paper and the AI Energy Score project. The work began with testing open-weight models to see how they perform in realistic deployment configurations and directly measuring their energy consumption. From there the researchers identified mathematical relationships between various open models’ energy use and other measurable statistics, such as their size.

The next step was to take those statistical relationships from the open-weight models and apply them to similarly-sized proprietary models. The problem is, no one knows how “big” proprietary models are. The size of an AI model usually refers to the number of parameters it contains, i.e. the quantity of numerical representations of what the model has learned that it uses to produce a response.

“When we have a closed model, we don't have the model size. We don't have the deployment conditions. We don't have anything, so we need to find things we can observe from this closed model that can reflect its size,” Jegham explained to me. One key discovery, he said, was that “knowledge retention,” or how well the model can remember factual information, is a strong predictor of model size. A company called Artificial Analysis tests models for knowledge retention, so the researchers compared those results to model size for open models and applied the same statistical relationship to estimate the size of closed models.

This is a simplified explanation — there were many other variables and data points that went into the Sustainable AI Group’s estimates. The researchers also had to develop a separate methodology to evaluate Google’s models, since those mostly run on the company’s proprietary “tensor processing units,” rather than the Nvidia chips the researchers’ initial measurements were based on.

The group’s main findings are based on a per-token estimate of each model’s energy use, i.e. the energy required to process the smallest units of data that an AI deals with. Every time you type a question into a chatbot, the model breaks down the words into smaller bits — i.e. tokens — each just a few characters long, usually. The model also first formulates its response in tokens before translating it to text, an image, or whatever you’re requesting; input tokens are less energy-intensive than the tokens the models spit out. The Sustainable AI Group reports each of its per-token estimates as a range to reflect uncertainty.

For now, the firm is keeping its per-token estimates behind a paywall, but it has already started to use them to advise corporate clients in estimating their AI-related emissions, Gamazaychikov said. For example, he mentioned working with Etsy to help the online retailer develop a “model router,” essentially some software that routes a given query to the most appropriate model for the task, taking into account carbon and cost. It’s also partnering with the corporate emissions accounting platform Watershed to explore how to integrate its model-specific energy numbers into Watershed’s system.

Instead of displaying per-token energy use, the Sustainable AI Group’s public dashboard ranks models’ energy intensity per “typical” session, whether chat or agentic. It defines a typical chat session as “a short back-and-forth” with “a question, an answer, and a follow-up or two to refine or clarify it,” whereas a typical agentic session is “an hour or two of the assistant reading files, making changes and checking its own work across a project.” There are also results for a “heavier” or “lighter” session — generally tasks that take more or less time or require greater or fewer back-and-forths with the AI.

The least efficient AI model for both a typical chat and agentic session, per the dashboard, is Anthropic’s Claude Fable 5. A typical agentic session uses 76 watt-hours, according to the Sustainable AI Group’s estimate, or about the amount of electricity it would take to charge four smartphones, per Department of Energy estimates. The most efficient model for a typical chat session was Claude Haiku 4.5, while the most efficient model for a typical agentic session was Open AI’s GPT-5 nano.

Jegham said the point of the dashboard is not to villainize particular companies or models or to argue that more efficient models are superior. He acknowledged that a more complex task may require a larger model, and a larger model is likely going to be more energy intensive than a smaller one.

The ranking is also flawed in that it assumes every model delivers responses with the same amount of verbosity. In reality, some models may use more words, and therefore more tokens, to answer the same question. Jegham gave the example of Anthropic’s Sonnet and Opus models: Sonnet is less energy intensive per token, but it typically requires more tokens for the same task, so sometimes it’s more energy intensive than Opus. The dashboard doesn’t reflect these differences.

While energy intensity is the core of the dashboard’s function, it also includes estimates of each model’s carbon emissions per session. That calculation opens up many more cans of worms, since actual emissions depend on where in the country the hardware that’s processing the AI session is located and what’s powering it. There’s no easy way to know which data center is processing a given AI request. Instead, the dashboard offers users the option to toggle between different emissions intensities to reflect different scenarios — a data center powered by behind-the-meter natural gas plants, for example, versus one located on a relatively clean grid.

A typical agentic session with Claude Fable 5 powered by a behind-the-meter gas plant emits roughly 52 grams of CO2, it says, while a heavy session emits just over 200 grams — equivalent to driving about half a mile in a gasoline-powered vehicle.

I reached out to OpenAI and Anthropic to ask why they don’t publish energy intensity data, whether there are barriers to doing so, and whether they have plans to do so in the future. A spokesperson from OpenAI told me the company relies “on infrastructure partners to operate the data centers that run our models, so we don’t directly collect the underlying energy data. That’s an important consideration in how we assess and provide this information.” Anthropic declined to comment.

Google, on the other hand, has published an energy use estimate for “the median Gemini Apps text prompt in May 2025,” but has not provided an update for subsequent model versions. In response to my request for comment, the company reiterated statements from Cooper Elsworth, a senior technical manager for AI energy, which Google shared with me for a previous story on Watershed’s efforts to calculate AI-related emissions. He said there is no industry consensus for how to measure and disclose the environmental footprint of frontier AI models. He also echoed OpenAI’s comments, noting that gathering accurate energy use data requires “highly advanced measurement infrastructure,” which not all AI providers have access to.

“We believe there is immense value in aligning the industry on comparable metrics to fairly compare and incentivize action,” he said.

Yellow

You’re out of free articles.

Subscribe to access Heatmap’s expert analysis of energy, climate change, and sustainability, including coverage of our regular survey research. Save $57 on an annual subscription, just $156 $99/year.
To continue reading
Create a free account or sign in to unlock more free articles.
or
Please enter an email address
By continuing, you agree to the Terms of Service and acknowledge our Privacy Policy
AM Briefing

Clinching Deals

On solar pivots, NYPA solar, and a receding Rhine

The Clinch River nuclear site.
Heatmap Illustration/Tennessee Valley Authority

Current conditions: Last weekend’s nor’easter caused up to $13 billion in damages across the Mid-Atlantic and Northeast regions of the United States • Hurricane Nolo shut down a major highway on Hawaii’s Big Island • A heat dome forming over eastern Africa is driving temperatures in Juba, the impoverished capital of South Sudan, past 100 degrees Fahrenheit.


THE TOP FIVE

1. The Senate reaches a permitting deal

At last, right after hopes dimmed, we have a deal. Senate negotiators reached a bipartisan agreement on a package of federal permitting reforms, locking in what Politico described as “the contours of long-sought legislation to speed up approvals for new energy projects in the U.S.” Democratic negotiators Senators Martin Heinrich of New Mexico and Sheldon Whitehouse of Rhode Island told the news outlet they were withholding endorsements of a final deal as “the last five yards” of the agreement are hammered out. Whitehouse cautioned that he needed “more clarity from the Trump administration” on what their easing of the blockade on wind and solar approvals would mean. Neither Democrats nor Republicans released text of the bill, which both parties said should come out this week.

Keep reading...Show less
Green
Daily Briefing

The Long-Term Risk of Trump’s Fuel Economy Rollback

By neutering the Corporate Average Fuel Economy standards, the Trump administration cements the country’s dependence on oil and liquid fuels.

Trump getting out of a Tesla.
Heatmap Illustration/Getty Images

This is Heatmap Daily, a weekday news digest written by our executive editor.

President Trump’s big fuel efficiency rollback is here. This afternoon, the Department of Transportation significantly weakened the Corporate Average Fuel Economy standards, the federal government’s rules that encourage new cars and trucks to get gradually more fuel-efficient over time. Instead of mandating that new cars and trucks hit a target of more than 50 miles per gallon, as the old Biden-era rules had required, new vehicles sold in the U.S. will now need to average only 34.9 miles per gallon.

Keep reading...Show less
Ideas

Should You Trust a Politician’s Pivot on Data Centers?

The cofounders of The Impact Project have a three-step test for voters.

Greg Abbott.
Heatmap Illustration/Getty Images

In November 2025, Texas Governor Greg Abbott announced a $40 billion Google investment in his state and declared, “Texas is the epicenter of AI development, where companies can pair innovation with expanding energy.” At a campaign stop in East Texas seven months later, he had a different message: “We must prohibit them from building AI data centers in rural Texas neighborhoods.” Last week, Abbott instructed Texas’ environmental agency to stop issuing permits to data center projects until the state’s grid operator completes an audit of all data centers in the interconnection process.

Abbott is not alone. In the past week, three other candidates for governor moved toward limits. On September 23, Maryland Governor Wes Moore, a Democrat, signed an executive order tying state incentives for large projects to a new review process, pledged that “the state will not go around a local community’s ‘no,’” and announced that he would ask lawmakers to repeal the state’s data center tax exemption, passed in 2020. The same day, Kansas Democratic nominee Cindy Holscher, who voted for data center tax incentives as a state senator and now backs a moratorium, said she “certainly would vote differently based on the information we have now.” Teri Ann Hourihan, Arizona’s No Labels candidate, also promised a “Day 1” moratorium on new data centers.

Keep reading...Show less