EcoInference.ai — edge AI research

The Regeneration Tax

What the Shift from Film to Digital Teaches Us About AI's Hidden Cost Multiplier

1. Two Photographers

I've spent a lot of years behind a camera — first film, then digital — and the two made me into two different photographers.

With film, I shot less and cared more. A roll gave you 24 or 36 exposures, and every one cost real money: the film, the processing, the prints. That cost put the discipline in front of the shutter, where it belonged. You metered the light. You waited for the moment. You thought about the picture before you took it, because a bad frame wasn't just a bad frame — it was money wasted.

Digital broke that discipline. With nothing riding on each frame, I stopped being careful, and so did everyone else. Now I shotgun my shots — dozens of frames of roughly the same thing, trusting volume to get me to a good one. And it works, sort of: I end up with more good pictures than I ever did on film. But it traded one problem for another. I have a hard drive full of near-duplicates and no patience to sort them. The fix is well known — serious photographers develop the discipline to shoot deliberately and cull ruthlessly, even when the frames are free. Most people never get there. I haven't.

I'm telling you this because it's not really a story about photography. It's a story about what happens to human behavior when cost disappears — and it's playing out again, at far larger scale, every time someone clicks "regenerate."

Generating an AI image or video feels like digital photography: free, instant, consequence-free. So people shotgun it — run the prompt again, run it differently, run it six more times until something looks right. But it only feels free. On the other end of that button is a data center drawing real power and evaporating real water on every attempt, whether you keep the result or throw it away. From your chair, it's digital. From the data center's chair, it's film — every frame has a cost. The difference: film's price was printed right on the box. The price of a generation is printed nowhere.

This paper is about that gap: what happens when the cost of a wasted attempt disappears from the view of the person making it — but doesn't disappear at all.

One thing to state plainly before we begin: this paper proposes no solutions. Decisions about pricing, interface design, and data center operations belong to people with expertise and data we do not have. What we do here is describe, in plain language, a problem that is real, large, measurable, and almost entirely missing from public conversation — then lay out the questions the people who do have the data should be asked, by the press, by policymakers, and by each other. Our single request of the industry is disclosure. The rest is a conversation we aim to start, not settle.

What this paper argues. Every published per-output cost figure for AI — including the ones in our three prior papers — is priced per kept output. Nobody keeps the first roll. The true cost of the image you posted, the clip you published, the reply you pasted is that cost times the number of attempts it took — and no one who knows that number will say it out loud.

2. A Few Terms, in Plain Language

This paper is written for non-technical readers — above all the reporters and policymakers we hope will pick it up. Everything that follows relies on the short list of terms below — and on nothing more technical than that.

Term What it means in this paper
ModelThe AI program itself — ChatGPT, Gemini, Veo, Midjourney. Software trained on vast amounts of data to produce text, images, video, or code on request.
PromptThe instruction a person gives the AI: the typed question, or the description of the image or video you want.
InferenceOne run of the model — the work of turning a prompt into an output, performed on servers in a data center. Every reply, image, or clip is one inference. (Training a model is a separate, one-time cost; this paper is about inference, which happens billions of times a day.)
Generation runOne complete inference producing one output. A "full generation run" means the data center did all the work and spent all the energy, whether or not anyone kept the result.
Regeneration / re-rollAsking the AI for another version of the same request. Each one is a new, full-price inference.
GPUGraphics processing unit — the specialized, power-hungry chip that does AI's mathematical heavy lifting. A data center runs thousands of them.
Data centerA warehouse-scale building full of servers. It consumes electricity to run them and, in most designs, water to cool them.
Watt-hour (Wh) / kilowatt-hour (kWh)Units of energy. Running a 10-watt LED bulb for an hour uses 10 Wh. A kilowatt-hour is 1,000 Wh; a typical U.S. household uses about 30 kWh per day.
JouleThe basic scientific unit of energy — 3,600 joules make one watt-hour. Used here simply to mean "every last bit of energy."
ModalityA category of AI output: text, images, video, or code.
StitchingJoining several short AI-generated clips into one longer video. AI models generate a few seconds at a time, so "a 30-second AI video" is really several clips stitched together — the subject of our previous paper.

3. Defining the Regeneration Tax

The regeneration tax is the ratio between the number of generation runs a system actually executes and the number of outputs a person actually keeps, uses, or publishes. If it takes four attempts to get an image worth posting, the regeneration tax on that image is 4x. Every joule and every drop of water spent on the three discarded attempts belongs on the bill of the one that survived.

Three properties make this tax more consequential than it first appears.

First, a regeneration is a full-price inference. There is no discount for rejection. The image you glance at and dismiss in half a second cost the data center exactly what the keeper cost — the same GPU seconds, the same watts, the same cooling. The model has no idea it's producing a reject, and the infrastructure bills accordingly.

Second, the tax compounds. Our previous paper showed that a "single" 30-second AI video is really 4 to 8 separately generated clips stitched together. The regeneration tax applies to each clip individually — the creator re-rolls clip three until the hands look right, re-rolls clip five until the camera move matches. A video that took four attempts per clip across seven clips is 28 full generation runs presented to the world as one output.

Third, the tax is invisible in every disclosure the industry has ever made. When Google reports that a median Gemini prompt consumes 0.24 watt-hours, that figure is per prompt — not per answer the user accepted. Every public number about AI's per-output footprint — ours included — is therefore a floor. The true cost of any output a human actually uses is that floor multiplied by a number no company has ever published.

Why this number matters more than any other. If the regeneration multiplier is 3x, every energy estimate in the public conversation is 3x too low — including ours. The best-documented professional case we found, examined in Section 5, ran above 20x.

4. The Silence, Again

Readers of our first three papers will recognize the shape of what follows. In our first paper, the missing number was energy per query. In our second, data center utilization. In our third, the energy cost of video generation. Each time, the companies measured the number precisely — and did not publish it. This paper makes it four for four: the regeneration ratio is tracked, dashboarded, and reviewed inside every company that operates a generative model, and it has never been disclosed by any of them, for any modality, ever.

This is not speculation: there is direct evidence the number exists and is watched.

When Forbes analyzed the operating cost of OpenAI's video generator Sora, its reporting noted that the estimates did not account for videos blocked before release or drafts that spend credits but never get posted — confirming that discarded generations are a large enough cost to change the math, and that OpenAI tracks kept versus discarded. You cannot exclude a category you don't measure.

In March 2026, Character.AI began rationing regenerations — "swipes," in its interface — citing infrastructure cost. Users widely reported a cap of about ten per day; the company never confirmed the number. Designing that cap required knowing exactly how often users swipe, what each swipe costs, and where the line had to be drawn to protect the margins. The company rationed the behavior. It still has not published the numbers that made the ration necessary.

And the regenerate button itself is one of the most instrumented controls in these products — placed, A/B tested, and optimized, its click rate a core engagement signal. A user who accepts the first output and leaves is, by the internal metrics, a worse user than one who stays and rolls six times. The regeneration ratio isn't merely known. It is being actively managed upward.

Set this against the old economics of a wasted frame. The price of film was printed on the box; the lab handed you an itemized receipt. Photography billed waste to the person who created it, immediately and legibly, and behavior followed. AI generation inverts that arrangement: a flat monthly subscription in front, a metered power feed into a building in someone else's county out back, and no receipt anywhere in between. The waste is not unmeasured. It is measured meticulously — and billed to the atmosphere.

The pattern, four papers running. Every number this series has needed exists inside the companies, at high precision, updated in real time. What has been published, in every case: nothing. At some point the silence stops being an oversight and becomes the disclosure.

5. Video: The Best-Documented Waste

Video is the natural place to start, because our previous paper established that it is already — per generation — the most expensive thing an ordinary person can ask an AI to do: roughly 30 to 3,000 times the energy of a text reply per clip, depending on the model.

The best data point in this entire paper comes from an advertisement. During the 2025 NBA Finals, the prediction market Kalshi aired a fully AI-generated commercial made with Google's Veo 3. The coverage focused on the price — about $2,000 in generation costs and two days of work, against the hundreds of thousands of dollars and months a conventional spot requires. But buried in NPR's reporting was the number that matters here. The ad's creator, veteran advertising director P.J. Accetturo, said plainly: "This took about 300–400 generations to get 15 usable clips."

That is a keep rate of roughly 4 percent — a regeneration tax of 20 to 27x. For every clip that aired, some twenty were generated at full energy cost and thrown away. And this was not a careless amateur burning free credits; it was an experienced professional with published prompts, working efficiently under a deadline. At current model quality, for professional standards, the waste is the workflow.

Put energy numbers on it. A commercial-grade generation call plausibly costs 300 watt-hours to 1 kilowatt-hour (an extrapolation from measured open models — Appendix B). The fifteen aired clips represent perhaps 4.5 to 15 kWh of "visible" generation; the 300–400 runs behind them, something like 90 to 400 kWh — several days to two weeks of a typical U.S. household's electricity, spent on one 30-second ad, with roughly 96 percent of it going to clips no one will ever see. Every published estimate of "what an AI video costs" counts the 4 percent and misses the 96.

Honesty requires the other end of the range too. For casual social content — where the bar is "good enough to post," not "good enough to air during the NBA Finals" — the limited data that exists suggests 1.2 to 1.7 attempts per kept clip. The population-wide multiplier for video therefore sits somewhere between roughly 1.5x and 27x, with essentially no public data describing the middle. That width is not a flaw in our analysis. It is the finding: a factor this large and this material is documented in public by exactly one on-the-record quote from one ad director.

The stitching math, revisited. Paper 3 showed a 30-second video is really 4–8 clips. This paper shows each of those clips may be attempt number twenty. The two multipliers stack — and nobody's estimate counts either one.

6. Code: The Only Modality with Vendor Numbers

Exactly one corner of the industry has published something close to a regeneration rate — almost by accident, in the service of marketing. GitHub, promoting Copilot, has published telemetry on how often developers accept the code suggestions it generates: roughly 26 to 34 percent. Read as a product statistic, that's a success story. Read as an energy statistic, it says that for every suggestion a developer accepts, two to four were generated, displayed, and discarded — a vendor-published regeneration tax of roughly 3–4x. It proves the ratio is knowable and publishable, which makes every other vendor's silence a choice rather than a limitation.

The surrounding evidence suggests the true waste runs higher, because acceptance is not the end of the story. GitClear's analysis of 211 million changed lines of code found "code churn" — code rewritten or deleted within two weeks of being committed — climbing in the AI-assistant era: 5.7 percent in 2024, up from a 3.3 percent baseline. And in Stack Overflow's 2025 survey, 66 percent of developers named "almost right, but not quite" AI solutions as their single biggest frustration, with 45 percent saying debugging AI-generated code costs them more time. Accepted code that gets rewritten two weeks later is a slow-motion regeneration: the inference was paid for, kept a while, then replaced — often by another inference.

The caveats, stated plainly: acceptance rate is a proxy, not a measurement; a rejected autocomplete suggestion costs far less than a rejected video; and the suggestions arrive unsolicited rather than being re-rolled by a user. But that is exactly why the modality matters — it shows the regeneration tax exists even where nobody is "shotgunning." It is built into how these tools are deployed. The deeper economics of agentic AI coding — tools that generate, test, discard, and regenerate entire changes autonomously — deserve a paper of their own, and we intend to write it.

7. Images and Text: Modeling in the Dark

For the two most widely used modalities, we found no vendor-published regeneration data at all — none from Midjourney, OpenAI, Stability, Adobe, or Google on images; none from any chatbot operator on text. This section is therefore different, and we flag it plainly: the per-generation energy science is solid; the multipliers are ours, labeled as scenarios, because the real ones are secrets.

Images. The energy cost of one image generation is well measured: a 2025 academic study (Luccioni et al.) benchmarked seventeen image models across more than nine thousand runs and found a range of about 0.09 to 4.1 watt-hours per image. What is not measured anywhere public is how many generations a user makes before keeping one. The behavioral research is small but telling: one 2025 study found image quality stops improving around the sixth or seventh attempt, while participants kept generating well past that point — the shotgun with no film counter, observed in a lab. And interface defaults set a floor before behavior even enters: Midjourney answers every prompt with a grid of four images, so a user who keeps one image from one prompt has already imposed a 4x tax by design.

Text. The weakest evidence base of the four: no chatbot vendor has published regenerate-button data. What exists is adjacent — studies find most task-oriented chatbot conversations involve multiple rounds of re-prompting and correction, regeneration's close cousin. Per-response energy, at least, is anchored by Google's disclosed 0.24 Wh per median Gemini prompt.

Since no measured multiplier exists for either modality, we model scenarios at 3x, 6x, and 10x — bands bracketed by the one vendor-published proxy (code, at 3–4x), the observed behavioral plateau (6–7 attempts), and a ceiling well below the documented professional video case. These are our assumptions, not measurements, and we would be delighted to see any vendor embarrass us with the real number.

Modality Measured per-generation energy At 3x At 6x At 10x
Text reply (Gemini, vendor figure)0.24 Wh0.72 Wh1.4 Wh2.4 Wh
Image (mid-range diffusion model)~3 Wh9 Wh18 Wh30 Wh
Image (grid-of-4 as the unit)~12 Wh per prompt36 Wh72 Wh120 Wh
Video clip (small open model)~80–90 Wh~250 Wh~500 Wh~850 Wh
Video clip (commercial-grade, extrapolated)~300–1,000 Wh~1–3 kWh~2–6 kWh~3–10 kWh

Table 1: Per-kept-output energy under scenario regeneration multipliers. Per-generation figures from sources in Appendix A; multipliers are the authors' scenario assumptions, not measurements. The video rows understate finished videos, which stitch multiple kept clips (Section 5).

The numbers look small until they meet the volume. In the nine days after OpenAI launched image generation in GPT-4o, users generated over 700 million images — announced proudly, with no estimate of how many were kept. At 3x, roughly 470 million of those images were waste heat with a preview thumbnail; at 10x, 630 million. Somebody at OpenAI knows which. That difference — between merely understated and wildly understated — is currently a number the public is not permitted to see.

A note on method. Where a number does not exist publicly, we model scenarios and label them as such. The companies that could replace our scenarios with measurements are welcome to do so at any time. Nothing would please us more than a correction accompanied by data.

8. The Tool Is Complicit

It would be easy to read this paper as a complaint about careless users. The photography story explains why that reading is wrong. Amateur photographers didn't become careless because their character changed between 1999 and 2005. They became careless because burst mode shipped. The tool taught the behavior.

The same teaching is happening now, deliberately. The regenerate button sits one click from every output — placed, tested, and optimized. "Re-roll" and "variations" are marketed as features, which they are; they are also energy multipliers with marketing budgets. Early research contains a telling result: users given one image at a time regenerate more than users shown a grid — the contact-sheet effect, rediscovered in a lab sixty years after every film photographer knew it instinctively. Interface choices have measurable energy consequences; to our knowledge, no AI product team has ever published an interface decision justified by energy reduction.

And the incentives point the other way. Flat-rate subscriptions make the tenth attempt feel identical to the first — a marginal price of zero that no other metered industrial process offers — while every re-roll registers internally as engagement, the metric product teams are paid to increase. This is not an accusation of intended harm. It is simpler: nobody in the loop has a reason to care. The user can't see the cost, the product team is rewarded for the behavior, and the utility bill is somebody else's line item. Systems shaped like that do not self-correct. They get asked questions from outside — which brings us to the questions.

9. The Questions That Deserve Answers

Our role is not to design fixes. We claim expertise adjacent to this topic — decades in large-scale software infrastructure, and three papers on AI's resource footprint — not in energy science, interface research, or data center operations. The people with the data and the real experts need to have this conversation with each other, in public. What we can supply is the questions — sorted by who should be answering, and who should be asking.

For the AI vendors

To be asked by reporters, by analysts on earnings calls, and by policymakers in hearings:

For the research community

For the press and for policymakers

None of these questions is hostile, and every one can be answered by people who already have the data. Our request is that they be asked — repeatedly, in public, by people whose questions cannot be ignored — until the answering starts.

10. My Actual Take

Since these papers have made a habit of ending with what I actually think, here it is.

The discipline of the film era wasn't virtue. It was priced-in scarcity — the cost of waste was visible, immediate, and billed to the person creating it, and behavior followed the price tag. Nobody needed a sustainability report from a film lab. The price on the box did the regulating.

The AI industry has built the precise inversion of that system. The most wasteful possible behavior — roll it again, roll it again — has been made the most frictionless, celebrated as a feature, priced at zero at the margin, and measured internally as success. And the companies operating the meters know the true number to two decimal places — they have simply never said it out loud. People shotgun because the tool told them shooting was free. It isn't. Somebody's aquifer is in the viewfinder.

And the sharpest point from our video paper belongs here too: the regeneration tax is steepest exactly where the output matters least. Nobody re-rolls a tumor-detection model twenty times chasing a funnier tumor. The twenty-seven-attempt workflows live in advertising, in content farms, in political rage-bait — the disposable end of the content spectrum, sold to us as democratized creativity at a 3-to-25x hidden markup.

To be clear about where this criticism comes from: not the anti-AI corner. I think AI is one of the great boons to humanity in my lifetime — I co-write these papers with one — and I want to see it used for the benefit of all. That is exactly why the waste matters. Every kilowatt-hour spent on the twenty-sixth discarded re-roll of a meme video burns social license — grid headroom, the patience of communities hosting data centers, goodwill in legislatures — that AI's genuinely transformative uses are going to need. The real threat to AI's future isn't the people pointing at the waste — it's the waste itself. Left invisible long enough, it will bring a backlash that takes down the good uses along with the bad. Transparency now is how this technology keeps its welcome.

I've spent forty-plus years building large-scale software systems, which qualifies me to recognize this problem and describe it honestly — not to design the fix. That belongs to the people inside these companies who already stare at the real numbers, the researchers who measure these systems for a living, and the policymakers whose job is to weigh public costs. My job, in these papers, is to make the problem legible enough that those people can no longer talk past it. What I want isn't my solution adopted. It's the conversation held, in public, with the data on the table.

11. Conclusion: The Contact Sheet

Film photography had an instrument of honesty called the contact sheet: every frame on the roll, keepers and mistakes alike, printed small on a single page. It existed for a practical reason — you had to see everything to choose anything — but it had a moral side effect. It confronted you with your own waste. Twelve nearly identical frames, eleven of them money wasted, staring back at you at once. You couldn't not know.

Every AI company has the contact sheet. It is reviewed weekly, in dashboards, by people whose job is to increase the number of frames — kept and discarded, by modality, by product, at a precision film photographers never dreamed of. The industry has simply decided that the rest of us — whose grids are being expanded, whose water is being signed away, whose per-query "transparency" figures are off by an undisclosed multiple — don't get to see it.

This paper had one purpose: to establish that the multiplier exists, that it is large, that it is known, and that it is withheld — and to put the right questions in the hands of the people positioned to ask them. Somewhere between 1.5x and 27x, everything currently believed about the per-output cost of generative AI is an undercount. Narrowing that range requires nothing to be invented. It requires only that someone who already has the number decide, or be pressed, to share it.

Show us the contact sheet.

Appendix A — Sources and Claims

Claims are graded in four tiers: [V] vendor-published, [R] on-record reporting, [A] academic (noting scale), [M] authors' modeling/extrapolation.

# Claim Tier Source
A1Kalshi NBA Finals ad took "about 300–400 generations to get 15 usable clips" (~4% keep rate)[R]P.J. Accetturo, quoted by NPR, June 23, 2025
A2The ad cost ~$2,000 in generation and took two days[R]NPR, June 23, 2025
A3Forbes' Sora cost analysis notes estimates exclude scrapped/unposted videos[R]Forbes reporting on Sora operating costs, Nov. 2025
A4Character.AI metered "swipes" (March 2026) citing infrastructure cost; ~10/day cap user-reported, never company-confirmed[R]Company statements / coverage, March 2026
A5GitHub Copilot suggestion acceptance rate ~26–34%[V]GitHub published research and telemetry
A6Code churn rose from ~3.3% baseline (2020–22) to 5.7% in 2024 (211M changed lines)[A]GitClear AI Copilot Code Quality report, 2025
A766% cite "almost right, but not quite" AI code as top frustration; 45% say debugging it takes more time[A]Stack Overflow Developer Survey 2025
A8Per-image energy 0.09–4.1 Wh across 17 diffusion models (46x spread); resolution multiplies cost up to ~5x[A]Luccioni et al., 2025 (arXiv:2506.17016; 9,000+ measured runs)
A9Image quality improvement plateaus around iteration 6–7 while users continue generating[A]2025 HCI study (20 participants — small-N, directional)
A10Single-image interfaces induce more regeneration than grid interfaces[A]2025 sustainability/HCI study (qualitative)
A11Google: median Gemini prompt = 0.24 Wh (incl. PUE and idle overhead)[V]Google technical disclosure, Aug 2025
A12Open-model video generation: ~3–360 Wh per clip depending on model; WAN2.1-1.3B ≈ 80–90 Wh[A]Measured open-model benchmarks (Paper 3, Appendix A)
A13Commercial-grade video generation ≈ 300 Wh–1 kWh per call[M]Authors' extrapolation from A12 (Paper 3 methodology)
A14Casual video generation ≈ 1.2–1.7 attempts per kept clip[A/M]Vendor benchmark blog (500 generations, 6 models) — low confidence
A15700M+ images generated by 130M+ users in ~9 days after GPT-4o image launch[V]OpenAI statements, reported April 2025
A163x / 6x / 10x multipliers for images and text[M]Authors' scenario bands; see Appendix B

Appendix B — Methodology

Scenario bands (3x/6x/10x). No measured regeneration ratio exists for images or text. Our bands are anchored at the low end by the one vendor-published proxy (Copilot acceptance, implying 3–4x), in the middle by the observed behavioral plateau (attempts 6–7) and interface floors (Midjourney's grid of four), and at the high end by a figure well below the only documented professional case (20–27x). We consider 3x conservative for any modality with a one-click regenerate control.

Video range endpoints. The 20–27x ceiling derives from A1 (300–400 generations ÷ 15 keepers). The ~1.5x casual floor derives from A14, which we grade low-confidence: it originates from a commercially interested party, at sample sizes we cannot audit. The population-wide mean is unknown; we deliberately decline to invent one.

Composition with stitching. Paper 3 established that a finished ~30-second video comprises 4–8 generation calls. The regeneration tax applies per call: total runs ≈ (clips stitched) × (attempts per kept clip). The Kalshi case is consistent — 15 kept clips × ~23 average attempts ≈ 345 generations, matching the reported 300–400.

Energy figures. Per-generation values in Table 1 carry over from Paper 3's Appendix A (open-model measurements; commercial extrapolation) and Luccioni et al. 2025 for images. Google's 0.24 Wh Gemini figure is used for text as the only vendor-published, full-overhead per-prompt number in existence. Household comparisons use standard U.S. averages.

What we did not do. We did not estimate a global aggregate energy cost of regeneration. That requires the population-wide multiplier — precisely the undisclosed number this paper is about. Producing a headline global figure from our own scenario assumptions would launder an assumption into a statistic, and we decline.

Appendix C — Limitations

The central limitation is the paper's subject: the regeneration ratio is unpublished for three of the four modalities discussed, and our scenario bands are placeholders for measurements the industry declines to release. Nothing in Section 7's table should be quoted as measured fact; the [M] tiers in Appendix A mark every number that is ours rather than the world's.

Specific cautions: the Kalshi data point (A1) is a single professional production at broadcast standards — an outlier by construction. Copilot acceptance (A5) measures unsolicited completion suggestions, which are cheaper per rejection than image or video generations. The behavioral studies (A9, A10) are small-N and lab-based. The casual-video figure (A14) comes from a commercially interested party. Google's 0.24 Wh (A11) is a median, and medians conceal the long tail of heavy use.

Finally, on expertise: the authors' background is in large-scale software systems and in the three preceding papers — not energy science, HCI research, or data center operations. This paper's claims about what those experts' data would show are questions, not predictions. The paper succeeds not if our scenario numbers prove right, but if the people holding the real ones are finally asked to produce them.

A note on this document: This white paper was researched and written collaboratively by Mark J. Divitt and Claude, an AI assistant made by Anthropic — a fact we disclose in every paper in this series, and which we note is itself an instance of the subject matter: this document took multiple drafts, revisions, and regenerations to produce. We have tried to spend them honestly. It is the fourth in a series, following The Case for Greener AI, AI Data Center Overbuild, and The Price of a Throwaway Video.

August 2026  |  Version 1.0  |  info@ecoinference.ai