EcoInference.ai — edge AI research

The Price of a Throwaway Video

What 30 Seconds of AI-Generated Video Actually Costs in Energy and Water — and Where Most of That Cost Is Wasted

1. Introduction

The first paper in this series, The Case for Greener AI, looked at the everyday case: a chat session, a quick question, the kind of AI use that happens billions of times a day. The second, AI Data Center Overbuild, looked at the infrastructure being built to serve all of it, and who ends up paying for capacity that may not be strictly necessary. This paper narrows the lens to a single, specific, and unusually expensive case: AI-generated video.

Video is worth its own paper because it breaks the pattern the rest of this series has established. A text query costs a fraction of a watt-hour. An image costs a few watt-hours. A single short AI-generated video clip — by the best available independent measurement — costs on the order of thirty times more energy than that image, and thousands of times more than the text query. And unlike a chatbot reply, which at minimum attempts to answer a question, much of the AI video being produced right now exists purely as disposable social media amusement or as ammunition in political arguments. That combination — the highest resource cost in mainstream generative AI, spent overwhelmingly on the lowest-stakes content — is what this paper is about.

A note from Mark: There is a second irony here, on top of the one I raised in the first paper. Producing this paper involved not just a handful of AI chat sessions, but a large automated research effort — on the order of a hundred coordinated AI sub-agents fanning out across the web, fetching sources, and cross-checking each other's claims before a word of this was written. That process has its own real energy cost, almost certainly larger than what it takes to research a typical white paper. I don't think that undercuts the argument; if anything it sharpens it. We used a meaningful amount of compute to produce a rigorously fact-checked case for using compute more carefully. I'd rather be honest about that trade than pretend it doesn't exist.

What this paper argues. No AI video vendor discloses per-video energy or water figures. The best independent evidence that does exist shows video generation costs one to several orders of magnitude more than text or image generation, that a realistic "30-second video" is actually several short clips stitched together rather than one continuous render, and that a large share of this expensive output is spent on content with no lasting value. We think that is worth saying plainly, not just measuring quietly.

2. The Disclosure Gap

Before any numbers can be discussed, one fact has to be established: essentially none of the companies running commercial AI video generation at scale — Google (Veo), OpenAI (Sora), Runway, or Kling — have published a per-video energy, water, or compute figure for their production models. This is a meaningfully different situation from text generation, where OpenAI and Google have both released per-query figures, however imperfect.

Three independent checks confirm the same gap from different angles:

Why the silence matters. When a company discloses a number, however imperfect, it invites scrutiny and comparison. When it discloses nothing, the absence itself becomes informative: it is far easier to stay quiet about a cost that would look bad than one that would look good. Every figure in the rest of this paper is a proxy for the real number — not because we chose to estimate, but because estimating is the only option available.

3. The Best Real Measurements We Have

The single strongest piece of evidence in this entire research area is a genuine academic study: "Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models" (Delavande, Pierrard, and Luccioni of Hugging Face; presented at a NeurIPS 2025 workshop). Unlike everything else available on this topic, it directly measured GPU energy on real hardware for seven actual video models running at their own default settings.

Model Measured GPU energy Generation time
AnimateDiff0.115 Wh0.68 sec
LTX-Video-0.9.7-dev3.16 Wh9.7 sec
CogVideoX-2b8.3 Wh50.6 sec
CogVideoX-5b21.6 Wh124 sec
Mochi-1-preview44.7 Wh263 sec
WAN2.1-T2V-1.3B78.8 Wh410 sec
WAN2.1-T2V-14B359.7 Wh1,875 sec

Table: Measured GPU energy and generation time for seven open text-to-video models at default settings. Source: Delavande, Pierrard & Luccioni, "Video Killed the Energy Budget" (arXiv:2509.19222, Sept. 2025).

That is nearly a 3,000-fold spread between the cheapest and most expensive model — at each model's own default settings, before anyone has touched the duration or resolution dials. Which model happens to be running under the hood matters as much as anything else in this paper.

Separately, MIT Technology Review's own direct measurement of CogVideoX (a different sub-study using the same CodeCarbon measurement approach, on NVIDIA H100 GPUs) found an older, lower-quality version of the model at roughly 30 Wh per video, and a newer, better version at roughly 942 Wh per video — a thirty-fold jump for what the reporters describe as a fairly modest quality improvement. The two studies used different settings and are not directly comparable figure-for-figure, but they agree on the same underlying lesson.

Key finding. Absolute energy cost for "an AI video" swings by one to three orders of magnitude depending on model size and quality settings alone — all before duration or resolution enters the picture. There is no such thing as a single trustworthy number for "what an AI video costs." There is only a range, and the range is enormous.

4. How Video Compares to Text and Image

This is the comparison at the heart of this paper, because the same Hugging Face study did it explicitly and put a number on it:

Direct quote from the study. "Generating a single short video with WAN2.1-T2V-1.3B consumes nearly ~90 Wh. This places video diffusion roughly 30× more costly than image generation, 2,000× than text generation, and 45,000× than text classification." (Note: the paper's own default-settings table lists this same model at 78.8 Wh; the ~90 Wh figure appears to reflect a slightly different configuration used specifically for this comparison. We report both rather than silently picking one.)
Task Typical energy per output What it means in practice
Text classification
(e.g., spam filter)
a small fraction of a Wh The cheapest common AI task there is
Text generation
(a chat reply)
0.047–0.24 Wh Google's own "full-system" Gemini figure (0.24 Wh) is higher because it deliberately counts idle capacity, cooling overhead, and water — not just the GPU
Image generation 0.086–4.08 Wh across 17 diffusion models; ~2.9 Wh baseline Order of magnitude: single digits of a watt-hour
Video generation
(small open model)
78.8–90 Wh per clip ~30× an image, ~2,000× a text reply, per the study's own comparison
Video generation
(larger open model)
359.7 Wh per clip ~125× an image, ~7,650× a text reply, by our own extension of the same baseline figures

The plain-language version of this table is the whole point of the paper: one AI-generated video clip — before anyone stitches several of them together to reach 30 seconds — costs roughly the same energy as somewhere between 30 and 7,000-plus individual text chatbot replies, depending on which video model and settings are being compared. And that comparison uses small, relatively cheap open-source models. Commercial closed models such as Sora and Veo are generally assumed to run larger and at higher resolution than these open baselines — which would push the true multiplier higher, not lower, though again, no vendor discloses enough to confirm this directly.

In plain terms. If a text chatbot reply is a light bulb left on for a few minutes, a single AI video clip is closer to running a load of laundry. Multiply either one by billions of uses a day and the comparison stops being cute and starts being a real line item on a real grid.

5. Duration, Resolution, and What "30 Seconds" Actually Requires

The same study pinned down how energy scales with the two dials that matter most, and the result reframes the entire question of what a 30-second video actually costs.

Duration scales worse than linearly

Doubling the frame count — that is, doubling duration at a fixed frame rate — roughly quadruples energy. Energy tracks closer to duration squared than to duration itself.

Resolution scales similarly

Doubling spatial resolution in each dimension (four times the total pixels) also roughly quadruples energy. Doing both at once compounds — doubling duration and resolution together multiplies energy by roughly sixteen.

Denoising steps scale linearly

Twice the quality-refinement steps, twice the energy — no surprise there, and no super-linear penalty.

Here is the part that actually matters for a real 30-second Facebook or YouTube clip: nobody generates one continuous 30-second video in a single pass. Every major commercial tool caps a single generation somewhere between four and ten seconds, then optionally "extends" it. Google's Veo 3.1 documents this precisely: each extension adds seven seconds, reading the final second of the prior clip as a seed and generating a genuinely new segment from there — a complete regeneration pass, not a cheap continuation. Reaching 30 seconds this way takes roughly four to five separate full generation calls. Reaching Veo's advertised maximum of 148 seconds takes 21 separate calls.

This produces a real, and somewhat counterintuitive, silver lining: because nobody actually renders one continuous 30-second clip, the worst of the quadratic-duration penalty above never gets to fully apply. A single real 30-second render, if anyone actually built one, could cost on the order of 36 times a 5-second clip (six times the duration, squared). Stitching several short clips instead means each call only pays the quadratic penalty on its own ~7 seconds — so the total cost across the whole video comes out closer to (number of clips) × (cost of one short clip): roughly linear in the number of stitches, not quadratic in the final length.

The stitching math. Applying that structure to Veo's own documented mechanics: if one 7-second, 1080p, commercial-quality generation call costs on the order of a few hundred watt-hours to roughly one kilowatt-hour — a reasoned extrapolation from the measured table in Section 3, not a disclosed figure — then a stitched 30-second video assembled from four to five calls lands in the range of roughly 2 to 10 kilowatt-hours, and a full 148-second video from 21 calls could run 6 to 20-plus kilowatt-hours. The structure of this math (linear in the number of stitches) is solid and directly sourced; the specific per-call anchor number is our extrapolation, clearly labeled as such.

One wrinkle worth stating plainly: the extend feature gets no hidden efficiency discount for reusing context. It is a complete regeneration pass every single time, seeded from the previous clip's final frame but not meaningfully cheaper because of it.

6. What People Actually Generate

It is worth being precise about what a "typical" AI video actually is, because both the default length and the default resolution are commonly misunderstood.

Length

Kling AI defaults to a 5-second generation, 10 seconds maximum per call. Google Veo 3.1's native single-clip lengths are 4, 6, or 8 seconds, extendable in 7-second increments. Runway's gen4_turbo is billed per second of output ($0.05/sec), which pushes users toward short clips to control cost. None of the major tools default anywhere near 30 seconds.

Resolution

720p is increasingly the bargain-bin tier, not the standard. Across pricing and quality comparisons for Kling, Runway, and Veo, 1080p is the reference tier that plans are built around; 720p is reserved mainly for free or watermarked output, and 4K is a premium add-on — Veo 3.1 charges an extra $0.60 per second for native 4K. Anyone actually posting a finished clip to Facebook or YouTube, rather than testing a free tier, is more realistically working in 1080p than 720p.

So the realistic mental model for "a 30-second AI video posted to social media" is: four to six separate model generation calls, at 1080p, each one a full independent compute job — not one continuous render. In one sense that is arguably worse for total resource use: there is no economy of scale, and the short-clip overhead is paid every time. In another it is better: it avoids the worst of the quadratic duration penalty a single continuous 30-second render would incur, as shown in Section 5.

7. Numbers That Did Not Survive Fact-Checking

In researching this paper, we passed every claim through an adversarial verification process: independent reviewers, working blind to each other's conclusions, tried to refute each figure against its original source before it was allowed into this document. A specific set of widely-circulated figures did not survive that process, and because they are likely to resurface elsewhere, we are naming them rather than quietly omitting them.

All three of the following come from a single July 2026 blog post, "Lights, Camera, Carbon," published by a for-profit AI-sustainability consultancy summarizing a paper still described as "undergoing peer review":

A note on method. We are flagging these not to single out one consultancy, but because the incentive problem is structural: firms that sell AI-sustainability advisory services have a business reason to produce dramatic, citable numbers, and dramatic numbers travel faster online than careful ones. The fix is not to stop citing industry estimates — it is to check them before repeating them, which is what we tried to do throughout this paper.

8. Hyperscaler Infrastructure Context

Since the video-generation figures above are raw compute energy, it is worth stating what sits on top of them in an actual hyperscaler facility.

Put together: total facility energy for a video-generation job is approximately (GPU and server energy) multiplied by PUE (roughly 1.1 to 1.2). Cooling water is roughly WUE multiplied by that total, plus the off-site "embedded" water consumed generating the electricity in the first place — and by Google's own comprehensive text-query methodology, that embedded share roughly doubles the water footprint beyond on-site cooling alone.

9. Where the Waste Actually Is

Everything to this point has been an attempt at careful measurement under real uncertainty. This section is not that. This is our judgment, and we think a white paper that hides behind neutral tone when the evidence points somewhere specific is doing its readers a disservice.

For the overwhelming majority of what is actually being generated — meme clips, "satisfying" filler content, engagement bait tuned for a recommendation algorithm, and the flood of AI-generated video that partisans on both ends of the political spectrum now use to make the other side look stupid or evil — this is a genuinely poor use of a real, finite resource. Each of these clips spends the energy-equivalent of dozens to thousands of ordinary chatbot exchanges, run on hardware that pulls real power off a real grid and real water for cooling, to produce something that gets watched once, maybe screenshotted for a caption, and forgotten by the next scroll. A chat reply at least attempts to answer a question. A throwaway video clip's entire value proposition is whether it earned a reaction in the three seconds someone happened to look at it — competing for that same three seconds against every other piece of content that cost almost nothing to make.

The political content is the part that troubles us most. A large share of current high-volume AI video use is rage bait: fabricated or exaggerated clips engineered to make one political tribe look foolish or dangerous to the other, optimized for shares rather than accuracy, produced by the left and the right in roughly equal measure — we are not letting either side off the hook here. That is not "using AI to solve a problem," and it is not even "using AI to be creative." It is spending real watts and real water to make people angrier at each other, more efficiently than before. There are real, defensible uses of this technology — accessibility, prototyping, film and game production, education — where the resource cost is a reasonable trade for value actually created. Disposable social media amusement and political rage content mostly are not that trade, and we do not think it is unreasonable to say so plainly, in a paper that otherwise tries hard to be measured.

The core problem. The highest per-output resource cost in mainstream generative AI is being spent, in large volume, on the lowest-stakes content it produces. That mismatch — not the existence of AI video generation itself — is the actual waste this paper is describing.

10. A Path Forward

None of this is an argument for banning AI video generation, and we are not making that argument. It is an argument for matching real resource cost to real value created — the same argument this series has made about chat and about data center construction, now applied to the most expensive mainstream case.

What follows from this

This is the same direction EcoInference.ai has argued from the start of this series: that a meaningful share of AI's footprint is not an inevitable cost of the technology, but a consequence of routing work to the heaviest available tool by default rather than by need. Video generation is simply where that gap between cost and value is currently at its widest. Closing it does not require anyone to stop using AI video — it requires being honest, as an industry and as individual users, about which 30 seconds are actually worth what they cost.

The technology to make video generation efficient enough to be nearly free may well arrive. The technology to make throwaway content worth generating in the first place will not. That is not an energy problem. That is a taste problem — and no efficiency gain will fix it for us.

Appendix A — Primary Sources

All figures reflect data available as of mid-2026 and should be re-checked against the latest releases.

#SourceData point used
A1Delavande, Pierrard & Luccioni, "Video Killed the Energy Budget," arXiv:2509.19222, Sept. 2025Per-model GPU energy/latency table; duration² and resolution² scaling laws; 30×/2,000×/45,000× video-vs-image-vs-text comparison
A2MIT Technology Review, "We did the math on AI's energy footprint," May 2025CogVideoX direct measurements (~30 Wh old version, ~942 Wh new version); "No Energy Data" for Sora, Veo 2, Adobe Firefly
A3MIT Technology Review, methodology companion piece, May 2025GPU energy as ~50% of total server energy demand, doubled to estimate full system cost
A4Hugging Face AI Energy Score (huggingface.github.io/AIEnergyScore)H100-only benchmark methodology; confirms no video-generation task category exists
A5Google Cloud Blog, "Measuring the environmental impact of AI inference," Aug. 20250.24 Wh per median Gemini text prompt (comprehensive full-system methodology); fleet PUE 1.09
A6Google 2025 Environmental ReportConfirms no per-query/per-video figures disclosed anywhere in the report; fleet PUE 1.09
A7Microsoft 2025 Environmental Sustainability Report; datacenters.microsoft.com efficiency pagePUE 1.17 (FY25); WUE 0.27 L/kWh (FY25); 125,000+ m³/year water savings from liquid cooling
A8NVIDIA H100 product brief (PB-11133-001)PCIe variant TDP 350W vs. SXM5 variant TDP up to 700W
A9Runway API pricing documentationgen4_turbo billed at $0.05 per second of output video
A10Kling AI video length documentationDefault generation length 5 seconds, 10-second maximum; free tier 720p, paid tiers 1080p
A11Google Veo 3.1 continuation/extend feature documentation7-second extension increments; seed-frame continuity mechanism; up to 20 extensions (148 seconds total)
A12AI video model comparison, 2026 (industry blog aggregating vendor pricing pages)1080p as standard/reference pricing tier across vendors; Veo 3.1 native 4K at $0.60/sec premium
A13Bertazzini et al., "The Hidden Cost of an Image," arXiv:2506.17016, June 2025Per-image energy range of 0.086–4.08 Wh across 17 diffusion models, measured on consumer (non-hyperscaler) GPU hardware
A14Luccioni et al., "Power Hungry Processing," arXiv:2311.16863, ACM FAccT 2024Single image generation ≈ energy to fully charge a smartphone; 1,000 SDXL images ≈ CO2 of driving 4.1 miles
A15Carbon Trust, "The carbon impact of AI video generation," June 2026AI video generation requires "at least two orders of magnitude more energy" than a text query; 50–100 gCO2e for one specific 5.4-second, 720p test case

Appendix B — Methodology and Derived Estimates

Video-vs-text/image comparison ratios

The 30×/2,000×/45,000× figures in Section 4 are quoted directly from [A1]. The additional ratio for WAN2.1-T2V-14B (~125× image, ~7,650× text) is our own extension, calculated by dividing that model's measured 359.7 Wh figure by the same paper's baseline text (0.047 Wh) and image (2.9 Wh) figures. We flag this explicitly as our derivation, not a quoted figure from the source.

The stitched 30-second estimate

Per-clip anchor: the measured open-model table in Section 3 spans 0.115–359.7 Wh depending on model and settings, none of which are directly representative of a commercial-quality, 1080p, several-second clip. We anchor our estimate at the higher end of that range — roughly 300 Wh to 1 kWh per 7-second call — reasoning that commercial closed models are generally understood to run larger and at higher resolution than the open models actually measured, though this remains an extrapolation, not a disclosed figure.

Number of calls: Google Veo 3.1's documented mechanics require one initial generation plus extensions in 7-second increments. Reaching 30 seconds requires roughly four to five total calls; reaching the 148-second advertised maximum requires 21 total calls (1 initial + 20 extensions).

Total estimate: (per-call energy) × (number of calls) = roughly 1.2–5 kWh (4–5 calls at 300 Wh–1 kWh) for a 30-second video, and roughly 6.3–21 kWh (21 calls) for a 148-second video. We rounded the 30-second range to 2–10 kWh in the body of this paper to avoid implying more precision than the underlying anchor supports.

Scaling law basis

Duration² and resolution² scaling (Section 5) are taken directly from [A1]'s reported findings: doubling frame count and doubling each spatial dimension each independently produce roughly a fourfold increase in energy, with combined scaling compounding toward roughly sixteenfold. Denoising step scaling is reported as linear in the same source.

Appendix C — Known Limitations and What We Do Not Know

No commercial model data exists

Every specific figure attached to Sora, Veo, Runway, or Kling in this paper is either a pricing signal (a business decision, not an energy measurement) or a third-party extrapolation. Nobody outside those companies knows their actual per-generation energy cost at real production settings.

Model-to-model variance is enormous

The measured table in Section 3 spans nearly three orders of magnitude between models at their own default settings. Any single-number estimate for "what an AI video costs" necessarily obscures this variance, and we have tried to show ranges rather than points wherever possible.

The stitching estimate is an extrapolation

The linear-in-number-of-calls structure of our 30-second estimate is well-supported by the scaling laws and Veo's documented mechanics. The specific per-call anchor (300 Wh–1 kWh) is a reasoned extrapolation from open-model data, not a measurement of any commercial system, and should be treated accordingly.

Retries are not counted

All estimates in this paper are per successful generation. In practice, people regenerate AI video clips multiple times before landing on a version they post. We found no published data quantifying this retry multiplier, so every estimate in this paper should be read as a floor on the true resource cost of a published video, not a ceiling.

Debunked figures are named, not erased

Section 7 names specific figures that failed our adversarial verification process rather than silently omitting them, because we expect readers to encounter them elsewhere. Their inclusion is not an endorsement.

Rapid pace of change

AI video model efficiency, default resolutions, and vendor disclosure practices are all moving quickly. The figures in this paper reflect conditions as of mid-2026 and should be re-checked against current releases before being relied upon for any specific decision.

A note on this document: This white paper was researched and written collaboratively by Mark J. Divitt and Claude, an AI assistant made by Anthropic. It is the third in a series, following The Case for Greener AI and AI Data Center Overbuild. All figures are estimates derived from public sources and reflect data available as of mid-2026; readers are encouraged to consult the primary sources in Appendix A and reach their own conclusions.

July 2026  |  Version 1.0  |  info@ecoinference.ai