The Price of a Throwaway Video
What 30 Seconds of AI-Generated Video Actually Costs in Energy and Water — and Where Most of That Cost Is Wasted
1. Introduction
The first paper in this series, The Case for Greener AI, looked at the everyday case: a chat session, a quick question, the kind of AI use that happens billions of times a day. The second, AI Data Center Overbuild, looked at the infrastructure being built to serve all of it, and who ends up paying for capacity that may not be strictly necessary. This paper narrows the lens to a single, specific, and unusually expensive case: AI-generated video.
Video is worth its own paper because it breaks the pattern the rest of this series has established. A text query costs a fraction of a watt-hour. An image costs a few watt-hours. A single short AI-generated video clip — by the best available independent measurement — costs on the order of thirty times more energy than that image, and thousands of times more than the text query. And unlike a chatbot reply, which at minimum attempts to answer a question, much of the AI video being produced right now exists purely as disposable social media amusement or as ammunition in political arguments. That combination — the highest resource cost in mainstream generative AI, spent overwhelmingly on the lowest-stakes content — is what this paper is about.
A note from Mark: There is a second irony here, on top of the one I raised in the first paper. Producing this paper involved not just a handful of AI chat sessions, but a large automated research effort — on the order of a hundred coordinated AI sub-agents fanning out across the web, fetching sources, and cross-checking each other's claims before a word of this was written. That process has its own real energy cost, almost certainly larger than what it takes to research a typical white paper. I don't think that undercuts the argument; if anything it sharpens it. We used a meaningful amount of compute to produce a rigorously fact-checked case for using compute more carefully. I'd rather be honest about that trade than pretend it doesn't exist.
2. The Disclosure Gap
Before any numbers can be discussed, one fact has to be established: essentially none of the companies running commercial AI video generation at scale — Google (Veo), OpenAI (Sora), Runway, or Kling — have published a per-video energy, water, or compute figure for their production models. This is a meaningfully different situation from text generation, where OpenAI and Google have both released per-query figures, however imperfect.
Three independent checks confirm the same gap from different angles:
- Hugging Face's AI Energy Score. The closest thing to an industry-standard, cross-model energy benchmark runs exclusively on NVIDIA H100 GPUs and covers text generation, reasoning, summarization, classification, image generation and captioning, and speech-to-text. It has no video generation category at all.
- Google's own 2025 Environmental Report. Google's tenth annual environmental report discloses fleet-wide aggregate energy, water, and emissions — but contains zero per-query, per-image, or per-video figures anywhere in the document. Veo is not mentioned.
- MIT Technology Review's 2025 investigation. The most rigorous independent attempt at this problem states plainly that closed commercial video models — Sora, Veo 2, Adobe Firefly — provide "No Energy Data." The reporters had to fall back to an open-source model to get any real measurement at all.
3. The Best Real Measurements We Have
The single strongest piece of evidence in this entire research area is a genuine academic study: "Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models" (Delavande, Pierrard, and Luccioni of Hugging Face; presented at a NeurIPS 2025 workshop). Unlike everything else available on this topic, it directly measured GPU energy on real hardware for seven actual video models running at their own default settings.
| Model | Measured GPU energy | Generation time |
|---|---|---|
| AnimateDiff | 0.115 Wh | 0.68 sec |
| LTX-Video-0.9.7-dev | 3.16 Wh | 9.7 sec |
| CogVideoX-2b | 8.3 Wh | 50.6 sec |
| CogVideoX-5b | 21.6 Wh | 124 sec |
| Mochi-1-preview | 44.7 Wh | 263 sec |
| WAN2.1-T2V-1.3B | 78.8 Wh | 410 sec |
| WAN2.1-T2V-14B | 359.7 Wh | 1,875 sec |
Table: Measured GPU energy and generation time for seven open text-to-video models at default settings. Source: Delavande, Pierrard & Luccioni, "Video Killed the Energy Budget" (arXiv:2509.19222, Sept. 2025).
That is nearly a 3,000-fold spread between the cheapest and most expensive model — at each model's own default settings, before anyone has touched the duration or resolution dials. Which model happens to be running under the hood matters as much as anything else in this paper.
Separately, MIT Technology Review's own direct measurement of CogVideoX (a different sub-study using the same CodeCarbon measurement approach, on NVIDIA H100 GPUs) found an older, lower-quality version of the model at roughly 30 Wh per video, and a newer, better version at roughly 942 Wh per video — a thirty-fold jump for what the reporters describe as a fairly modest quality improvement. The two studies used different settings and are not directly comparable figure-for-figure, but they agree on the same underlying lesson.
4. How Video Compares to Text and Image
This is the comparison at the heart of this paper, because the same Hugging Face study did it explicitly and put a number on it:
| Task | Typical energy per output | What it means in practice |
|---|---|---|
| Text classification (e.g., spam filter) |
a small fraction of a Wh | The cheapest common AI task there is |
| Text generation (a chat reply) |
0.047–0.24 Wh | Google's own "full-system" Gemini figure (0.24 Wh) is higher because it deliberately counts idle capacity, cooling overhead, and water — not just the GPU |
| Image generation | 0.086–4.08 Wh across 17 diffusion models; ~2.9 Wh baseline | Order of magnitude: single digits of a watt-hour |
| Video generation (small open model) |
78.8–90 Wh per clip | ~30× an image, ~2,000× a text reply, per the study's own comparison |
| Video generation (larger open model) |
359.7 Wh per clip | ~125× an image, ~7,650× a text reply, by our own extension of the same baseline figures |
The plain-language version of this table is the whole point of the paper: one AI-generated video clip — before anyone stitches several of them together to reach 30 seconds — costs roughly the same energy as somewhere between 30 and 7,000-plus individual text chatbot replies, depending on which video model and settings are being compared. And that comparison uses small, relatively cheap open-source models. Commercial closed models such as Sora and Veo are generally assumed to run larger and at higher resolution than these open baselines — which would push the true multiplier higher, not lower, though again, no vendor discloses enough to confirm this directly.
5. Duration, Resolution, and What "30 Seconds" Actually Requires
The same study pinned down how energy scales with the two dials that matter most, and the result reframes the entire question of what a 30-second video actually costs.
Duration scales worse than linearly
Doubling the frame count — that is, doubling duration at a fixed frame rate — roughly quadruples energy. Energy tracks closer to duration squared than to duration itself.
Resolution scales similarly
Doubling spatial resolution in each dimension (four times the total pixels) also roughly quadruples energy. Doing both at once compounds — doubling duration and resolution together multiplies energy by roughly sixteen.
Denoising steps scale linearly
Twice the quality-refinement steps, twice the energy — no surprise there, and no super-linear penalty.
Here is the part that actually matters for a real 30-second Facebook or YouTube clip: nobody generates one continuous 30-second video in a single pass. Every major commercial tool caps a single generation somewhere between four and ten seconds, then optionally "extends" it. Google's Veo 3.1 documents this precisely: each extension adds seven seconds, reading the final second of the prior clip as a seed and generating a genuinely new segment from there — a complete regeneration pass, not a cheap continuation. Reaching 30 seconds this way takes roughly four to five separate full generation calls. Reaching Veo's advertised maximum of 148 seconds takes 21 separate calls.
This produces a real, and somewhat counterintuitive, silver lining: because nobody actually renders one continuous 30-second clip, the worst of the quadratic-duration penalty above never gets to fully apply. A single real 30-second render, if anyone actually built one, could cost on the order of 36 times a 5-second clip (six times the duration, squared). Stitching several short clips instead means each call only pays the quadratic penalty on its own ~7 seconds — so the total cost across the whole video comes out closer to (number of clips) × (cost of one short clip): roughly linear in the number of stitches, not quadratic in the final length.
One wrinkle worth stating plainly: the extend feature gets no hidden efficiency discount for reusing context. It is a complete regeneration pass every single time, seeded from the previous clip's final frame but not meaningfully cheaper because of it.
6. What People Actually Generate
It is worth being precise about what a "typical" AI video actually is, because both the default length and the default resolution are commonly misunderstood.
Length
Kling AI defaults to a 5-second generation, 10 seconds maximum per call. Google Veo 3.1's native single-clip lengths are 4, 6, or 8 seconds, extendable in 7-second increments. Runway's gen4_turbo is billed per second of output ($0.05/sec), which pushes users toward short clips to control cost. None of the major tools default anywhere near 30 seconds.
Resolution
720p is increasingly the bargain-bin tier, not the standard. Across pricing and quality comparisons for Kling, Runway, and Veo, 1080p is the reference tier that plans are built around; 720p is reserved mainly for free or watermarked output, and 4K is a premium add-on — Veo 3.1 charges an extra $0.60 per second for native 4K. Anyone actually posting a finished clip to Facebook or YouTube, rather than testing a free tier, is more realistically working in 1080p than 720p.
So the realistic mental model for "a 30-second AI video posted to social media" is: four to six separate model generation calls, at 1080p, each one a full independent compute job — not one continuous render. In one sense that is arguably worse for total resource use: there is no economy of scale, and the short-clip overhead is paid every time. In another it is better: it avoids the worst of the quadratic duration penalty a single continuous 30-second render would incur, as shown in Section 5.
7. Numbers That Did Not Survive Fact-Checking
In researching this paper, we passed every claim through an adversarial verification process: independent reviewers, working blind to each other's conclusions, tried to refute each figure against its original source before it was allowed into this document. A specific set of widely-circulated figures did not survive that process, and because they are likely to resurface elsewhere, we are naming them rather than quietly omitting them.
All three of the following come from a single July 2026 blog post, "Lights, Camera, Carbon," published by a for-profit AI-sustainability consultancy summarizing a paper still described as "undergoing peer review":
- "A standard 5-second AI video uses 57.5–114.8 Wh" (compared by the source to running a kitchen air fryer for five minutes) — rejected unanimously; the comparison could not be reproduced from any measured source.
- "An 8-second 720p video on Veo 3, Runway Gen-4.5, or Seedance costs 200–400 Wh" — rejected; the figure as characterized could not be located verbatim in the cited source.
- "A 12-second 1080p Sora 2.0 Pro video costs 1,313 Wh" — rejected; the source's own authors label this figure "estimated, not measured," derived from indirect modeling rather than an actual measurement.
8. Hyperscaler Infrastructure Context
Since the video-generation figures above are raw compute energy, it is worth stating what sits on top of them in an actual hyperscaler facility.
- Google's global data center fleet runs at a Power Usage Effectiveness (PUE) of 1.09 — the best in six years — meaning roughly 9% overhead energy for cooling and power distribution beyond the compute itself.
- Microsoft's global fleet runs at a PUE of 1.17 (FY25) and a Water Usage Effectiveness (WUE) of 0.27 liters per kilowatt-hour (FY25, down from 0.30 the prior year). New liquid-cooled facilities save more than 125,000 cubic meters of water per site per year versus older evaporative designs.
- A common and consequential mix-up: NVIDIA's H100 PCIe card draws 350 watts — a figure that circulates constantly online — but the SXM5 variant, which is what hyperscalers actually deploy in multi-GPU AI clusters, draws up to 700 watts, double the PCIe figure. Any energy estimate anchored to "H100 = 350W" is quietly using the wrong hardware for this context.
Put together: total facility energy for a video-generation job is approximately (GPU and server energy) multiplied by PUE (roughly 1.1 to 1.2). Cooling water is roughly WUE multiplied by that total, plus the off-site "embedded" water consumed generating the electricity in the first place — and by Google's own comprehensive text-query methodology, that embedded share roughly doubles the water footprint beyond on-site cooling alone.
9. Where the Waste Actually Is
Everything to this point has been an attempt at careful measurement under real uncertainty. This section is not that. This is our judgment, and we think a white paper that hides behind neutral tone when the evidence points somewhere specific is doing its readers a disservice.
For the overwhelming majority of what is actually being generated — meme clips, "satisfying" filler content, engagement bait tuned for a recommendation algorithm, and the flood of AI-generated video that partisans on both ends of the political spectrum now use to make the other side look stupid or evil — this is a genuinely poor use of a real, finite resource. Each of these clips spends the energy-equivalent of dozens to thousands of ordinary chatbot exchanges, run on hardware that pulls real power off a real grid and real water for cooling, to produce something that gets watched once, maybe screenshotted for a caption, and forgotten by the next scroll. A chat reply at least attempts to answer a question. A throwaway video clip's entire value proposition is whether it earned a reaction in the three seconds someone happened to look at it — competing for that same three seconds against every other piece of content that cost almost nothing to make.
The political content is the part that troubles us most. A large share of current high-volume AI video use is rage bait: fabricated or exaggerated clips engineered to make one political tribe look foolish or dangerous to the other, optimized for shares rather than accuracy, produced by the left and the right in roughly equal measure — we are not letting either side off the hook here. That is not "using AI to solve a problem," and it is not even "using AI to be creative." It is spending real watts and real water to make people angrier at each other, more efficiently than before. There are real, defensible uses of this technology — accessibility, prototyping, film and game production, education — where the resource cost is a reasonable trade for value actually created. Disposable social media amusement and political rage content mostly are not that trade, and we do not think it is unreasonable to say so plainly, in a paper that otherwise tries hard to be measured.
10. A Path Forward
None of this is an argument for banning AI video generation, and we are not making that argument. It is an argument for matching real resource cost to real value created — the same argument this series has made about chat and about data center construction, now applied to the most expensive mainstream case.
What follows from this
- Ask for video the same transparency text already got. OpenAI and Google both eventually published per-query text energy figures. No major vendor has done the equivalent for video, despite video costing far more per output. That should change, and there is no principled reason it hasn't.
- Match the tool to the job — including for video. The heaviest generative tool available is not the correct default for the lightest-stakes content. A meme clip does not need the same model and resolution as a serious creative or commercial project.
- Be intentional, not repetitive. Every regenerated take or near-duplicate prompt is another full generation call, and this paper's own math shows those calls do not get cheaper by being repeated.
- Extend disclosure standards designed for text to video. The Hugging Face AI Energy Score already rates models like appliance efficiency labels for text and image tasks. Extending an equivalent standard to video generation is a solvable problem, not a hypothetical one — the measurement methodology in Section 3 of this paper already exists.
This is the same direction EcoInference.ai has argued from the start of this series: that a meaningful share of AI's footprint is not an inevitable cost of the technology, but a consequence of routing work to the heaviest available tool by default rather than by need. Video generation is simply where that gap between cost and value is currently at its widest. Closing it does not require anyone to stop using AI video — it requires being honest, as an industry and as individual users, about which 30 seconds are actually worth what they cost.
The technology to make video generation efficient enough to be nearly free may well arrive. The technology to make throwaway content worth generating in the first place will not. That is not an energy problem. That is a taste problem — and no efficiency gain will fix it for us.
Appendix A — Primary Sources
All figures reflect data available as of mid-2026 and should be re-checked against the latest releases.
| # | Source | Data point used |
|---|---|---|
| A1 | Delavande, Pierrard & Luccioni, "Video Killed the Energy Budget," arXiv:2509.19222, Sept. 2025 | Per-model GPU energy/latency table; duration² and resolution² scaling laws; 30×/2,000×/45,000× video-vs-image-vs-text comparison |
| A2 | MIT Technology Review, "We did the math on AI's energy footprint," May 2025 | CogVideoX direct measurements (~30 Wh old version, ~942 Wh new version); "No Energy Data" for Sora, Veo 2, Adobe Firefly |
| A3 | MIT Technology Review, methodology companion piece, May 2025 | GPU energy as ~50% of total server energy demand, doubled to estimate full system cost |
| A4 | Hugging Face AI Energy Score (huggingface.github.io/AIEnergyScore) | H100-only benchmark methodology; confirms no video-generation task category exists |
| A5 | Google Cloud Blog, "Measuring the environmental impact of AI inference," Aug. 2025 | 0.24 Wh per median Gemini text prompt (comprehensive full-system methodology); fleet PUE 1.09 |
| A6 | Google 2025 Environmental Report | Confirms no per-query/per-video figures disclosed anywhere in the report; fleet PUE 1.09 |
| A7 | Microsoft 2025 Environmental Sustainability Report; datacenters.microsoft.com efficiency page | PUE 1.17 (FY25); WUE 0.27 L/kWh (FY25); 125,000+ m³/year water savings from liquid cooling |
| A8 | NVIDIA H100 product brief (PB-11133-001) | PCIe variant TDP 350W vs. SXM5 variant TDP up to 700W |
| A9 | Runway API pricing documentation | gen4_turbo billed at $0.05 per second of output video |
| A10 | Kling AI video length documentation | Default generation length 5 seconds, 10-second maximum; free tier 720p, paid tiers 1080p |
| A11 | Google Veo 3.1 continuation/extend feature documentation | 7-second extension increments; seed-frame continuity mechanism; up to 20 extensions (148 seconds total) |
| A12 | AI video model comparison, 2026 (industry blog aggregating vendor pricing pages) | 1080p as standard/reference pricing tier across vendors; Veo 3.1 native 4K at $0.60/sec premium |
| A13 | Bertazzini et al., "The Hidden Cost of an Image," arXiv:2506.17016, June 2025 | Per-image energy range of 0.086–4.08 Wh across 17 diffusion models, measured on consumer (non-hyperscaler) GPU hardware |
| A14 | Luccioni et al., "Power Hungry Processing," arXiv:2311.16863, ACM FAccT 2024 | Single image generation ≈ energy to fully charge a smartphone; 1,000 SDXL images ≈ CO2 of driving 4.1 miles |
| A15 | Carbon Trust, "The carbon impact of AI video generation," June 2026 | AI video generation requires "at least two orders of magnitude more energy" than a text query; 50–100 gCO2e for one specific 5.4-second, 720p test case |
Appendix B — Methodology and Derived Estimates
Video-vs-text/image comparison ratios
The 30×/2,000×/45,000× figures in Section 4 are quoted directly from [A1]. The additional ratio for WAN2.1-T2V-14B (~125× image, ~7,650× text) is our own extension, calculated by dividing that model's measured 359.7 Wh figure by the same paper's baseline text (0.047 Wh) and image (2.9 Wh) figures. We flag this explicitly as our derivation, not a quoted figure from the source.
The stitched 30-second estimate
Per-clip anchor: the measured open-model table in Section 3 spans 0.115–359.7 Wh depending on model and settings, none of which are directly representative of a commercial-quality, 1080p, several-second clip. We anchor our estimate at the higher end of that range — roughly 300 Wh to 1 kWh per 7-second call — reasoning that commercial closed models are generally understood to run larger and at higher resolution than the open models actually measured, though this remains an extrapolation, not a disclosed figure.
Number of calls: Google Veo 3.1's documented mechanics require one initial generation plus extensions in 7-second increments. Reaching 30 seconds requires roughly four to five total calls; reaching the 148-second advertised maximum requires 21 total calls (1 initial + 20 extensions).
Total estimate: (per-call energy) × (number of calls) = roughly 1.2–5 kWh (4–5 calls at 300 Wh–1 kWh) for a 30-second video, and roughly 6.3–21 kWh (21 calls) for a 148-second video. We rounded the 30-second range to 2–10 kWh in the body of this paper to avoid implying more precision than the underlying anchor supports.
Scaling law basis
Duration² and resolution² scaling (Section 5) are taken directly from [A1]'s reported findings: doubling frame count and doubling each spatial dimension each independently produce roughly a fourfold increase in energy, with combined scaling compounding toward roughly sixteenfold. Denoising step scaling is reported as linear in the same source.
Appendix C — Known Limitations and What We Do Not Know
No commercial model data exists
Every specific figure attached to Sora, Veo, Runway, or Kling in this paper is either a pricing signal (a business decision, not an energy measurement) or a third-party extrapolation. Nobody outside those companies knows their actual per-generation energy cost at real production settings.
Model-to-model variance is enormous
The measured table in Section 3 spans nearly three orders of magnitude between models at their own default settings. Any single-number estimate for "what an AI video costs" necessarily obscures this variance, and we have tried to show ranges rather than points wherever possible.
The stitching estimate is an extrapolation
The linear-in-number-of-calls structure of our 30-second estimate is well-supported by the scaling laws and Veo's documented mechanics. The specific per-call anchor (300 Wh–1 kWh) is a reasoned extrapolation from open-model data, not a measurement of any commercial system, and should be treated accordingly.
Retries are not counted
All estimates in this paper are per successful generation. In practice, people regenerate AI video clips multiple times before landing on a version they post. We found no published data quantifying this retry multiplier, so every estimate in this paper should be read as a floor on the true resource cost of a published video, not a ceiling.
Debunked figures are named, not erased
Section 7 names specific figures that failed our adversarial verification process rather than silently omitting them, because we expect readers to encounter them elsewhere. Their inclusion is not an endorsement.
Rapid pace of change
AI video model efficiency, default resolutions, and vendor disclosure practices are all moving quickly. The figures in this paper reflect conditions as of mid-2026 and should be re-checked against current releases before being relied upon for any specific decision.
July 2026 | Version 1.0 | info@ecoinference.ai