How Much To Run AI

Video APIs charge per second. Past a certain monthly volume, buying your own GPUs gets cheaper — where is that point?

At what volume does self-hosting video generation start saving money?

Enter your workload to estimate the total cost of ownership and break-even point of running open-weight video generation models on your own GPUs, compared against official and hosted API per-second pricing.

Workload & assumptions

Cost assumptions

Get the cost of one second of video, then decide how to run it

Turn the cost of self-hosted video generation into a number your team can check, compare, and use to discuss building versus calling an API.

Open weights are free to download, but running them is not. GPUs, power, and day-to-day operations all land on the bill, while API paths grow linearly with every output second.

Use this estimate to judge whether self-hosting belongs in your video pipeline plan — then take the budget to your team.

The total is only the start. Results list the upfront investment, monthly cost, per-output-second cost, and break-even months, so you can explain where the budget goes.

What determines the cost of one second of video?

This estimate looks at four things: resolution, workload, hardware profile, and the operational work needed to keep a deployment running.

Resolution tier

The open checkpoints self-host only at native 768p; 480p and 2K/4K outputs are API-only, and per-second pricing differs by tier.

Expected workload

Clip length × clips per month sets the GPU-hours required. The higher the volume, the thinner the fixed hardware investment spreads.

Hardware profile

Throughput samples come from 4× H100 80GB (BF16, SGLang). A different tier changes generation speed, GPU-hours, and utilization.

Operations and power

Deploying, monitoring, upgrading, and handling failures all take people; electricity and PUE add to the monthly bill based on your region.

You can adjust the exchange rate, electricity price, salary range, FTE ratio, PUE, system cost multiplier, and depreciation period.

Advanced capacity inputs like multi-node scaling and 2K/4K self-hosting are still on the roadmap.

Method & assumptions

How the estimate works — the assumptions behind every number on this page.

  • Workload is measured in seconds of output video: clip duration × clips per month. API prices are published per output second by resolution tier.
  • Self-host throughput is an estimate: public samples report roughly 13–74 seconds to generate a 5-second 768p clip on 4× H100 80GB (BF16, SGLang). We use the conservative end-to-end figure and assume generation time scales approximately linearly with output duration at fixed resolution and step count. Treat all throughput-derived numbers as directional estimates.
  • Local purchase cost = one-time hardware (× system multiplier) depreciated over the chosen period + electricity (GPU-hours × power × PUE × rate) + operations labor. Cloud rental bills only the GPU-hours actually spent generating.
  • Hosted API cost = monthly output seconds × the per-second price of the selected resolution tier. Reference input materials (images, reference video) may be billed separately by the provider.
  • Break-even months ≈ one-time local investment ÷ (cheapest API monthly cost − local monthly operating cost), shown only when the difference is positive. Directional only.
  • License restrictions are displayed as facts (territory, revenue threshold, open checkpoint scope) and are not legal advice. They do not block the estimate.

What this estimate does not include

This is a planning estimate. Below are the production costs and quality factors the calculator does not fully cover yet.

  • This is a planning estimate, not a vendor quote.
  • Generation quality, prompt adherence, and aesthetics are outside the price estimate — the post-training gap of fal H3 Max cannot be measured in price alone.
  • Content moderation, safety, and compliance costs are not included.
  • Video storage, delivery bandwidth, and asset management may add to the monthly bill.
  • Compute burned on failed generations and retries is not modeled separately.
  • Throughput is based on sparse public samples; real speed varies with steps, framework, and load.

How to estimate the cost of self-hosted video generation

Set your workload, pick a hardware profile and assumptions, and get a shareable number you can take to a budget discussion.

  1. 01

    Set the workload

    Enter clip length, clips per month, and target resolution to fix your monthly output seconds.

  2. 02

    Pick hardware and assumptions

    Choose owned or rented GPU profiles, then adjust electricity, labor, and other cost assumptions for your region.

  3. 03

    Compare the four paths

    See monthly and per-output-second costs for local purchase, cloud rental, the official API, and fal H3 Max — then share the URL with your team.

Should you self-host video models?

The calculator shows you the self-hosting side of the bill first; the final call also depends on your volume, budget, region, and whether the team can run the infrastructure.

Self-hosting may fit when:

  • Your monthly volume is stable and large enough to amortize the hardware.
  • You need full control over data and the generation pipeline.
  • Native 768p output meets your quality requirements.
  • The team can deploy and maintain GPU inference services.
  • Your region is within the MiniMax H3 license territory (which excludes the US, EU, UK, and South Korea).

A hosted API may fit when:

  • You want to ship fast without operating GPU infrastructure.
  • Volume is low, uncertain, or highly variable.
  • You need 2K/4K output or API-only components like Context-IR.
  • You need fal H3 Max's post-trained quality, and its weights are not released.
  • Your region falls outside the H3 license's applicable territory.

Both approaches can be right. The calculator places the self-hosting estimate next to the published prices of the official MiniMax API and fal H3 Max, and shows break-even months against the cheapest API path.

Supported video models

Open-weight video generation models you can self-host, with their license terms and open checkpoint scope.

MiniMax H3MiniMax
~69.2B (33B DiT + Qwen3-VL-32B encoder + VAE) · 768p

Open checkpoints: FL2VA, Ref2VA

API-only components: H3-Context-IR, H3-Regenerate-2K

MiniMax H3 Community License (not OSI-approved): the applicable territory excludes the US, EU, UK and South Korea — self-hosting there requires separate written authorization. Commercial use is free below USD 20M annual revenue with attribution; above that threshold, written authorization is required. Only H3-Base checkpoints (FL2VA, Ref2VA) are open-weight.

Data sources

Every price and throughput figure on this page traces back to a public source.

Hosted API list prices, manually verified: Last verified 2026-09-03.

Throughput estimate sources:

GPU purchase and rental prices come from the same governed price-observation pipeline as the LLM calculator; the exact source, region, and observation date behind the current estimate are listed in the calculator result's "Hardware & cloud rental price sources" card.

Prices are list prices for budget estimation only; promotions, discounts, and regional terms may differ. Confirm with the provider.

FAQ

Common questions about self-hosting video models and how this estimate works.