Your boss asks, "Should we run AI ourselves?" See not only how much it costs, but what the money is paying for.
How much does it cost to run AI yourself?
Pick a model, deployment option, and usage level to estimate the cost of running open models such as Kimi, GLM, DeepSeek, and Qwen on your own hardware.
Found incorrect data? Report it to the administrator
How this estimate is calculated
The result shows the VRAM, hardware, electricity, labor, and source assumptions behind the estimate, so you can check what went into it.
VRAM Formula
VRAM ≈ params × bytes/param × runtime headroom (default ×1.2; adjustable for low/standard/high scenarios). FP16 = 2 bytes/param, INT4 = 0.5 bytes/param. For MoE models, all expert weights must reside in VRAM, so total parameter count is used.
Hardware Matching
We start with the smallest GPU that fits the required VRAM. Manual model-to-hardware recommendations override the automatic match when one is available.
Electricity Calculation
Monthly electricity = power (kW) × running hours × rate × PUE. Power defaults to the single-GPU TDP reference; for multi-GPU or full-system rigs, enter measured wall power in the calculator. Default values are shown; adjust them to match your region and facility. Cloud rental already includes electricity.
Labor Estimate
We estimate monthly labor as a monthly salary range for an AI operations engineer × an FTE share. Default values are shown; adjust them to match your team and region. This is a planning estimate.
Data Sources & Updates
Model parameters, architecture, and model-card information come from official materials and Hugging Face; GPU specifications come from NVIDIA and other hardware vendors. Cloud and local-hardware prices use published public prices, channel or market references, or estimates on an entry-by-entry basis. Each price shows its update time, conditions, and source link when available; estimates without a verifiable link are marked as having no recorded source link.
The current public snapshot contains cloud prices from AWS, Google Cloud, and Microsoft Azure, plus manually checked local channel and market references. Model and hardware information is checked against official sources and Hugging Face. Use the provider, update time, and source link shown for each entry.
Start with one number, then decide how to run AI
Turn the cost of running AI yourself into a number your team can check, compare, and use in a build-or-buy discussion.
Free model weights do not make a deployment free. The bill can include hardware, power, and the people who keep the system running. It also depends on the model, capacity, and deployment choice.
Use the estimate to decide whether self-hosting belongs in your AI plan, then walk your team through the budget.
The total is only the starting point. The result shows the cost lines and assumptions behind it, so you can explain where the budget goes.
What drives the cost of running AI?
The estimate starts with four things: the model, how you deploy it, how much you use it, and the work required to keep it running.
Model and size
A model's size determines its VRAM, compute, and hardware needs.
Deployment option
Cloud rental, purchased hardware, and existing infrastructure lead to different bills.
Expected usage
Higher usage usually means more capacity. This version does not yet model throughput, concurrency, or latency targets.
Local operations cost
Someone still has to deploy, monitor, upgrade, and respond to incidents. The cost depends on your team and region.
You can adjust exchange rate, electricity rate, salary range, FTE share, PUE, system multiplier, and depreciation period.
Advanced capacity inputs such as context length and concurrency are planned.
How the estimate is built
The calculator combines the selected model's infrastructure needs with your deployment choice, expected usage, and operating assumptions. Each one affects the final number.
- Model and size data
- Parameter counts, architecture, and model sizes come from official documentation and Hugging Face.
- Price data
- GPU, hardware, and cloud prices are dated, manually checked references. Individual entries may come from official price pages, public product pages, channels, or third-party markets; the result shows each entry's update time, conditions, and source link when available.
- Cloud billing model
- Cloud prices use pay-as-you-go hourly rates. Reserved capacity, negotiated rates, and volume discounts are not included.
- Running time
- Light, medium, and heavy usage correspond to roughly 120, 360, and 720 hours per month.
- Operations cost
- Monthly labor uses a local monthly salary multiplied by an FTE share. It is not a fixed vendor fee.
- Exchange rate
- The default is 1 USD = 7.2 CNY. Change it in the calculator if your planning currency differs.
- Adjustable inputs
- Exchange rate, electricity rate, labor salary, FTE share, PUE, system cost multiplier, and depreciation period can all be adjusted to match your region and facility.
What this estimate leaves out
This is a planning estimate. These are the production costs and capacity questions it does not fully capture.
- This is a planning estimate, not a vendor quote.
- Provider, region, contract, and availability can change the price.
- Real capacity also depends on throughput, concurrency, context length, quantization, and latency targets.
- Model quality is not part of the price estimate.
- Migration, fine-tuning, security, compliance, and application development can add cost.
- Production requirements such as redundancy and incident response can change the total.
How to estimate the cost of running AI yourself
Choose a model, describe the setup, and leave with a shareable number for the budget conversation.
- 01
Choose a model
Pick the model and size closest to what you plan to run.
- 02
Describe the setup
Choose cloud rental or local hardware, then set expected usage and tune the cost assumptions to match your region.
- 03
Check and share the estimate
Review the total and each cost line, then share the URL with your team.
Should you run AI yourself?
The calculator covers the self-hosting side. Your workload, budget, and team's ability to operate the system decide which option fits.
Running AI yourself may fit better when:
- You need more control over your data and infrastructure.
- Your usage is stable enough to plan capacity.
- An open model meets your quality needs.
- Your team can deploy and maintain the system.
- Private deployment or customization has clear value.
A managed API may fit better when:
- You want to ship quickly.
- Usage is low, uncertain, or highly variable.
- Your team does not want to run GPU infrastructure.
- You need a proprietary frontier model.
- You would rather pay for what you use.
Both can be sensible. When a current public price is available, the calculator places the self-hosting estimate alongside the original model vendor's API reference price.
What the calculator gives you
Use How Much To Run AI as a starting point: see the one-time and monthly costs before deciding whether self-hosting deserves a closer look.
- Estimated one-time deployment cost range
- Estimated monthly operating cost
- Breakdown for compute, electricity, and operations labor
- The model and deployment assumptions used
- Adjustable assumptions: exchange rate, electricity rate, labor cost, PUE, system multiplier, and depreciation period
- Shareable link: your setup is stored in the URL, so others see the same result. You can also export the result as a PNG or PDF report.
Open models currently covered
These are the open model families currently covered, so you can compare self-hosting costs across model sizes.
Frequently asked questions
These answers cover the questions that change the estimate from one setup to another.