> Source: https://meetrix.io/blogs/self-hosted-llms-enterprise-cost/
> Markdown copy of that page. Cite the URL above, not this file.

LLMs & APIs

# Self-Hosted LLMs: When They Beat API Costs for Enterprises

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 29, 2026 • 10 min read

Self-hosting an LLM is cheaper than an API once your traffic is steady enough to keep a GPU busy, and not before. That's the honest answer, and it depends on a number most teams never calculate: how much of every GPU-hour is actually doing work. Get that number wrong and you can burn more on an idle GPU than you'd have spent on the API you were trying to avoid.

Below is the actual math, with API prices checked on September 29, 2026, and GPU prices from our own [GPU cost breakdown for AWS and GCP](https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/). I'm not going to hand you someone else's "breakeven point," because I don't trust the suspiciously precise numbers other sites publish on this topic. Here's the formula instead, so you can run it on your own traffic.

The Quick Verdict

Steady, high volume

Self-host. Once a GPU clears roughly a tenth of its theoretical capacity, it undercuts a mid-tier API on raw token cost.

Spiky or low volume

Stay on the API. An idle GPU is the single fastest way to lose this comparison.

Data can't leave your network

Self-host regardless of cost. Compliance and contract terms override the token math entirely.

Need frontier-level reasoning

Keep an API in the loop. Open-weight models are closing the gap, but the hardest tasks still favor Claude, GPT or Gemini's top tier.

~9% GPU utilization where an AWS H100 matches a mid-tier API's blended token price, on our placeholder throughput

$5,022 What one always-on H100 costs per month on AWS, whether it serves one request or a million

3 to 5x The multiplier several vendor cost-of-ownership writeups put on top of raw GPU rental for labor and idle capacity. Their number, not independently verified.

## What "Cheaper" Actually Means

Every self-host-versus-API comparison that only compares the GPU's hourly price against the API's per-token price is measuring the wrong thing. The GPU bills by the hour whether it's busy or not. The API bills by the token, full stop. That asymmetry is the entire question, and it means the honest comparison isn't "GPU price vs. token price," it's "GPU price divided by how many tokens you actually push through it."

A busy GPU behind a well-batched server can beat almost any API on cost per token. An idle one is just an expensive way to heat a data center. Utilization is the variable that decides which one you're running.

## API Pricing Right Now

Checked against each vendor's own pricing page on September 29, 2026. Prices move: OpenAI alone has shipped two pricing generations since the spring, so treat this table as a snapshot, not a promise.

| Model | Input, per million tokens | Output, per million tokens | Tier |
| --- | --- | --- | --- |
| GPT-6 Astra | $10.00 | $50.00 | Flagship |
| GPT-6 Sol | $2.00 | $10.00 | Mid |
| GPT-6 Luna | $0.10 | $0.50 | Budget |
| Claude Fable 5.1 | $10.00 | $50.00 | Flagship |
| Claude Opus 5.5 | $4.00 | $20.00 | High |
| Claude Sonnet 5.5 | $2.00 | $10.00 | Mid |
| Claude Haiku 4.5 | $1.00 | $5.00 | Budget |
| Gemini 3.1 Pro | $2.00 | $12.00 | High |
| Gemini 3.8 Flash | $0.75 | $3.75 | Mid |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Budget |

From the [OpenAI pricing page](https://developers.openai.com/api/docs/pricing), [Claude's pricing page](https://claude.com/pricing) and [Gemini's API pricing page](https://ai.google.dev/gemini-api/docs/pricing), September 29, 2026. Gemini's ≤200k-token rate is shown, and its Flash tiers are already scheduled to roughly double on January 1, 2027. GPT-6 Sol also has an unpublished long-context threshold past which input and output both cost more. All three vendors offer roughly 50% off through a batch API and a steep discount on cached input, so a heavily-cached or batched workload will beat these headline numbers.

Notice the spread inside a single vendor. GPT-6 Astra to GPT-6 Luna is a 100x difference on input tokens. Before you price self-hosting against "the API," decide which tier you're actually comparing against, because the answer changes completely depending on whether your current spend is on a flagship model or a budget one.

## Self-Hosted Cost Per Token

We worked through this in detail in the [GPU costs article](https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/), so here's just the part that matters for this comparison: on a placeholder throughput of 2,500 output tokens per second, one H100 on AWS costs $0.76 per million tokens at 100% utilization, $1.53 at 50%, and $3.06 at 25%. On GCP the same three numbers are $1.22, $2.44 and $4.88.

Every one of those numbers beats GPT-6 Sol or Claude Sonnet 5.5's blended price of $8.40 per million tokens (worked out below), even at a quarter utilization. Against the budget tier, GPT-6 Luna's blended $0.42 per million, the comparison flips: only a very busy GPU beats it, and you're also comparing a small open model against a small hosted one, not against the flagship you might actually be replacing.

A note on that 2,500 tok/s placeholder: it's deliberately conservative, and real hardware can do a lot better on a small model. Third-party H100 benchmarks of Llama 3.1 8B in offline batch inference put SGLang and LMDeploy around 16,200 tokens per second and vLLM around 12,550, a gap the benchmark's authors trace to scheduling overhead rather than raw compute. Those numbers are for an 8B model running in a best-case batch mode, not the 70B-class model this article's GPU pricing assumes, and I haven't reproduced them myself, so treat them as an upper bound on what's possible, not a number to plan a fleet around. Measure your own model.

## The Breakeven Formula

Here's the calculation behind the "~9%" figure in the stat row above. It compares one always-on H100 against a mid-tier API, for a chat-style workload where four out of five tokens are output rather than input, which is typical for a conversational assistant.

```plaintext
# Steady, always-on H100 vs. a mid-tier hosted API
gpu_price_per_hour = 6.88          # AWS p5.4xlarge, on demand
hours_per_month = 730
monthly_gpu_cost = gpu_price_per_hour * hours_per_month   # $5,022

# Mid-tier API, blended for a chat-style 20% input / 80% output mix
api_input_per_million = 2.00       # e.g. GPT-6 Sol or Claude Sonnet 5.5
api_output_per_million = 10.00
api_blended_per_million = 0.2 * api_input_per_million + 0.8 * api_output_per_million  # $8.40

# Tokens the API spend of $5,022 would buy
breakeven_tokens_millions = monthly_gpu_cost / api_blended_per_million   # about 598 million

# Tokens the GPU can produce in a month, if it ran flat out
tokens_per_second = 2500           # placeholder: measure this on your own model
max_tokens_millions = tokens_per_second * 3600 * 24 * 30 / 1_000_000    # about 6,480 million

utilization_needed = breakeven_tokens_millions / max_tokens_millions
print(f"{utilization_needed:.1%} utilization needed to break even")   # about 9.2%
```

Read that as: if your GPU is busy more than about a tenth of the time, at this placeholder throughput, it's already cheaper than paying GPT-6 Sol or Claude Sonnet prices for the same volume. The throughput number is the one you have to replace with your own measurement. Halve it and the utilization threshold roughly doubles; double it and the threshold roughly halves. The shape of the conclusion doesn't move much. The exact crossover point does, so run this on your own model before you commit budget.

Run the same formula against GPT-6 Astra or Claude Fable's flagship pricing and the self-hosted GPU wins at well under 2% utilization, because those models cost six times more per token. Run it against GPT-6 Luna's budget pricing and you need well over half the GPU's capacity just to break even, and that's before asking whether Luna and your open model are even comparable in quality.

## Costs the API Hides

-   **Rate limits.** Grow past your tier and you either wait, pay for a higher tier, or get throttled during your busiest hour.
-   **Price volatility.** OpenAI's flagship pricing has already changed once this year, GPT-5.6 to GPT-6 in early September. Your cost model from six months ago may already be wrong.
-   **Model deprecation.** The exact model your prompts are tuned for gets retired on the vendor's schedule, not yours, and the replacement doesn't always behave the same way.
-   **Context overage.** GPT-6 Sol doubles its input and output price past an unpublished long-context threshold, and Gemini's Flash tiers are already scheduled to roughly double on January 1, 2027. A workload that creeps past either line gets quietly more expensive.
-   **Data terms.** Free and some lower tiers use your content to improve the vendor's models. Enterprise agreements usually opt out of this, but check the actual contract, not the marketing page.

## Costs Self-Hosting Hides

We covered GPU quota, idle time, storage and egress in the [GPU costs article](https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/), so I won't repeat the list here. The one worth adding for an enterprise specifically is staffing: someone has to own version upgrades, patch security issues in the serving stack, and answer the pager when inference latency spikes at 2 a.m. Several vendor writeups on this topic put the all-in cost of self-hosting at three to five times the raw GPU rental once labor is counted, and while I can't verify their specific hourly rates, the direction is right: budget for a person, not just a GPU.

There's also a model-quality tax that's easy to skip past. Open-weight models are genuinely good now, our [roundup of the best self-hosted LLMs](https://meetrix.io/blogs/best-self-hosted-llm/) covers which ones, but "cheaper per token" only matters if the cheaper model actually does the job. Swapping GPT-6 Astra for a 4-bit open model that gets the task wrong twice as often isn't a cost saving, it's a different cost.

## When It's Not About Cost

Some enterprises self-host an LLM even when the token math favors the API, and they're not wrong to. In March 2023, three separate Samsung Semiconductor engineers pasted proprietary material into ChatGPT within about three weeks of each other: source code while debugging, an internal meeting transcript, and a test sequence for chip defect detection. Samsung banned the company's staff from using ChatGPT and similar tools that May. That's the plainest example of why "the data leaves the building" is a real risk, not a theoretical one. If a contract, a regulator or your own security policy says customer data can't reach a third party, that's the whole decision. No breakeven formula overrides it.

The regulatory picture is moving too, and it's more layered than a single date. Under the EU's Digital Omnibus, which entered into force in July 2026, high-risk obligations for stand-alone systems (Annex III: hiring, credit scoring, education and similar) are now due December 2, 2027, and obligations for high-risk AI embedded in already-regulated products (Annex I: medical devices, machinery) follow on August 2, 2028. The Article 50 transparency rules, disclosing AI-generated content and telling people they're talking to an AI, were not delayed and still apply from August 2, 2026. None of that changes the eventual data governance and human-oversight obligations for systems in scope, and those are easier to demonstrate when the model runs on infrastructure you control. If you're anywhere near this scope, talk to counsel before treating cost as the only variable. This is a moving regulation and my summary of it is not legal advice.

Latency is the other non-cost reason. A model running in your own VPC, next to the service that calls it, skips a round trip to a third-party API that a public internet path can make unpredictable. For anything closer to real time than a chat reply, that matters more than the per-token price.

## The Hybrid Middle Ground

-   Self-hosted vLLM for routine traffic
-   Frontier API for hard cases
-   Same OpenAI-compatible format on both

Most enterprises that self-host don't go all in. They route the high-volume, routine traffic, classification, extraction, routine drafting, to a self-hosted open model, and keep a frontier API for the requests that genuinely need it. A model gateway such as LiteLLM sits in front of both, exposing one OpenAI-compatible endpoint and routing each request to vLLM on your own GPU or out to Claude, GPT or Gemini by data sensitivity or task difficulty. AWS's own sample toolkits and Solutions Library guidance for GenAI on EKS are built around exactly this pattern: LiteLLM as the gateway, vLLM or SGLang doing the self-hosted serving behind it, on GPU nodes that [Kubernetes schedules](https://meetrix.io/blogs/gpu-scheduling-kubernetes/) like any other resource. Because both sides speak the same OpenAI-compatible API shape, adding a route is a model name in a config file, not a rewrite. Our guide to [self-hosting an OpenAI-compatible API](https://meetrix.io/blogs/openai-compatible-api-self-hosted/) covers the swap itself, and [deploying vLLM on Amazon EKS with GPUs](https://meetrix.io/blogs/vllm-eks-gpu-deployment/) covers running the self-hosted side on Kubernetes rather than a single instance.

We package [vLLM](https://meetrix.io/store/vllm/) as a pre-configured image for AWS and Google Cloud, with drivers, CUDA and the OpenAI-compatible endpoint already set up, so the self-hosted side of a hybrid setup doesn't need a from-scratch build. The [AWS developer guide](https://meetrix.io/blogs/vllm-developer-guide/) and [GCP version](https://meetrix.io/blogs/vllm-gcp-developer-guide/) walk through the launch.

## Where I'd Start

Piloting or under a few hundred million tokens a month

Stay on the API. Building GPU infrastructure to save money on traffic you haven't proven yet is the classic way to lose this bet.

Steady, predictable, high volume

Run the breakeven formula above on your real throughput. If your GPU would clear that utilization threshold, price a self-hosted pilot against your current API bill for a month.

Data can't leave your network

Self-host. Confirm the compliance requirement first, then use the cost math to pick the cheapest infrastructure that satisfies it, not to decide whether to satisfy it.

Mixed workload, some routine and some hard

Go hybrid. Self-host the bulk, high-volume traffic and keep an API for the requests where model quality genuinely matters.

If you only take one number from this article, take the formula, not the "9%." Your throughput, your model, your token mix and your region all move that number. Run it before you buy hardware, and rerun it every time a vendor changes their prices, because in 2026 that's happening more than once a quarter.

## Frequently Asked Questions

Is self-hosting an LLM cheaper than the OpenAI or Claude API for an enterprise?

Only past a real volume threshold, and only if you count engineering time and idle GPU hours, not just the instance price. Against a mid-tier hosted model, an always-on H100 on AWS breaks even at roughly a quarter of its capacity. Below that, the API is cheaper.

How many tokens per month before self-hosting an LLM pays off?

Enough that your GPU spends more time processing than idling. In our worked example, an AWS H100 matches a mid-tier API's blended token price at around 9 to 10% utilization on a placeholder throughput, which most steady enterprise traffic clears easily. Measure your own tokens per second before trusting any number, including ours.

What hidden costs does self-hosting an LLM add beyond the GPU price?

Engineering time to patch, monitor and upgrade the serving stack; idle GPU hours outside peak traffic; storage and load time for model weights; and GPU quota requests that can take days. None of these show up on the instance price.

What hidden costs does an LLM API add beyond the token price?

Rate limits that force you onto a higher paid tier, price changes you don't control, a shrinking discount as your usage grows past their volume tiers, and the ongoing risk that your vendor deprecates the exact model your prompts were tuned for.

Do open-weight models match GPT or Claude for enterprise use?

On many business tasks, close enough. On the hardest reasoning and coding benchmarks, the frontier API models are usually still ahead. Check our roundup of the best self-hosted LLMs against your actual task before assuming a smaller open model is good enough.

Why would an enterprise self-host an LLM even if the API is cheaper?

Data residency and compliance are the usual reasons. If a contract, a regulator, or your own security policy says customer data cannot leave your network, self-hosting is not a cost decision, it's a requirement, regardless of what the math says. The EU AI Act's phased deadlines are pushing more enterprises to ask this question earlier.

Can an enterprise mix self-hosted and API models?

Yes, and most that self-host end up doing this. A gateway like LiteLLM routes routine, high-volume traffic to a self-hosted open model and sends the hard cases to a frontier API. Both speak the same OpenAI-compatible format, so adding a route is a model name, not a rewrite.

## Test the Self-Hosted Side of the Math

A pre-configured vLLM inference server for AWS and Google Cloud, with an OpenAI-compatible API, so you can price a real pilot against your current API bill.

[Launch vLLM on Your GPU](https://meetrix.io/store/vllm/)

Meetrix Store

vLLM

Serve open models on your own GPU

[Deploy it](https://meetrix.io/store/vllm/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [vLLM Serve open models on your own GPU](https://meetrix.io/store/vllm/)
-    [OpenWebUI A private ChatGPT-style assistant](https://meetrix.io/store/openwebui/)
-    [Llama 4 Scout Mixture-of-Experts, 17B active parameters](https://meetrix.io/store/llama-4-scout/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)

[Browse all products](https://meetrix.io/store/)
