> Source: https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/
> Markdown copy of that page. Cite the URL above, not this file.

LLMs & APIs

# GPU Costs for LLM Inference: AWS vs GCP (2026)

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 28, 2026 • 12 min read

For LLM inference on list prices, AWS is cheaper if you need an H100 or an A100, and GCP is a little cheaper if a single L4 is enough. That's the short answer, and it's not the one that will decide your bill. An H100 costs $6.88 an hour on AWS and $10.98 on GCP, but the same GPU costs four times more per token at 25% utilization than at full load, on either cloud.

Prices below were checked on September 28, 2026, for on-demand instances in US East (AWS) and us-central1 (GCP). I haven't benchmarked throughput for this article, so where a tokens-per-second figure appears it's a labelled placeholder you should replace with your own measurement. The sections cover the prices, how to turn them into cost per million tokens, Spot and commitment discounts, which GPU fits which model, and the costs that never show up on the instance price.

The Quick Verdict

H100 or A100 workloads

AWS on demand. About 37% less per H100 and 25% less per A100 than GCP list price.

Small models on one L4

GCP, by roughly 12 to 13%. The gap is real but small, so region and quota matter more.

Spiky or batch traffic

Spot on AWS, on H100 or A10G, where the discount is 53 to 62% today. Not on L40S, where it's 1%.

Steady, always-on traffic

Fix utilization first, then commit. A 3-year reservation saves 54 to 57% on AWS, but GPUs age quickly.

37% Cheaper per H100 on AWS than on GCP, on demand

4x Cost per token swing between 100% and 25% utilization, same GPU

$5,022 One H100 on AWS, on demand, running all month (730 hours)

## GPU Prices Side by Side

These are per-GPU hourly prices, which is the fair way to compare, because the clouds package GPUs differently. AWS sells a single H100 as `p5.4xlarge` and eight as `p5.48xlarge`. GCP's H100 lives in the `a3-highgpu` family, and its 8-GPU shape is the one I could price reliably.

| GPU | AWS instance | AWS per GPU-hour | GCP instance | GCP per GPU-hour | Cheaper |
| --- | --- | --- | --- | --- | --- |
| L4, 24 GB | g6.xlarge | $0.805 | g2-standard-4 | about $0.70 | GCP |
| A10G, 24 GB | g5.xlarge | $1.006 | No equivalent | n/a | n/a |
| L40S, 48 GB | g6e.xlarge | $1.861 | Not compared | n/a | n/a |
| A100, 40 GB | p4d.24xlarge (8 GPUs) | $2.75 ($21.96 for 8) | a2-highgpu-1g | $3.67 | AWS but 8-GPU minimum |
| H100, 80 GB | p5.4xlarge (1 GPU) | $6.88 | a3-highgpu-8g (8 GPUs) | $10.98 ($87.83 for 8) | AWS |

AWS figures from [Vantage's EC2 tracker](https://instances.vantage.sh/aws/ec2/p5.48xlarge), updated September 28, 2026. Google's own [GPU pricing page](https://cloud.google.com/compute/gpus-pricing) renders its table client-side and I couldn't read it programmatically, so the GCP figures are the ones price trackers publish, and three of them agreed on $87.83 for the a3-highgpu-8g (a fourth listed $88.49). Check your region in the pricing calculator before you commit budget.

37%AWS H100 advantage

If you remember AWS as the expensive option for GPUs, that was true until mid-2025. AWS [cut on-demand prices](https://aws.amazon.com/blogs/aws/announcing-up-to-45-price-reduction-for-amazon-ec2-nvidia-gpu-accelerated-instances/) by 33% for P4d and P4de and by 44% for P5, measured from its May 31, 2025 prices, with Savings Plan discounts up to 45% on P5 as well.

One thing that did go up: EC2 Capacity Blocks for ML, where you reserve GPUs for a fixed window. [The Register reported a rise of about 15%](https://www.theregister.com/2026/01/05/aws_price_increase/) in January 2026, with p5e going from $34.61 to $39.80 an hour. That is a separate product from the on-demand prices in the table.

Two more things the table hides. A single L4 is where the GPU choice barely matters, because at $0.805 versus about $0.70 you're arguing over roughly $75 a month. And the A100 row comes with a catch: AWS only sells it in 8-GPU shapes (`p4d.24xlarge` with 40 GB cards, `p4de.24xlarge` with 80 GB), so if you need one A100, GCP's single-GPU `a2-highgpu-1g` is the practical option even though each GPU costs more.

GCP's smaller A3 shapes are the other unknown. One tracker lists the single-GPU `a3-highgpu-1g` as Spot or Flex-start only, so confirm in your console that you can get one on demand before you build a plan around it.

## Cost Per Million Tokens

An hourly price tells you almost nothing about inference cost. What you pay per token depends on how many tokens the GPU produces per hour, and that depends on your model, your batch size and, above all, how often the GPU is busy. The formula is short:

```plaintext
price_per_hour = 6.88         # one H100 on AWS p5.4xlarge, on demand
tokens_per_second = 2500      # placeholder: measure this on your own model and traffic
utilization = 0.5             # share of the hour the GPU is actually busy

tokens_per_hour = tokens_per_second * 3600 * utilization
print(price_per_hour / (tokens_per_hour / 1_000_000))   # about 1.53 dollars per million tokens
```

To show the shape of it, here is the same placeholder throughput of 2,500 tokens per second on one H100, at three utilization levels. The throughput is an assumption, not a measurement. The ratios between the rows don't depend on it.

| GPU utilization | AWS H100 ($6.88/hr) | GCP H100 ($10.98/hr) |
| --- | --- | --- |
| 100% busy | $0.76 per million | $1.22 per million |
| 50% busy | $1.53 per million | $2.44 per million |
| 25% busy | $3.06 per million | $4.88 per million |

Illustrative arithmetic on a placeholder of 2,500 output tokens per second. Real throughput varies widely with model size, quantization, context length and batch size.

Read down a column, not across. Going from 100% to 25% busy multiplies your cost per token by four, and no cloud switch gives you that. Going from GCP to AWS on the same H100 saves about 37%, which is worth having, but it's smaller than the swing from an idle afternoon.

This is also why the self-hosting question ("is it cheaper than an API?") has no general answer. A busy GPU behind a well-batched server, such as the [vLLM or SGLang](https://meetrix.io/blogs/vllm-vs-sglang/) options, can be very cheap per token. A GPU that sits idle overnight is paying $5,022 a month on AWS or $8,015 on GCP for a hobby. Measure your actual utilization first. If it's under a third, ask whether a smaller GPU or a Spot fleet fits better than a bigger box.

## Spot and Commitments

Here's what the same AWS instances cost with a discount attached. Both are snapshots from the tracker: Spot prices move constantly.

| AWS instance | On demand | Spot | 3-year reserved |
| --- | --- | --- | --- |
| g6.xlarge (L4) | $0.805 | $0.604 (25% off) | $0.369 (54% off) |
| g5.xlarge (A10G) | $1.006 | $0.473 (53% off) | $0.435 (57% off) |
| g6e.xlarge (L40S) | $1.861 | $1.836 (1% off) | $0.804 (57% off) |
| p4d.24xlarge (8x A100) | $21.96 | $17.54 (20% off) | $9.37 (57% off) |
| p5.4xlarge (1x H100) | $6.88 | $2.623 (62% off) | $2.972 (57% off) |

AWS US East, from Vantage's tracker on September 28, 2026. Reserved prices are the tracker's 3-year rates.

The last row is the one to stare at: today's Spot price for an H100 is lower than the 3-year reserved price. That won't last, and Spot capacity can be reclaimed at short notice, but it shows how much of a GPU's list price is a premium for certainty. The L40S row shows the opposite: Spot is 1% off, so there's nothing to gain.

GCP's Spot discount is deep as well, but the H100 figures I found ranged from about $1.09 to $3.69 per GPU-hour depending on region and source, which is too wide to print as a fact. Look at your region in the console.

My opinion on commitments: use them for the floor of your traffic and nothing more. A 3-year reservation on a GPU generation is a bet that you won't want the next one, and inference hardware has been turning over faster than that. A 1-year term, or a reservation sized to your quietest week, gets you most of the benefit. Then send the peaks to on-demand or Spot.

For Spot to work with inference, the serving layer has to tolerate losing a node: two or more replicas behind a load balancer, model weights on fast storage or baked into the image, and clients that retry. A single Spot endpoint that customers depend on is a bad idea.

## Which GPU for Which Model

Weights take roughly the parameter count times the bytes per parameter. That's about 2 bytes at FP16, 1 at FP8 and half a byte at 4-bit. A 70-billion-parameter model is therefore around 140 GB, 70 GB and 35 GB. The KV cache, which grows with context length and concurrent requests, comes on top, and it's usually why a model that "fits" on paper falls over in practice.

| Model size | Fits on | AWS | GCP |
| --- | --- | --- | --- |
| Up to about 8B, FP16 | One 24 GB GPU (L4 or A10G) | g6.xlarge or g5.xlarge | g2-standard-4 |
| About 30B at 4-bit, or 70B at 4-bit with little headroom | One 48 GB GPU (L40S) | g6e.xlarge | Not checked here |
| 70B at FP8 | One 80 GB H100, tight, or two GPUs | p5.4xlarge | a3-highgpu shapes |
| 100B and up, or large mixture-of-experts | Eight 80 GB GPUs | p5.48xlarge ($55.04/hr) | a3-highgpu-8g ($87.83/hr) |

A rule of thumb from parameter count, not a benchmark. Test your exact model, quantization and context length before you size a fleet. Our [list of open-source LLMs to self-host](https://meetrix.io/blogs/best-open-source-llms-self-hosted-2026/) covers what's worth running.

The cheapest mistake to avoid is buying an H100 for a model that a quantized build fits on an L4 or L40S. The second cheapest is the opposite: squeezing a model onto a card with no room for KV cache, then adding GPUs anyway when latency collapses under concurrency. Both are visible in a one-day test.

## Costs Beyond the GPU

The instance price is most of the bill, not all of it.

-   Quota. Both clouds start new accounts with little or no GPU quota, and increases can take days. On AWS the request is a vCPU quota for the instance family, and we've written up [how to increase your AWS vCPU quota](https://meetrix.io/blogs/increase-aws-vcpu-quota/). Ask before you plan a launch date.
-   Idle time. A GPU bills while it waits. An always-on H100 is about $5,022 a month on AWS and $8,015 on GCP, and an always-on L4 is about $588 and $515. If you're running on Kubernetes, tight [GPU scheduling](https://meetrix.io/blogs/gpu-scheduling-kubernetes/) is a big part of avoiding this, not just an instance-shutdown habit.
-   Storage. Model weights are large (140 GB at FP16 for a 70B model), and you pay for the disk, plus the time it takes to copy them onto a fresh node every time a Spot instance is replaced.
-   Data transfer out. Text responses are small, so egress rarely dominates, but a busy public API does add up, and both clouds bill it. Check the network pricing for your volume.
-   Engineering time. Inference servers release often and sometimes break things. You're paying someone to pin versions, test upgrades and answer the pager.

Quota and idle time are the two I'd check first. The quota one is just annoying. The idle one is a habit: set an auto-shutdown or a schedule for anything that isn't serving production traffic.

## Running vLLM on Either Cloud

-   vLLM on AWS and Google Cloud
-   NVIDIA drivers and CUDA included
-   OpenAI-compatible API

Since the answer to "AWS or GCP?" depends on your GPU, region and quota, it helps not to be locked to either. We package the same open-source vLLM server for both clouds: the [vLLM image](https://meetrix.io/store/vllm/) launches on a GPU instance with drivers, CUDA and an OpenAI-compatible API ready, and there's no per-token fee. The GPU rates above are instance prices only. Marketplace listings can add an hourly line on top, and our [AWS listing article](https://meetrix.io/blogs/vllm-aws-marketplace/) quotes $0.019 to $0.048 an hour depending on the instance, so check the current listing before you budget. The setup guides are the [vLLM developer guide for AWS](https://meetrix.io/blogs/vllm-developer-guide/) and the [GCP version](https://meetrix.io/blogs/vllm-gcp-developer-guide/).

If you're not sure vLLM is the right engine, read [Ollama vs vLLM](https://meetrix.io/blogs/ollama-vs-vllm/) for the single-user case, and [how to self-host an OpenAI-compatible API](https://meetrix.io/blogs/openai-compatible-api-self-hosted/) for the client side. Because the API is the same on both clouds, you can run one short test on each and compare real cost per token. Running this at fleet scale rather than on a single instance? Our guide to [deploying vLLM on Amazon EKS with GPUs](https://meetrix.io/blogs/vllm-eks-gpu-deployment/) covers the Kubernetes path, including the autoscaling that keeps an idle GPU fleet from billing like a busy one.

## Where I'd Start

Testing a model

Use **one L4 or A10G on demand**, on whichever cloud gives you quota first. Cost differences here are pocket change. Learn your tokens per second.

Production, 70B class

Price **AWS p5.4xlarge first**. On these list prices it's about 37% cheaper per GPU, with Spot for the peaks.

Lots of small models

**L4s on GCP** are the cheapest per GPU by a small margin. Watch quota per region before you plan a fleet.

Steady 24/7 load

Measure utilization for two weeks, then **commit to the floor**, on a 1-year term, and keep the rest on demand.

Already running on one cloud and it's working? Moving for a 37% price difference on one GPU shape costs you a migration, new quota requests and a retest. It only pays if your GPU bill is large or your utilization is already high. If it's not, raising utilization on the cloud you have is usually the better first move.

## Frequently Asked Questions

Is AWS or GCP cheaper for LLM inference?

On list prices as of September 28, 2026, AWS is cheaper for H100 and A100 GPUs, and GCP is slightly cheaper for a single L4. An H100 costs about $6.88 per GPU-hour on AWS and $10.98 on GCP. Utilization moves your bill more than the cloud does.

How much does an H100 cost per hour on AWS and GCP?

On demand, AWS charges $6.88 per hour for one H100 on p5.4xlarge, or $55.04 for eight on p5.48xlarge. GCP's a3-highgpu-8g lists at about $87.83 for eight, roughly $10.98 per GPU, in us-central1.

How do I calculate cost per million tokens for a self-hosted LLM?

Divide the GPU's hourly price by the millions of tokens it produces per hour. That is tokens per second times 3,600, times the share of the hour the GPU is actually busy. Measure tokens per second on your own model and traffic.

Is self-hosting an LLM cheaper than paying per token?

Only when the GPU stays busy. In our example an H100 costs $0.76 per million tokens at full load and $3.06 at 25% utilization on AWS. Compare that against your API's price for the same model, with your real traffic pattern.

Are Spot GPUs safe for LLM inference?

For stateless, replicated inference behind a load balancer, yes, if you can lose a node at short notice. For a single production endpoint, no. Keep an on-demand baseline and let Spot absorb the extra traffic.

What GPU do I need to run a 70B model?

Weights take roughly the parameter count times bytes per parameter: about 140 GB at FP16, 70 GB at FP8 and 35 GB at 4-bit. Add KV cache on top. A 4-bit build is tight on 48 GB, and FP8 is tight on one 80 GB H100.

Do I need to request GPU quota on AWS and GCP?

Usually, yes. New accounts often start with little or no GPU quota on either cloud, and requests can take a day or more. On AWS it is a vCPU quota for the GPU instance family. Ask before you plan a launch date.

Can I run the same vLLM setup on AWS and GCP?

Yes. vLLM is open source and cloud-agnostic, and the OpenAI-compatible API is identical on both. Meetrix packages the same vLLM server for AWS and Google Cloud, so you can price-test each cloud before you commit.

## Run vLLM on Your Own GPU

A pre-configured vLLM inference server for AWS and Google Cloud, with NVIDIA drivers, CUDA and an OpenAI-compatible API ready to go.

[Launch vLLM on Your GPU](https://meetrix.io/store/vllm/)

Meetrix Store

vLLM

Serve open models on your own GPU

[Deploy it](https://meetrix.io/store/vllm/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [vLLM Serve open models on your own GPU](https://meetrix.io/store/vllm/)
-    [OpenWebUI A private ChatGPT-style assistant](https://meetrix.io/store/openwebui/)
-    [Llama 4 Scout Mixture-of-Experts, 17B active parameters](https://meetrix.io/store/llama-4-scout/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)

[Browse all products](https://meetrix.io/store/)
