> Source: https://meetrix.io/blogs/open-weight-llms-gke/
> Markdown copy of that page. Cite the URL above, not this file.

LLMs & APIs

# Deploying Open-Weight LLMs on Google Cloud (GKE)

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 29, 2026 • 10 min read

GKE gets one thing right by default that most Kubernetes GPU guides spend a paragraph warning you about: as long as you let it install the driver automatically, the NVIDIA device plugin comes along with it, on both Standard and Autopilot. No separate DaemonSet, no node stuck advertising zero GPUs because a plugin never registered. That's the opposite of what we found writing the equivalent [EKS guide](https://meetrix.io/blogs/vllm-eks-gpu-deployment/), where one of two AMI families skips the device plugin entirely.

This covers Standard versus Autopilot, a working `vllm-openai` Deployment, autoscaling, and multi-node serving for models too big for one GPU. If you want the mechanism underneath all of this explained without a specific cloud attached, read [GPU scheduling on Kubernetes](https://meetrix.io/blogs/gpu-scheduling-kubernetes/) first. And if you don't need Kubernetes at all, our [vLLM developer guide for GCP](https://meetrix.io/blogs/vllm-gcp-developer-guide/) gets you a single GPU instance running with Deployment Manager instead.

The Quick Verdict

Testing a model

Autopilot on a single L4. No node pool to manage, and GKE handles the driver and device plugin for you.

Production, exact config needed

Standard mode. Pick the machine type, GPU and driver version yourself instead of working within Autopilot's supported set.

Model too big for one GPU

LeaderWorkerSet across a Standard node pool, with Hyperdisk ML if weight-loading time matters.

Steady, high-utilization traffic

Standard, tuned to run hot. Autopilot's per-pod convenience costs more once a node pool is already busy.

0 Separate device plugin steps needed on GKE, versus one AMI family on EKS that requires it

11.9x Faster model weight loading with Hyperdisk ML versus pulling straight from a model registry

$0.10/hr GKE cluster management fee, identical on Standard and Autopilot, before any GPU compute

Prerequisites

-   An existing GKE cluster (Standard or Autopilot), plus `gcloud`, `kubectl` and `helm` installed locally.
-   GPU quota in the region you're launching into. New GCP projects often start with little or none, so request an increase from the Quotas page before you plan a launch date. If cluster autoscaling is on, your quota needs to cover the maximum node count times GPUs per node, not just what you launch with.
-   A Hugging Face token if the model you're serving is gated, stored as a Kubernetes Secret.

## Standard or Autopilot?

**Autopilot** is the hands-off mode. You describe the pod, including the GPU type and count, and GKE provisions a matching node, installs the driver, and manages the device plugin and node health, all without you touching a node pool. Each GPU has its own minimum GKE version before Autopilot will schedule it:

| GPU | Minimum GKE version |
| --- | --- |
| B200 (180 GB) | 1.32.2-gke.1422000 |
| H200 (141 GB) | 1.31.4-gke.1183000 |
| H100 Mega (80 GB) | 1.28.9-gke.1250000, or 1.29.4-gke.1542000 |
| H100 (80 GB) | 1.28.6-gke.1369000, or 1.29.1-gke.1575000 |
| A100, L4, T4 | All supported versions |

$0.10per hour, either mode

Two billing changes worth knowing before you launch. From October 1, 2026, Autopilot adds a node management premium on top of the usual charges specifically for RTX PRO 6000 (G4) GPUs. And more broadly, the billing model itself is version-gated: from GKE 1.29.4-gke.1427000, every GPU pod on Autopilot uses node-based billing automatically, Compute Engine hardware plus a management premium, without you adding anything to the manifest.

Between 1.28.9-gke.1069000 and that version, you get node-based billing only if you explicitly add the `cloud.google.com/compute-class: Accelerator` selector; skip it and you're on the older per-pod billing instead. Below 1.28.9-gke.1069000, Autopilot doesn't support the Accelerator compute class at all, and GPU pods bill like any other Autopilot pod.

**Standard** gives you the node pool directly: you pick the machine type, the exact GPU and the driver version, and GKE still auto-installs the driver and device plugin by default (since GKE 1.30.1, you don't even need to pass `gpu-driver-version`, it defaults to on). The trade-off is the one you'd expect: you're managing the node pool, not handing it to Google.

Neither is objectively better. Autopilot suits a team that wants to stop thinking about nodes and whose GKE version and GPU choice line up with what it supports. Standard suits a team that wants a configuration Autopilot doesn't offer yet, or needs machine type control for cost or performance reasons.

## Create the GPU Node Pool (Standard)

An L4, enough to serve an 8B-class model for testing, with automatic driver installation:

```plaintext
gcloud container node-pools create gpu-workers \
  --cluster your-cluster \
  --accelerator type=nvidia-l4,count=1,gpu-driver-version=default \
  --machine-type g2-standard-8 \
  --num-nodes 1 --min-nodes 0 --max-nodes 3 \
  --node-labels "cloud.google.com/gke-accelerator=nvidia-l4"
```

That's the whole node-side setup. There's no separate device plugin step to remember, and no taint to add by hand the way EKS needs one, since GKE's scheduler already keeps ordinary workloads off GPU-labelled nodes through the node label and selector pairing above.

Want the NVIDIA GPU Operator instead of GKE's managed device plugin, say, for its feature set across multiple cloud providers? Create the node pool with `gpu-driver-version=disabled` so GKE doesn't install its own driver, and add the node label `gke-no-default-nvidia-gpu-device-plugin=true` to stop GKE's device plugin DaemonSet from running, so the Operator's plugin is the only thing claiming `nvidia.com/gpu`. Running both plugins at once produces inconsistent, hard-to-diagnose GPU counts. Autopilot clusters don't support the GPU Operator at all.

## Deploy vLLM

A minimal Deployment and Service for a single-GPU model on Standard mode, following the same shape as [Google's own vLLM serving tutorial](https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-gemma-gpu-vllm): the `cloud.google.com/gke-accelerator` node selector to land on the right GPU type, a PVC for the Hugging Face cache, and `/dev/shm` mounted as memory-backed for tensor-parallel setups.

```plaintext
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-server
spec:
  replicas: 1
  selector:
    matchLabels: { app: vllm-server }
  template:
    metadata:
      labels: { app: vllm-server }
    spec:
      nodeSelector:
        cloud.google.com/gke-accelerator: nvidia-l4
      containers:
        - name: vllm
          image: vllm/vllm-openai:latest
          args:
            - "--model=meta-llama/Llama-3.1-8B-Instruct"
            - "--max-model-len=8192"
          ports:
            - containerPort: 8000
          resources:
            limits: { nvidia.com/gpu: "1" }
            requests: { nvidia.com/gpu: "1" }
          volumeMounts:
            - { name: hf-cache, mountPath: /root/.cache/huggingface }
            - { name: shm, mountPath: /dev/shm }
          startupProbe:
            httpGet: { path: /v1/models, port: 8000 }
            initialDelaySeconds: 15
            periodSeconds: 30
            failureThreshold: 60
          livenessProbe:
            httpGet: { path: /health, port: 8000 }
            periodSeconds: 10
          readinessProbe:
            httpGet: { path: /v1/models, port: 8000 }
            periodSeconds: 5
            failureThreshold: 3
      volumes:
        - name: hf-cache
          persistentVolumeClaim: { claimName: hf-cache-pvc }
        - name: shm
          emptyDir: { medium: Memory, sizeLimit: 4Gi }
---
apiVersion: v1
kind: Service
metadata:
  name: vllm-server
spec:
  selector: { app: vllm-server }
  ports:
    - { port: 8000, targetPort: 8000 }
```

Three probes instead of two, and that's deliberate. vLLM's `/health` endpoint returns 200 as soon as the server process binds to the port, not once the model has actually loaded, so pointing a readiness probe at it can add the pod to the Service before it can serve a request. `/v1/models` is model-aware: it fails or refuses the connection while loading, and returns 200 with model metadata once the model is actually ready. The `startupProbe` on `/v1/models` gives the model room to load without the liveness probe killing the container mid-download; `livenessProbe` on `/health` just confirms the process itself is alive; `readinessProbe` on `/v1/models` is what actually controls traffic routing.

Create the PVC before applying this, then check the pod landed on the GPU node and passed its probes:

```plaintext
kubectl get pods -o wide
kubectl logs -f deployment/vllm-server
```

## The Same Deployment on Autopilot

The container spec doesn't change. What changes is the node selector, and whether you need the compute-class selector depends entirely on your GKE version, per the billing breakdown above.

On GKE 1.29.4-gke.1427000 and later, just request the GPU:

```plaintext
      nodeSelector:
        cloud.google.com/gke-accelerator: nvidia-l4   # GKE 1.29.4-gke.1427000+, nothing else needed
```

On the transitional versions, 1.28.9-gke.1069000 up to but not including 1.29.4-gke.1427000, add the compute class explicitly, or you get pod-based billing instead of node-based, and on some versions in that range Autopilot rejects the pod outright without it:

```plaintext
      nodeSelector:
        cloud.google.com/compute-class: Accelerator   # required on GKE < 1.29.4-gke.1427000
        cloud.google.com/gke-accelerator: nvidia-l4
```

## Expose and Test the API

The Service above is `ClusterIP`, fine for testing through a port-forward. For real traffic, put an Ingress in front of it, or an internal `LoadBalancer` Service if every caller is inside your VPC.

```plaintext
kubectl port-forward svc/vllm-server 8000:8000

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}]
  }'
```

vLLM speaks the OpenAI chat completions format on both clouds, so anything built against the OpenAI SDK works here by changing the `base_url`. Our guide to [self-hosting an OpenAI-compatible API](https://meetrix.io/blogs/openai-compatible-api-self-hosted/) covers that swap.

## Autoscaling GPU Capacity

This is the one area where GKE and EKS genuinely diverge in tooling, not just naming. EKS increasingly runs on Karpenter. GKE runs Cluster Autoscaler, extended by Node Auto-Provisioning, which creates whole new node pools when nothing existing fits a pending pod, and Compute Classes, which let you rank a preference across machine types and let GKE pick. As of early 2026 there's no `karpenter-provider-gcp`, so if you're moving a Karpenter-based setup from AWS, budget time to rebuild it around GKE's own tools rather than expecting a drop-in port.

A Compute Class groups several node pools under one class with a priority order; when the autoscaler sees pods it can't schedule under that class, it provisions from the node pools in that exact order. For GPU workloads where falling back to a different instance type matters, that's worth setting up on Standard mode.

On top of node-level autoscaling, the same HPA-versus-KEDA split applies as on any Kubernetes cluster: HPA can scale vLLM replicas using its Prometheus metrics but can't scale to zero, and KEDA can, if you have idle windows worth the setup. We cover this in more depth, vendor-neutral, in [GPU scheduling on Kubernetes](https://meetrix.io/blogs/gpu-scheduling-kubernetes/).

If hand-tuning autoscaling thresholds isn't how you want to spend an afternoon, Google's own [GKE Inference Quickstart](https://cloud.google.com/blog/products/ai-machine-learning/gke-inference-gateway-and-quickstart-are-ga) analyzes your model and traffic pattern and generates manifests, HPA thresholds included, built on the [llm-d](https://developers.redhat.com/articles/2025/10/07/master-kv-cache-aware-routing-llm-d-efficient-ai-inference) serving stack, which runs vLLM underneath with a routing layer on top. The GKE Inference Gateway that comes with it routes by KV cache state rather than round robin, which matters once you're running more than one replica.

## Serving Models Too Big for One GPU

Google's own tutorials for serving DeepSeek-R1 671B and Llama 3.1 405B on GKE use the same [LeaderWorkerSet (LWS)](https://lws.sigs.k8s.io/) API vLLM documents for any Kubernetes cluster: one leader pod running `vllm serve`, the rest running headless workers, tensor and pipeline parallelism spanning the group.

One GKE-specific trick worth knowing if you go this far: downloading a 405B-class model fresh on every pod can take around 90 minutes. [Hyperdisk ML](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/persistent-volumes/hyperdisk-ml) in read-only-many mode, mounted into every pod instead of downloading fresh, is rated up to 11.9 times faster than loading straight from a model registry, scaling to 2,500 concurrent nodes at 1.2 TiB/s. In practice that cuts a 405B model's load time to somewhere around 20 minutes.

Same advice as on any cloud: don't reach for this until a single node's GPUs genuinely can't fit the model. It's a real jump in operational complexity for a real class of problem, not a default.

## What This Costs

Both modes carry the same GKE cluster management fee: $0.10 an hour, about $73 a month, billed in one-second increments, whether you run Standard or Autopilot. The free tier covers exactly one cluster with $74.40 in monthly credit.

On top of that, Standard bills per node through ordinary Compute Engine pricing: a `g2-standard-8` (one L4, 8 vCPUs, 32 GB) runs about $0.85 an hour on demand, roughly $621 a month. Autopilot's billing depends on your GKE version and GPU, per the table earlier in this article: node-based billing (hardware plus the management premium) once you're on a version and configuration that qualifies, pod-based billing (CPU, memory and GPU the pod actually requests) otherwise. Autopilot tends to win when utilization is uneven; Standard wins once you've tuned a node pool to run hot. For the GPU hourly prices themselves, A100, L4 and H100 on GCP, our [GPU costs for LLM inference](https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/) article has the numbers side by side with AWS.

## Troubleshooting

-   **Pod stuck in Pending on Standard.** The node pool's `cloud.google.com/gke-accelerator` label doesn't match the pod's node selector, or there's no node pool with that accelerator at all. Check `kubectl get nodes --show-labels`.
-   **Pod rejected outright on Autopilot.** Check your GKE version against the table earlier in this article. On the transitional versions you need the Accelerator compute class selector; miss it and some versions reject the pod rather than leaving it pending.
-   **CrashLoopBackOff on a fresh deploy.** The pod is still downloading the model when the liveness probe kills it. Switch to the startup-probe pattern above rather than raising a single probe's delay.
-   **OOMKilled or a CUDA out-of-memory error.** The PVC for the Hugging Face cache is too small, or `/dev/shm` is undersized for tensor-parallel. Both show up as the pod dying partway through startup.
-   **Works in one region, fails in another.** GPU availability and quota are regional. Confirm both before assuming the manifest is wrong.
-   **Running the GPU Operator, but GPU counts look wrong.** GKE's own device plugin is probably still running alongside it. Confirm the node has `gke-no-default-nvidia-gpu-device-plugin=true` and was created with `gpu-driver-version=disabled`.

## Where I'd Start

Testing a model

Autopilot, a single L4. Confirm your GKE version is current so you're not fighting the transitional compute-class rules.

Production, single node fits

Standard, with the three-probe pattern above. Simpler to operate than Autopilot once you know your exact configuration.

Model needs multiple nodes

LeaderWorkerSet on Standard, and price a Hyperdisk ML volume before you accept a 90-minute cold start as normal.

Don't want to hand-tune any of this

Start with GKE Inference Quickstart. It generates the manifests and HPA thresholds for you, built on the same vLLM this article uses.

Already comfortable operating Kubernetes on AWS? The mechanism is identical, but the tooling underneath isn't. Budget time for Karpenter's absence on GKE specifically, not just a general "it's all Kubernetes" assumption.

## Frequently Asked Questions

Does GKE install the NVIDIA device plugin automatically?

Yes. GKE's managed device plugin deploys as an add-on and advertises nvidia.com/gpu on GPU nodes by default, on both Standard and Autopilot. Swapping in the NVIDIA GPU Operator instead means disabling GKE's plugin yourself, and only on Standard mode; Autopilot doesn't support the Operator at all.

Which GPUs can I use on GKE?

Autopilot reaches B200, H200, H100 and H100 Mega, A100, L4 and T4, each gated to a minimum GKE patch version. Standard mode reaches a wider range through Compute Engine machine types directly, without that version gating.

Autopilot or Standard for serving an open-weight model?

Autopilot if you want to stop thinking about nodes and your GPU and GKE version fit what Autopilot supports. Standard if you need a specific machine type, driver version or configuration Autopilot doesn't offer, or want to tune a node pool to run hot for cost reasons.

Why is my GPU pod rejected on GKE Autopilot?

Almost always a version mismatch. On GKE versions before 1.29.4-gke.1427000, you need cloud.google.com/compute-class: Accelerator alongside your GPU selectors, or Autopilot rejects the pod. From 1.29.4-gke.1427000 on, requesting the GPU directly is enough.

How much does a GKE cluster cost before you even add GPUs?

$0.10 per cluster per hour, about $73 a month, for both Standard and Autopilot, billed in one-second increments. The free tier covers exactly one cluster with $74.40 in monthly credit.

How do I serve a model too large for one GPU on GKE?

The same LeaderWorkerSet (LWS) API vLLM documents for any Kubernetes cluster. For very large models, mounting weights from a Hyperdisk ML volume instead of downloading them fresh to every pod is the trick Google's own tutorials use to cut load time dramatically.

Is GPU scheduling different on GKE compared to EKS?

The core mechanism, device plugins and extended resources, is identical. The autoscaler differs: EKS increasingly runs Karpenter, GKE runs Cluster Autoscaler with Node Auto-Provisioning and Compute Classes. There's no karpenter-provider-gcp, so a Karpenter setup doesn't port over directly.

## Don't Need Kubernetes? Deploy vLLM on One GPU Instance

If GKE is more than your workload needs, our Deployment Manager-based vLLM developer guide gets an OpenAI-compatible endpoint running on a single Google Cloud GPU instance in minutes.

[See the Single-Instance Guide](https://meetrix.io/blogs/vllm-gcp-developer-guide/)

Meetrix Store

vLLM

Serve open models on your own GPU

[Deploy it](https://meetrix.io/store/vllm/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [vLLM Serve open models on your own GPU](https://meetrix.io/store/vllm/)
-    [OpenWebUI A private ChatGPT-style assistant](https://meetrix.io/store/openwebui/)
-    [Llama 4 Scout Mixture-of-Experts, 17B active parameters](https://meetrix.io/store/llama-4-scout/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)

[Browse all products](https://meetrix.io/store/)
