> Source: https://meetrix.io/store/llama-4-scout/
> Markdown copy of that page. Cite the URL above, not this file.

[Store](https://meetrix.io/store/)  Llama 4 Scout

Open-source LLM

# Self-hosted Llama 4 Scout on AWS

Run Meta's Llama 4 Scout, a Mixture-of-Experts model with 17B active parameters, behind an OpenAI-compatible API in your own AWS account. No dependency wrangling or GPU driver setup.

 [Launch on AWS](https://aws.amazon.com/marketplace/pp/prodview-iu7qh5qnnbjt6)

-   17B active, 109B total
-   OpenAI-compatible API
-   Pay per hour

![Llama 4 Scout deployed on AWS by Meetrix](https://meetrix.io/assets/store/llama-4-scout/llama-4-scout.webp)

## What is Llama 4 Scout?

Llama 4 Scout is part of Meta's Llama 4 open model generation. It uses a Mixture-of-Experts architecture with 17 billion active parameters out of 109 billion in total, for strong text generation, chat and embeddings.

-   ### Mixture of Experts

    17B active parameters per token, 16 experts.

-   ### Text generation

    Writing, translation and question answering.

-   ### Code writing

    Generate and explain code.

-   ### OpenAI-compatible API

    Chat, completions, embeddings and model listing.

-   ### API docs included

    Test endpoints from the dashboard URL.

-   ### Private by default

    Prompts and outputs stay in your own account.

## What's in the Meetrix Llama 4 Scout AMI

The AMI ships Llama 4 Scout and an OpenAI-compatible API server on Ubuntu 22.04, with GPU drivers ready.

-   Llama 4 Scout with an OpenAI-compatible API
-   A llama systemd service you can restart if it hangs
-   Automatic SSL for your domain
-   A certificate script for issuing SSL manually otherwise
-   A CloudFormation template for your own AWS account
-   Email support from Meetrix at aws@meetrix.io

## How to set up a self-hosted Llama 4 Scout API

1.  ### Check your GPU quota

    Make sure your account has vCPU quota for g4dn instances in your region.

2.  ### Launch the stack

    Subscribe on AWS Marketplace and launch the stack. Double-check the admin email; it cannot be changed later.

3.  ### Point your domain at it

    After 5-10 minutes, create a DNS record with the PublicIp from the stack outputs.

4.  ### Call the API

    Open the DashboardUrl to test the endpoints, then point your OpenAI client at your server.

## Deploy on AWS

### Llama 4 Scout on AWS

An AMI on AWS Marketplace, launched on a g4dn GPU instance in your own AWS account with a CloudFormation stack.

Deploys with

CloudFormation

Recommended size

g4dn.metal

OS

Ubuntu 22.04

#### Setup guides

-   [**Developer guide** Launch, SSL, API, testing](https://meetrix.io/blogs/llama-4-developer-guide/)
-   [**Llama 4 Scout on AWS Marketplace** Use cases and technical highlights](https://meetrix.io/blogs/llama-4-scout-aws-meetrix/)
-   [**Video walkthrough** Watch the deployment on YouTube](https://www.youtube.com/watch?v=zvRbySK6Q6w)

 [Launch on AWS](https://aws.amazon.com/marketplace/pp/prodview-iu7qh5qnnbjt6)

## Llama 4 Scout API endpoints

| Endpoint | Purpose |
| --- | --- |
| /v1/chat/completions | Chat completions from a list of messages |
| /v1/completions | Completions from a prompt |
| /v1/embeddings | Embeddings for input text |
| /v1/models | List available models |

The developer guide includes a small Node.js script that checks every endpoint. [Read the Llama 4 developer guide →](https://meetrix.io/blogs/llama-4-developer-guide/)

## Video: Llama AMI installation walkthrough (shown with Llama 3)

## Llama 4 Scout FAQ

What is Llama 4 Scout?

A Meta Llama 4 model with a Mixture-of-Experts architecture: 17 billion active parameters, 109 billion in total.

Which instance should I use?

The AWS listing recommends g4dn.metal.

Is it compatible with the OpenAI API?

Yes. It serves the OpenAI chat, completions, embeddings and models endpoints.

Can I change the admin email later?

No. It is used for the SSL certificate and cannot be changed after the stack is created.

What if the Llama service hangs?

SSH in and restart the llama service, then check the disk is not full.

What if I see a 502 Bad Gateway error?

The model is still loading. Wait about 5 minutes and refresh.

How do I upgrade?

Back up your server data, remove the old stack and launch the new version from AWS Marketplace.

## Llama guides and articles

### [Llama 4 Scout by Meetrix on AWS Marketplace](https://meetrix.io/blogs/llama-4-scout-aws-meetrix/)

Use cases, technical highlights and the Meetrix AMI.

### [Best Open Source LLMs to Self-Host in 2026](https://meetrix.io/blogs/best-open-source-llms-self-hosted-2026/)

How Llama 4 compares with other open models.

### [Llama 4: Developer Guide](https://meetrix.io/blogs/llama-4-developer-guide/)

Launch the stack, set up SSL and test the API.

## Need a hand with Llama 4?

We build and run self-hosted AI infrastructure every day, from model selection to GPU sizing. Tell us what you need.

[Contact us](https://meetrix.io/contact-us/)
