> Source: https://meetrix.io/store/llama-3/
> Markdown copy of that page. Cite the URL above, not this file.

[Store](https://meetrix.io/store/)  Llama 3

Open-source LLM

# Self-hosted Llama 3 on AWS

Run Meta's Llama 3, in 8B or 70B, behind an OpenAI-compatible API on a GPU instance in your own AWS account. Point your OpenAI tools at your own endpoint and keep your data private.

 [Browse Meetrix on AWS](https://aws.amazon.com/marketplace/seller-profile?id=62b53dd8-45d0-4b50-a654-bd6993167486)

-   8B and 70B models
-   OpenAI-compatible API
-   Pay per hour

![Llama 3 deployed on AWS by Meetrix](https://meetrix.io/assets/store/llama-3/llama-3.webp)

## What is Llama 3?

Llama 3 is Meta's open large language model family, with refined post-training and better scalability than Llama 2. It handles language understanding, translation, dialogue, reasoning and code generation.

-   ### Two sizes

    8B for cost and speed, 70B for harder tasks.

-   ### Dialogue and reasoning

    Assistants, Q&A and multi-step answers.

-   ### Code generation

    Write and explain code.

-   ### OpenAI-compatible API

    Chat, completions, embeddings and model listing endpoints.

-   ### Switch models

    Pick a model id from /v1/models and pass it in your request.

-   ### Private by default

    Prompts and outputs stay in your own account.

## What's in the Meetrix Llama 3 AMIs

Each AMI ships the model and an OpenAI-compatible API server on Ubuntu 22.04, ready on a g4dn GPU instance.

-   Llama 3 8B or 70B with an OpenAI-compatible API
-   A llama systemd service you can restart if it hangs
-   Automatic SSL when your domain is hosted on Route 53
-   A certificate script for issuing SSL manually otherwise
-   A CloudFormation template for your own AWS account
-   Email support from Meetrix at aws@meetrix.io

## How to set up a self-hosted Llama 3 API

1.  ### Pick a size

    Choose 8B or 70B and check your account has vCPU quota for the matching g4dn instance.

2.  ### Launch the stack

    Subscribe on AWS Marketplace and launch the CloudFormation stack with your domain and admin email.

3.  ### Point your domain at it

    After 5-10 minutes, create a DNS record with the PublicIp from the stack outputs.

4.  ### Call the API

    Open the DashboardUrl for the API docs, then point your OpenAI client at your server.

## Deploy on AWS

### Llama 3 on AWS

Two AMIs on AWS Marketplace, each launched on a g4dn GPU instance in your own AWS account with CloudFormation.

Deploys with

CloudFormation

8B instance

g4dn.xlarge

70B instance

g4dn.metal

#### Setup guides

-   [**Developer guide** Launch, SSL, API, switching models](https://meetrix.io/blogs/llama3-developer-guide/)
-   [**Install Llama 3 on AWS** Setup, comparison and best practices](https://meetrix.io/blogs/how-to-install-llama-3/)
-   [**Video walkthrough** Watch the deployment on YouTube](https://www.youtube.com/watch?v=zvRbySK6Q6w)

 [Browse Meetrix on AWS](https://aws.amazon.com/marketplace/seller-profile?id=62b53dd8-45d0-4b50-a654-bd6993167486)

## Llama 3 models

| Model | Recommended instance | AWS |
| --- | --- | --- |
| Llama 3 8B | g4dn.xlarge | [Launch on AWS](https://aws.amazon.com/marketplace/pp/prodview-pc5rinshvui7o) |
| Llama 3 70B | g4dn.metal | [Launch on AWS](https://aws.amazon.com/marketplace/pp/prodview-f7225w6emwfu6) |

The 8B listing supports g4dn instances from xlarge up to metal.

## Llama 3 API endpoints

| Endpoint | Purpose |
| --- | --- |
| /v1/chat/completions | Chat completions from a list of messages |
| /v1/completions | Completions from a prompt |
| /v1/embeddings | Embeddings for input text |
| /v1/models | List available models |

Switching models takes a little longer on the first response. [Read the Llama 3 developer guide →](https://meetrix.io/blogs/llama3-developer-guide/)

## Video: deploy Llama 3 on AWS

## Llama 3 FAQ

What is Llama 3?

Meta's large language model family with improved post-training, for understanding, translation, dialogue, reasoning and code generation.

Which size should I choose?

8B on g4dn.xlarge for lower cost; 70B on g4dn.metal for more capable output.

How do I switch models?

Call /v1/models, copy the id you want and use it as the model value in your request. The first response after a switch is slower.

What if the Llama service hangs?

SSH in and run sudo systemctl restart llama.service, wait a few minutes and reload the dashboard.

Is it compatible with the OpenAI API?

Yes. OpenAI SDKs work with your server as the base URL.

What if SSL does not set up automatically?

SSH in and run sudo /root/certificate\_generate\_standalone.sh with your admin email.

How do I upgrade?

Back up your server data, remove the old stack and launch the new version from AWS Marketplace.

## Llama guides and articles

### [How to Install Llama 3 on AWS via Pre-configured AMI](https://meetrix.io/blogs/how-to-install-llama-3/)

Setup steps, comparisons and cost management.

### [Best Open Source LLMs to Self-Host in 2026](https://meetrix.io/blogs/best-open-source-llms-self-hosted-2026/)

How Llama compares with other open models.

### [Self-Host an OpenAI-Compatible LLM API on AWS](https://meetrix.io/blogs/openai-compatible-api-self-hosted/)

Swap the OpenAI base URL without rewriting app code.

## Need a hand with Llama?

We build and run self-hosted AI infrastructure every day, from model selection to GPU sizing. Tell us what you need.

[Contact us](https://meetrix.io/contact-us/)
