Self-Hosting an Open-Weight Model for Cortega
Guide for platform administrators who want to run their own open-weight LLM (Qwen, Llama, Mistral, GLM, DeepSeek, Kimi K2, and similar) behind Cortega, instead of — or alongside — a hosted provider like OpenAI or Anthropic.
Overview
Cortega can govern traffic to any OpenAI-compatible model endpoint, including
one you run yourself. The deploy/self-hosted-llm/ wizard handles the parts
that are otherwise tedious to get right by hand, in one command:
- Sizing — give it a model, it proposes a GPU instance type and cost estimate; you confirm or change it before anything is provisioned.
- Provisioning and deploy — creates the EC2 host (or deploys straight onto a machine you already have), installs vLLM, and waits for it to come up.
- Verification — runs a smoke-test chat completion automatically.
- Multiple models — run several self-hosted models side by side, each independently manageable and independently disposable.
Registering the resulting endpoint with Cortega (creating the provider and model) is a short manual step at the end — the wizard prints the exact JSON to paste in, but does not call the Cortega API for you.
Before you start
Download the install-package archive cortega-<version>.tar.gz from
Cortega releases,
not GitHub's Source code archives. Extract it and enter the package directory:
tar -xzf cortega-<version>.tar.gz
cd cortega-<version>
Replace <version> with your downloaded version. The same package includes
the Cortega platform installer and deploy/self-hosted-llm/. Run all commands
below from this directory; use the included README.md to install Cortega.
You'll need:
- An AWS account with GPU instance quota (see Instance quota below), or an existing GPU machine (on-prem, another cloud, an EC2 instance you already manage) reachable over SSH.
awsCLI v2,ssh,scp,openssl,curl, andjqinstalled wherever you run the wizard.- A model in mind. If you don't have one yet, Model Intel is a separate, purely informational recommender —
industry/application questions, real benchmark scores across a large
open-weight catalog, and hardware sizing for cloud and local hosting. It
has no config to import into this wizard; once you've picked a model
there, just give its name to
questionnaire.shbelow. If you'd rather let this wizard suggest something from its own small catalog, leave the model prompt blank and it'll ask for your budget, expected concurrent users, and use case instead.
One command: model in, running endpoint out
./deploy/self-hosted-llm/questionnaire.sh my-model --model qwen2.5-14b-awq
or a full HuggingFace repo not in this repo's own catalog:
./deploy/self-hosted-llm/questionnaire.sh my-model --model Qwen/Qwen2.5-32B-Instruct-AWQ
or leave --model off entirely for an interactive prompt (which also lets
you fall back to a budget/use-case shortlist if you don't know which model
you want):
./deploy/self-hosted-llm/questionnaire.sh my-model
It resolves the model (an exact or fuzzy match against catalog.json, or a
sizing estimate from the model's inferred parameter count if it isn't
catalogued), then prints the proposed instance type, GPU, quantization,
context window, and estimated cost:
Proposed hardware (from catalog.json):
Instance type: g5.2xlarge (1x A10G)
Tensor parallel size: 1
Context window: 32768
Root volume: 150GB
Quantization: auto-detect from checkpoint
Est. cost: $1.21/hr on-demand (~$883/mo if left running)
Use this configuration? [Y/n]
Say yes to accept it, or no to override any field (instance type, tensor parallel size, root volume, context window, quantization) before continuing.
Confirm and it asks one more question — provision and deploy now? — then
(unless you say no) runs setup.sh and deploy.sh for you, waits for the
model to finish loading, runs a smoke-test chat completion, and prints the
Cortega provider/model registration JSON. One command, no manual
stitching-together of separate steps.
Already have a GPU machine? Skip AWS provisioning: open
deploy/self-hosted-llm/deployments/my-model/.env and set EXISTING_HOST
(its IP or hostname) and SSH_KEY_PATH before answering "yes" to deploy —
the wizard detects it and skips setup.sh automatically. No AWS resources
are created in this path.
Instance quota
New AWS accounts often start with a 0-vCPU quota for GPU instance families.
setup.sh checks this before launching and fails fast with a clear message
naming the quota to raise, rather than a cryptic launch error. If you hit
this, request the increase in the AWS Console under Service Quotas ▸
Amazon EC2 and rerun setup.sh once it's approved (approval can take from
minutes to a day depending on the instance family and your account history).
Register with Cortega
The wizard's last step prints two JSON blocks. In Cortega, go to Admin ▸ Providers ▸ Add Provider, choose Custom, and paste the first block's fields. Once the provider is saved, add the model under it using the second block's fields. Your self-hosted model now appears anywhere Cortega lets you pick a model — team routing, virtual models, budgets, guardrails — exactly like a hosted provider's model.
Re-print it any time without redeploying:
./deploy/self-hosted-llm/manage.sh my-model cortega-payload
Verifying it works
The wizard already ran one smoke test automatically. To check again, or run a few latency samples:
./deploy/self-hosted-llm/manage.sh my-model smoke # one test chat completion, direct to the host
./deploy/self-hosted-llm/manage.sh my-model bench 5 # 5 latency samples
Then send a real request through Cortega itself (via a virtual key, to the model name you registered) to confirm it's reachable through the gateway, not just directly.
Running more than one model
Every command takes the deployment name as its first argument, so models run independently:
./deploy/self-hosted-llm/manage.sh list
shows every deployment you've created, its model, and its host.
Changing your mind / starting over
./deploy/self-hosted-llm/teardown.sh my-model
Terminates the EC2 instance and its security group (for the new-infrastructure path — nothing is torn down for an existing-machine deployment beyond local bookkeeping). The name is immediately free to reuse:
./deploy/self-hosted-llm/questionnaire.sh my-model --model <a-different-model>
Remove the old provider/model in Cortega's Admin ▸ Providers if you're not immediately replacing it with the new one under the same name.
Running unattended
For a fully hands-off run (no terminal prompts at any step), pass
--config to the questionnaire with a small file. Set MODEL_QUERY to skip
straight to the hardware proposal, or leave it unset and set
BUDGET_USD_MONTH/CONCURRENT_USERS/USE_CASE/SELECT_INDEX for the
budget-shortlist path instead. CONFIRM_HARDWARE and AUTO_DEPLOY (y/n,
default y) skip the two confirmation prompts:
cat > my-answers.env <<'EOF'
MODEL_QUERY=qwen2.5-14b-awq
CONFIRM_HARDWARE=y
AUTO_DEPLOY=y
EOF
./deploy/self-hosted-llm/questionnaire.sh my-model --config my-answers.env
runs the whole thing — resolve, propose, provision, deploy, smoke test, print the Cortega payload — end to end with no terminal attention.
Choosing a model
See deploy/self-hosted-llm/catalog.json for the full list with cost and
sizing detail. It spans:
| Tier | Examples | Typical host |
|---|---|---|
| Budget | Qwen2.5 7B, Llama 3.1 8B, GLM-4 9B | 1x A10G |
| Balanced | Qwen2.5 14B, Mistral Small 24B, DeepSeek-R1-Distill 32B | 1x A10G/L40S |
| Performance | Qwen2.5 72B, Llama 3.1 70B, GLM-4.5-Air | 4x L40S |
| Frontier | DeepSeek-V3, GLM-4.6, Kimi K2 | 8x H100/H200 |
The catalog is a hand-maintained snapshot refreshed periodically against public benchmarks and current AWS pricing — treat its cost estimates as directional, and verify pricing for your account/region before committing to the frontier tier in particular.