Skip to main content

Self-Hosting an Open-Weight Model for Cortega

Guide for platform administrators who want to run their own open-weight LLM (Qwen, Llama, Mistral, GLM, DeepSeek, Kimi K2, and similar) behind Cortega, instead of — or alongside — a hosted provider like OpenAI or Anthropic.


Overview​

Cortega can govern traffic to any OpenAI-compatible model endpoint, including one you run yourself. The deploy/self-hosted-llm/ wizard handles the parts that are otherwise tedious to get right by hand, in one command:

  • Sizing — give it a model, it proposes a GPU instance type and cost estimate; you confirm or change it before anything is provisioned.
  • Provisioning and deploy — creates the EC2 host (or deploys straight onto a machine you already have), installs vLLM, and waits for it to come up.
  • Verification — runs a smoke-test chat completion automatically.
  • Multiple models — run several self-hosted models side by side, each independently manageable and independently disposable.

Registering the resulting endpoint with Cortega (creating the provider and model) is a short manual step at the end — the wizard prints the exact JSON to paste in, but does not call the Cortega API for you.

Before you start​

Download the install-package archive cortega-<version>.tar.gz from Cortega releases, not GitHub's Source code archives. Extract it and enter the package directory:

tar -xzf cortega-<version>.tar.gz
cd cortega-<version>

Replace <version> with your downloaded version. The same package includes the Cortega platform installer and deploy/self-hosted-llm/. Run all commands below from this directory; use the included README.md to install Cortega.

You'll need:

  • An AWS account with GPU instance quota (see Instance quota below), or an existing GPU machine (on-prem, another cloud, an EC2 instance you already manage) reachable over SSH.
  • aws CLI v2, ssh, scp, openssl, curl, and jq installed wherever you run the wizard.
  • A model in mind. If you don't have one yet, Model Intel is a separate, purely informational recommender — industry/application questions, real benchmark scores across a large open-weight catalog, and hardware sizing for cloud and local hosting. It has no config to import into this wizard; once you've picked a model there, just give its name to questionnaire.sh below. If you'd rather let this wizard suggest something from its own small catalog, leave the model prompt blank and it'll ask for your budget, expected concurrent users, and use case instead.

One command: model in, running endpoint out​

./deploy/self-hosted-llm/questionnaire.sh my-model --model qwen2.5-14b-awq

or a full HuggingFace repo not in this repo's own catalog:

./deploy/self-hosted-llm/questionnaire.sh my-model --model Qwen/Qwen2.5-32B-Instruct-AWQ

or leave --model off entirely for an interactive prompt (which also lets you fall back to a budget/use-case shortlist if you don't know which model you want):

./deploy/self-hosted-llm/questionnaire.sh my-model

It resolves the model (an exact or fuzzy match against catalog.json, or a sizing estimate from the model's inferred parameter count if it isn't catalogued), then prints the proposed instance type, GPU, quantization, context window, and estimated cost:

Proposed hardware (from catalog.json):
Instance type: g5.2xlarge (1x A10G)
Tensor parallel size: 1
Context window: 32768
Root volume: 150GB
Quantization: auto-detect from checkpoint
Est. cost: $1.21/hr on-demand (~$883/mo if left running)

Use this configuration? [Y/n]

Say yes to accept it, or no to override any field (instance type, tensor parallel size, root volume, context window, quantization) before continuing.

Confirm and it asks one more question — provision and deploy now? — then (unless you say no) runs setup.sh and deploy.sh for you, waits for the model to finish loading, runs a smoke-test chat completion, and prints the Cortega provider/model registration JSON. One command, no manual stitching-together of separate steps.

Already have a GPU machine? Skip AWS provisioning: open deploy/self-hosted-llm/deployments/my-model/.env and set EXISTING_HOST (its IP or hostname) and SSH_KEY_PATH before answering "yes" to deploy — the wizard detects it and skips setup.sh automatically. No AWS resources are created in this path.

Instance quota​

New AWS accounts often start with a 0-vCPU quota for GPU instance families. setup.sh checks this before launching and fails fast with a clear message naming the quota to raise, rather than a cryptic launch error. If you hit this, request the increase in the AWS Console under Service Quotas ▸ Amazon EC2 and rerun setup.sh once it's approved (approval can take from minutes to a day depending on the instance family and your account history).

Register with Cortega​

The wizard's last step prints two JSON blocks. In Cortega, go to Admin ▸ Providers ▸ Add Provider, choose Custom, and paste the first block's fields. Once the provider is saved, add the model under it using the second block's fields. Your self-hosted model now appears anywhere Cortega lets you pick a model — team routing, virtual models, budgets, guardrails — exactly like a hosted provider's model.

Re-print it any time without redeploying:

./deploy/self-hosted-llm/manage.sh my-model cortega-payload

Verifying it works​

The wizard already ran one smoke test automatically. To check again, or run a few latency samples:

./deploy/self-hosted-llm/manage.sh my-model smoke # one test chat completion, direct to the host
./deploy/self-hosted-llm/manage.sh my-model bench 5 # 5 latency samples

Then send a real request through Cortega itself (via a virtual key, to the model name you registered) to confirm it's reachable through the gateway, not just directly.

Running more than one model​

Every command takes the deployment name as its first argument, so models run independently:

./deploy/self-hosted-llm/manage.sh list

shows every deployment you've created, its model, and its host.

Changing your mind / starting over​

./deploy/self-hosted-llm/teardown.sh my-model

Terminates the EC2 instance and its security group (for the new-infrastructure path — nothing is torn down for an existing-machine deployment beyond local bookkeeping). The name is immediately free to reuse:

./deploy/self-hosted-llm/questionnaire.sh my-model --model <a-different-model>

Remove the old provider/model in Cortega's Admin ▸ Providers if you're not immediately replacing it with the new one under the same name.

Running unattended​

For a fully hands-off run (no terminal prompts at any step), pass --config to the questionnaire with a small file. Set MODEL_QUERY to skip straight to the hardware proposal, or leave it unset and set BUDGET_USD_MONTH/CONCURRENT_USERS/USE_CASE/SELECT_INDEX for the budget-shortlist path instead. CONFIRM_HARDWARE and AUTO_DEPLOY (y/n, default y) skip the two confirmation prompts:

cat > my-answers.env <<'EOF'
MODEL_QUERY=qwen2.5-14b-awq
CONFIRM_HARDWARE=y
AUTO_DEPLOY=y
EOF
./deploy/self-hosted-llm/questionnaire.sh my-model --config my-answers.env

runs the whole thing — resolve, propose, provision, deploy, smoke test, print the Cortega payload — end to end with no terminal attention.

Choosing a model​

See deploy/self-hosted-llm/catalog.json for the full list with cost and sizing detail. It spans:

TierExamplesTypical host
BudgetQwen2.5 7B, Llama 3.1 8B, GLM-4 9B1x A10G
BalancedQwen2.5 14B, Mistral Small 24B, DeepSeek-R1-Distill 32B1x A10G/L40S
PerformanceQwen2.5 72B, Llama 3.1 70B, GLM-4.5-Air4x L40S
FrontierDeepSeek-V3, GLM-4.6, Kimi K28x H100/H200

The catalog is a hand-maintained snapshot refreshed periodically against public benchmarks and current AWS pricing — treat its cost estimates as directional, and verify pricing for your account/region before committing to the frontier tier in particular.