Skip to main content

Build your private AI infrastructure

Run your models on infrastructure you control, govern access through AI Border Gateway (AIBG), and connect your developer tools to that gateway. This guide covers a personal machine, an existing GPU server, and cloud infrastructure in your own account.

Here, private means model inference and gateway data stay within your chosen infrastructure. Downloading images and model weights still needs network access. Agent telemetry, external tools, web search, and hosted fallback models need separate review; this setup alone does not make every application offline.

Before you start — Download the Cortega install package​

Download cortega-<version>.tar.gz from Cortega releases and extract it. Choose the install-package asset, rather than GitHub's Source code archives. Replace <version> below with the downloaded version:

tar -xzf cortega-<version>.tar.gz
cd cortega-<version>

This single package includes the Cortega platform installer and the deploy/self-hosted-llm/ scripts. Use the extracted package directory for all commands in this guide; no source checkout is needed. Its README.md explains the platform installation options.

Step 1 — Choose and run your LLMs​

Start at Model Intelligence. Enter your workloads, usage, and priorities, select Only open-weight models, and generate recommendations. Check each model's license, context requirements, and tool-use capabilities before choosing it. Open weights do not necessarily mean an unrestricted open-source license.

The cost modes answer different questions:

ModelInt cost modeHow to use it for this guide
On Prem (Cortega + LLM)Compare self-hosted inference costs. Use the chosen model name with the packaged installer.
Locally Hosted (Buy Your Own Hardware)Compare hardware purchases and sizing for local inference. These recommendations are sizing guidance, not a universal installer.
On Prem (Cortega only)Uses external model APIs; does not meet the private-inference goal here.
Hosted (Cortega.AI)A hosted service option, not an LLM deployment into your own cloud account.

Choose a runtime path below. Each ends with an OpenAI-compatible endpoint that AIBG can reach.

Option A — Personal use with Ollama​

Install Ollama, then find the model and exact tag in the Ollama model library. Check the available variants against your machine's memory; a ModelInt recommendation does not guarantee that an equivalent Ollama package exists.

Run the local model you selected:

ollama run <local-model-tag>
curl http://localhost:11434/v1/models

Use a local model, not a cloud variant. Set OLLAMA_NO_CLOUD=1 for the Ollama server and restart it to disable cloud features, following the Ollama configuration instructions.

The local OpenAI-compatible base URL is http://localhost:11434/v1. Ollama ignores the API key on this local interface; use a placeholder such as ollama if provider configuration requires one. See Ollama API compatibility.

For a containerized or remote AIBG gateway, localhost points to the gateway itself. Configure a reachable host address and Ollama's bind address as needed, and restrict access to your gateway's network. The local Ollama endpoint does not authenticate requests.

Option B — vLLM on an existing GPU server​

Use ModelInt's BYOH mode to estimate the hardware you need, whether you already own it or plan to buy it. The Cortega deployment scripts currently target Linux NVIDIA GPU hosts with working GPU drivers, SSH, sudo, and an apt-based installation environment. They do not install all the Intel, AMD, or Apple hardware configurations shown by ModelInt.

Use ModelInt to choose a model, then pass its catalog ID or full Hugging Face repository name to the installer. ModelInt provides recommendations and sizing; there is no configuration file to download or import.

From the extracted Cortega package directory, run:

./deploy/self-hosted-llm/questionnaire.sh private-chat --model <catalog-id-or-hf-repo>

Review the proposed configuration and choose no when asked to provision and deploy now, so you can configure your existing host first. The scripts require ssh, scp, openssl, curl, and jq (plus AWS CLI v2 for AWS provisioning).

Edit deploy/self-hosted-llm/deployments/private-chat/.env:

EXISTING_HOST=<gpu-host-address>
SSH_KEY_PATH=<absolute-path-to-your-ssh-key>
SSH_USER=ubuntu
VLLM_PROVIDER_BASE_URL=http://<private-address-reachable-from-aibg>:8000/v1

Verify the proposed GPU count, tensor parallelism, context size, and disk needs against the actual machine. Set HF_TOKEN for gated model repositories, and pin VLLM_IMAGE to a version you have validated. Then deploy:

./deploy/self-hosted-llm/deploy.sh private-chat
./deploy/self-hosted-llm/manage.sh private-chat smoke
./deploy/self-hosted-llm/manage.sh private-chat cortega-payload

The deployment installs the container runtime components and starts vLLM. Keep the printed provider credentials secure for Step 2. If you have not chosen a model, omit --model to use the wizard's budget and workload questionnaire.

Option C — vLLM in your cloud account​

For new AWS infrastructure, run the same questionnaire with your model name. Review the proposed hardware and cost, then confirm provisioning and deployment:

./deploy/self-hosted-llm/questionnaire.sh private-chat --model <catalog-id-or-hf-repo>

The wizard runs setup, deployment, and a smoke test, then prints the provider registration fields. You can decline deployment to edit the generated .env first, then run setup.sh private-chat and deploy.sh private-chat from deploy/self-hosted-llm/.

Set your AWS region, confirm GPU quota, and restrict VLLM_ALLOWED_CIDR to the network that needs to reach inference. AWS provisioning creates billable resources. ModelInt costs are estimates; check the selected instance and storage pricing before launch.

For another cloud, provision a compatible NVIDIA GPU host yourself, then use Option B. Automatic infrastructure provisioning currently supports AWS; the existing-host path can deploy to a compatible machine in another cloud.

For multiple models, repeat with distinct deployment names and enough GPU capacity. These scripts manage single-host deployments, not distributed multi-node model serving. See the self-hosted LLM walkthrough for operations and teardown commands.

Checkpoint: your model answers a direct test request, and its private endpoint is reachable from the network where AIBG will run.

Step 2 — Download AIBG and connect your models​

Use the same extracted Cortega install package to deploy AI Border Gateway. Follow its README.md and the installation guides. Choose a local trial, an existing environment, or AWS to match your topology. Place the gateway and its data stores inside your chosen privacy boundary, with network access to the model hosts.

  1. In Admin → Providers, add a Custom provider for your OpenAI-compatible endpoint. For the vLLM script path, use the fields printed by manage.sh private-chat cortega-payload; registration is manual.
  2. Register the served model under that provider. For Ollama, use the exact local model tag; for vLLM, use the generated model payload.
  3. Configure the team's model access and issue a virtual key (ck_...). Keep upstream credentials in AIBG; clients use the virtual key.
  4. Test the model in the console's Model Playground with the intended caller identity. Confirm routing selects your private provider.

The user guide walks through providers, models, teams, and keys. AIBG capabilities explains routing, access controls, budgets, guardrails, and observability. For private inference, keep every permitted route and fallback on private providers, including any separately configured models used by auxiliary AI features.

Checkpoint: a governed request through AIBG reaches your self-hosted model.

Step 3 — Connect your developer tools and agents​

Install and enroll Endpoint Guard, then open Endpoint Guard → AI Assistants:

  1. Enable Route AI assistant traffic through AIBG.
  2. Choose the default team for users without an assigned team. Ensure the user's actual team routes to the private models configured in Step 2.
  3. Check the supported assistants you want Endpoint Guard to configure: Claude Code, Codex CLI, OpenCode, Pi, Cursor, Goose, and Aider.
  4. Keep the Endpoint Guard native app running. It detects installed assistants and applies their gateway settings using one Cortega API key per device, shared across the enabled assistants. Let the device receive the settings, then restart tools that need to reload them; Cursor requires a restart.
  5. Test a representative task, including streaming and tool calls, and verify that AIBG routes it to your private model.

Endpoint Guard maintains this configuration as the selection changes, reducing the need to configure each user's tools manually. Unchecking a tool removes its managed Cortega configuration. Disabling the master switch revokes existing device assistant keys tenant-wide; enabling it again reactivates them.

See AI Assistants setup for the full workflow. GitHub Copilot Chat and Windsurf are not automatically configured. Check each VS Code extension separately; not every AI feature in an editor supports a custom model endpoint.

Alternative — Configure clients manually​

For supported clients you prefer to manage yourself, open AI Border Gateway → Client Setup. Copy the client-specific endpoint and settings, provide a virtual key (ck_...), and select the registered model. Follow Connecting clients. Manual settings need to be maintained across users, machines, and tool updates.

For either setup method, verify model compatibility: configuring a gateway URL does not make every private model compatible with every agent. Codex requires streaming tool calls on the Responses API, which not every self-hosted endpoint supports. Test the actual model and route before rollout.

Govern other AI traffic​

Endpoint Guard's Apps policies separately observe, police, or block watched traffic to its original destination. That proxy path does not replace external SaaS models with your private model. Review browser assistants, external tools, proxy coverage, and fail-open settings as part of your privacy boundary. See Endpoint Guard for both capabilities.

Checkpoint: a real request from each enabled tool appears in AIBG observability and is attributed to the intended caller and private model.

Verify your privacy boundary​

Before using sensitive data, check the complete request path:

  • Confirm inference runs on your own model hosts, and test both normal routing and failure behavior. An unavailable private model must not silently route to an external provider.
  • Check agent web search, MCP servers, extensions, telemetry, and background services separately. These can communicate externally even when inference is private.
  • Review where prompts, responses, traces, backups, and operational logs are stored, who can access them, and how long they are retained.
  • Verify network controls restrict direct model access and undesired outbound traffic. Separate installation-time downloads from runtime requirements.
  • Repeat the client test after changing models, routes, tools, or device policy.

The result is a verified private inference path with centrally governed access, plus an explicit account of any remaining external dependencies.