Model selection and updates¶
Cloud routes (Milestone 1)¶
The gateway supports openai, anthropic, azure-openai, bedrock, and vertex
alongside ollama and vllm. Cloud templates in the existing model catalog are
proposed, with placeholder model IDs, limits, and prices. They enable no external
traffic. Review provider terms, region, retention, model limits, and input/output
prices, then use the existing promotion workflow. MODEL_ROUTING_POLICY_PATH accepts
a ModelCatalog (only status: approved entries) or the existing ModelRoutingPolicy.
The catalog checker carries connection, pricing, and routing fields into the
approved environment's routing.policy.models.
An operator's Helm/GitOps values can contain:
runtime:
modelId: approved-cloud
allowedModels: [approved-cloud]
providerCredentials:
- env: OPENAI_API_KEY
secretName: model-provider-credentials
secretKey: openai-api-key
routing:
policy:
enabled: true
models:
- id: approved-cloud
backend: openai
connection:
baseUrl: https://api.openai.com/v1
model: YOUR_APPROVED_MODEL_ID
credentialEnv: OPENAI_API_KEY
pricing:
inputUsdPer1kTokens: 0 # replace with your contracted rate
outputUsdPer1kTokens: 0 # replace with your contracted rate
Create the referenced Secret using your existing secret backend; do not put credential
values in Git, catalog entries, Helm values, or client requests. Each connection reads
only its named environment variable. Credential-bearing URLs and inline credential
fields are rejected. Use HTTPS for real providers. The gateway's default NetworkPolicy
still denies external egress: provision an approved HTTPS egress policy for the
gateway namespace through the existing egress catalog, or route through an internal
egress proxy permitted by networkPolicy.runtimeEgress. Workspaces retain their
default-deny boundary and receive no provider credentials.
| Backend | connection.baseUrl |
connection.model / credential environment |
|---|---|---|
openai |
https://api.openai.com/v1 |
Approved model ID / OPENAI_API_KEY |
anthropic |
https://api.anthropic.com/v1 |
Approved Claude model ID / ANTHROPIC_API_KEY |
azure-openai |
https://RESOURCE.openai.azure.com/openai/v1 |
Azure deployment name / AZURE_OPENAI_API_KEY |
bedrock |
https://bedrock-runtime.REGION.amazonaws.com |
Bedrock model ID or inference profile / AWS_BEARER_TOKEN_BEDROCK |
vertex |
https://REGION-aiplatform.googleapis.com/v1/projects/PROJECT/locations/REGION/endpoints/openapi |
google/APPROVED_GEMINI_MODEL / VERTEX_ACCESS_TOKEN |
These adapters use OpenAI Chat Completions, Anthropic Messages, Azure's v1 API, Bedrock Converse with a Bedrock bearer API key, and Vertex's OpenAI-compatible API. Vertex needs an OAuth access token; the operator must refresh it and roll the gateway when an environment-backed Secret changes. Automatic credential acquisition/refresh and AWS SigV4 are not implemented. Readiness checks cloud credential presence, not provider availability or account permissions; actual calls use the shared circuit breaker and retry policy.
All five support chat, function tools, and streaming through Chat and Messages; Responses and synchronous batch items retain non-streaming behavior. Anthropic and Bedrock adapters accept text and function tools; unsupported modalities/parameters receive an explicit, receipted 400. OpenAI and Azure also proxy embeddings and legacy completions. The other adapters reject those endpoints explicitly. Model-specific API restrictions still apply. No live cloud credential or paid call is needed for the provider tests or Compose walkthrough.
fallbacks is an ordered preference: put a local route first for local preference,
or a cloud route first for cloud preference. Chat, Messages, Responses, and synchronous
batch items try the next permitted route on upstream overload (429/5xx), connection
failure, or an open circuit. Ordinary 4xx errors never fall back; streaming fallback
stops once output begins. Gateway-wide load shedding still returns 503 before routing.
Embeddings and legacy completions retain single-route behavior.
Set data_classification: confidential in an inference body, or send
X-Data-Classification: confidential. Tenant dataClassification is a floor: neither
a body nor a header can lower it. Values are public, internal, confidential, and
restricted; the last two permit only local routes, including fallback, canary,
shadow, and cache selection. If no eligible local route can serve the request, the
gateway returns 403 data_classification_denied with a hash-chained refusal receipt.
Stored Responses and uploaded batch files/jobs preserve the floor during chaining and
replay. Bind tenant identity to a key record or verified JWT claim for this to be a
tenant security boundary.
Model metadata was reviewed against the publishers' repositories on 2026-09-23 for the v0.2.0 release. The catalog separates models approved for the existing lab profiles from newer candidates that still need evaluation.
Models used by the shipped profiles¶
| Purpose | Model | Release behavior |
|---|---|---|
| Local CPU smoke tests | qwen2.5:0.5b |
Retains the small, non-reasoning model used by the quickstart. Its Ollama weight-layer digest was reverified. |
| Customer CPU lab | qwen3.5:0.8b |
Retains the customer reasoning model. Its Ollama weight-layer digest was reverified. |
| Customer coding agents | Qwen/Qwen3-Coder-Next |
Now pins the upstream commit and all safetensors shard checksums in a reproducible inventory. |
| Customer RAG embeddings | BAAI/bge-small-en-v1.5 |
Now pins the upstream commit and weight inventory; the embedding deployment uses the same revision. |
These approvals describe lab profiles. They do not establish production suitability for a customer's workloads. Read the model cards for the existing evidence and limitations.
Current GPU candidates¶
All rows below have status: proposed and are absent from the gateway allowlists.
Context sizes are upstream configuration values, not measured capacity or enabled
gateway limits. Candidate licenses and revision links are recorded in
platform/model-catalog/models.yaml.
| Candidate | Upstream license | Context tokens | Evaluation focus |
|---|---|---|---|
| Qwen3.8-27B | Apache-2.0 | 262,144 | Dense model for coding and agent workloads; first candidate to compare with the current coding profile. |
| Qwen3.8-Flash-Next | Qwen Community 1.0 | 262,144 | Larger architecture with custom license terms and model-specific serving requirements. |
| GLM-5.3-Flash | MIT | 1,048,576 | Large multi-GPU candidate; validate the publisher's serving recipe and hardware requirements. |
| DeepSeek-V4.1-Flash | MIT | 1,048,576 | Large model with a specialized architecture; validate runtime support before capacity testing. |
| Qwen3.6-35B-A3B | Apache-2.0 | 262,144 | Retained as a comparison candidate with verified upstream context metadata. |
Qwen3.8-Flash-Next is not Apache-2.0. Its publisher uses the
Qwen Community 1.0 license.
The catalog records LicenseRef-Qwen-Community-1.0 so its terms cannot be confused
with the permissive license of Qwen3.8-27B.
The newer models include upstream multimodal capabilities. Their presence in the catalog does not add image or video support to the gateway. The pinned runtime image and existing GPU profiles have not been validated with these candidates. Use the upstream recipe, then run AgentWorkflows' real-model evals and load tests before changing a serving profile or approving a model.
Reproduce the approved model metadata¶
The first two commands validate local catalog, manifest, and profile consistency. The third contacts the Ollama registry and Hugging Face metadata API. It reproduces the approved Ollama layer digests and Hugging Face weight inventories without downloading model weights.
A manifest digest identifies the inventory of weight filenames, byte sizes, and upstream SHA-256 checksums at one immutable commit. It is not a checksum of the entire model, and it does not verify bytes already installed in a customer model store. Compare those files against the inventory during model-store ingestion.
See the provenance runbook for the manifest format and update commands.
Self-hosted model revisions¶
- The default vLLM chart and approved customer profiles now set
model.revision. Custom overlays that changemodel.namemust also set the matching revision. - The embedding profile explicitly overrides the coding-model pin with the BGE revision. Keep these revisions distinct.
- The AWQ example is an operator template. It now names an explicit customer checkpoint placeholder instead of implying that an official Qwen AWQ repository exists. Supply an approved checkpoint and revision before using it.
- No new candidate is automatically deployed or promoted. Follow the model governance runbook to collect evaluation, load, and security evidence for a promotion.