Production Readiness Matrix¶
AgentWorkflows is cloud-first. The local trial demonstrates its governed provider path; production deployments must supply identity, persistence, and current evidence. The Kubernetes profiles are optional deployment templates. Customer clusters should keep the same interfaces and replace only the platform services they already operate, such as ingress, storage, secrets, logging, and GPU node pools.
Required Controls¶
| Area | Local lab implementation | Customer cluster expectation | Validation |
|---|---|---|---|
| Runtime isolation | Ollama in ollama, optional vLLM in vllm, gateway in inference |
Dedicated namespaces and runtime service accounts | make smoke RUNTIME_BACKEND=ollama |
| Accelerator portability | vLLM profiles for CPU-off local lab, NVIDIA, and AMD ROCm | Customer clusters expose nvidia.com/gpu or amd.com/gpu |
make production-check |
| Runtime high availability | Gateway replicas, vLLM replicas, HPA, PDBs, and topology spread | Size min/max replicas to SLOs and GPU inventory | helm template in make validate |
| Traceability | X-Request-ID, X-Sandbox-ID, optional traceparent, JSON audit events |
Forward the same headers through ingress and log pipeline | make smoke, make trace-smoke, gateway tests |
| API authentication | Gateway and RAG business endpoints require API keys in local/customer values | Back hashes with customer secret manager and rotate keys through External Secrets | gateway/RAG auth tests, make smoke, make rag-smoke |
| API contracts | platform/api-contracts/ stores OpenAPI snapshots for gateway and RAG with stable operation IDs and auth declarations |
Review contract diffs before changing customer-facing routes, request schemas, or auth semantics | make api-contract, make api-contract-update |
| Configuration contracts | platform/config-contracts/ stores service runtime env snapshots and checks them against Python settings, Helm templates, and chart defaults |
Review config diffs before changing customer overlays, secrets, budgets, retrieval settings, or runtime endpoints | make config-contract, make config-contract-update |
| Prompt privacy | Audit logs include prompt length and SHA-256 only | Do not log raw prompt text by default | test_audit_log_redacts_prompt_content |
| Data retention | platform/governance/data-retention.yaml covers audit logs, generated evidence, RAG knowledge, agent workspace data, and model governance records |
Align retention days and classification to customer policy | make retention-check, make retention-report |
| Model governance | Gateway ALLOWED_MODELS rejects unapproved model IDs |
Maintain an approved model catalog per environment | test_chat_completion_rejects_disallowed_model |
| Model catalog | platform/model-catalog/models.yaml and cluster ConfigMap |
Treat model additions as reviewed changes | make production-check |
| Model lifecycle | Approved models require promotion requests, evidence references, runtime metadata, and approved-only allowlists | Review promotion requests before adding models to gateway allowlists | make model-check, make model-report |
| Model provenance | Approved Hugging Face models use pinned commits and safetensors checksum inventories; Ollama records registry weight-layer digests | Verify downloaded weights against the inventory or registry digest before production use; record customer model-store changes in provenance | make model-provenance-check, make model-provenance-verify, make model-provenance-report |
| Admission control | Gateway caps message count, prompt size, completion tokens, temperature, and streaming | Tune limits by sandbox and runtime capacity | test_admission_policy_rejects_unsafe_or_expensive_requests |
| Prompt secret detection | Gateway rejects prompts that match configured credential patterns before runtime forwarding | Keep enabled for coding-agent workspaces and tune patterns only after review | gateway guardrail tests, make production-check |
| Output guardrail | Gateway inspects model completions for leaked credentials/PII/blocked content and flags, redacts, or blocks before return (OWASP LLM02:2025/LLM05:2025) | Enable guardrails.outputGuardrail; use non-streaming for hard redact/block enforcement |
gateway output-guardrail tests |
| Per-tenant RAG isolation | Retrieval is scoped to the caller's tenant via the owner payload field stamped at ingest; enforced by default on both backends and fail-closed |
Keep retrieval.tenantIsolation enabled for multi-tenant corpora (chart local single-tenant profile turns it off); stamp each source owner with the tenant id; front the service with the gateway or a header-stamping proxy since the tenant id is header-asserted |
RAG tenant-isolation tests |
| Validation toolchain | platform/tools/validation-toolchain.yaml declares validate, local, and strict profiles with a pinned Linux/CI installer |
Install the strict profile before customer handoff or production-readiness sign-off | make toolchain-install, make toolchain-doctor, make validate-full |
| SLO and error budget | platform/slo/objectives.yaml defines inference, eval, restore, and coding-agent platform objectives with alert references |
Align targets to the customer's contract and review burn-rate alerts | make slo-check, make slo-report |
| Sandbox budgets | Gateway enforces request, prompt-character, and estimated-token ceilings by X-Sandbox-ID |
Size limits by tenant and review overage events | test_sandbox_budget_status_and_request_limit_rejection |
| Shared budget backend | Local/customer values use Redis-compatible shared counters for multi-replica gateways; budgets fail closed on a Redis outage while the rate limiter can opt into fail-open (rateLimit.failOpen, default closed) |
Replace bundled Redis with external managed/Sentinel/Cluster Redis (runbook); choose the rate-limit fail policy per availability target | test_redis_budget_tracker_shares_usage_across_tracker_instances, test_rate_limit_fail_open_admits_when_backend_down |
| Quota and chargeback | platform/governance/quota-plans.yaml connects tenant quotas, gateway budgets, workspace sizing, and chargeback labels |
Align quota plans to customer showback or chargeback policy before onboarding tenants | make quota-check, make quota-report |
| Sandbox isolation | ai-sandbox namespace, quota, limits, default-deny network policy |
Per-team sandbox namespaces with quotas and egress allowlists | make trace-smoke |
| Tenant labs | make tenant-up and make tenant-smoke create team namespaces with quota, RBAC, trace contract, and network controls |
One namespace per team or approved experiment boundary | make tenant-smoke |
| Tenant onboarding | TenantOnboarding spec renders tenant controls and matching agent workspace values |
Review generated quota, RBAC, PVC, storage, and egress before apply | make tenant-onboard, scripts/tenant-onboard.py --check |
| Regulated offline tenant profile | tenants/onboarding/regulated-offline-coding-agents.yaml renders confidential, no-external-egress agent controls |
Use for offline or regulated teams and add egress only through reviewed catalog-backed changes | make tenant-onboard-regulated, make production-check |
| RAG service | Local retrieval service returns platform context and OpenAI-compatible grounded messages | Replace or extend the knowledge set with customer-approved internal docs | make rag-smoke, RAG service tests |
| Vector RAG profile | deploy/charts/qdrant-vector-store and customer RAG values provide a persistent Qdrant backend |
Size storage, vector dimensions, and ingestion to the customer's embedding model and approved document pipeline | make production-check, Qdrant/RAG Helm renders |
| Agent workspaces | agent-workspace chart creates a locked-down namespace, PVC, RBAC, trace contract, and approved egress for coding agents |
One workspace per team, project, or agent boundary with customer-approved external egress | make agent-smoke |
| Egress governance | platform/network/egress-catalog.yaml requires external agent egress to reference approved catalog entries |
Review and expire Git, package mirror, artifact, and ticketing egress entries | make egress-check, make egress-report |
| Chaos drills | Safe rollout drills for gateway, budget Redis, Ollama, RAG, Qdrant, vLLM, and GPU capacity preflight | Run after platform upgrades and before customer demos or maintenance windows | make chaos-drill, DRILL=gpu-capacity-preflight RUN_SMOKE=0 make chaos-drill |
| Evaluation harness | platform/evals/smoke-suite.yaml, platform/evals/coding-agent-suite.yaml, and make eval for repeatable prompt and coding-agent checks |
Maintain environment-specific suites and keep summaries as release evidence | scripts/eval-suite.py --check-config |
| Adversarial safety eval | platform/evals/safety-suite.yaml red-team battery (prompt-injection, jailbreak, data-exfiltration, unsafe-tool-use, and bias/fairness cases) gated by a safety release gate (minRefusalRate over a minCases floor) |
Extend the red-team suite per model and require it before promotion | SUITE=platform/evals/safety-suite.yaml make eval, make release-gate |
| RAG grounding eval | scripts/rag-eval.py scores retrieval plus RAGAS-style context precision and answer faithfulness |
Add ground-truth answers; gate on minFaithfulness/minContextPrecision |
make rag-eval-check |
| Model quality drift | Proxy alerts (governance.rules) plus scheduled eval comparison |
Schedule evals and alert routing; roll back on threshold breach | runbooks/model-drift-monitoring.md |
| Release gates | platform/slo/release-gates.yaml enforces eval, load, restore, strict toolchain, SLO, governance, supply-chain, and evidence-pack thresholds |
Run strict gates before demos, releases, restore reviews, and production-readiness handoff so checked-in sample evidence cannot pass | make release-gate, make release-gate-strict, make release-report-strict |
| Autoscaling | KEDA ScaledObject from Prometheus request rate | Tune thresholds to customer SLOs and GPU capacity | helm template in make validate |
| Observability | Prometheus metrics, Grafana dashboards, Loki + Promtail log pipeline | Centralize metrics, logs, and alerts | make validate, dashboards in deploy/observability/ |
| Distributed tracing | Gateway/RAG export OTLP spans to a Tempo backend wired as a Grafana datasource | Set observability.tracing endpoint; size Tempo retention/storage |
make dashboard-check, deploy/observability/applications.yaml |
| Cost and FinOps | OpenCost app, gateway estimated_cost_usd metric + /v1/usage, cost dashboards |
Map cost-center labels to chargeback; review cost dashboards | make dashboard-check, cost panels in inference/opencost dashboards |
| Encryption in transit | Plaintext HTTP data plane by default; opt-in mTLS/cert-manager overlay provided | Enable a CNI/mesh mTLS or cert-manager TLS control before regulated multi-tenant use | deploy/clusters/customer/mtls/ |
| Runtime threat detection | Optional Falco/Tetragon detective layer (opt-in Argo app) | Deploy into a Kyverno-excluded namespace; route alerts to the log pipeline | deploy/observability/runtime-security.yaml, runbook |
| Policy as code | Kyverno required labels, resources, pod hardening, read-only root filesystems, image signature audit | Enforce on AI namespaces and exclude platform operators | make policy-test when Kyverno CLI is installed |
| Cost controls | Required owner/cost/environment/sandbox labels and OpenCost app | Map labels to chargeback/showback taxonomy | make validate YAML checks |
| Secret handling | External Secrets examples and no committed runtime tokens | Replace local Kubernetes provider with enterprise backend | deploy/clusters/customer/external-secrets.yaml |
| Supply chain | Pinned Alpine runtime images, hashed Python dependency locks, runtime-only Python dependencies, high/critical Trivy image and repo failure gates, local SBOM/SARIF/checksum evidence, Cosign digest signing and release asset upload in the tag-only release workflow | Deploy only immutable signed image digests that you have scanned | make dependency-lock-check, make repo-security-scan, make image-scan, make supply-chain-check, GitHub Actions release workflow |
| Backup and restore | restore-drill application-data validation and Velero examples |
Run scheduled restore evidence for each critical data store | make restore-drill, make backup-drill |
| Disaster recovery | Single-cluster DR posture with named RPO/RTO and a whole-platform restore order | Provision an off-cluster backup target; design secondary-cluster/multi-region if required | runbooks/disaster-recovery.md |
| Model cards | Each approved model ships a card/datasheet referenced from the catalog | Keep cards current with promotion; treat as a review artifact | make model-check |
| Load testing | k6 chat-completion scenario with sandbox tags, live-gateway mode, and self-contained local gateway-path mode | Store summaries and compare against SLOs | make loadtest, make loadtest-local |
| Evidence pack | Static customer handoff report plus optional live Kubernetes readiness checks | Attach reports to release, demo, restore drill, or incident review evidence | make evidence, make evidence LIVE=1 |
Stateful stores: dev/reference footprints and their HA path¶
Four bundled stateful stores ship as single-node reference footprints so a laptop lab and a fresh cluster start with no external dependencies. They are deliberately not production topologies. State this plainly to any operator sizing a production environment, and swap each to its external/HA path before a regulated or multi-tenant handoff. The full opt-in procedure (with rollback) is in the external / managed stores runbook.
| Bundled store | Reference footprint | Production / HA path |
|---|---|---|
Temporal PostgreSQL (deploy/charts/workflows) |
1 replica and a persistent volume for history and visibility | Configure both Temporal SQL datastores for managed HA Postgres; back up both databases and validate restore before switching traffic |
Budget / response-cache Redis (deploy/charts/budget-redis) |
1 replica, AOF and PVC persistence, minAvailable: 0 (disk loss can lose counters; an outage fails budgets closed) |
Point budget.redisUrl / responseCache.redisUrl at an external managed Redis, Redis Sentinel failover pair, or Redis Cluster and stop syncing the bundled Application. Budgets stay fail-closed on outage; the rate limiter can opt into fail-open (rateLimit.failOpen) as an availability-vs-enforcement tradeoff |
Qdrant vector store (deploy/charts/qdrant-vector-store) |
Single-instance, schema-enforced (replicaCount max 1) on one RWO PVC |
Use an external managed Qdrant or a Qdrant cluster (sharded/replicated) and point retrieval.vectorStore.url at it; the bundled chart intentionally does not model clustering |
Loki (deploy/observability/applications.yaml) |
SingleBinary, replication_factor: 1, filesystem storage |
Move to a scalable/distributed Loki mode with object storage and replication; forward the tamper-evident audit receipts onward to a SIEM for durable long-term hold |
These stores are the "external HA stores" operator-owned item tracked in Scope and non-goals; the bundled charts are working references and are never removed, so rolling back to the reference footprint for a demo is a one-line values change.
Workflow upgrades and recovery¶
The 0.5.1 gateway uses Redis, not PostgreSQL, for run indexes, step timelines and
team/run budgets. Temporal owns workflow history in temporal and SQL visibility in
temporal_visibility, both on the existing PostgreSQL service. Back up all three stores
together with the complete receipt export. No gateway PostgreSQL database or new service
is introduced by this milestone.
Gateway startup applies numbered Lua migrations from app/migrations atomically with
the schema marker in the existing budget Redis. Migration 001 adopts the unversioned
0.2.0 keys without rewriting counters, run IDs, receipts or expiry times. Repeated and
concurrent startups are safe. Unknown/newer schema versions prevent startup; restore a
pre-upgrade backup when rolling back across an incompatible future migration.
Temporal uses its upstream versioned SQL migrations for both databases. Compose pins
server/auto-setup to 1.29.1; the existing Helm dependency uses server/admin-tools 1.32.0,
now explicitly pinned together. These are separate upgrade tracks: do not move a Compose
database directly across those server minors. Follow the
Temporal upgrade sequence,
apply schema updates before rolling servers, and never change numHistoryShards on an
existing database. The chart explicitly declares the visibility store and schema management.
For the first Helm install, temporal.schema.useHelmHooks=false lets the bundled Postgres
start before the schema Job. For upgrades on an existing release, run:
helm dependency build deploy/charts/workflows
helm upgrade workflows deploy/charts/workflows -n workflows -f your-workflow-values.yaml \
--set temporal.schema.useHelmHooks=true --wait --wait-for-jobs --timeout 10m
Keep the same release, namespace, database names, secrets and PVCs. For GitOps, order the schema Job before the server rollout using the controller's sync phases. Keep old worker workflow code compatible with retained histories; this migration does not rewrite histories.
Run the isolated fixture drills (Docker, Git and Python 3.12+; no provider credentials):
make workflow-upgrade-test
make workflow-restore-drill
make workflow-helm-upgrade-test # also requires kind, Helm and kubectl
On Windows, the first two entry points also run directly in PowerShell as
python scripts/workflow-recovery.py upgrade and python scripts/workflow-recovery.py drill.
The Helm test uses native Helm/kubectl and the gateway's Python environment (PyYAML),
with kind and Docker in WSL Ubuntu.
The upgrade drill builds the 0.2.0 release commit f17850aa91091c306921db577c18c15df7da3ef0,
starts a run and waits for approval, then replaces images while retaining volumes. It checks
all retained gateway values, both SQL schema versions, run listing, receipt IDs and budgets,
then approves the original execution. The restore drill dumps both Postgres databases,
snapshots Redis, destroys only its own disposable volumes, restores and resumes that run.
Both verify the receipt chain and terminate a worker during an active model call, checking
that the call completes once. Reports and backups go under .out/aw-hardening-*; all drill
containers and volumes are removed even on failure. Builds are sequential, workers have
two activity slots, and Windows invokes Docker through wsl.exe -d Ubuntu -e docker.
The Helm test creates a separate kind cluster, caps its node at two CPUs and 4 GiB, installs
the previous charts/images, then upgrades the same releases to two gateway and worker
replicas. It checks retained PVC identities, run state, budgets and receipts, resumes the
waiting run and deletes the cluster. It uses an isolated kubeconfig, local fixtures and
temporary loopback ports. This single-node drill tests upgrades, not multi-node failover.
Seeding the old Helm release supplies the required connectProtocol: tcp operator override
for both SQL stores; current chart defaults include it. The database contents and previous
application images remain the 0.2.0 baseline.
For your Compose trial, take a quiesced backup and restore into a new project:
python scripts/workflow-recovery.py backup --project agentworkflows \
--directory .out/workflow-backup --receipts /path/to/complete-retained-receipts.jsonl
# Stop the source before reusing its published ports, or select unused Compose ports.
python scripts/workflow-recovery.py restore --project agentworkflows-restored \
--directory .out/workflow-backup
Omit --receipts only for a first container lifetime with no earlier receipts. Backup stops
workers, drains the gateway, stops Temporal, then takes pg_dump -Fc of each database and
a Redis RDB snapshot; it restarts the source afterward. Restore refuses existing Temporal
tables or nonempty Redis. It restores the archived receipts alongside the backup and
verifies them with audit-verify --anchor --strict-continuity, including new restart links.
Checksum or chain errors fail before database writes. Copy the backup and its manifest/head
anchors to separately controlled storage; checksums alone do not authenticate a rewritten
backup. Retain policy files, credentials and exact deployment values through your existing
secret/configuration backup process. A completed backup defines the recovery point; writes
after it are outside that backup. Measure recovery time on your own data volume.
For Helm/managed stores, use the same maintenance order and database-native pg_dump -Fc
and pg_restore --create --exit-on-error for both Temporal databases, plus a persistence
snapshot of every gateway Redis database and the SIEM receipt export. Restore to isolated
databases first, verify schema versions and receipts with the included verifier, point a
test gateway/worker at those stores, and resume a waiting run before switching traffic.
The Compose drill is a reproducible application recovery test, not a managed-database
failover or Kubernetes storage certification.
Multiple replicas and shutdown¶
Merge the chart's values-ha.yaml after your reviewed deployment values:
helm upgrade inference-gateway deploy/charts/inference-gateway -n inference \
-f your-gateway-values.yaml -f deploy/charts/inference-gateway/values-ha.yaml
helm upgrade workflows deploy/charts/workflows -n workflows \
-f your-workflow-values.yaml -f deploy/charts/workflows/values-ha.yaml \
--set temporal.schema.useHelmHooks=true --wait --wait-for-jobs --timeout 10m
The gateway and its bundled console run in two replicas; the console has no separate server or in-memory session to synchronize. The worker has two replicas on the same team task queue. Each deployment keeps at least one pod available during voluntary disruptions and rolls with zero unavailable pods. Topology spread distributes pods when nodes permit. KEDA's minimum is two in the gateway HA values. Temporal's example uses three replicas per server role with PDBs keeping two. Use enough nodes to realize those guarantees. The bundled Postgres and Redis remain single-node references; application replication does not make storage highly available. Point the existing store settings at your managed HA stores and configure persistence, backups and networking there.
/healthz checks the running event loop and intentionally stays live during dependency
outages. Gateway /readyz checks required Redis stores, local runtimes, cloud credential configuration and
Temporal when workflows.temporalAddress/TEMPORAL_ADDRESS is set. Worker port 8081
/readyz requires an active worker, Temporal and gateway readiness; /healthz checks
its event loop. Dependency outages withdraw readiness without causing restart storms.
SIGTERM stops worker polling and lets active activities finish for 180 seconds before Temporal's shutdown cancellation. Pods and Compose allow 210 seconds; the gateway drains requests for 180 seconds, with a five-second Kubernetes pre-stop delay for endpoint removal. Set the pod grace above your longest custom activity and SDK shutdown timeout when extending the supplied three-minute activity limit. Forced kills or exhausted grace may retry an activity: Temporal retains the step, but arbitrary external side effects cannot be promised exactly once. Tools must honor their idempotency key; container steps deliberately do not automatically retry. Retain every replica's audit stream and external head anchors; the shared head key alone cannot prove completeness of concurrent replica lifetimes.
Live providers and concurrent runs¶
make test excludes the paid acceptance directory. Even explicit acceptance runs skip
unless AGENTWORKFLOWS_LIVE_PROVIDERS=1, and always skip when CI is set. Each provider
also requires its own real environment credential; absent credentials skip that provider.
The suite covers chat, streamed content and usage, forced tool calls, a deterministic
primary outage followed by a real-provider fallback, and receipt/provider cost accounting.
Do not place credentials in tracked files or command arguments.
| Provider | Credential variable | Configuration prefix |
|---|---|---|
| OpenAI | OPENAI_API_KEY |
LIVE_OPENAI |
| Anthropic | ANTHROPIC_API_KEY |
LIVE_ANTHROPIC |
| Azure OpenAI | AZURE_OPENAI_API_KEY |
LIVE_AZURE_OPENAI |
| Bedrock bearer-token API | AWS_BEARER_TOKEN_BEDROCK |
LIVE_BEDROCK |
| Vertex OpenAI-compatible endpoint | VERTEX_ACCESS_TOKEN |
LIVE_VERTEX |
For each enabled provider set <PREFIX>_MODEL, <PREFIX>_INPUT_USD_PER_1K, and
<PREFIX>_OUTPUT_USD_PER_1K to the approved model and current prices. Set
<PREFIX>_BASE_URL for Azure, Bedrock and Vertex; OpenAI and Anthropic use their standard
API URLs. Choose models supporting tools and reported streaming usage. Cost assertions
use configured prices, not invoices. With your environment already provisioned:
No real-provider calls were made for this hardening change. The default test gate checks the same adapters with local protocol fixtures.
With the fake Compose trial running, make workflow-loadtest submits 20 workflow runs at
concurrency two, checks their results, receipt IDs and per-run budgets, and asserts the
aggregate team spend is exactly $0.220 in synthetic configured prices. It writes throughput,
median/p95 latency and failures to results/loadtest/workflow-runs.json. This small test
checks concurrent accounting and completion; it does not establish provider latency or
production capacity.
Local result on 2026-10-07 (Windows client, WSL Docker, fake Compose providers, one worker with two activity slots): 20/20 completed, 0 failures, 3.476 seconds total, 5.754 runs/s, 0.337 s median and 0.353 s p95 end-to-end latency. All 20 receipt IDs and run budgets verified; shared spend increased by $0.220. This was a short functional load check on a shared machine, with warm containers and synthetic provider responses.
Promotion Review¶
Use the matrix above as the source of truth, then review this shorter sequence before a customer handoff or production-style demo:
- Run
make validate-full,make api-contract,make config-contract, andmake release-gate-strictagainst current evidence. - Confirm auth, prompt redaction, model allowlists, sandbox budgets, quota labels, and NetworkPolicies match the target environment.
- Review tenant onboarding output before applying it; regulated/offline tenants must keep external CIDR egress disabled.
- Verify RAG knowledge, vector-store dimensions, GPU resource names, runtime replicas, HPA/PDB settings, and topology spread against customer capacity.
- Run smoke, RAG, agent, eval, load, restore, and chaos evidence paths that are relevant to the handoff.
- Confirm image scan, SBOM, checksum, signature, and repo security evidence exists for the images being promoted.
- Generate an Evidence pack and attach the Markdown report to the handoff notes; retain JSON evidence with the release or drill record.