Skip to main content
The unified Runlayer operator reconciles two custom resources in one controller: Use this path for one or many isolated tenants on a Kubernetes cluster (EKS or GKE).
Do not install the unified operator on a cluster that already runs runlayer-deploy-operator. Both reconcile MCPServer CRs.

End-to-end flow

  1. Provision cluster features and external per-tenant dependencies (Postgres, Redis, buckets, IAM / Workload Identity, streams).
  2. Install the runlayer-operator Helm chart once per cluster.
  3. Create (or let the operator create) the platform namespace and seed required Secrets.
  4. Apply a RunlayerInstance (raw YAML, optional runlayer-instance Helm chart, or Terraform on AWS).
  5. Wait for status.phase: Ready, point DNS at the load balancer, then create MCPServer CRs if using Runlayer Deploy.
Supporting guides (read before install):

1. External dependencies

The operator does not provision cloud infrastructure. Each tenant needs the following out-of-band (Terraform, console, or your platform).

Required for every tenant

Network: cluster nodes (or pods via CNI) must reach Postgres and Redis on private connectivity (VPC, Private Service Connect, peering). Tighten egress with spec.security.networkPolicy.additionalEgressRules once CIDRs are known.

Optional (feature-dependent)

Sample Terraform for Postgres, Redis, buckets, streams, and identity is embedded in External dependencies and Workload identity (AWS and GCP side by side). Adapt the samples to your IaC — Runlayer does not require a specific Terraform module.

Shared GCP project

When any component uses streamBackend: pubsub, set one project for the instance:
Do not put projectId on per-component pubsub blocks (subscription / DLQ only).

Identity model (IRSA / Workload Identity)

The operator creates dedicated ServiceAccounts per workload and never embeds long-lived cloud keys. Annotate via spec.serviceAccount: AWS (IRSA):
GCP (Workload Identity):
Prefer WI/IRSA over ECS-style assume-role env vars (AUDIT_LOG_SIEM_EXPORT_ASSUME_ROLE_ARN is not the K8s path).

Saved Bedrock providers for monitor consumers

The EKS application policy already grants scoped saved-Bedrock role assumption for separately provisioned monitor consumers. Additional ServiceAccounts need explicit application_irsa_additional_service_accounts trust and matching role annotations. The saved provider role must trust that AWS identity with its configured external ID and grant model access; GKE consumers also need an AWS credential source because GCP Workload Identity does not authenticate to AWS.

2. Cluster requirements

Kubernetes version

Shared features (EKS and GKE)

EKS-specific

See Kubernetes prerequisites for the full cluster contract.

GKE-specific

Namespace layout

Override with spec.deploy.platformNamespace / spec.deploy.mcpNamespace.

3. Install and upgrade the operator

Install (once per cluster)

Copy and customize values:
OCI (Customer Distribution — production): versions declared production-ready via the runlayer-operator-promote-customer-distribution workflow.
OCI (staging / pre-promote): continuous push on main in account 412677576004. Override image.repository to the staging operator image when using those charts.
Use the Customer Distribution OCI chart for production installs. Runlayer can share a chart tarball if OCI pull is blocked in your environment. CRDs ship in the chart crds/ directory and are applied on first helm install only:
  • platform.runlayer.com/v1alpha1 — RunlayerInstance
  • deploy.runlayer.com/v1alpha1 — MCPServer

Common operator values

Full values reference ships with the Helm chart you install from Customer Distribution ECR (values.example.yaml in the chart package).

Upgrading the operator (manual CRD apply)

Helm does not upgrade CRDs on subsequent releases. When a release notes CRD schema changes, apply CRDs before upgrading the chart:
If CRDs were never Helm-managed, plain kubectl apply -f ... is enough.
Always apply CRDs from the same chart version you are upgrading to. Skipping CRD apply can leave the controller unable to reconcile new fields.

4. Prerequisites: namespace and Secrets

Install order for Secrets

The operator never creates, reads, or syncs Secret data (no secrets write RBAC). You must create Secrets in the platform namespace. Two workable sequences: Many teams use sequence A with their own Terraform syncing Secrets Manager / Secret Manager into the platform namespace.

Required Secrets (platform namespace)

Example (sequence A):

MCP Catalog (required)

Onboarding calls the MCP Catalog API. Without a key, GET /api/v1/catalog/ returns 503 and setup shows “Couldn’t load setup status”. Contact Runlayer for the key — see Runlayer-provided inputs.

AI Watch binary distribution (when using AI Watch)

Installer resolution needs a Runlayer download token, a customer-owned object-storage cache, and workload permission to that cache. Without them the UI reports “No AI Watch installer is currently resolved”. After setting these, restart backend/worker and run release discovery (Check now or the scheduled poll). Details: External dependencies and Workload identity.

Optional component Secrets

App secret

Platform Secret named runlayer-app by convention (spec.appSecretRef.name). Mounted with envFrom on backend and worker (and related backend-image jobs). Every key in the Secret becomes an environment variable.

Required keys

Common optional keys

AUTH_API_KEY may also appear here; it is also injected from spec.auth.apiKeyFrom (duplicate is harmless).

Common non-secret configuration via the same Secret

Because envFrom loads all keys, teams often put non-sensitive settings in runlayer-app as well (or use a separate ConfigMap pattern outside the operator). Typical production keys: Runlayer Assistant inference. BEDROCK_ENABLED=true (typical AWS): new/unbound assistants use platform-billed Bedrock (RUNLAYER_ASSISTANT_BEDROCK_MODEL; IRSA/WI or BEDROCK_ACCESS_KEY_ID / BEDROCK_SECRET_ACCESS_KEY). BEDROCK_ENABLED=false (typical GCP/GKE): the assistant binds an org AI provider instead — recommended model if one is set, else the custom-agent default, else the oldest provider’s first model. Until an admin adds a provider in Settings → AI Providers, the assistant is still provisioned but chat is blocked. An admin can change the assistant’s model/provider in the agent panel (MANAGE_ORG_SETTINGS); a bound org provider is sticky and is not reset when Bedrock is later enabled. Prompt, name, connectors, and skills stay managed. Any backend Settings field (see application config) can be supplied this way. Prefer CR fields for connectivity that the operator already knows how to wire (database, redis, auth, component toggles). On AWS, when this deployment invokes Anthropic models on Bedrock (BEDROCK_ENABLED=true, or Bedrock-backed custom agent models), submit the one-time form in the account those calls come from:
Anthropic model access. Anthropic requires a use case details form before its models can be invoked in an AWS account, plus an accepted model agreement. You need this when the deployment invokes Anthropic models on Bedrock — typically BEDROCK_ENABLED=true (Runlayer Assistant platform inference and optional Bedrock agent models). With BEDROCK_ENABLED=false, Runlayer Assistant binds an org AI provider instead; skip this step unless you also enable Bedrock inference elsewhere (for example BEDROCK_ANTHROPIC_MODELS_ENABLED for custom agents).The backend does both for you at startup and on login, in the account the Bedrock calls come from, using the model-access IAM permissions your deployment grants. Four settings describe the deployment on the form — BEDROCK_MODEL_ACCESS_COMPANY_NAME (defaults to Runlayer, the operator of the software making the calls), BEDROCK_MODEL_ACCESS_COMPANY_WEBSITE, BEDROCK_MODEL_ACCESS_INDUSTRY and BEDROCK_MODEL_ACCESS_INTENDED_USERS. Override them if you run this install yourself. They are never derived from your data: the form is submit-once, so a guessed answer would be permanent.Set BEDROCK_AUTO_MODEL_ACCESS=false to turn this off and do it by hand instead: Bedrock console → Model catalog → any Anthropic model → Submit use case details. Do the same if your deployment role predates the model-access permissions — the backend logs exactly which actions it is missing. AWS accepts the form once per account, or once at your AWS Organization’s management account, which member accounts inherit.Deploying outside the US? Also set RUNLAYER_ASSISTANT_BEDROCK_MODEL to your geography’s inference profile, for example au.anthropic.claude-opus-4-6-v1 in ap-southeast-2. The risk-categorization explainer follows BEDROCK_REGION on its own and uses Claude Haiku 4.5 outside the US, so the Anthropic form above covers it; only set RISK_CATEGORIZATION_MODEL if you want a different model, and pick one served in your Region. If Bedrock rejects the primary model it falls back to Amazon Nova Lite, which needs no access request.Access takes up to 15 minutes to propagate. Until it does — or if this is skipped entirely — Bedrock-backed agents fail with Model use case details have not been submitted for this account; see troubleshooting for the CLI checks and Region caveats.

5. Deploy a RunlayerInstance

Required spec fields

Optional highlights: database.port (5432), database.name (runlayer), redis.port / tls / credentials, components, deploy, serviceAccount, workloadScheduling, security, trustedCa, autoscaling, podDisruptionBudget, ingress, otelCollector, auditSpool, deploymentTelemetry, extraEnv (extra environment variables for every backend-image workload; unlike the ECS module’s additional_backend_env_vars, a name the operator sets itself is rejected rather than overridden). Field-level detail is covered in the sections below and in the supporting guides linked at the top of this page.

Minimal example (EKS)

Minimal example (GKE)

Optional components (snippet)

When enabling an optional component, set its images.* field (and any required Secret refs). The operator rejects the CR if those are missing. status.phase: Ready / Available wait for every enabled optional Deployment (not only backend/frontend/worker). Pull credentials for platform components: the operator sets no imagePullSecrets on platform component pods (spec.deploy.defaultImagePullSecrets applies to MCPServer pods only — the Helm chart’s toolguard.image.pullSecrets knob has no operator equivalent). Customer Distribution ECR (088332244652) supports cross-account pulls when Runlayer allowlists your AWS account or IAM pull principal during onboarding — use those image URIs directly on EKS once allowlisted. Prefer a GAR / private mirror for GKE, air-gapped networks, or orgs that require an internal registry copy.
For streamBackend: kinesis on auditConsumer / siemExport / sessionMaterializer, replicas must be ≤ 1 (CEL). Session materializer is always singleton (CEL) for both backends — compaction safety. It also injects SESSION_* / HOOK_EVENTS_* into backend/worker (payload read + hook publish). For Pub/Sub, provision the hook-events subscription with message ordering enabled; publishers use ordering_key=session_id. Publish RPC timeout defaults to 5s (HOOK_EVENTS_GCP_PUBSUB_PUBLISH_TIMEOUT_SECONDS); override via that env or hookEvents.publishTimeoutSeconds on the CR. Infra (streams, checkpoints, buckets, ordered subscription) is out-of-band; see External dependencies. spec.agents env (RUNLAYER_AGENT_SANDBOX*) is injected onto backend and worker (and onto agentEventTrigger when that component is enabled) — see External dependencies. Procedures validation runs in the same sandbox runtime, so Procedures need spec.agents configured too; without it, creating or validating a Procedure returns 503. agents.filesBucket is a bucket you own, so two things on it are yours to set:
  • The backend role needs s3:PutObjectTagging alongside s3:PutObject. Files the backend hands a run (Slack attachments) are uploaded with a runlayer-kind=run-input tag, and a tagged PutObject is denied without that action — the upload fails and the attachment is silently dropped. Grant s3:PutObjectTagging on the agent-files bucket in the backend role (see Workload identity).
  • Those run inputs live under sessions/{agent_id}/workspace/.runlayer-inputs/{run_scope}/ inside the durable per-agent prefix, and nothing deletes them. They are only ever readable by the run they were delivered to, but to stop them accumulating add a lifecycle rule expiring objects tagged runlayer-kind=run-input after ~7 days (plus noncurrent_version_expiration if the bucket is versioned). Scope the rule to the sessions/ prefix as well as the tag: published artifacts are copies that live under artifacts/, and a tag-only rule would reach them too. The ECS module ships this rule as expire-run-inputs.

K8s Agent Sandbox (sandboxMode: k8s)

In-cluster runtime using kubernetes-sigs/agent-sandbox (SandboxClaim + SandboxWarmPool). The operator only wires env; it does not install the controller, Template, WarmPool, or isolation runtime. 1. Install the controller 2. Isolation on the nodes (gVisor or Kata — pick one) Sandbox pods must not use the host runc kernel. Nodes that schedule them need a RuntimeClass (kubectl get runtimeclass): GKE add-on admission also requires automountServiceAccountToken: false, runAsNonRoot, drop ALL capabilities, CPU+memory limits, and nodeSelector/toleration sandbox.gke.io/runtime=gvisor. 3. Template + WarmPool + RBAC, then set spec.agents.sandboxMode: k8s + spec.agents.k8s.warmPool. Ask Runlayer for Agent Sandbox Template / WarmPool examples if you choose sandboxMode: k8s.

Deploy (spec.deploy)

When enabled, backend/worker get RUNLAYER_DEPLOY=K8S and related env; frontend gets PUBLIC_RUNLAYER_DEPLOY=K8S.

Wait for Ready

Phases: Pending → Installing → Migrating / Upgrading → Ready (or Degraded / Failed). Available=True only when every enabled core and optional component Deployment is Ready (for example ToolguardReady, TopicCpuReady, TopicCpuFastReady, LLMGatewayReady). Disabled optionals are ignored. Ingress is separate: when Ingress is enabled, also watch IngressReady for traffic (on GCE, that includes the llm-gateway GLB backend when the Ingress routes to it). Then create MCP env Secrets in the MCP namespace and apply MCPServer CRs (create MCPServer CRs in the MCP namespace after Ready).

6. Optional: runlayer-instance Helm chart

The optional runlayer-instance Helm chart (same Customer Distribution OCI registry as the operator) renders one RunlayerInstance CR per Helm release. Prefer it for GitOps when you do not apply the CR with kubectl/Terraform. v1 owns only the CR — Secrets and (when not using spec.ingress) Ingress stay outside Helm.
Values: name, namespace, labels, annotations, and spec (full RunlayerInstanceSpec passthrough). Do not put tenant CRs in the operator chart — operator upgrades must not re-render tenants.

7. Ingress, NetworkPolicy, scheduling, resources

Ingress

Opt-in (spec.ingress.enabled, default false). When enabled, the operator reconciles path routing: On gce / gce-internal, the operator creates BackendConfigs (backend, frontend, and llm-gateway when enabled) before annotating Services / applying Ingress so GKE GLBC does not stick on missing BackendConfig. Set controllerNamespace so NetworkPolicies allow the ingress controller → backend (AWS LBC, nginx). GKE also uses GCP health-check CIDR defaults. Do not enable certManager and gcp.managedCertificate on the same instance.

NetworkPolicies

Defaults isolate tenants and allow required DNS / platform / MCP paths. Extensions are append-only:
When deploy.enabled: true, platform policies come from RunlayerInstance; MCP policies come from each MCPServer.spec.security.networkPolicy.

Workload scheduling

  • spec.workloadScheduling — default nodeSelector / affinity / tolerations for platform pods + migration Job
  • spec.components.<name>.scheduling — per-component override (e.g. ToolGuard → GPU pool)

Default container resources

When components.<name>.resources is omitted, the operator applies these defaults: ¹ backend also carries a 2Gi ephemeral-storage request (no limit) for the audit-log spool; it is kept even when components.backend.resources is overridden, unless the override sets an explicit ephemeral-storage request. Override per component for smaller clusters. Topic CPU also enforces ECS-parity memory floors vs gunicornWorkers / fast.gunicornWorkers. If you set deploy.resourceQuota.hard, ensure requests fit the quota.

Autoscaling and PDBs

  • spec.autoscaling.<component> — HPA when enabled: true and a CPU or memory target is set (backend, frontend, worker, sentryRelay, llmGateway)
  • llmGateway gets a default HPA (70% CPU, min 1, max 3) when enabled and autoscaling.llmGateway is unset — set autoscaling.llmGateway.enabled: false to disable
  • Default HPA behavior (when unset): scale-down waits for a quiet fleet — 1800s for backend (keeps replicas added by a spike until it is over), 300s for other components — and removes at most 10% per minute; scale-up waits 60s and adds up to 50% or 2 pods per minute. Set behavior to replace the whole block.
  • spec.podDisruptionBudget — optional PDBs for core Deployments

Trusted CA

spec.trustedCa mounts a customer CA ConfigMap and sets SSL_CERT_FILE / REQUESTS_CA_BUNDLE / NODE_EXTRA_CA_CERTS for enterprise TLS inspection.

OTEL collector sidecar

spec.otelCollector.enabled injects a native sidecar on the backend Deployment (Kubernetes ≥ 1.29). Requires reachable OTLP endpoints; hostname endpoints may need extra NetworkPolicy CIDRs. Configure reachable OTLP endpoints; see observability notes below. Name the Runlayer-operated endpoint runlayer-… and give it protocol: http: the sidecar exports its own health metrics (export queue depth, send failures, receiver refusals — no application data) to endpoints with that prefix and only those, which is how Runlayer monitors your deployment’s telemetry export.

Audit log spool

The backend’s local audit-log spool is enabled by default: audit records are appended to a spool on the pod’s local disk (a dedicated emptyDir at /var/audit-spool) and delivered asynchronously by an in-process tailer, so MCP proxy requests do not wait on audit delivery. Records are typically delivered within seconds, and delivery is at-least-once with deduplication. If the spool cannot keep up or its disk budget fills, individual records fall back to inline delivery on the request path, so nothing is lost to backpressure.
  • Node-disk encryption is required. The spool briefly holds pre-redaction audit payloads at rest on the node’s disk. On GKE, node disks are encrypted by default with Google-managed keys (CMEK optional). On EKS, encrypted EBS node volumes must be explicitly enabled. The operator cannot verify either.
  • The backend pod’s shutdown drain window for the spool is 110 seconds (grace period 140s minus 30s headroom); records not delivered before the window expires are lost with the pod.
  • After an outage of a delivery target (database or Kinesis), expect elevated write load while each backend pod drains its accumulated backlog. On small database instances, AUDIT_LOG_SPOOL_DELIVERY_CONCURRENCY (default 8, per pod) can be lowered via the app secret to throttle the catch-up.
  • Each backend pod can hold up to ~512MB of spool data on node disk by default (AUDIT_LOG_SPOOL_MAX_BYTES, settable via spec.auditSpool.maxBytes; dead-letter bytes count against the same budget — a soft cap that worker processes re-check every 5 seconds, so bursts can briefly overshoot it). Budget node ephemeral storage accordingly when running many instances per cluster. The spool volume is bounded by an emptyDir sizeLimit derived as max(2 × maxBytes, 1Gi) — raising maxBytes automatically keeps the backstop above the budget — and backend pods carry an ephemeral-storage request (2Gi, no limit) so the scheduler accounts for the disk footprint. If spool usage ever exceeds the sizeLimit, the kubelet evicts the backend pod and its undelivered spool is discarded — spec.auditSpool.volumeSizeLimit overrides the derivation; if you set it, keep it comfortably above the maxBytes budget.
  • Alert on spool health. A spool that persistently rejects appends degrades silently to inline delivery (a latency regression, not data loss), and nothing alerts on it out of the box. See Monitoring the audit log spool below.
To revert to inline audit delivery, set spec.auditSpool.enabled: false. New pods write audit logs inline again, and each old pod drains its residual spool within the shutdown window as it stops. No data migration is needed in either direction. Disabling requires the operator 0.2.22 CRD: helm upgrade does not apply crds/, and against an older CRD the API server silently prunes spec.auditSpool, leaving the spool on. Apply the updated CRD first — see Upgrading the operator.

Monitoring the audit log spool

The spool degrades gracefully — records that it rejects are delivered inline instead of lost — so a persistently unhealthy spool is invisible unless you alert on it. The operator does not ship a spool alert; use the metrics and alert rules below to build one.

Prerequisite: metrics export

The backend exports spool metrics over OTLP. spec.otelCollector.enabled (above) injects a collector sidecar and wires the backend’s OTLP/gRPC endpoint for you — but at least one configured endpoint must carry metrics: an endpoint restricted to signals: [traces] (or a configMapRef collector config without a metrics exporter) leaves the sidecar’s metrics pipeline on the nop exporter, and the spool metrics go nowhere. Each backend worker process exports its own cumulative counters, tagged with a unique service.instance.id resource attribute. The operator’s generated sidecar config promotes that identity onto every datapoint, so the per-process series survive into your storage. If you replace the config via configMapRef, replicate that promotion (or enable resource_to_telemetry_conversion on a Prometheus/remote-write exporter): if the workers collapse into one series, the merged counters zigzag, every dip reads as a counter reset, and rate() / increase() overcount by orders of magnitude — the alerts below will page on fabricated spikes. Metric names below are the Prometheus-translated forms (OTLP ingestion appends the _total / _seconds / _milliseconds / _bytes suffixes), as stored by Prometheus-compatible backends such as Mimir.

Fallback reasons

Every rejected append increments runlayer_audit_spool_fallback_count_total with a reason label. Each reason has a different operational meaning:

Prometheus alert rules

A ready-to-adapt PrometheusRule (kube-prometheus-stack shape — adapt the labels/selectors to your stack, or lift the PromQL into whatever rule format your Prometheus uses). The operator does not ship these as manifests:
One residual false-fire mode: after a metrics-ingestion gap of 15 minutes or more (collector or scrape-pipeline outage), the dead-letters rule’s birth branches can re-fire on pre-existing series when ingestion resumes — silence the alert on pipeline recovery rather than treating it as new loss. One loss path is not covered by these rules: when the tailer fails to write a poison record into dead/ in the first place, the record is dropped outright. The matching counter (runlayer_audit_spool_dead_letter_write_failed_count_total) is disabled by default on self-hosted deployments, so a metric alert on it would silently never fire. If your stack supports log-based alerting, alert on the audit_spool_dead_letter_write_failed backend log line instead — it fires unconditionally.

Kubernetes prerequisites

Cluster requirements for EKS and GKE

External dependencies

Postgres, Redis, streams, buckets + Terraform samples

Runlayer-provided inputs

Catalog key, WorkOS, download token, images

Workload identity

IRSA / GKE Workload Identity matrix

Deployment overview

Account strategy and isolation models

Network firewall

Client egress allowlists for the tenant hostname