Skip to main content
Select required features before the first apply so their credentials, permissions, quotas, and network access are ready. Each HCL snippet shows module inputs; merge it into module "runlayer_infrastructure". Examples enable features and do not represent the defaults.

Bedrock and Anthropic

On this ECS deployment path, Runlayer Assistant uses Bedrock in the deployment account (default model global.anthropic.claude-opus-4-8). It ignores Settings → AI Providers. Having “some Anthropic model” enabled in the console is not enough — that specific model (or its base) must be available. Module v28.1.x+ attaches the model-access IAM actions to the ECS task role (GetFoundationModelAvailability, PutUseCaseForModelAccess, Marketplace subscribe, etc.). You do not add those by hand unless an SCP / boundary strips them.
If Assistant chat returns does not have access to the Anthropic model yet, see troubleshooting.

Distribution configuration

The deployment guide configures Distribution with all three inputs below. Obtain a registered reader key from Runlayer before the first apply:
The module stores the key in Secrets Manager and injects the URL as non-secret runtime configuration. It rejects flagd without both reader inputs. The application default remains WorkOS when the provider input is omitted; setting the key and URL alone does not switch providers. Existing WorkOS deployments can stage the reader credentials before explicitly switching to flagd. Hook-body compression (gzip-hooks, zstd-hooks) is on by default for Distribution-backed deployments. WorkOS-provider deployments opt in until they move to Distribution; FF_DISABLE wins if both name a flag:
If Runlayer has registered your account for self-managed Distribution registration, also pass the IDs Runlayer support provides and the module manages your registry manifest as a Terraform resource:
Rotate keys with zero gap by setting distribution_api_key_previous to the outgoing key, applying, waiting for ECS services to reach steady state, then clearing it and applying again. Set distribution_status = "disabled" to revoke access without destroying the stack; move openfeature_provider off flagd first.

Optional feature flags

Enable only the features you intend to operate. ToolGuard needs GPU quota and bootstrap egress. With WAF IP allowlisting, configure AgentCore VPC mode and PrivateLink before enabling agents. Runlayer Deploy is always enabled by the ECS module.
Saved Bedrock providers for Slack monitors. The agent-event-trigger consumer uses its dedicated ECS task role to assume the saved provider role. The module grants sts:AssumeRole only on arn:aws:iam::*:role/runlayer-bedrock-* and arn:aws:iam::*:role/RunlayerBedrock*. The saved role must trust the caller’s AWS account/identity with the saved external ID and grant the required model access. When both BEDROCK_ENABLED and BEDROCK_ANTHROPIC_MODELS_ENABLED are enabled, the consumer also receives nonstreaming bedrock:InvokeModel access to Claude foundation models and the deployment account’s Claude inference profiles for the legacy platform fallback. Upgrade and apply the Terraform module alongside the backend/consumer image; updating the image alone does not install this permission. After IAM propagation, verify an authorized monitor canary and its dead-letter alarm.

ToolGuard capacity and scaling

ToolGuard scale-out speed. A new ToolGuard GPU host installs NVIDIA GRID drivers at boot and takes 10–15 minutes before it serves traffic, so a scale-out decision is only as fast as the spare hosts already running. That boot path also needs outbound HTTPS to Ubuntu, NVIDIA GRID S3, Docker, and related hosts — see Self-Hosted Egress Requirements → ToolGuard. runlayer_tool_guard_scaling_profile trades idle GPU cost for time-to-capacity: balanced and aggressive keep one to two extra g6f.large instances running at all times (billed even when idle) so the next task starts immediately. Scale-in waits 15 minutes on every profile. The dashboard’s Tasks: Desired vs Running panel shows how long scale-outs wait on new hosts. The runlayer-tool-guard-scale-out-lag alarm fires only when desired tasks exceed running tasks for 20 consecutive minutes — longer than a cold GPU host boot — so a routine scale-out does not trip it; when it does fire, new hosts are not arriving (check the Auto Scaling group and capacity provider). GPU signal on aggressive. CPU is a weak load signal on a GPU inference host, so aggressive adds a second target-tracking policy on fleet-average GPU utilization (target 60%). The policy tracks the existing RunlayerToolGuard/GPUUtilization metric (Environment dimension, Average statistic): every host whose ToolGuard answers /health 200 publishes one sample per minute, so the per-minute average is the mean across serving GPUs. Hosts that are not serving (capacity-provider spares, hosts still installing drivers or loading the model, hosts in the recycle drain) publish nothing and are left out, so the spare pool cannot hold the average under the target while serving GPUs are busy. No per-host series is added, so CloudWatch metric cardinality is unchanged. Scaling takes the higher of the CPU and GPU decisions, so the GPU policy never forces the fleet below what the CPU or predictive policies request. Hosts publish GPU metrics from their boot script every 60 seconds, so changes to the ToolGuard launch template trigger an ASG rolling instance refresh that replaces the hosts: tasks restart on new hosts as old hosts drain (each restart reloads the model, so expect a capacity dip per batch). Until the refresh replaces them, hosts from the previous launch template still publish every 5 minutes and so land in only one of every five 1-minute averages. Predictive scaling. Because a reactive scale-out cannot beat the 10–15 minute host boot, the module also creates an ECS predictive scaling policy on the ToolGuard service. It forecasts the next 48 hours of demand hourly from up to 14 days of ToolGuard predict request history and raises the task count runlayer_tool_guard_predictive_scheduling_buffer_seconds (default 1200) before each forecast hour, so GPU hosts are booting before the load arrives. conservative runs it in forecast_only mode (forecasts are recorded but never acted on); balanced and aggressive run forecast_and_scale. Override the profile default with runlayer_tool_guard_predictive_scaling_mode (disabled, forecast_only, forecast_and_scale). The first forecast appears after 24 hours of metrics; review it in the ECS console (service → Service auto scaling) or with aws application-autoscaling get-predictive-scaling-forecast, then flip forecast_only → forecast_and_scale once the curve matches your traffic. Predictive scaling only scales out and never raises the task count above runlayer_tool_guard_max_count; the CPU policy still handles scale-in and unforecast bursts. In regions where ECS predictive scaling is unavailable (for example il-central-1) the policy is skipped automatically; the runlayer_tool_guard_predictive_scaling_mode_effective output reports the mode actually applied. Why the load is floored. The forecast learns daily and weekly patterns from the load metric, so a few days with scanning switched off would teach it that those weekdays are idle and under-provision them for the next two weeks. Before forecasting, weekday hours (Mon–Fri UTC) below runlayer_tool_guard_predictive_load_floor_fraction (default 0.5) of the window’s mean hourly request count are raised to that floor, so off days look like quiet business days instead of zero; weekends forecast from actual traffic (runlayer_tool_guard_predictive_load_floor_weekdays_only = false floors every day, 0 disables the floor). Capacity is sized so each task receives runlayer_tool_guard_predictive_target_requests_per_task_per_minute (default 50 per minute, i.e. 3000 per hour). The default keeps per-task load in the range where ToolGuard predict p95 stays near its floor; latency rises steeply once a task serves more than roughly twice that rate. Raise the target to trade latency for fewer hosts, lower it for more headroom. Running tasks are read from Container Insights (RunningTaskCount) when it is enabled, otherwise from the GPU Auto Scaling group’s in-service instance count (one task per host). The Auto Scaling group count also includes the spare hosts balanced and aggressive keep warm (roughly 20% and 30% of the fleet), so without Container Insights requests per task read low by that share and the forecast under-provisions. Pair forecast_and_scale with container_insights = "enabled"; terraform plan warns when it is not.

Agent Sandbox (Optional)

Procedures also need the agent sandbox: their code is validated and test-run inside it. Without an agent runtime (enable_runlayer_agents with AgentCore), creating or validating a Procedure returns 503. Select a network configuration that allows AgentCore to reach the application before applying.
With the egress proxy enabled, every sandbox request to the public internet is authenticated per run and checked against the agent’s network access policy. Runlayer traffic stays on PrivateLink and is never proxied.
Network access policies are enforced only by this ECS / AgentCore sandbox runtime. The in-cluster Kubernetes agent sandbox (sandboxMode: k8s) does not enforce them yet: an agent’s policy is stored but has no effect there.
By default, Agent Sandbox routes LLM calls through the backend Anthropic proxy. To use OpenAI or an OpenAI-compatible gateway for custom agents, add a provider in Settings → AI Providers (no Terraform change needed). This Terraform path provisions the ECS / AgentCore sandbox. In-cluster kubernetes-sigs Agent Sandbox (sandboxMode: k8s, gVisor or Kata on the nodes) is operator/GKE — see Runlayer Operator — K8s Agent Sandbox. Runlayer Assistant always runs on Bedrock in this account — it ignores Settings → AI Providers by design, so it works without configuring a provider. Model access is handled as described above; this module exposes it as bedrock_auto_model_access, bedrock_model_access_company_name, bedrock_model_access_company_website, bedrock_model_access_industry and bedrock_model_access_intended_users. Backend egress for WorkOS AuthKit, Bedrock, and AgentCore is listed in Self-Hosted Egress Requirements.

LLM gateway

Microsoft Teams

Optional. Off unless all five values are set — the module variable is the client secret, the rest ride additional_backend_env_vars. Provisioning creates a per-agent Entra app registration and Azure Bot Service resource in your Entra tenant and Azure subscription, so a self-managed install needs its own multi-tenant app registration (Graph application permissions Application.ReadWrite.All, AppCatalog.ReadWrite.All, TeamsAppInstallation.ReadWriteForUser.All, admin-consented), an Azure resource group whose IAM grants that app Azure Bot Service Contributor, and the Microsoft.BotService provider registered on the subscription.
MS_TEAMS_APP_URL defaults to APP_URL (https://<domain_name>). Register https://<domain_name>/api/v1/integrations/ms-teams/admin-consent-callback as a Redirect URI (Authentication → Web) on that app registration, or admin consent fails with AADSTS500113. Each tenant that installs an agent’s bot must also enable Upload custom apps in its Teams Admin Center setup policy; without it, publishing to the tenant’s app catalog returns 403 no matter how many times it is retried. domain_name is effectively immutable once agents are provisioned: each agent’s Azure Bot resource stores the messaging endpoint captured at creation time and is never rewritten, so changing the domain orphans existing bots until each is deleted and recreated.

Audit SIEM, sessions, and hook events

Binary packages and downloads

Customer OpenTelemetry

Export metrics from the backend, and/or run an OTEL collector sidecar that ships to your collector:

Secrets, Sentry, and catalog

Deployment identity

For Sentry Relay credentials, follow Sentry Relay Credentials from WorkOS Vault. For software upgrades and rollback, follow Updates.