> ## Documentation Index
> Fetch the complete documentation index at: https://docs.runlayer.com/llms.txt
> Use this file to discover all available pages before exploring further.

# ECS Optional Features

> Prepare and configure agents, ToolGuard, Distribution, observability, and integrations on ECS

Select required features **before the first apply** so their credentials, permissions, quotas, and network access are ready. Each HCL snippet shows module inputs; merge it into `module "runlayer_infrastructure"`. Examples enable features and do not represent the defaults.

## Bedrock and Anthropic

On this ECS deployment path, Runlayer Assistant uses Bedrock in the deployment account (default model
`global.anthropic.claude-opus-4-8`). It ignores **Settings → AI Providers**.
Having “some Anthropic model” enabled in the console is not enough — that
specific model (or its base) must be available.

Module `v28.1.x+` attaches the model-access IAM actions to the ECS task role
(`GetFoundationModelAvailability`, `PutUseCaseForModelAccess`, Marketplace
subscribe, etc.). You do not add those by hand unless an SCP / boundary strips
them.

```hcl theme={null}
bedrock_auto_model_access = true
bedrock_region            = "" # empty → deployment region

# Self-managed installs: override before first successful auto-submit
# (form is submit-once and permanent). Defaults name "Runlayer".
bedrock_model_access_company_name    = "Example Corp"
bedrock_model_access_company_website = "https://example.com"
bedrock_model_access_industry        = "Technology"
bedrock_model_access_intended_users  = 100

# Custom agents only — not required for Runlayer Assistant
enable_bedrock_anthropic_models = false
# anthropic_api_key = var.anthropic_api_key  # via TF_VAR_*
# platform_llm_trial_default_budget_usd = "50"
```

If Assistant chat returns `does not have access to the Anthropic model yet`,
see [troubleshooting](/operations/troubleshooting#runlayer-assistant-anthropic-bedrock-model-access).

## Distribution configuration

The [deployment guide](/deployment/terraform) configures Distribution with all
three inputs below. Obtain a registered reader key from Runlayer before the
first apply:

```hcl theme={null}
distribution_api_key = var.distribution_api_key
distribution_api_url = "https://distribution.prod.runlayer.com"
openfeature_provider = "flagd"
```

The module stores the key in Secrets Manager and injects the URL as non-secret
runtime configuration. It rejects `flagd` without both reader inputs. The
application default remains WorkOS when the provider input is omitted; setting
the key and URL alone does not switch providers. Existing WorkOS deployments
can stage the reader credentials before explicitly switching to `flagd`.

Hook-body compression (`gzip-hooks`, `zstd-hooks`) is on by default for
Distribution-backed deployments. WorkOS-provider deployments opt in until they
move to Distribution; `FF_DISABLE` wins if both name a flag:

```hcl theme={null}
additional_backend_env_vars = {
  FF_ENABLE = "gzip-hooks,zstd-hooks"
}
```

If Runlayer has registered your account for self-managed Distribution
registration, also pass the IDs Runlayer support provides and the module manages
your registry manifest as a Terraform resource:

```hcl theme={null}
distribution_registry_bucket  = "mcp-catalog"
distribution_registry_source  = "runlayer-customer-deployments"
distribution_target_id        = "managed-aws:<your-account-id>"
distribution_configuration_id = "<uuid-from-runlayer>"
```

Rotate keys with zero gap by setting `distribution_api_key_previous` to the
outgoing key, applying, waiting for ECS services to reach steady state, then
clearing it and applying again. Set `distribution_status = "disabled"` to revoke
access without destroying the stack; move `openfeature_provider` off `flagd` first.

## Optional feature flags

Enable only the features you intend to operate. ToolGuard needs GPU quota and [bootstrap egress](/deployment/egress-requirements#toolguard-gpu-bootstrap). With WAF IP allowlisting, configure [AgentCore VPC mode and PrivateLink](/deployment/ecs-networking#security-configuration) before enabling agents.

Runlayer Deploy is always enabled by the ECS module.

```hcl theme={null}
enable_runlayer_agents             = true
enable_runlayer_agentcore_runtime  = true
enable_runlayer_tool_guard         = true
enable_runlayer_topic_cpu          = false
enable_runlayer_rev                = false # separate GPU decision-model service
enable_runlayer_cimd               = true
enable_runlayer_deploy_healthcheck = false
# runlayer_deploy_use_shared_sg = false
# runlayer_deploy_use_shared_task_role = false

# ToolGuard sizing (when enable_runlayer_tool_guard = true)
# runlayer_tool_guard_desired_count = 2
# runlayer_tool_guard_max_count = 4
# runlayer_tool_guard_timeout = 30   # deprecated for enforcement scans (use the workspace scan latency budget); still sets the tool-list/analyze/models timeout
# runlayer_tool_guard_gpu_resource_value = "ALL"  # required for g6f.* fractional GPUs
# runlayer_tool_guard_scaling_profile = "conservative" # balanced | aggressive
# runlayer_tool_guard_predictive_scaling_mode = null # disabled | forecast_only | forecast_and_scale (default follows profile)

enable_audit_log_kinesis    = true
audit_log_write_target      = "kinesis"
enable_audit_log_s3_archive = true
# S3 archive alone keeps Aurora hot-window size bounded (default
# audit_log_archive_hot_window_days = 90). Kinesis is only for SIEM export of
# full audit metadata — enable when you have a consumer ready.
audit_log_kinesis_consumer_config = {
  enabled       = true
  desired_count = 1
}

enable_audit_payload_offload = true
enable_audit_log_spool       = true
enable_toolguard_feedback    = true
enable_tool_output_offload   = false
```

**Saved Bedrock providers for Slack monitors.** The agent-event-trigger consumer
uses its dedicated ECS task role to assume the saved provider role. The module
grants `sts:AssumeRole` only on `arn:aws:iam::*:role/runlayer-bedrock-*` and
`arn:aws:iam::*:role/RunlayerBedrock*`. The saved role must trust the caller's AWS
account/identity with the saved external ID and grant the required model access.
When both `BEDROCK_ENABLED` and `BEDROCK_ANTHROPIC_MODELS_ENABLED` are enabled,
the consumer also receives nonstreaming `bedrock:InvokeModel` access to Claude
foundation models and the deployment account's Claude inference profiles for
the legacy platform fallback.
Upgrade and apply the Terraform module alongside the backend/consumer image;
updating the image alone does not install this permission. After IAM propagation,
verify an authorized monitor canary and its dead-letter alarm.

## ToolGuard capacity and scaling

**ToolGuard scale-out speed.** A new ToolGuard GPU host installs NVIDIA GRID
drivers at boot and takes 10–15 minutes before it serves traffic, so a
scale-out decision is only as fast as the spare hosts already running.
That boot path also needs outbound HTTPS to Ubuntu, NVIDIA GRID S3, Docker,
and related hosts — see [Self-Hosted Egress Requirements → ToolGuard](/deployment/egress-requirements#toolguard-gpu-bootstrap).
`runlayer_tool_guard_scaling_profile` trades idle GPU cost for
time-to-capacity:

| Profile | Scale-out trigger | Spare EC2 headroom | Instances per step | Scale-out cooldown |
| - | - | - | - | - |
| `conservative` (default) | 50% avg CPU | none (hosts launch after tasks are already pending) | `runlayer_tool_guard_desired_count` | 60s |
| `balanced` | 35% avg CPU | \~20% (capacity provider target 80) | 2 | 30s |
| `aggressive` | 25% avg CPU **or** 60% fleet-average GPU | \~30% (capacity provider target 70) | up to `runlayer_tool_guard_max_count` | 30s (GPU policy: 60s) |

`balanced` and `aggressive` keep one to two extra `g6f.large` instances
running at all times (billed even when idle) so the next task starts
immediately. Scale-in waits 15 minutes on every profile. The dashboard's
**Tasks: Desired vs Running** panel shows how long scale-outs wait on new
hosts. The `runlayer-tool-guard-scale-out-lag` alarm fires only when desired
tasks exceed running tasks for 20 consecutive minutes — longer than a cold
GPU host boot — so a routine scale-out does not trip it; when it does fire,
new hosts are not arriving (check the Auto Scaling group and capacity
provider).

**GPU signal on `aggressive`.** CPU is a weak load signal on a GPU inference
host, so `aggressive` adds a second target-tracking policy on fleet-average
GPU utilization (target 60%). The policy tracks the existing
`RunlayerToolGuard/GPUUtilization` metric (`Environment` dimension,
`Average` statistic): every host whose ToolGuard answers `/health` 200
publishes one sample per minute, so the per-minute average is the mean across
serving GPUs. Hosts that are not serving (capacity-provider spares, hosts
still installing drivers or loading the model, hosts in the recycle drain)
publish nothing and are left out, so the spare pool cannot hold the average
under the target while serving GPUs are busy. No per-host series is added, so
CloudWatch metric cardinality is unchanged.
Scaling takes the higher of the CPU and GPU decisions, so the GPU policy
never forces the fleet below what the CPU or predictive policies request.
Hosts publish GPU metrics from their boot script every 60 seconds, so changes to the ToolGuard launch
template trigger an ASG rolling instance refresh that replaces the
hosts: tasks restart on new hosts as old hosts drain (each restart reloads the
model, so expect a capacity dip per batch). Until the refresh replaces them,
hosts from the previous launch template still publish every 5 minutes and so
land in only one of every five 1-minute averages.

**Predictive scaling.** Because a reactive scale-out cannot beat the 10–15
minute host boot, the module also creates an ECS predictive scaling policy
on the ToolGuard service. It forecasts the next 48 hours of demand hourly
from up to 14 days of ToolGuard `predict` request history and raises the
task count `runlayer_tool_guard_predictive_scheduling_buffer_seconds`
(default 1200) before each forecast hour, so GPU hosts are booting before
the load arrives. `conservative` runs it in `forecast_only` mode (forecasts
are recorded but never acted on); `balanced` and `aggressive` run
`forecast_and_scale`. Override the profile default with
`runlayer_tool_guard_predictive_scaling_mode` (`disabled`, `forecast_only`,
`forecast_and_scale`). The first forecast appears after 24 hours of metrics;
review it in the ECS console (service → **Service auto scaling**) or with
`aws application-autoscaling get-predictive-scaling-forecast`, then flip
`forecast_only` → `forecast_and_scale` once the curve matches your traffic.
Predictive scaling only scales out and never raises the task count above
`runlayer_tool_guard_max_count`; the CPU policy still handles scale-in and
unforecast bursts. In regions where ECS predictive scaling is unavailable
(for example `il-central-1`) the policy is skipped automatically; the
`runlayer_tool_guard_predictive_scaling_mode_effective` output reports the
mode actually applied.

**Why the load is floored.** The forecast learns daily and weekly patterns
from the load metric, so a few days with scanning switched off would teach
it that those weekdays are idle and under-provision them for the next two
weeks. Before forecasting, weekday hours (Mon–Fri UTC) below
`runlayer_tool_guard_predictive_load_floor_fraction` (default `0.5`) of the
window's mean hourly request count are raised to that floor, so off days
look like quiet business days instead of zero; weekends forecast from actual
traffic (`runlayer_tool_guard_predictive_load_floor_weekdays_only = false`
floors every day, `0` disables the floor). Capacity is sized so each task
receives `runlayer_tool_guard_predictive_target_requests_per_task_per_minute`
(default `50` per minute, i.e. 3000 per hour). The default keeps per-task
load in the range where ToolGuard
`predict` p95 stays near its floor; latency rises steeply once a task serves more
than roughly twice that rate. Raise the target to trade latency for fewer hosts,
lower it for more headroom.

Running tasks are read from Container Insights (`RunningTaskCount`) when it
is enabled, otherwise from the GPU Auto Scaling group's in-service instance
count (one task per host). The Auto Scaling group count also includes the
spare hosts `balanced` and `aggressive` keep warm (roughly 20% and 30% of
the fleet), so without Container Insights requests per task read low by that
share and the forecast under-provisions. Pair `forecast_and_scale` with
`container_insights = "enabled"`; `terraform plan` warns when it is not.

## Agent Sandbox (Optional)

Procedures also need the agent sandbox: their code is validated and test-run
inside it. Without an agent runtime (`enable_runlayer_agents` with AgentCore),
creating or validating a Procedure returns 503. Select a [network configuration](/deployment/ecs-networking#security-configuration) that allows AgentCore to reach the application before applying.

```hcl theme={null}
# Separate Sentry DSN for sandbox errors (does not reuse backend sentry_dsn)
# Pass via TF_VAR / secrets backend — never commit
runlayer_agent_sandbox_sentry_dsn = var.agent_sandbox_sentry_dsn

# Optional image overrides
# runlayer_agent_sandbox_image_uri = "123456789012.dkr.ecr.us-east-1.amazonaws.com/runlayer/agent-sandbox:1.28.162"
# With WAF IP allowlisting, use VPC + PrivateLink (not PUBLIC):
# enable_privatelink              = true
# runlayer_agentcore_network_mode = "VPC"
# runlayer_agentcore_consumer_vpc_id = "vpc-xxxxx"
# runlayer_agentcore_consumer_private_subnet_ids = ["subnet-xxxxx", "subnet-yyyyy"]
# runlayer_agentcore_max_lifetime_seconds = 3600
# runlayer_agentcore_idle_timeout_seconds = 900

# Per-agent network access policies (requires VPC mode). Roll out in order:
# 1. VPC mode with direct egress left on (the defaults above).
# 2. sandbox_egress_mode = "proxy"                  # proxy runs, nothing routes through it yet
# 3. Turn the sandbox-egress-proxy feature flag on for every organization, for example
#    additional_backend_env_vars = { FF_ENABLE = "<flags already forced>,sandbox-egress-proxy" }
#    and wait for the backend rollout to finish before step 4.
# 4. sandbox_egress_mode = "proxy_only"             # proxy becomes the only path for every run; skipping
#    step 3 costs organizations without the flag a few minutes of sandbox internet while the backend rolls out
# sandbox_egress_proxy_egress_ports = [443, 8443]    # 443 is always open; add ports that host:port policies may use
# sandbox_egress_proxy_image_uri    = "…/runlayer/sandbox-egress-proxy:1.28.500"
```

With the egress proxy enabled, every sandbox request to the public internet is
authenticated per run and checked against the agent's
[network access policy](/platform-agents#network-access-policy). Runlayer
traffic stays on PrivateLink and is never proxied.

<Warning>
  Network access policies are enforced only by this ECS / AgentCore sandbox
  runtime. The in-cluster Kubernetes agent sandbox (`sandboxMode: k8s`) does not
  enforce them yet: an agent's policy is stored but has no effect there.
</Warning>

By default, Agent Sandbox routes LLM calls through the backend Anthropic proxy. To use OpenAI or an OpenAI-compatible gateway for **custom** agents, add a provider in **Settings → AI Providers** (no Terraform change needed).

This Terraform path provisions the **ECS / AgentCore** sandbox. In-cluster [kubernetes-sigs Agent Sandbox](https://agent-sandbox.sigs.k8s.io/) (`sandboxMode: k8s`, gVisor or Kata on the nodes) is operator/GKE — see [Runlayer Operator — K8s Agent Sandbox](/deployment/runlayer-operator#k8s-agent-sandbox-sandboxmode-k8s).

Runlayer Assistant always runs on Bedrock in this account — it ignores **Settings → AI Providers** by design, so it works without configuring a provider. Model access is handled as described above; this module exposes it as `bedrock_auto_model_access`, `bedrock_model_access_company_name`, `bedrock_model_access_company_website`, `bedrock_model_access_industry` and `bedrock_model_access_intended_users`.

Backend egress for WorkOS AuthKit, Bedrock, and AgentCore is listed in [Self-Hosted Egress Requirements](/deployment/egress-requirements).

## LLM gateway

```hcl theme={null}
enable_llm_gateway            = true
llm_gateway_listener_priority = 12
# llm_gateway_image_uri = "123456789012.dkr.ecr.us-east-1.amazonaws.com/runlayer/llm-gateway:1.28.162"
# Pass via TF_VAR_llm_gateway_runlayer_api_key — never commit
# llm_gateway_runlayer_api_key = var.llm_gateway_runlayer_api_key
```

## Microsoft Teams

Optional. Off unless all five values are set — the module variable is the client
secret, the rest ride `additional_backend_env_vars`.

Provisioning creates a per-agent Entra app registration and Azure Bot Service
resource in **your** Entra tenant and Azure subscription, so a self-managed
install needs its own multi-tenant app registration (Graph application
permissions `Application.ReadWrite.All`, `AppCatalog.ReadWrite.All`,
`TeamsAppInstallation.ReadWriteForUser.All`, admin-consented), an Azure
resource group whose IAM grants that app **Azure Bot Service Contributor**, and
the `Microsoft.BotService` provider registered on the subscription.

```hcl theme={null}
# ms_teams_app_client_secret = var.ms_teams_app_client_secret  # via TF_VAR_*

additional_backend_env_vars = {
  MS_TEAMS_APP_CLIENT_ID         = "00000000-0000-0000-0000-000000000000"
  MS_TEAMS_RUNLAYER_TENANT_ID    = "00000000-0000-0000-0000-000000000000"
  MS_TEAMS_AZURE_SUBSCRIPTION_ID = "00000000-0000-0000-0000-000000000000"
  MS_TEAMS_AZURE_RESOURCE_GROUP  = "runlayer-ms-teams-bots"
}
```

`MS_TEAMS_APP_URL` defaults to `APP_URL` (`https://<domain_name>`). Register
`https://<domain_name>/api/v1/integrations/ms-teams/admin-consent-callback` as a
Redirect URI (Authentication → Web) on that app registration, or admin consent
fails with `AADSTS500113`.

Each tenant that installs an agent's bot must also enable **Upload custom apps**
in its Teams Admin Center setup policy; without it, publishing to the tenant's
app catalog returns 403 no matter how many times it is retried.

`domain_name` is effectively immutable once agents are provisioned: each agent's
Azure Bot resource stores the messaging endpoint captured at creation time and
is never rewritten, so changing the domain orphans existing bots until each is
deleted and recreated.

## Audit SIEM, sessions, and hook events

```hcl theme={null}
enable_hook_events_pipeline = true
# hook_events_alert_email = "ops@example.com"

enable_audit_log_siem_export = true
# Write to your bucket (or omit external_* to use the Runlayer-managed bucket)
# audit_log_siem_export_external_s3_bucket = "acme-audit-siem"
# audit_log_siem_export_external_s3_region = "us-east-1"
# audit_log_siem_export_assume_role_arn    = "arn:aws:iam::123456789012:role/siem-writer"
# audit_log_siem_export_external_id        = "REPLACE_ME"
# audit_log_siem_export_s3_prefix          = "runlayer/audit"
# audit_log_siem_export_read_role_arns     = ["arn:aws:iam::123456789012:role/siem-reader"]

# session_siem_export_kinesis_consumer_config = { enabled = true, desired_count = 1 }
# session_siem_export_external_s3_bucket = "acme-session-siem"
# sessions_read_from_postgres = false
# session_materializer_config = { enabled = true, desired_count = 1 }

# data_export_bucket_arn = "arn:aws:s3:::acme-hook-events-export"
```

## Binary packages and downloads

```hcl theme={null}
enable_binary_package_distribution = true
# runlayer_download_token = var.runlayer_download_token  # via TF_VAR_*
# runlayer_downloads_base_url = "https://downloads.example.com"
```

## Customer OpenTelemetry

Export metrics from the backend, and/or run an OTEL collector sidecar that ships to **your** collector:

```hcl theme={null}
otel_exporter_otlp_metrics_endpoint = "https://otel.example.com:4317"
otel_exporter_otlp_metrics_protocol = "grpc"
otel_service_name                   = "runlayer-backend"

otel_collector_config = {
  enabled         = true
  export_endpoint = "otel-collector.example.com:4317" # your collector; not localhost
  export_insecure = false
}

flow_metrics_collection_enabled      = true
flow_trace_adaptive_sampling_enabled = true
```

## Secrets, Sentry, and catalog

```hcl theme={null}
# Prefer TF_VAR_* / a secrets backend — do not commit values
# secret_key  = var.secret_key   # must pair with master_salt, or omit both
# master_salt = var.master_salt
# sentry_dsn  = var.sentry_dsn
mcp_catalog_api_key = var.mcp_catalog_api_key # required for catalog / onboarding
# distribution_api_key = var.distribution_api_key

mcp_catalog_api_url = "" # empty → backend default

frontend_sentry_dsn                = ""
frontend_sentry_traces_sample_rate = "0.1"
sentry_audit_events_enabled        = true
sentry_relay_enabled               = true
# sentry_relay_upstream = "https://sentry.example.com/"
# sentry_relay_config   = { ... }

oauth_broker_url = "" # empty → module default when broker is used
# oauth_client_registration_rate_limit_per_hour = 1000
```

## Deployment identity

```hcl theme={null}
customer_id     = "acme"
deployment_name = "staging"
infra_version   = "v33.1.0"
```

For Sentry Relay credentials, follow [Sentry Relay Credentials from WorkOS Vault](/deployment/sentry-relay-vault). For software upgrades and rollback, follow [Updates](/operations/updates).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.