module "runlayer_infrastructure". Examples enable features and do not represent the defaults.
Bedrock and Anthropic
On this ECS deployment path, Runlayer Assistant uses Bedrock in the deployment account (default modelglobal.anthropic.claude-opus-4-8). It ignores Settings → AI Providers.
Having “some Anthropic model” enabled in the console is not enough — that
specific model (or its base) must be available.
Module v28.1.x+ attaches the model-access IAM actions to the ECS task role
(GetFoundationModelAvailability, PutUseCaseForModelAccess, Marketplace
subscribe, etc.). You do not add those by hand unless an SCP / boundary strips
them.
does not have access to the Anthropic model yet,
see troubleshooting.
Distribution configuration
The deployment guide configures Distribution with all three inputs below. Obtain a registered reader key from Runlayer before the first apply:flagd without both reader inputs. The
application default remains WorkOS when the provider input is omitted; setting
the key and URL alone does not switch providers. Existing WorkOS deployments
can stage the reader credentials before explicitly switching to flagd.
Hook-body compression (gzip-hooks, zstd-hooks) is on by default for
Distribution-backed deployments. WorkOS-provider deployments opt in until they
move to Distribution; FF_DISABLE wins if both name a flag:
distribution_api_key_previous to the
outgoing key, applying, waiting for ECS services to reach steady state, then
clearing it and applying again. Set distribution_status = "disabled" to revoke
access without destroying the stack; move openfeature_provider off flagd first.
Optional feature flags
Enable only the features you intend to operate. ToolGuard needs GPU quota and bootstrap egress. With WAF IP allowlisting, configure AgentCore VPC mode and PrivateLink before enabling agents. Runlayer Deploy is always enabled by the ECS module.sts:AssumeRole only on arn:aws:iam::*:role/runlayer-bedrock-* and
arn:aws:iam::*:role/RunlayerBedrock*. The saved role must trust the caller’s AWS
account/identity with the saved external ID and grant the required model access.
When both BEDROCK_ENABLED and BEDROCK_ANTHROPIC_MODELS_ENABLED are enabled,
the consumer also receives nonstreaming bedrock:InvokeModel access to Claude
foundation models and the deployment account’s Claude inference profiles for
the legacy platform fallback.
Upgrade and apply the Terraform module alongside the backend/consumer image;
updating the image alone does not install this permission. After IAM propagation,
verify an authorized monitor canary and its dead-letter alarm.
ToolGuard capacity and scaling
ToolGuard scale-out speed. A new ToolGuard GPU host installs NVIDIA GRID drivers at boot and takes 10–15 minutes before it serves traffic, so a scale-out decision is only as fast as the spare hosts already running. That boot path also needs outbound HTTPS to Ubuntu, NVIDIA GRID S3, Docker, and related hosts — see Self-Hosted Egress Requirements → ToolGuard.runlayer_tool_guard_scaling_profile trades idle GPU cost for
time-to-capacity:
balanced and aggressive keep one to two extra g6f.large instances
running at all times (billed even when idle) so the next task starts
immediately. Scale-in waits 15 minutes on every profile. The dashboard’s
Tasks: Desired vs Running panel shows how long scale-outs wait on new
hosts. The runlayer-tool-guard-scale-out-lag alarm fires only when desired
tasks exceed running tasks for 20 consecutive minutes — longer than a cold
GPU host boot — so a routine scale-out does not trip it; when it does fire,
new hosts are not arriving (check the Auto Scaling group and capacity
provider).
GPU signal on aggressive. CPU is a weak load signal on a GPU inference
host, so aggressive adds a second target-tracking policy on fleet-average
GPU utilization (target 60%). The policy tracks the existing
RunlayerToolGuard/GPUUtilization metric (Environment dimension,
Average statistic): every host whose ToolGuard answers /health 200
publishes one sample per minute, so the per-minute average is the mean across
serving GPUs. Hosts that are not serving (capacity-provider spares, hosts
still installing drivers or loading the model, hosts in the recycle drain)
publish nothing and are left out, so the spare pool cannot hold the average
under the target while serving GPUs are busy. No per-host series is added, so
CloudWatch metric cardinality is unchanged.
Scaling takes the higher of the CPU and GPU decisions, so the GPU policy
never forces the fleet below what the CPU or predictive policies request.
Hosts publish GPU metrics from their boot script every 60 seconds, so changes to the ToolGuard launch
template trigger an ASG rolling instance refresh that replaces the
hosts: tasks restart on new hosts as old hosts drain (each restart reloads the
model, so expect a capacity dip per batch). Until the refresh replaces them,
hosts from the previous launch template still publish every 5 minutes and so
land in only one of every five 1-minute averages.
Predictive scaling. Because a reactive scale-out cannot beat the 10–15
minute host boot, the module also creates an ECS predictive scaling policy
on the ToolGuard service. It forecasts the next 48 hours of demand hourly
from up to 14 days of ToolGuard predict request history and raises the
task count runlayer_tool_guard_predictive_scheduling_buffer_seconds
(default 1200) before each forecast hour, so GPU hosts are booting before
the load arrives. conservative runs it in forecast_only mode (forecasts
are recorded but never acted on); balanced and aggressive run
forecast_and_scale. Override the profile default with
runlayer_tool_guard_predictive_scaling_mode (disabled, forecast_only,
forecast_and_scale). The first forecast appears after 24 hours of metrics;
review it in the ECS console (service → Service auto scaling) or with
aws application-autoscaling get-predictive-scaling-forecast, then flip
forecast_only → forecast_and_scale once the curve matches your traffic.
Predictive scaling only scales out and never raises the task count above
runlayer_tool_guard_max_count; the CPU policy still handles scale-in and
unforecast bursts. In regions where ECS predictive scaling is unavailable
(for example il-central-1) the policy is skipped automatically; the
runlayer_tool_guard_predictive_scaling_mode_effective output reports the
mode actually applied.
Why the load is floored. The forecast learns daily and weekly patterns
from the load metric, so a few days with scanning switched off would teach
it that those weekdays are idle and under-provision them for the next two
weeks. Before forecasting, weekday hours (Mon–Fri UTC) below
runlayer_tool_guard_predictive_load_floor_fraction (default 0.5) of the
window’s mean hourly request count are raised to that floor, so off days
look like quiet business days instead of zero; weekends forecast from actual
traffic (runlayer_tool_guard_predictive_load_floor_weekdays_only = false
floors every day, 0 disables the floor). Capacity is sized so each task
receives runlayer_tool_guard_predictive_target_requests_per_task_per_minute
(default 50 per minute, i.e. 3000 per hour). The default keeps per-task
load in the range where ToolGuard
predict p95 stays near its floor; latency rises steeply once a task serves more
than roughly twice that rate. Raise the target to trade latency for fewer hosts,
lower it for more headroom.
Running tasks are read from Container Insights (RunningTaskCount) when it
is enabled, otherwise from the GPU Auto Scaling group’s in-service instance
count (one task per host). The Auto Scaling group count also includes the
spare hosts balanced and aggressive keep warm (roughly 20% and 30% of
the fleet), so without Container Insights requests per task read low by that
share and the forecast under-provisions. Pair forecast_and_scale with
container_insights = "enabled"; terraform plan warns when it is not.
Agent Sandbox (Optional)
Procedures also need the agent sandbox: their code is validated and test-run inside it. Without an agent runtime (enable_runlayer_agents with AgentCore),
creating or validating a Procedure returns 503. Select a network configuration that allows AgentCore to reach the application before applying.
sandboxMode: k8s, gVisor or Kata on the nodes) is operator/GKE — see Runlayer Operator — K8s Agent Sandbox.
Runlayer Assistant always runs on Bedrock in this account — it ignores Settings → AI Providers by design, so it works without configuring a provider. Model access is handled as described above; this module exposes it as bedrock_auto_model_access, bedrock_model_access_company_name, bedrock_model_access_company_website, bedrock_model_access_industry and bedrock_model_access_intended_users.
Backend egress for WorkOS AuthKit, Bedrock, and AgentCore is listed in Self-Hosted Egress Requirements.
LLM gateway
Microsoft Teams
Optional. Off unless all five values are set — the module variable is the client secret, the rest rideadditional_backend_env_vars.
Provisioning creates a per-agent Entra app registration and Azure Bot Service
resource in your Entra tenant and Azure subscription, so a self-managed
install needs its own multi-tenant app registration (Graph application
permissions Application.ReadWrite.All, AppCatalog.ReadWrite.All,
TeamsAppInstallation.ReadWriteForUser.All, admin-consented), an Azure
resource group whose IAM grants that app Azure Bot Service Contributor, and
the Microsoft.BotService provider registered on the subscription.
MS_TEAMS_APP_URL defaults to APP_URL (https://<domain_name>). Register
https://<domain_name>/api/v1/integrations/ms-teams/admin-consent-callback as a
Redirect URI (Authentication → Web) on that app registration, or admin consent
fails with AADSTS500113.
Each tenant that installs an agent’s bot must also enable Upload custom apps
in its Teams Admin Center setup policy; without it, publishing to the tenant’s
app catalog returns 403 no matter how many times it is retried.
domain_name is effectively immutable once agents are provisioned: each agent’s
Azure Bot resource stores the messaging endpoint captured at creation time and
is never rewritten, so changing the domain orphans existing bots until each is
deleted and recreated.