> ## Documentation Index
> Fetch the complete documentation index at: https://docs.runlayer.com/llms.txt
> Use this file to discover all available pages before exploring further.

# ECS Configuration and Operations

> Database, cache, scaling, monitoring, secrets, and outputs for ECS deployments

Use these examples when preparing your [ECS deployment](/deployment/terraform), or when updating an existing installation. All input snippets belong inside the module block unless marked as root outputs. Merge examples into your existing configuration.

The [inputs reference](/deployment/terraform-ecs-inputs) is the full variable contract. [Networking](/deployment/ecs-networking) and [optional features](/deployment/ecs-features) have their own guides.

## Naming constraints

These values are used in AWS resource names (ElastiCache replication group IDs, secrets, IAM, etc.) and are validated by the module:

| Variable | Constraint | Default |
| - | - | - |
| `project` | **1–10 characters**; letters, numbers, and hyphens only (`^[a-zA-Z0-9-]+$`) | `"anysource"` |
| `environment` | One of: `production`, `staging`, `development` | `"production"` |

Generated names follow `<project>-<environment>-…`. Example: `project = "acme"`, `environment = "staging"` → Redis replication group `acme-staging-redis`.

## Database Configuration

Managed Aurora is always **PostgreSQL Serverless v2** (`db.serverless`). There is no provisioned instance-class input (`db.r6g.large`, etc.). Size the cluster with Aurora Capacity Units (ACUs):

| Field | Default | Meaning |
| - | - | - |
| `database_config.min_capacity` | `2` | Minimum ACUs (floor while the cluster is running) |
| `database_config.max_capacity` | `16` | Maximum ACUs (ceiling the cluster can scale to) |
| `database_config.enable_ops_reader` | `false` | Extra Serverless v2 reader + custom endpoint for ops/analytics |

```hcl theme={null}
database_name     = "anysource"
database_username = "postgres"
database_config = {
  engine_version      = "16"
  min_capacity        = 4 # ACUs, not an instance class
  max_capacity        = 32
  publicly_accessible = false
  backup_retention    = 30
}
```

Pick ACU values [supported for Aurora PostgreSQL Serverless v2](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/aurora-serverless-v2.setting-capacity.html) in your region. For a database you operate yourself (any size or engine), use `external_database` instead.

## Bring Your Own Database / Cache

Skip Aurora and/or ElastiCache when you operate PostgreSQL and Redis/Valkey yourself. Postgres and Redis can be configured independently.

```hcl theme={null}
external_database = {
  host        = "postgres.internal.example.com"
  reader_host = "postgres-reader.internal.example.com" # optional
  port        = 5432
  database    = "runlayer"
  username    = "runlayer"
  ssl_mode    = "require"
}
external_database_password = var.external_db_password

external_redis = {
  host        = "valkey.internal.example.com"
  port        = 6379
  tls_enabled = true
}
external_redis_password = var.external_redis_password
```

ECS tasks receive `POSTGRES_*` env vars and Secrets Manager entries for `PLATFORM_DB_PASSWORD` and `REDIS_URL`. Ensure private subnets can reach your endpoints. For Kubernetes, see the [Runlayer Operator database and Redis contract](/deployment/runlayer-operator#5-deploy-a-runlayerinstance).

## Redis / ElastiCache

Managed Redis (skipped when `external_redis` is set):

| Setting | Default | Notes |
| - | - | - |
| Replication group ID | `<project>-<environment>-redis` | Derived; not an input |
| `redis_node_type` | `"cache.t3.medium"` | Any [ElastiCache node type](https://docs.aws.amazon.com/AmazonElastiCache/latest/dg/CacheNodes.SupportedTypes.html) available in your region |
| `redis_num_cache_clusters` | `1` | Use `2+` with failover for HA |
| `redis_automatic_failover_enabled` | `false` | Requires 2+ nodes |
| `redis_multi_az_enabled` | `false` | Requires 2+ nodes and failover |
| `redis_snapshot_retention_limit` | `null` | Resolves to `1` in production, `0` otherwise; set `0`–`35` to override |

```hcl theme={null}
redis_node_type                  = "cache.r7g.large"
redis_num_cache_clusters         = 2
redis_automatic_failover_enabled = true
redis_multi_az_enabled           = true
redis_snapshot_retention_limit   = 7

# Dedicated noeviction Redis (REDIS_CRITICAL_URL)
# Group name: <project>-<environment>-redis-critical
enable_critical_redis             = true
critical_redis_node_type          = "cache.t4g.small"
critical_redis_num_cache_clusters = 2
critical_redis_multi_az_enabled   = true
```

## Service Scaling

Defaults already set production-ready CPU/memory. To override, replace the full `services_configurations` map (partial maps are not merged):

```hcl theme={null}
services_configurations = {
  "backend" = {
    name                              = "backend"
    path_pattern                      = ["/api/*", "/docs*", "/redoc*", "/.well-known*"]
    health_check_path                 = "/api/v1/utils/health-check/"
    container_port                    = 8000
    host_port                         = 8000
    port                              = 8000
    priority                          = 10
    min_capacity                      = 3  # runtime floor (autoscaling target)
    max_capacity                      = 10 # runtime ceiling
    cpu                               = 2048
    memory                            = 4096
    cpu_auto_scalling_target_value    = 70
    memory_auto_scalling_target_value = 80
  }

  "frontend" = {
    name              = "frontend"
    path_pattern      = ["/*"]
    health_check_path = "/"
    container_port    = 80
    host_port         = 80
    port              = 80
    priority          = 20
    min_capacity      = 2
    max_capacity      = 10
    cpu               = 512
    memory            = 1024
  }
}
```

`min_capacity` / `max_capacity` control the running fleet: Application Auto Scaling owns each service's live `desired_count`, and Terraform ignores it after the service is created, so `desired_count` only seeds a brand-new service. To change the floor of an existing service, raise `min_capacity` (or `backend_min_capacity`).

Reserve `backend.priority + 1` for the `/mcp` alias listener rule (module validation enforces this).

## ECS services, workers, and migrations

```hcl theme={null}
backend_min_capacity               = 3 # runtime floor; autoscaling owns the live count
backend_ephemeral_storage_size_gib = 40
backend_stop_timeout_seconds       = 120
health_check_grace_period_seconds  = 60
target_group_deregistration_delay  = 30
wait_for_ecs_steady_state          = true
ecs_service_wait_timeout           = "30m"
enable_ecs_exec                    = false
enable_read_only_root_filesystem   = true
deletion_protection                = true

workers = 4
worker_config = {
  min_capacity = 2 # runtime floor; autoscaling owns the live count
  cpu          = 1024
  memory       = 2048
}

additional_backend_env_vars = {
  LOG_LEVEL = "INFO"
}

# Restore Aurora from snapshot at create time only
# snapshot_identifier = "rds:acme-staging-YYYY-MM-DD"

# migration_lock_timeout_ms = 0
# migration_statement_timeout_ms = 0
# migration_retry_count = 0
# backend_migration_wait_timeout_seconds = 1800
# backend_migration_aws_role = { arn = "arn:aws:iam::123456789012:role/migration-runner" }
```

## Monitoring & Alerting

```hcl theme={null}
enable_monitoring = true

# Default true: mission-critical CloudWatch alarms → SNS topic
# ${project}-${environment}-critical-alerts
critical_alerts_enabled = true

# Fan alarms to an existing same-region SNS topic (e.g. your Slack/Chatbot relay).
# The module does not create Slack Chatbot config; wire Slack on that topic yourself.
critical_alerts_extra_sns_topic_arns = [
  "arn:aws:sns:us-east-1:123456789012:my-ops-alerts",
]
# ToolGuard's actionable alarms fan out to these topics too (see the module README).
critical_alerts_create_sqs_subscription = true

enable_db_health_alarms    = true
enable_agent_run_alarms    = true
enable_worker_queue_alarms = true
enable_dashboard           = true

# ECS Container Insights: null uses environment default
# (production → "enabled", staging/development → "disabled")
container_insights = "enabled"
```

**Automatic CloudWatch alarms cover:** ECS CPU/memory, RDS, Redis, ALB latency/unhealthy targets/5XX, and VPC flow logs.

The Redis per-task connection alarm below requires a module release that includes `redis_connections_per_backend_task_threshold`. For the deployment guide's `v33.1.0`, omit that input and use `redis_curr_connections_threshold = 2000` instead.

```hcl theme={null}
rds_alarm_config = {
  DiskQueueDepth = { period = 300, threshold = 5, unit = "Count" }
  WriteIOPS      = { period = 300, threshold = 1000, unit = "Count" }
  ReadIOPS       = { period = 300, threshold = 1000, unit = "Count" }
  Storage        = { period = 300, threshold = 107374182400, unit = "Bytes" }
}

alb_5xx_alarm_period    = 300
alb_5xx_alarm_threshold = 1

redis_curr_connections_threshold             = 20000
redis_connections_per_backend_task_threshold = 750
redis_evictions_threshold                    = 1
redis_burst_evictions_threshold              = 5000
redis_sustained_evictions_threshold          = 500
rds_database_connections_threshold           = 2700
rds_writer_cpu_utilization_threshold         = 85
rds_reader_cpu_utilization_threshold         = 90
rds_writer_acu_utilization_threshold         = 90
db_idle_in_txn_max_duration_threshold        = 300
db_idle_in_txn_count_threshold               = 5

agent_run_stall_threshold_minutes        = 60
agent_run_stalled_count_threshold        = 0
agent_run_failure_rate_threshold_percent = 5
agent_run_failure_rate_min_volume        = 20
```

## Resource tags

The `additional_tags` input accepts a string map and defaults to `{}`. It applies
deployment metadata to directly managed workload resources. Resource identity tags
take precedence. No partner attribution is enabled by default. Verify product code,
resource eligibility and existing tag ownership before opt-in. See the module README
for the coverage available in your published module version.

## Secrets management

All application secrets live in **AWS Secrets Manager in your account**. The module creates two primary secrets (plus an optional LLM gateway secret):

| Secret / key | Where stored | Source |
| - | - | - |
| `SECRET_KEY` | Restricted app-keys secret (`restricted/...-app-keys-<suffix>`) | Auto-generated (or from `secret_key`) |
| `MASTER_SALT` | Restricted app-keys secret | Auto-generated (or from `master_salt`) |
| `PLATFORM_DB_PASSWORD` | App secrets secret (`<project>-<env>-<suffix>`) | Auto-generated (or from `external_database_password`) |
| `REDIS_URL` | App secrets secret | Auto-generated (or from external Redis settings) |
| `AUTH_API_KEY` | App secrets secret | From `auth_api_key` / `TF_VAR_auth_api_key` |
| `MCP_CATALOG_API_KEY` | App secrets secret | From `mcp_catalog_api_key` (required for catalog and onboarding) |
| `DISTRIBUTION_API_KEY` | App secrets secret | From `distribution_api_key` / `TF_VAR_distribution_api_key` (sensitive; required when `openfeature_provider = "flagd"`, as in the deployment guide) |
| `SENTRY_*` | App secrets secret | Optional (`sentry_dsn` or WorkOS Vault via Relay) |
| `LLM_GATEWAY_RUNLAYER_API_KEY` | Separate `restricted/...-llm-gateway-<suffix>` secret | From `llm_gateway_runlayer_api_key` when `enable_llm_gateway = true` |

**Design notes:**

* `SECRET_KEY` and `MASTER_SALT` are isolated in the restricted secret so only privileged task roles can read them.
* The app secret and restricted secret share the same random name suffix to avoid recreate collisions during Secrets Manager recovery windows.
* Prefer `TF_VAR_*` / a secrets backend for sensitive inputs — do not commit API keys to git.
* Rotate `DISTRIBUTION_API_KEY` through the Terraform input and coordinate registration with Runlayer; follow the [Distribution key rotation procedure](/deployment/ecs-features#distribution-configuration) so the new key is accepted before tasks use it.
* After deploy, rotate other non-key values with Secrets Manager (then force a new ECS deployment so tasks pick up the new version). Rotate `SECRET_KEY`/`MASTER_SALT` via `secret_key`/`master_salt` instead — Terraform reverts out-of-band edits to the restricted secret on the next apply.

The deployment guide exports `ecs_cluster_name`. To inspect secret names with Terraform, also add these **root outputs**:

```hcl theme={null}
output "app_secrets_name" {
  value = module.runlayer_infrastructure.app_secrets_name
}

output "restricted_app_keys_secret_name" {
  value = module.runlayer_infrastructure.restricted_app_keys_secret_name
}
```

Apply the output-only change, then inspect names with `terraform output -raw app_secrets_name` or `terraform output -raw restricted_app_keys_secret_name`. These outputs contain names, not secret values.

Use your approved secret-management workflow to update non-key values. Preserve all other keys in the JSON secret. Restart every service that consumes the changed value; for the core services:

```bash theme={null}
for service in backend-service worker-service mcp-exec-worker-service; do
  aws ecs update-service \
    --cluster "$(terraform output -raw ecs_cluster_name)" \
    --service "$service" --force-new-deployment \
    --query 'service.{name:serviceName,status:status}' --output table
done
```

Also restart any enabled optional consumers using that secret, then repeat the [deployment verification](/deployment/terraform#verify).

## Module outputs (attributes)

These are child-module attributes. `terraform output` reads only outputs declared in your root configuration; export an attribute there before using it from the CLI. Useful attributes include:

| Output | Description |
| - | - |
| `application_url` | Primary `https://` URL |
| `alb_dns_name` / `alb_zone_id` | Public ALB DNS + hosted-zone ID |
| `public_alb_dns_name` / `internal_alb_dns_name` | Dual-ALB DNS names (when enabled) |
| `privatelink_service_name` | VPC endpoint service name (when PrivateLink enabled) |
| `ecs_cluster_arn` / `ecs_cluster_name` | ECS cluster identifiers |
| `vpc_id` / `private_subnet_ids` | Network IDs |
| `task_role_arn` / `ecs_task_execution_role_arn` | IAM roles |
| `app_secrets_name` | Application Secrets Manager secret name |
| `restricted_app_keys_secret_name` / `restricted_app_keys_secret_arn` | Restricted key material |
| `llm_gateway_base_url` | LLM gateway base URL (when enabled) |

```bash theme={null}
terraform output application_url
curl "$(terraform output -raw application_url)/api/v1/utils/health-check/"
```

## Common use cases

### Private enterprise deployment

```hcl theme={null}
alb_access_type   = "private"
alb_allowed_cidrs = ["10.0.0.0/8"]

waf_enable_ip_allowlisting = true
waf_allowlist_ipv4_cidrs = [
  "10.100.0.0/16",
  "10.200.0.0/16",
]

database_config = {
  publicly_accessible = false
  backup_retention    = 30
}
```

### High availability production

```hcl theme={null}
# Aurora Serverless v2 ACUs (not db.r* instance classes)
database_config = {
  min_capacity = 8
  max_capacity = 64
}

redis_num_cache_clusters         = 2
redis_automatic_failover_enabled = true
redis_multi_az_enabled           = true
```

### Development environment

```hcl theme={null}
environment = "development"
database_config = {
  min_capacity     = 2
  max_capacity     = 4
  backup_retention = 1
}
enable_monitoring = false
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.