Skip to main content

Troubleshooting

This guide covers common issues and troubleshooting procedures for Runlayer. For complex issues or enterprise support, contact our technical team.

Quick Diagnostics

System Health Check

Kubernetes (Helm/EKS)
ECS (Terraform)

Database Connectivity

Common Issues

Blocked egress (self-hosted)

If login, agent runs, ToolGuard, or Bedrock fail behind a corporate SWG / URL allowlist, check Self-Hosted Egress Requirements (includes a symptom → host table). Device/client allowlists are on Network & Firewall Requirements.

Service Won’t Start

Symptoms: Services fail to start or immediately exit Solutions:
  1. Kubernetes: Inspect pod status and events
    • kubectl describe pod <pod> -n anysource
    • kubectl logs <pod> -n anysource
  2. ECS: Inspect stopped tasks and CloudWatch logs
    • aws ecs describe-tasks --cluster <cluster> --tasks <task-id>
    • Review the task’s CloudWatch log group
  3. Verify configuration values and secrets (Kubernetes Secrets or AWS SSM/Secrets Manager)
  4. Check security groups and network connectivity between services
  5. For ToolGuard EMPTY CAPACITY PROVIDER or AuthKit agent-token failures, see Self-Hosted Egress Requirements

Database Connection Issues

Symptoms: Application cannot connect to database Solutions:
  1. Verify the database endpoint and credentials
  2. Check security groups / network policies for database access
  3. Review database logs in AWS (RDS logs / CloudWatch)
  4. Validate environment variables or secret values used by the backend

Performance Issues

Symptoms: Slow response times or high resource usage Solutions:
  1. Check resource utilization (CloudWatch metrics or Kubernetes metrics)
  2. Review database performance and slow query logs
  3. Verify Redis cache connectivity and hit rate
  4. Inspect application logs for errors and timeouts

ACM Certificate Issues

Symptoms: Terraform fails looking up certificate or ACM certificates remain in PENDING_VALIDATION. Wildcard lookup fails (default behavior): The ECS module derives a wildcard domain from your domain (e.g., ecs.staging.runlayer.com → *.staging.runlayer.com) and looks up an existing certificate.
  1. Verify a wildcard certificate exists: aws acm list-certificates --query "CertificateSummaryList[?contains(DomainName, '*')]"
  2. If no wildcard certificate exists, either create one manually or set enable_acm_dns_validation = true to have Terraform create it.
DNS validation fails (when creating new certificates):
  1. Confirm the Route53 hosted zone exists in the AWS account running Terraform.
  2. Ensure hosted_zone_name matches the zone name (e.g., staging.runlayer.com).
  3. Re-run terraform apply so the _acme-challenge CNAME records are created automatically.

Authentication Problems

Symptoms: Users cannot log in or access resources Solutions:
  1. Verify authentication configuration
  2. Check external identity provider connectivity
  3. Review user permissions and roles
  4. Check JWT token configuration

Runlayer Assistant: Anthropic Bedrock model access

Symptoms: Runlayer Assistant (or platform Bedrock inference) replies with an error instead of an answer. Two common shapes:
Cause: Anthropic requires a one-time use case details form (and model agreement) before its models can be invoked in an AWS account. Runlayer Assistant always routes through Bedrock using your deployment’s own AWS credentials / task role, so the form and the specific model (default global.anthropic.claude-opus-4-8) must be available in the deployment account and Region. Enabling a different Claude model in the console is not enough. The IAM action list in the Forbidden message is a hint for old task roles. Current ECS module releases (including v28.1.x) already attach those actions to the ECS task role when Agents/Bedrock are in play — you should not need to paste them into a hand-written policy unless SCPs / permission boundaries strip them, or you are on a pre-model-access module version. With bedrock_auto_model_access = true (default), the backend submits the form at startup/login using bedrock_model_access_*. Defaults name Runlayer as the company — for a self-managed install, override those fields before first successful submit (the form is submit-once and permanent). Solutions:
  1. Check whether the form was already submitted, and whether the assistant model is available:
    Requires AWS CLI 2.27.42+ and sufficient Bedrock permissions on the calling principal.
  2. Confirm the ECS task role actually has the model-access actions (module-managed policy) and that org SCPs do not deny bedrock:PutUseCaseForModelAccess / aws-marketplace:Subscribe. Backend logs bedrock_model_access_iam_denied when the auto-grant path hits AccessDenied.
  3. If the form was not submitted, either let auto-grant run (with correct bedrock_model_access_* overrides) or submit by hand — Bedrock console → Model catalog → any Anthropic model → Submit use case details. Allow up to 15 minutes to propagate.
  4. AWS accepts the submission once per account, or once at your AWS Organization’s management account, with member accounts inheriting it. Check the management account before submitting again. Opt-in Regions require their own submission.
  5. Confirm the model is served from your Region. The default global. inference profile is not offered from every BEDROCK_REGION — set RUNLAYER_ASSISTANT_BEDROCK_MODEL to your geography’s profile (for example au.anthropic.claude-opus-4-6-v1 in ap-southeast-2). The risk-categorization explainer picks its model from BEDROCK_REGION automatically (Claude Haiku 4.5 on your geography’s profile outside the US); if you pinned RISK_CATEGORIZATION_MODEL, make sure that id is served in your Region too. A wrong id shows up as a single risk_categorization_breaker_opened warning naming the model and Region, after which categorization switches to the fallback model (Amazon Nova Lite on your geography’s us./eu./apac. profile; the warning’s fallback_model_id field names it) and retries the primary every RISK_CATEGORIZATION_BREAKER_SECONDS (default 300). If the fallback is rejected too, or your Region has none, categorization is suspended between retries and violations keep their heuristic reason.
  6. If your organization restricts Regions with SCPs, allow bedrock:InvokeModel* in every destination Region of the inference profile.
See Bedrock and Anthropic (configuration accordion) and Self-Hosted Egress Requirements for related networking.

Log Analysis

Application Logs

Kubernetes (Helm/EKS)
ECS (Terraform)

Database Logs

Review database logs in AWS (RDS logs / CloudWatch).

Enterprise Support

For complex issues, performance optimization, or enterprise-level troubleshooting:

Enterprise Technical Support

Contact our technical team for advanced troubleshooting and 24/7 support

Support Information

When contacting support, please include:
  • System Information: OS, deployment method, AWS region
  • Error Messages: Complete error messages and stack traces
  • Log Files: Relevant application and system logs
  • Configuration: Sanitized configuration files (remove secrets)
  • Steps to Reproduce: Detailed steps that led to the issue

Escalation Process

  1. Level 1: Basic troubleshooting (this guide)
  2. Level 2: Advanced diagnostics (contact support)
  3. Level 3: Engineering escalation (critical issues)

Preventive Measures

  • Regular Monitoring: Set up health checks and alerting
  • Log Rotation: Configure proper log management
  • Resource Monitoring: Monitor CPU, memory, and disk usage
  • Backup Verification: Regularly test backup and restore procedures
Contact our support team for comprehensive monitoring setup and proactive issue prevention.