Troubleshooting
This guide covers common issues and troubleshooting procedures for Runlayer. For complex issues or enterprise support, contact our technical team.Quick Diagnostics
System Health Check
Kubernetes (Helm/EKS)Database Connectivity
Common Issues
Blocked egress (self-hosted)
If login, agent runs, ToolGuard, or Bedrock fail behind a corporate SWG / URL allowlist, check Self-Hosted Egress Requirements (includes a symptom → host table). Device/client allowlists are on Network & Firewall Requirements.Service Won’t Start
Symptoms: Services fail to start or immediately exit Solutions:- Kubernetes: Inspect pod status and events
kubectl describe pod <pod> -n anysourcekubectl logs <pod> -n anysource
- ECS: Inspect stopped tasks and CloudWatch logs
aws ecs describe-tasks --cluster <cluster> --tasks <task-id>- Review the task’s CloudWatch log group
- Verify configuration values and secrets (Kubernetes Secrets or AWS SSM/Secrets Manager)
- Check security groups and network connectivity between services
- For ToolGuard
EMPTY CAPACITY PROVIDERor AuthKit agent-token failures, see Self-Hosted Egress Requirements
Database Connection Issues
Symptoms: Application cannot connect to database Solutions:- Verify the database endpoint and credentials
- Check security groups / network policies for database access
- Review database logs in AWS (RDS logs / CloudWatch)
- Validate environment variables or secret values used by the backend
Performance Issues
Symptoms: Slow response times or high resource usage Solutions:- Check resource utilization (CloudWatch metrics or Kubernetes metrics)
- Review database performance and slow query logs
- Verify Redis cache connectivity and hit rate
- Inspect application logs for errors and timeouts
ACM Certificate Issues
Symptoms: Terraform fails looking up certificate or ACM certificates remain inPENDING_VALIDATION.
Wildcard lookup fails (default behavior):
The ECS module derives a wildcard domain from your domain (e.g., ecs.staging.runlayer.com → *.staging.runlayer.com) and looks up an existing certificate.
- Verify a wildcard certificate exists:
aws acm list-certificates --query "CertificateSummaryList[?contains(DomainName, '*')]" - If no wildcard certificate exists, either create one manually or set
enable_acm_dns_validation = trueto have Terraform create it.
- Confirm the Route53 hosted zone exists in the AWS account running Terraform.
- Ensure
hosted_zone_namematches the zone name (e.g.,staging.runlayer.com). - Re-run
terraform applyso the_acme-challengeCNAME records are created automatically.
Authentication Problems
Symptoms: Users cannot log in or access resources Solutions:- Verify authentication configuration
- Check external identity provider connectivity
- Review user permissions and roles
- Check JWT token configuration
Runlayer Assistant: Anthropic Bedrock model access
Symptoms: Runlayer Assistant (or platform Bedrock inference) replies with an error instead of an answer. Two common shapes:global.anthropic.claude-opus-4-8) must be available in the deployment account and Region. Enabling a different Claude model in the console is not enough.
The IAM action list in the Forbidden message is a hint for old task roles. Current ECS module releases (including v28.1.x) already attach those actions to the ECS task role when Agents/Bedrock are in play — you should not need to paste them into a hand-written policy unless SCPs / permission boundaries strip them, or you are on a pre-model-access module version.
With bedrock_auto_model_access = true (default), the backend submits the form at startup/login using bedrock_model_access_*. Defaults name Runlayer as the company — for a self-managed install, override those fields before first successful submit (the form is submit-once and permanent).
Solutions:
-
Check whether the form was already submitted, and whether the assistant model is available:
Requires AWS CLI 2.27.42+ and sufficient Bedrock permissions on the calling principal.
-
Confirm the ECS task role actually has the model-access actions (module-managed policy) and that org SCPs do not deny
bedrock:PutUseCaseForModelAccess/aws-marketplace:Subscribe. Backend logsbedrock_model_access_iam_deniedwhen the auto-grant path hits AccessDenied. -
If the form was not submitted, either let auto-grant run (with correct
bedrock_model_access_*overrides) or submit by hand — Bedrock console → Model catalog → any Anthropic model → Submit use case details. Allow up to 15 minutes to propagate. - AWS accepts the submission once per account, or once at your AWS Organization’s management account, with member accounts inheriting it. Check the management account before submitting again. Opt-in Regions require their own submission.
-
Confirm the model is served from your Region. The default
global.inference profile is not offered from everyBEDROCK_REGION— setRUNLAYER_ASSISTANT_BEDROCK_MODELto your geography’s profile (for exampleau.anthropic.claude-opus-4-6-v1inap-southeast-2). The risk-categorization explainer picks its model fromBEDROCK_REGIONautomatically (Claude Haiku 4.5 on your geography’s profile outside the US); if you pinnedRISK_CATEGORIZATION_MODEL, make sure that id is served in your Region too. A wrong id shows up as a singlerisk_categorization_breaker_openedwarning naming the model and Region, after which categorization switches to the fallback model (Amazon Nova Lite on your geography’sus./eu./apac.profile; the warning’sfallback_model_idfield names it) and retries the primary everyRISK_CATEGORIZATION_BREAKER_SECONDS(default 300). If the fallback is rejected too, or your Region has none, categorization is suspended between retries and violations keep their heuristic reason. -
If your organization restricts Regions with SCPs, allow
bedrock:InvokeModel*in every destination Region of the inference profile.
Log Analysis
Application Logs
Kubernetes (Helm/EKS)Database Logs
Review database logs in AWS (RDS logs / CloudWatch).Enterprise Support
For complex issues, performance optimization, or enterprise-level troubleshooting:Enterprise Technical Support
Contact our technical team for advanced troubleshooting and 24/7 support
Support Information
When contacting support, please include:- System Information: OS, deployment method, AWS region
- Error Messages: Complete error messages and stack traces
- Log Files: Relevant application and system logs
- Configuration: Sanitized configuration files (remove secrets)
- Steps to Reproduce: Detailed steps that led to the issue
Escalation Process
- Level 1: Basic troubleshooting (this guide)
- Level 2: Advanced diagnostics (contact support)
- Level 3: Engineering escalation (critical issues)
Preventive Measures
- Regular Monitoring: Set up health checks and alerting
- Log Rotation: Configure proper log management
- Resource Monitoring: Monitor CPU, memory, and disk usage
- Backup Verification: Regularly test backup and restore procedures