AWS
This guide provides detailed, step-by-step instructions for installing and configuring Obsrv on AWS, utilizing Terraform, Terragrunt, and Helm.
Infrastructure Requirements
1. System Specifications
CPU Requirements:
Minimum: 19 CPUs.
Optimal Configuration: 5 nodes with 4 cores each, totaling 80GB of RAM.
The installation package includes both lakehouse and real-time OLAP storage by default. If the lakehouse component is not required, only the real-time OLAP storage can be installed, reducing requirements to 16 CPUs and 64GB of RAM.
In this case, we recommend using 2 nodes with 8 cores each, totaling 64GB of RAM, by selecting the t2.2xlarge AWS instance type.
Availability Zones: All instances should be within the same availability zone to minimize cross-zone data transfer costs. The Obsrv installer will automatically create the EKS (Elastic Kubernetes Service) cluster for you.
2. Networking Setup
CIDR Block: Use a
/23CIDR range (512 IPs) for your environment.Example: A VPC with
10.0.0.0/23provides IPs from10.0.0.0to10.0.1.255.
Subnets: Ensure subnets are created in all availability zones within your AWS region.
Prerequisites
Before beginning the installation, make sure the following tools are installed on your Linux-based system:
Tool
Version
Installation Command
Official Documentation
Terraform
1.5.x or earlier
curl "https://releases.hashicorp.com/terraform/1.5.2/terraform_1.5.2_linux_amd64.zip" -o terraform.zip && unzip terraform.zip && sudo mv terraform /usr/local/bin/ && rm terraform.zip
Terragrunt
0.48 or later
curl -OL https://github.com/gruntwork-io/terragrunt/releases/download/v0.49.0/terragrunt_linux_amd64 && sudo mv terragrunt_linux_amd64 /usr/local/bin/terragrunt && sudo chmod +x /usr/local/bin/terragrunt
Helm
3.10.2 or later
curl https://get.helm.sh/helm-v3.10.2-linux-amd64.tar.gz -o helm.tar.gz && tar -zxvf helm.tar.gz && sudo mv linux-amd64/helm /usr/local/bin/
AWS CLI
2.10 or later
curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip" && unzip awscliv2.zip && sudo ./aws/install
Installation Steps
1. Clone the Obsrv Repository
Start by cloning the Obsrv automation repository and checkout to either the latest release tag or main.
2. Configure the Kubernetes Cluster
By executing the following commands which will bring up the kubernetes cluster in the AWS environment of configured region.
Navigate to the Configuration Directory:
Update Configuration Files:
Open
cluster_overides.tfand modify the configuration values to match your environment.
Configure S3 for Cluster State:
Open
obsrv.confand update your AWS credentials and bucket names.
3. Run the Installation Script
Navigate to the Infra Setup Directory:
Make the Script Executable:
Run the Installation:
To start the installation, run the script:
If you want the installer to automatically handle dependencies, set
install_dependencies=true.
4. Verify the Cluster
Once the installation completes, verify that your Kubernetes cluster is up and running:
This should show the nodes in your Kubernetes cluster.
Helm Chart Configuration
1. Navigate to the Helm Chart Directory
2. Update AWS Cloud Configuration
Note:
global-cloud-values-aws.yamlis auto-generated by Terraform (theaws_cloud_valuesmodule) during the./obsrv.sh installstep in Installation Steps, using the values fromcluster_overrides.tfvarsand resources Terraform creates (S3 bucket names, region, IAM role ARNs, Elastic IP, etc). It gets overwritten every time you runobsrv.sh install— do not edit it manually. Shown below just so you know what Terraform fills in:
3. Update Domain Configuration
In global-values.yaml, replace <domain> with your actual domain or Elastic IP:
4. Install Obsrv
Make the script executable and set the environment variables and run the installation:
core-setup installs the bootstrap CRDs, prerequisites, core database, and Kafka — all (migrations, monitoring, oauth, coreinfra, obsrvapis, obsrvtools, additional) depends on these being in place first, so run them in this order.
Post-Installation Verification
After completing the installation, follow these steps to verify that all components are running correctly:
1. Check Kubernetes Components
Verify all pods are running:
All pods should be in
Runningstate. Common namespaces to check:flink: Core Pipelinemonitoring: Monitoring stackdataset-api: Dataset APIsweb-console: Dataset Management console
Check Services:
Verify that essential services have external IPs assigned, particularly the Kong service.
If any component fails these checks, refer to the component-specific logs:
Troubleshooting
"Unreadable module directory" / lstat ../modules: no such file or directory on terragrunt init
Terragrunt copies the source directory into .terragrunt-cache before running. If terraform/aws/terragrunt.hcl has no explicit source block, newer Terragrunt versions (v1.x) copy only the aws directory, breaking the ../modules relative paths used by main.tf. Older Terragrunt versions (~0.45) didn't hit this because they ran in place.
Fix: ensure terraform/aws/terragrunt.hcl has a terraform { source = ... } block pointing at the repo root with the //terraform/aws working-dir suffix, so both ../modules and ../../helmcharts resolve correctly from the cache copy.
Failed to resolve provider packages: locked provider ... does not match configured version constraint
terraform/aws/.terraform.lock.hcl can drift out of sync with the provider version constraints in main.tf (e.g. main.tf requiring ~> 6.0 while the lock file is still pinned to a 4.x version). This happens when the constraint is bumped in a commit without regenerating the lock file.
Fix:
Commit the regenerated .terraform.lock.hcl.
S3 bucket does not exist on terragrunt init
The remote state bucket configured in obsrv.conf (AWS_TERRAFORM_BACKEND_BUCKET_NAME) must exist before terragrunt init runs. Either create it manually with aws s3api create-bucket, or pass --backend-bootstrap to auto-provision it (requires Terragrunt with backend-bootstrap support, e.g. v1.x):
Certificate shows valid in cert-manager, but browser still says "Not Secure"
kubectl get certificate reporting Ready: True only means cert-manager issued the cert into its Kubernetes Secret — it doesn't guarantee every service actually serving TLS on your domain can read that secret. If a service's Ingress lives in a namespace other than where the cert's Secret was created (commonly web-console), the kubernetes-reflector controller must be explicitly allowed to copy the secret into that namespace via the reflector.v1.k8s.emberstack.com/reflection-allowed-namespaces annotation on the Certificate (set in helmcharts/services/letsencrypt-ssl/templates/issuer-and-certs.yaml). A namespace missing from that list (e.g. dataset-api) causes the Kong ingress controller to fail fetching the secret for that Ingress, which can break Kong's config sync entirely and make it fall back to its default self-signed certificate for every host, not just the affected one.
Diagnose — check what cert is actually being served on the wire, independent of Kubernetes state:
If the issuer shows O=Kong, CN=localhost instead of Let's Encrypt, check the Kong ingress controller logs for failed to fetch the secret errors:
Fix: add the missing namespace to the reflection-allowed-namespaces annotation, then re-run the migrations helm bundle and restart Kong:
Upgrade Steps
Pull the Latest Code:
Update Configurations: Review and update configuration values as needed.
Run Terraform for Upgrade:
Upgrade with Updated Cloud Values:
By following these steps, you will ensure a successful installation and configuration of Obsrv on AWS.
Last updated
Was this helpful?