Skip to content
English
  • There are no suggestions because the search field is empty.

Configuring Datadog for Amazon EKS Recovery

This article describes the options available for controlling where your recovered EKS cluster sends Datadog telemetry, and helps you choose the right one for your environment.

When Arpio recovers an Amazon EKS cluster, the recovered cluster includes the same Datadog configuration as the source cluster, including its production Datadog API key. If you make no changes, logs, metrics, and traces from your recovery environment can flow into your production Datadog organization. This can clutter production dashboards and trigger production monitors and alerts during a test or recovery.

Recommended for most customers: If you want to keep observability during a test or recovery while sending telemetry to a different Datadog destination, use the EKS translation ConfigMap (Option 1).

Options at a glance

Option Best for Main limitation
1. EKS translation ConfigMap Redirecting the Datadog Agent to a different Datadog organization or API key while keeping observability The ConfigMap stores values in plaintext, so it should remap secret references rather than hold secret values
2. Network Sandbox domain allowlist Blocking all Datadog traffic from the recovery environment during a test Applies to every application in the same region pair
3. SSM Parameter Store or Secrets Manager Datadog configuration that is sourced from AWS rather than hardcoded Does not apply to native Kubernetes secrets
4. Lifecycle event automation Hardcoded values that cannot easily be refactored Telemetry may reach production briefly before the automation runs

Before you choose: review your Datadog setup

The right option depends on how Datadog is installed in your cluster and where its configuration lives. Before you choose, answer the following questions.

How is Datadog set up?

  • Do you use the Datadog Operator EKS add-on, or a self-installed Datadog Agent deployed through Helm?
  • Do you send telemetry to standard Datadog endpoints or to custom endpoints?
  • Do you have a separate Datadog organization for recovery that should use a different API key, application key, or bearer token?
  • Do you need a tagging scheme so production and recovery telemetry can be told apart?

Where does the Datadog configuration live?

  • Is the API key stored as a native Kubernetes secret, an external secret, or an AWS Secrets Manager secret?
  • Is the endpoint or key hardcoded in values.yaml, or referenced from SSM Parameter Store or Secrets Manager?
  • Is the value set on a container environment variable, or at the deployment or Helm chart level?

The last question matters most. Values set at the chart level, or referenced through secretKeyRef, are usually good candidates for the translation ConfigMap. Values hardcoded deep in a Helm chart usually point toward lifecycle event automation.

Choosing an option

Work through the questions below in order.

  1. Does Datadog need to keep working during the test or recovery?
    • No – use the Network Sandbox (Option 2) to block Datadog traffic entirely.
    • Yes – continue to step 2.
  2. Is the Datadog configuration addressable in your Kubernetes manifests? For example, is the API key referenced with secretKeyRef.name, or set as a chart-level environment value?
  3. Is the configuration stored in, or can it be refactored into, SSM Parameter Store or AWS Secrets Manager?
  4. Use lifecycle event automation (Option 4). If the brief window before the automation runs is not acceptable, and you don't need observability during the test, combine it with the Network Sandbox (Option 2).
Your situation Recommended option
You need observability, and the value is set at the chart level or through secretKeyRef Option 1: EKS translation ConfigMap
You want to suppress Datadog telemetry entirely during a test Option 2: Network Sandbox
The key already lives in Secrets Manager or SSM, or the configuration can be moved there Option 3: SSM Parameter Store or Secrets Manager
Values are hardcoded in the Helm chart and can't be refactored Option 4: Lifecycle event automation

Note: Option 3 applies only to values stored in AWS Secrets Manager or SSM Parameter Store. Native Kubernetes secrets are not supported by this option.

Option 1: EKS translation ConfigMap

Arpio can rewrite values in your Kubernetes manifests as resources are restored into the recovery environment. This is the most effective way to keep observability active while pointing the Datadog Agent somewhere else.

For full details, see EKS Translation ConfigMap.

Recommended pattern: Remap the secret reference, not the secret value. This avoids placing a replacement Datadog API key in plaintext in the ConfigMap.

source:
  k8sTypes:
    - DaemonSet
    - Deployment
  paths:
    - $.spec.template.spec.containers.[*].env.[*].valueFrom.secretKeyRef.name
targets:
  - type: static
    matcher: full
    valueMap:
      datadog-agent: datadog-agent-dr

Option 1: EKS Translation configMap

How it works

  1. Create two secrets in your source cluster: datadog-agent for production and datadog-agent-dr for recovery.
  2. During restore, the translation ConfigMap rewrites the secretKeyRef.name from datadog-agent to datadog-agent-dr.
  3. When the recovered Datadog DaemonSet starts, DD_API_KEY resolves from the datadog-agent-dr secret instead of the production secret.

Other values you can translate

  • Datadog endpoint
  • DD_ENV
  • DD_SITE
  • Tags
  • Container image references
  • Account IDs, ARNs, or other manifest values addressable through JSONPath

Things to keep in mind

  • Don't store secrets in the ConfigMap. ConfigMap values are stored in plaintext. Use translation to remap references to secrets rather than to hold secret material directly.
  • Secret-to-secret value linking is not supported. You can replace one secretKeyRef with another, but you can't translate a secret's value so that it's sourced from a different secret.
  • A missing target secret can prevent pods from starting. If datadog-agent-dr doesn't exist in the recovery cluster, the Datadog pods may fail to initialize. Some customers use this deliberately, so they can verify secrets before enabling services.
  • Scaling the agent to zero is not supported. Disabling the Datadog Agent through replica count or scheduling changes is not a supported workflow for translation.
  • Keep translations narrow. Broad translations, such as rewriting region names everywhere, can unintentionally change recovery resource names. Scope each translation to the specific paths and values you need.

Option 2: Network Sandbox domain allowlist

The Network Sandbox blocks outbound traffic from your recovery environment at a shared AWS Network Firewall. You can use it to prevent the recovered cluster from sending any telemetry to Datadog during a test.

For full details, see Advanced Network Sandbox for AWS.

To block Datadog, leave Datadog domains out of the sandbox allowlist.

To allow Datadog, add these entries to the allowlist:

.datadoghq.com 
.datadoghq.eu

The leading dot permits all subdomains, including api., trace.agent., agent-intake.logs., and app., without listing each one individually.

Things to keep in mind

  • The firewall is shared across the region pair. Allowlist settings apply to every application that recovers into the same region pair.
  • Suggested Domains are updated with a delay of about five minutes. The list of blocked traffic shown in Arpio is near-real-time, so blocked Datadog traffic from a first test appears after roughly five minutes.
  • Updating sandbox settings while a recovery point is being applied. If one application shows that Arpio is applying the latest recovery point, you can open Sandbox Settings from another application in the same region pair, because the firewall is shared. To verify changes immediately, open the recovery account in the AWS console and check VPC > Network Firewall > the domain allowlist rule group.

Option 3: SSM Parameter Store or AWS Secrets Manager

If your Datadog key or endpoint is sourced from AWS rather than hardcoded, you can use Arpio configuration tags to give the recovery environment a different value.

For an overview, see Arpio Configuration Tags.

SSM Parameter Store

Configure Datadog to read its value from an SSM parameter, include that parameter in your Arpio application, and tag it with arpio-config:recovery-value so recovery uses a different value.

For details, see arpio-config:recovery-value.

AWS Secrets Manager

Tag the source secret with arpio-config:self-managed. Arpio creates the corresponding recovery secret and manages its resource policy, but never sets or overwrites the secret value.

  1. Make sure your source EKS cluster references the secret, and include the source secret in your Arpio application.
  2. On the first capture, Arpio creates the recovery secret with an empty value and maintains it in standby.
  3. In the recovery account, set the recovery value once.
  4. Future tests and recoveries continue to use that value. It persists in standby and is not removed when a test ends.

For details, see arpio-config:self-managed.

Important: This option applies to AWS Secrets Manager secrets only. It does not apply to native Kubernetes secrets.


Option 4: Lifecycle event automation

Arpio lifecycle event notifications can trigger an AWS Lambda function, an SSM document, or a kubectl workflow after restore. You can use this automation to rewrite the Datadog configuration or disable the agent in the recovered cluster.

For details, see Arpio Lifecycle Event Notifications.

This option is best when the Datadog API key and endpoint are hardcoded in a Helm chart and moving them into SSM Parameter Store or Secrets Manager isn't practical.

Important: There is a short window between when the restore completes and when your automation runs. During that time, the Datadog Agent may send telemetry to your production Datadog organization using production credentials.

If that window isn't acceptable and you don't need observability during the test, combine this option with the Network Sandbox.


Separating recovery telemetry in Datadog

Whichever option you choose, you can also make recovery telemetry easier to identify on the Datadog side. These are Datadog settings rather than Arpio controls.

  • Instance identity: Recovered instances have different identifiers from production, which can help separate host-level data.
  • Tagging: Define a clear production-versus-recovery tagging scheme (for example, with DD_ENV or custom tags) so dashboards, monitors, and queries can be scoped appropriately. You can set recovery-specific tag values with the EKS translation ConfigMap.

Arpio does not control how Datadog ingests or filters data. If you want to filter or drop recovery telemetry before it reaches your production Datadog organization (for example, by source account, VPC, or network), check with Datadog for the options available on your plan.


Related articles