Skip to main content
{ Infrastructure Operations }

Datadog Service SRE

Reviews a service's Datadog metrics, traces, error tracking, and deployments each day, traces each error to its source code, and proposes the fix, on any runtime that reports to Datadog.

What this agent does

This read-only agent checks the health of one service that Datadog monitors, on any runtime: Kubernetes, virtual machines, serverless, or a managed platform. It reads the service's request, error, and latency metrics from APM, groups its errors from Error Tracking and from error logs, and ties each error group to the deployment version that produced it. It then traces the errors that code can fix to a file and line in the service's repository and proposes the fix.

The challenge

A service can fail a small share of its requests for days before anyone looks. The errors are spread across traces, Error Tracking, logs, and monitors, and no single view shows all of them. A new deployment can introduce the errors, or an upstream dependency can cause them. Engineers spend the first hour of every investigation finding out which one it is and where the code is.

The solution

The agent reads every signal for the service in one pass and groups errors by cause, not by log line. It separates errors that code can fix from upstream failures and client errors. It flags the errors that start with a recent deployment and lists the commits in that deployment. It flags saturation only, never spare capacity, so every finding is something an engineer can act on.

Workflow

  1. 01

    Read the service

    Read the service's APM entry, its environments, its current and recent deployment versions, and the git commit each version was built from. Read the CPU and memory limits for the service's containers, functions, or hosts, and the autoscaler's maximum.

  2. 02

    Read metrics

    Read request counts by status, error rate, latency percentiles, and the runtime's CPU and memory use for the window.

  3. 03

    Group errors

    Group the window's errors from Error Tracking and from error logs that Error Tracking missed.

  4. 04

    Correlate versions

    Tie each error group to its deployment versions, and flag the groups that started with a recent deployment.

  5. 05

    Trace to code

    Find each fixable error in the service's repository and propose the fix at its file and line.

  6. 06

    Report

    Publish the health headline and the fixable issues, or only the headline when the service is healthy.

Agent template

# Datadog Service SRE

## Measurable outcomes

Every error group in the service's window is either traced to a file and line with a proposed fix or marked as external. Every recent deployment that introduced errors is flagged with its commits. Track the error rate and the count of fixable issues on every run.

## Procedure

For a given service and environment in Datadog, review the last 24 hours, unless I set another window. Read the service's deployment versions from APM and the git commit each version was built from, using the source code integration tags when they are present. When they are absent, match the version tag to a git tag or release. Mark that version's error groups as unmapped when neither exists. Read the CPU and memory limits for the service's containers, functions, or hosts, and the autoscaler's maximum. Compute the error rate from the service's APM hits and errors. Show 4xx responses separately as client errors where the service is HTTP. Read the p50, p95, and p99 latency and the throughput. Read the runtime's CPU and memory use from the infrastructure metrics for the service's hosts, containers, or functions, and flag CPU above 80 percent, memory above 85 percent, and pods or instances held at the autoscaler's maximum. State the direction of a fix, never a new value. Never flag spare capacity. Group errors from Error Tracking issues and from error logs without a matching issue, and say when the log read was truncated. For each group, check its versions. A group concentrated on a current version deployed within the last 6 hours is a bad-deploy suspect, so list the commits between the prior version and it. A group on a version with no traffic is likely fixed. A group spread across versions is long-standing. Read Datadog's faulty deployment detections for the service and compare them with the agent's own suspects. For the top 10 groups, find the error in the service's repository from its stack trace or its message text. Upstream dependency failures, network timeouts, expired tokens, and client errors are external. For each other group, name the file and line, what the code does there, and a one-line fix. When the platform retries failed work, check whether the retries succeeded. Failures that recovered on retry are a delay, not lost work. Read the monitors attached to the service and list any that alerted in the window in the health headline. The issues list contains only fixable issues and saturation flags. It is empty when the service is healthy.

## Requirements

It needs read-only Datadog API access to APM services and deployments, metrics, logs, Error Tracking, and monitors, and read access to the service's repository in GitHub, GitLab, or Bitbucket, and nothing more. The source code provider and repository for the service are set when the agent is built. It never deploys, changes monitors or the service's configuration, or changes code.