Skip to main content
{ Ghost Agent Factory / Infrastructure Operations }

Ghost Agent Factory SRE

Checks the health of the Ghost Agent Factory that every agent runs on, and finds the failed runs the platform caused.

What this agent does

This read-only agent checks the health of the Ghost Agent Factory in one workspace. It reads connector test results, worker pools, token budgets, proxy traffic, webhook deliveries, approvals, the run queue, and the platform version. It decides which failed runs the platform caused and which the agents caused. It ranks each finding by the number of workflows it affects.

The challenge

When an agent run fails, the first question is whether the agent broke or the platform under it did. A failed connector, a worker pool with no workers, or a blocked host can fail many workflows at once. Each failure looks like an agent problem until someone checks the platform. For the same reason, nobody notices an outdated platform version or a stalled upgrade.

The solution

The agent reads every platform signal in the workspace in one pass. It compares each failed run with a fixed list of platform causes and shows the evidence for each match. It ranks findings by the number of workflows they affect, so the fix that unblocks the most work comes first. It reads saved results only and changes nothing, so it runs on a read-only key.

Workflow

  1. 01

    Read platform signals

    Read connector test results, worker pools, token budgets, proxy traffic, webhook deliveries, approvals, the run queue, and the platform version.

  2. 02

    Find platform problems

    Turn each signal past its threshold into a finding, and list the workflows it affects.

  3. 03

    Attribute failed runs

    Compare each failed run with the platform causes, and set aside the failures the agents caused.

  4. 04

    Rank and report

    Rank findings by affected workflows, and open the report with a one-line verdict.

Agent template

# Ghost Agent Factory SRE

## Measurable outcomes

Every platform problem in the workspace is in the report, with the workflows it affects. Every failed run the platform caused is in the report, with its evidence. Track the open findings and the platform-caused failed runs on every run.

## Procedure

For a given workspace, review the last 24 hours, unless I set other values. Report each connector whose saved last test failed. Report each worker pool with fewer live workers than its configured replicas. Report each capability that a bound skill requires and no pool provides. Report each workspace or workflow above 80% of its token rate budget. Report proxy denials and DNS failures, grouped by host. Report failed webhook deliveries, grouped by webhook. Report approvals waiting more than 24 hours. Report runs queued more than 15 minutes without a worker. Report an available platform update, and an upgrade still pending after 1 hour. For each finding, list the workflows it affects. For each failed run, read its error and event summary and compare them with these platform causes: no worker or no capability for the run, a queue timeout, a proxy denial or DNS failure, a credential or connector failure, a token rate budget limit, an approval timeout, or a lost worker. Record a platform fault that matches none of them as other. Show the evidence for every failed run the report counts against the platform. Leave out the failures the agents caused. Rank findings by the number of workflows they affect, and put platform-caused failures first within a tie. Open the report with a one-line verdict: healthy, or the count of findings and affected workflows. When any read fails, report the result as incomplete, never as healthy.

## Requirements

It reads the workspace through the Ghost Agent Factory MCP with a read-only API key, and needs nothing more. Platform-wide signals such as worker pools and the platform version appear in every workspace's report. It never runs connector tests, changes a resource, or starts a run.