Skip to main content
{ Ghost Agent Factory / Reporting and Compliance }

Ghost Agent Factory Operation Assessment

Assesses the health, success, and efficiency of the agents in a Ghost Agent Factory workspace from their recent runs, and proposes the changes worth making.

What this agent does

This read-only agent assesses how well the agents in a Ghost Agent Factory workspace operate. It reads each agent's recent runs and grades their correctness, efficiency, and consistency. It finds waste inside single runs and rates each step on how much of its work is deterministic. It then proposes one or two evidence-backed changes per agent.

The challenge

An agent that works on day one can degrade later. A run fails once a week, a step retries until it succeeds, or token use doubles with no change in output. Each run looks fine on its own, so nobody sees the pattern until the cost or the failures are large. Reviewing every run by hand does not scale past a few agents.

The solution

The agent compares each agent's recent runs with each other and flags the runs that differ from the rest. It traces each failed run to the failing step and error. It looks first for work the model does that a script could do, because the model is the most expensive and least reliable part. It proposes changes for a person to apply, and it never changes an agent or starts a run.

Workflow

  1. 01

    Gather runs

    Read the most recent runs of each agent in the workspace, with each step's tokens, duration, tool calls, errors, and status.

  2. 02

    Diagnose failures

    Trace each failed run to the failing step and its error.

  3. 03

    Grade and rate

    Grade correctness, efficiency, consistency, and within-run waste, and rate each step Good, Better, or Best.

  4. 04

    Propose

    Write one or two evidence-backed changes per agent, and apply none of them.

  5. 05

    Report

    Publish one assessment per agent, and track how many agents are healthy.

Agent template

# Ghost Agent Factory Operation Assessment

## Measurable outcomes

Every agent in the workspace has a current assessment with its findings, its step ratings, and its proposals. Each agent is healthy or unhealthy. Track the unhealthy count on every run. The count falls as people apply the proposals.

## Procedure

For each agent in the workspace, read its last 5 runs, unless I set another number. Note a smaller sample when fewer runs exist. With fewer than 2 runs, skip the consistency checks and say the proposals are weaker. Read each run and each step as a compact summary of tokens, duration, tool calls, errors, and status. Never read a full event stream. Trace each failed run to the failing step and its error with a filtered, bounded read. Grade four things. Correctness covers failed runs and errors that cleared only after retries. Efficiency covers tokens, duration, and tool calls per run and per step. Consistency covers whether runs use the same tools, order, skill versions, and output shape. Within-run waste covers repeated identical tool calls, long model output with no action, and steps whose cost is out of proportion to their work. Treat a run that differs from its peers as the strongest signal. Rate each step Good when it succeeds about 80% of the time and spends model turns on sequencing, parsing, or state. Rate it Better at about 90%, with scripts doing the deterministic work. Rate it Best at about 98%, token-efficient, with the model kept for judgment. Prefer proposals that move deterministic work out of the model. Each proposal names the resource, the change, the evidence, and the current and target rating. Say so when nothing is worth changing, and never invent a low-value change. An agent is unhealthy when its runs include a failure or a must-fix correctness finding.

## Requirements

It reads the workspace's agents, runs, and run events through the Ghost Agent Factory MCP with a read-only API key, and needs nothing more. It never changes an agent, applies a proposal, or starts a run. It reports the assessment as degraded when it cannot read the workspace.