Skip to main content
Signals are the entrance and a deployed version is the exit. Here’s what happens in between.

The path

Every stage has a checkpoint where a human decides.

Stage by stage

1. Detection

When a run executes, the Failure Agent detects signals automatically. A failing tool or an answer that conflicts with its sources creates a signal in the queue (see Sources). Detection itself is controlled by the Failure detection toggle in Settings → Loop.

2. Confirm the failure

Review the signal detail and press the [Confirm failure] button or [Not a failure] button (see Confirming failures). Not-a-failure ends here: false alarms go no further.

3. Cases accumulate in the cluster

Confirmed failures are grouped into a cluster by cause, and the input that produced each signal becomes a case (a wrong-output example). The more expected answers and reasons you fill in, the better experiments can score candidates (see Cases).

4. Experiment

Start from Run simulation on a cluster’s detail page, or New experiment on the Simulation screen. A new experiment has three steps.
  1. Pick a cluster which cluster should we fix?
  2. Experiment setup choose how many fix candidates (2/3/4/6), the Budget (USD), and options: Cost optimization, consistency measurement, and Hold-out regression (Off / Light / Full). You can also steer it in natural language, e.g. “Keep the prompt as-is and try improving it by tweaking guardrails, memory, or retrieval instead.” Reference folders lets candidates be scored against a golden set you curated.
  3. Run check the summary and start the run.
The experiment runs in the background and shows up as a Measuring card in the experiment column of the Simulation board.

5. Review candidates

When the experiment finishes, the candidates come back ranked. Each candidate shows:
  • Results the gain, broken cases, cases tested, and a score
  • What actually changes which surface (prompt, guardrail, memory, retrieval, …) changes and how, with before → after
  • Per-case results each case marked improved / regressed / kept / unresolved
  • Deep verification accuracy & safety, cost, reliability, regression check
The best candidate is marked. See Simulation.

6. Deploy

Deploy the candidate you like. Only owners and managers can deploy. Approving an improvement ships a numbered version to the Versions → Improvements tab, with its diff, author, eval score, and one-click rollback (see Versions).

7. Watch

After deploying, watch whether signals matching the same pattern come back. The Revalidation tab tells you whether a banned capability has come back to life.

Next steps

Simulation

Experiment on your cases before anything reaches production.

Versions

Diffs and rollback for deployed improvements.