The path
Stage by stage
1. Detection
When a run executes, the Failure Agent detects signals automatically. A failing tool or an answer that conflicts with its sources creates a signal in the queue (see Sources). Detection itself is controlled by the Failure detection toggle in Settings → Loop.2. Confirm the failure
Review the signal detail and press the [Confirm failure] button or [Not a failure] button (see Confirming failures). Not-a-failure ends here: false alarms go no further.3. Cases accumulate in the cluster
Confirmed failures are grouped into a cluster by cause, and the input that produced each signal becomes a case (a wrong-output example). The more expected answers and reasons you fill in, the better experiments can score candidates (see Cases).4. Experiment
Start from Run simulation on a cluster’s detail page, or New experiment on the Simulation screen. A new experiment has three steps.- Pick a cluster which cluster should we fix?
- Experiment setup choose how many fix candidates (2/3/4/6), the Budget (USD), and options: Cost optimization, consistency measurement, and Hold-out regression (Off / Light / Full). You can also steer it in natural language, e.g. “Keep the prompt as-is and try improving it by tweaking guardrails, memory, or retrieval instead.” Reference folders lets candidates be scored against a golden set you curated.
- Run check the summary and start the run.
5. Review candidates
When the experiment finishes, the candidates come back ranked. Each candidate shows:- Results the gain, broken cases, cases tested, and a score
- What actually changes which surface (prompt, guardrail, memory, retrieval, …) changes and how, with before → after
- Per-case results each case marked improved / regressed / kept / unresolved
- Deep verification accuracy & safety, cost, reliability, regression check
6. Deploy
Deploy the candidate you like. Only owners and managers can deploy. Approving an improvement ships a numbered version to the Versions → Improvements tab, with its diff, author, eval score, and one-click rollback (see Versions).7. Watch
After deploying, watch whether signals matching the same pattern come back. The Revalidation tab tells you whether a banned capability has come back to life.Next steps
Simulation
Experiment on your cases before anything reaches production.
Versions
Diffs and rollback for deployed improvements.