Skip to main content
Simulation is where you pick a cluster and experiment with fix candidates. Instead of testing on production traffic first, candidates are scored against the cases accumulated in the cluster, so you see how much better they are, and what they break, in advance.

The board

One experiment is one card. Cards live in one of four columns.
  • Experiment currently running; the card shows a measuring state
  • Optimization experiments with a confirmed improvement
  • Regression experiments with regressions left, or no confirmed improvement
  • Archive experiments that have been put away
A card shows the change-type tag (prompt_replace, combo, and so on), the cluster name, the gain (%), the tested-case count and spend (like n=5 · $1.34), and the candidate count. When an experiment finishes, its card moves to the column the results call for.

Views and filters

  • The toggle at the top right switches between Kanban view and List view. List columns are Started, Status, Cluster, Best gain, Broke, Candidates, and Spend.
  • Search evaluations finds cards; Any date narrows the range.

The flow

  1. Pick a cluster and set the options in New experiment.
  2. Read the candidates, gains, and reasoning trace in Experiment results.
  3. Deep-test a candidate and publish it in Verify and publish.

Learn more

New experiment

Select cluster, configure options, launch. Three steps.

Experiment results

Candidates, gains, spend, and the reasoning trace.

Verify and publish

The four deep tests, then publish.

Guarding against regressions

The mechanisms that keep fixes from breaking things.