Skip to main content
An answer without evidence is just an assertion. Verification on a trace splits the answer into individual claims, checks each one against the evidence inside the run, and records a verdict per claim.

Running verification

In the trace detail panel, press the [Verify] button in the Verification section. Even on a trace you have not verified yet, the Claim Decomposition · Evidence Retrieval · Verification stages are already listed; click one and it tells you no artifact has been captured for that stage yet. Once you press the button, a running badge shows while it works, then flips to the result badge (for example, Not supported). Each stage fills in its duration on the right.

Verification stages

  1. Claim Decomposition The answer is split into individual claims. Each claim carries the quoted passage it came from plus category and importance tags (for example, Factual, High).
  2. Evidence Retrieval Evidence for each claim is gathered from the run. Click the expand icon on the left to reveal the Sentence splitting · Extract entities · Match evidence substeps.
  3. Verification Each claim gets a verdict against the evidence, with step-by-step reasoning.

Opening a stage artifact

Click a stage or substep to see what it actually produced on the right. Pick a substep and its parent stage name appears above the title. The Formatted / JSON toggle at the top right switches between a readable rendering and the raw JSON. Each claim card shows a character count on the right and can be collapsed.

Reading verdicts

  • Every claim gets a verdict badge. A claim the evidence cannot back is Not supported.
  • reasoning walks through which evidence was considered and how the verdict was reached. The last line carries the verdict label.
  • The gauge under the Verification section shows the share of supported claims for the whole trace. The same gauge appears in the Verdict column of the list.
  • The verdict filter in the list (Support · Partial · Not support · Unverified) narrows traces by verdict level. Traces you have not verified yet are Unverified.
A low support share is the strongest signal of subtle hallucination: the Agent did read the sources, but drew a stronger conclusion than they back. When entity_overlap or tfidf_score is low and the verdict is Not supported, the answer most likely invented a value that was never in the run at all. High overlap with a Not supported verdict points the other way: the same material was on the table and only the conclusion diverged.

Verdicts become signals

Claims that fail verification are automatically registered as evidence-check signals. From the signal detail you can jump straight to the trace and to the cluster it belongs to. See the signals overview. To run verification automatically on every run by placing a verification block inside the Flow, see Answer verification.

The same check scores your fixes

Verification is not only a detection step. The criteria in Settings → Review & Scores and the same claim-level verdicts are what Simulation uses to score a candidate fix: the Accuracy & safety and Holdout deep tests re-run this check over past cases. A candidate wins because it raises the share of supported claims, not because it merely looks better. See Verify and publish.