UNDERSTAND THE RUN
Follow the evidence.
Find the steps worth a closer look.
Review a completed agent run as a readable timeline. Structural checks highlight observations, with evidence and room for uncertainty.
Open a completed run
Trace Check JSON v1 · up to 2 MiB · 2,000 steps
Start with an example
One-minute walkthrough: open the illustrative suggestion, inspect a rule observation, enable local review, follow an action back to its timeline, then export the separate review notes. Deliberately selected handwritten examples demonstrate behavior, not unseen accuracy.
Logs stay in this browser tab. Nothing is uploaded or saved by this app. Likely tokens, passwords, and email addresses are redacted before display and export. Detection is incomplete: review sensitive logs before importing.
Supported format and review limits
Download sample, open it in a text editor, and replace its task and steps with your completed run. Provider-specific logs and JSONL need conversion to this format before import. Tool arguments and results belong in content as strings; preserve their order and use the same call_id for a call and its result. Multiple calls may have results in a different order.
Use schema_version 1, a run_id, task, status "completed", and a steps array. Each step has a unique id, kind, and content. Supported kinds: task, assistant, tool_call, tool_result. Optional fields: tool, call_id, status (ok or error). Link calls and results using call_id. Unknown fields are rejected.
Only linked calls can be checked for missing results. Flags describe structural observations, not proven mistakes. Expected failures and neutral exploration may be flagged. Rules run by default. Optional experimental ML review requires explicit original message grouping and per-run opt-in; no probability of error is claimed.
COMPLETED RUN
Report v1 includes the full redacted run and all observations, including steps hidden by filters. Check it before sharing. To review it again, save its run object as a separate JSON file.
OPTIONAL LOCAL EXPERIMENT
Experimental ML review suggestions
A frozen TF-IDF classifier can surface actions for manual review. It never changes rule flags, executes tools, suppresses findings, or certifies a run as correct.
Off for this run.
Model, coverage, and evaluation limits
20,000 unigram/bigram features → frozen three-class logistic head → fixed 0.70 review cutoff. The action uses its original message plus two preceding message slots. Fields are limited to 4,000 characters, then the joined context to its last 12,000 characters. Non-ASCII features and unknown vocabulary abstain. The final assistant message is excluded.
This is a deterministic reconstruction of an earlier experiment, not fresh held-out evidence. On previously examined development validation, 878 of 1,390 actions were supported; 50 suggestions contained 47 labeled mistakes and 3 false positives. Compared with rules, it added 21 true positives and 3 false positives, failing the zero-additional-false-positive gate. TEST remains excluded. A later semantic model also failed that gate and is not enabled.
Local worker · 15-second model load/verification limit, then 2-second inference limit · 2 MiB input/model limits · no uploads. Scores are uncalibrated. Vocabulary contains source-derived terms and is not anonymous. Model source, hash, and license notice. Example provenance and selection. Third-party training-source notices.