Skip to content
Raw Markdown

Evals

Read an eval run's status and scored per-case results. Run ids come from issue occurrences sourced from failing evals, or from runs started in CI or the web app.

Parameters: project_id, run_id.

FieldNotes
status"running", "completed", "failed", or "cancelled"
labelThe run’s label on the Evals page. CI runs are labeled with the branch name
total_cases, completed_casesProgress counters
summaryPassed and failed counts plus accuracy. Null until the run finishes
results[{question, status, duration_seconds, error, sql?}]
errorSet instead of the keys above when the request fails

Results are written as the run progresses, so partial results appear before the status turns terminal. A case’s generated sql is included only for cases that did not pass, since that is what needs diagnosing.

What it cannot do

There is no tool to start a run, add a case, or delete one. Runs start from CI or a checkout with cassis eval run, or from the Evals page in the app. An agent that has just fixed an issue therefore runs the CLI in its checkout, and uses this tool to read a run someone else started.

An eval run scores an ontology. The ontology an agent is changing lives in its checkout, so the CLI starts that run.