Evals
Read an eval run's status and scored per-case results. Run ids come from issue occurrences sourced from failing evals, or from runs started in CI or the web app.
Parameters: project_id (optional with an API key scoped to one project, which it defaults to), run_id.
| Field | Notes |
|---|---|
status | "running", "completed", "failed", or "cancelled" |
label | The run’s label on the Evals page. CI runs are labeled with the branch name |
total_cases, completed_cases | Progress counters |
summary | Passed and failed counts plus accuracy. Null until the run finishes |
results | [{question, status, duration_seconds, error, sql?}] |
error | Set instead of the keys above when the request fails |
Results are written together when the run finishes. Poll status while it is running; do not wait for partial results.
Each non-passing case includes the generated sql, gold_sql, expected_row_count, actual_row_count, missing_concepts, judge_verdict, and judge_reasoning when available. Result row values and the process log are not included. The question and gold SQL come from what the run scored; older results fall back to the current test case.
What it cannot do
There is no tool to start a run, add a case, or delete one. Runs start from CI or a checkout with cassis eval run, or from the Evals page in the app. An agent that has just fixed an issue therefore runs the CLI in its checkout, and uses this tool to read a run someone else started.
An eval run scores a version of the context. The context an agent is changing lives in its checkout, so the CLI starts that run.