Skip to content
Raw Markdown

Build and run evals

Validation proves the context can be imported. Evals prove a change still answers important business questions correctly.

Add a trusted case

Start with a real question whose correct behavior matters. Fix the context, then propose the question and its expected SQL together with the business assumptions that SQL encodes.

Have a domain owner or designated reviewer validate both the interpretation and the SQL before adding the case. SQL that runs is not necessarily SQL that expresses the right business rule.

Test Cases12
Runs
4 runs ▶ Run evaluation
Date Status Accuracy Version
Jun 6, 11:45 Completed 92% v14
Jun 5, 15:20 Completed 83% v13
Jun 2, 10:30 Completed 75% v12

How cases are judged

Warehouse-connected
Cassis executes the generated and expected SQL, then compares their result rows deterministically, with tolerance for insignificant numeric differences.
Schema-only
Neither query can run. An LLM equivalence judge compares the SQL logic and explains whether it matches.

Run before merge

In the web app
Run the suite against the published context or an app branch.
From a checkout
cassis eval run tests the local context without publishing it.
In CI
Run the same command on pull requests so regressions block the merge.

Failures return to the Review queue with their evidence attached. Open a result to compare generated and gold SQL, row counts, and the judge’s explanation. Each result preserves the question and gold SQL it scored. Older results that use the current case’s SQL are labeled accordingly.

Remove or update a case when the approved business definition changes. A stale expected query makes the suite noisy and easier to ignore.

Command syntax lives in the eval reference.