EH
Eval Healthcheckbeta
Pre-release health check

Check whether your AI eval signal is trustworthy before release.

Upload an eval dataset, run health checks on your judge and dataset, and get a release-readiness report your team can act on.

Built for AI engineers running evals before release.

Sample report verdict
WARN82 / 100 ยท Mostly Stable
Review needed before using this eval as a release gate.
6 of 30 samples produced unstable verdicts across 5 runs and the dataset has 3 rows with missing expected values.
What gets checked
  • Verdict stability across N runsLLM-assisted
  • Required dataset columnsDeterministic
  • Missing or empty expected valuesDeterministic
  • Duplicate rowsDeterministic
  • Basic length outliersDeterministic

Start with judge and dataset health

3 live
โš—
AI-assisted
01Judge Reliability
Is your LLM judge consistent enough to trust? We re-run it multiple times per sample and surface unstable verdicts.
๐Ÿ—„
Deterministic
02Dataset Health
Is your eval dataset structurally usable and obviously risky? Deterministic checks for missing fields, duplicates, and shape.
๐Ÿ“ˆ
Deterministic
03Score Health
Is score movement meaningful or likely noise? Compare eval score changes against judge reliability.