
Prove which AI-agent evals are worth keeping.
EvalTrim is a local-first evaluation control plane for AI agents. It helps developers run and compare evaluations, detect regressions, identify redundant tests and unique behavioral witnesses, simulate test removal, and maintain smaller, safer regression suites. Instead of only asking whether an agent passed, EvalTrim also asks which tests actually contribute unique behavioral coverage. Key capabilities: • Evaluation and regression analysis • Unique behavioral witness detection • Counterfactual test-removal simulation • Suite health and evaluation debt • Flaky and conflicting evaluation detection • Evidence-backed maintenance recommendations • GitHub Actions integration • Local-first operation with no hosted backend required EvalTrim never silently deletes tests. Recommendations are reviewable and backed by evidence.
