Compare two model outputs against a reference rubric using coverage, structure, concision, uncertainty, and lexical-overlap signals.
Scores are deterministic review signals, not a substitute for representative human or model-graded evals.
Compare two AI responses using reference coverage, structure, concision, uncertainty, token estimates, and lexical-overlap signals.