Evaluation

Quality measurement for AI-generated analyses

JRSigned in as Jordan Reyes
What this page will measure
Once analyses run against a live model, this page will score every output against a labelled review set: whether the themes match what a human reviewer would identify, whether each quote actually appears in the submitted feedback, whether the response conforms to the expected schema, and whether the product manager who ran it found it useful. The numbers below are placeholder values from a sample review set.

Theme accuracy

88%

Target 90% · below target

Share of generated themes a human reviewer accepts as a real, distinct pattern.

Evidence grounding

94%

Target 95% · below target

Share of quotes that appear verbatim in the submitted feedback.

Output validity

99%

Target 99% · meeting target

Responses that parse against the expected schema on the first attempt.

Helpfulness

92%

Target 85% · meeting target

Analyses rated helpful by the product manager who ran them.

Evaluation cases
CaseTheme accuracyValidity
Enterprise NPS — 120 verbatims91%Pass
Support tickets — mixed languages78%Pass
Short-form app reviews84%Pass
Sales call notes — unstructured69%Fail
Single long transcript90%Pass