What this page will measure
Once analyses run against a live model, this page will score every output against a labelled review set: whether the themes match what a human reviewer would identify, whether each quote actually appears in the submitted feedback, whether the response conforms to the expected schema, and whether the product manager who ran it found it useful. The numbers below are placeholder values from a sample review set.
Theme accuracy
88%
Target 90% · below target
Share of generated themes a human reviewer accepts as a real, distinct pattern.
Evidence grounding
94%
Target 95% · below target
Share of quotes that appear verbatim in the submitted feedback.
Output validity
99%
Target 99% · meeting target
Responses that parse against the expected schema on the first attempt.
Helpfulness
92%
Target 85% · meeting target
Analyses rated helpful by the product manager who ran them.
Evaluation cases
| Case | Theme accuracy | Validity |
|---|---|---|
| Enterprise NPS — 120 verbatims | 91% | Pass |
| Support tickets — mixed languages | 78% | Pass |
| Short-form app reviews | 84% | Pass |
| Sales call notes — unstructured | 69% | Fail |
| Single long transcript | 90% | Pass |