Built-in evaluation metrics
Choose the behavior you need to check, then open the metric’s page. Each page explains when to use it, why, how to configure it, and what its result means, with its own TypeScript example.
All 22 metrics are exported from @anvia/core/evals and run inside runEvalSuite().
Where to start
- A fixed value or phrase: start with exactMatch or contains.
- An answer based on documents: start with faithfulness, then add answerRelevancy.
- A JSON contract: use jsonCorrectness.
- Your own quality rule: use llmJudge for pass/fail or llmScore for a graded score.
- A conversation: use turnRelevancy and knowledgeRetention.
Start with one clear requirement. Add another metric when it catches a different failure that matters to your users.
Read Tips for how to choose a focused set of metrics for your cases.
Text and structure
| Metric | What it checks |
|---|---|
| exactMatch | Checks whether the output equals the expected value. |
| contains | Checks whether the answer includes a required phrase or pattern. |
| notContains | Checks whether the answer avoids a forbidden phrase or pattern. |
| containsAll | Checks whether every required phrase or pattern appears in the answer. |
| containsAny | Checks whether at least one accepted phrase or pattern appears in the answer. |
| matches | Checks whether output text matches a regular expression. |
| doesNotMatch | Checks whether output text avoids a forbidden regular expression. |
| maxLength | Checks whether output text stays within a character limit. |
| requiredFields | Checks whether an object contains all required top-level fields. |
| jsonCorrectness | Checks whether output text is valid JSON and satisfies a Zod schema. |
Meaning and custom judgments
| Metric | What it checks |
|---|---|
| semanticSimilarity | Compares the meaning of an answer with a reference answer using embeddings. |
| llmJudge | Uses a model to produce a structured judgment and applies your pass/fail rule. |
| llmScore | Uses a model to score an answer against your criteria and return feedback. |
| gEval | Uses explicit evaluation steps or criteria to score a custom quality requirement. |
Answer quality and grounding
| Metric | What it checks |
|---|---|
| answerRelevancy | Checks whether an answer stays relevant to the user’s question. |
| promptAlignment | Checks whether an answer follows the instructions you specify. |
| hallucination | Measures whether an answer contradicts trusted context passages. |
| faithfulness | Checks whether factual claims in the answer are supported by retrieved evidence. |
| abstention | Checks whether the assistant appropriately answers or declines to answer. |
| summarization | Checks whether a summary preserves important source facts and stays grounded. |
Conversations
| Metric | What it checks |
|---|---|
| turnRelevancy | Checks whether assistant replies stay relevant across a conversation. |
| knowledgeRetention | Checks whether the assistant preserves information the user supplied earlier. |
Continue with What to evaluate to design cases or Run evaluations to execute a suite.