Skip to content

Built-in evaluation metrics ​

Choose the behavior you need to check, then open the metric’s page. Each page explains when to use it, why, how to configure it, and what its result means, with its own TypeScript example.

All 22 metrics are exported from @anvia/core/evals and run inside runEvalSuite().

Where to start ​

Start with one clear requirement. Add another metric when it catches a different failure that matters to your users.

Read Tips for how to choose a focused set of metrics for your cases.

Text and structure ​

MetricWhat it checks
exactMatchChecks whether the output equals the expected value.
containsChecks whether the answer includes a required phrase or pattern.
notContainsChecks whether the answer avoids a forbidden phrase or pattern.
containsAllChecks whether every required phrase or pattern appears in the answer.
containsAnyChecks whether at least one accepted phrase or pattern appears in the answer.
matchesChecks whether output text matches a regular expression.
doesNotMatchChecks whether output text avoids a forbidden regular expression.
maxLengthChecks whether output text stays within a character limit.
requiredFieldsChecks whether an object contains all required top-level fields.
jsonCorrectnessChecks whether output text is valid JSON and satisfies a Zod schema.

Meaning and custom judgments ​

MetricWhat it checks
semanticSimilarityCompares the meaning of an answer with a reference answer using embeddings.
llmJudgeUses a model to produce a structured judgment and applies your pass/fail rule.
llmScoreUses a model to score an answer against your criteria and return feedback.
gEvalUses explicit evaluation steps or criteria to score a custom quality requirement.

Answer quality and grounding ​

MetricWhat it checks
answerRelevancyChecks whether an answer stays relevant to the user’s question.
promptAlignmentChecks whether an answer follows the instructions you specify.
hallucinationMeasures whether an answer contradicts trusted context passages.
faithfulnessChecks whether factual claims in the answer are supported by retrieved evidence.
abstentionChecks whether the assistant appropriately answers or declines to answer.
summarizationChecks whether a summary preserves important source facts and stays grounded.

Conversations ​

MetricWhat it checks
turnRelevancyChecks whether assistant replies stay relevant across a conversation.
knowledgeRetentionChecks whether the assistant preserves information the user supplied earlier.

Continue with What to evaluate to design cases or Run evaluations to execute a suite.

Built for Anvia.