Test set
Representative questions, user groups, documents, edge cases, adversarial cases, and expected evidence.
Evidence:Question taxonomy, source mapping, inclusion rationaleEvaluation framework
A useful knowledge assistant must retrieve the right evidence, respect permissions, support its claims, abstain when needed, perform reliably, and fit a human workflow.
Evaluation matrix
Representative questions, user groups, documents, edge cases, adversarial cases, and expected evidence.
Evidence:Question taxonomy, source mapping, inclusion rationaleWhether relevant, permitted, sufficiently complete evidence appears in the retrieved context.
Evidence:Recall/precision variants, rank review, failure examplesWhether important answer claims are supported by the supplied sources and citations point to the right material.
Evidence:Claim-source checks, citation correctness, unsupported-claim rateWhether the system recognizes missing evidence, ambiguity, conflicts, and questions outside the approved scope.
Evidence:Refusal/clarification cases, false-confidence reviewLatency, cost, availability, logging, source freshness, change control, monitoring, and incident response.
Evidence:Service measures, runbook, alerts, ownershipWhether users understand citations, limitations, review duties, escalation, and the decisions they still own.
Evidence:Task testing, adoption feedback, override and escalation recordsRelease decision
A high overall score can hide permission leakage, unsupported consequential claims, or failure for a key user group. Release gates should include non-negotiable conditions as well as aggregate measures.
Minimum evaluation record
This framework is an educational working method, not a certification, legal opinion, security assessment, or guarantee that a system is safe or compliant.