Evaluation & validation

How RafiHive answers are tested

Before a release, every answer is run against a golden set of regulatory questions with known expected sources. The numbers below come straight from the last harness run on 2026-09-07; nothing here is typed in by hand.

Retrieval hit rate100%
Answer pass rate75%
Fabricated citations0
Superseded or unapproved sources retrieved0

What the golden set covers

16 questions across variations, dossier structure, labelling, ICH quality and blue box topics. Retrieval is measured with top-8 results; answers are generated by eu.anthropic.claude-sonnet-4-6.

What counts as a pass

An answer passes only if every expected source is cited, no citation is invented, the required concepts are present and no forbidden claim appears. A single missing citation fails the whole case, which is why the answer pass rate sits below the retrieval hit rate.

Named RA sign-off

0 of 16 cases carry named RA sign-off; the rest are internal drafts. The release gate (npm run eval:kb:release) refuses to pass until every case has a named approver. Harness and cases live in the repository under scripts/kb.

Limits

A golden set is a regression guard, not proof of correctness on your question. Every RafiHive output still requires qualified RA review before it is used. See AI governance for the human-review boundary and sample output for what a reviewed answer looks like.