Part VI · Proof
Four suites. One result
is a refusal.
Scores are the easy part. The result worth pointing at is the question it declined to answer.
Grounding behaviour
The prompt refuses what the corpus does not state, and the refusal is exact. Runs locally, with no network and no key.
Retrieval, end to end
Real embeddings, clause chunking, top-K 4 — including the question whose answer spans two chunks, and four out-of-corpus probes.
The eligibility tool
Every branch of the verdict, including the refusal to compute from a question that carries no numbers at all.
Clause coverage
Every clause is reached by at least one question. An uncovered clause fails the build rather than going quietly untested.
Why a refusal is the stronger result
A test suite proves the system does what we told it to do. The refusal in Part I proves something harder: that it declines a question it could have answered correctly from training data, because the answer was not in the documents the business supplied.The same probe is in the demo set in Part II. It is the one question in this project that must come back empty.
A system that is right for the wrong reason is not safe to put in front of creators. That is the thing this was built against.
What we would do next
Version the corpus, so an answer can state what the rule was on a given date rather than only what it is now. Log every refusal — each one is a gap in the policy documentation, and that log is a product in itself. Escalate to a human with the retrieved clauses already attached, so the agent starts where the assistant stopped.