Part IV · Architecture
Two branches. One never
touches the model.
A classifier reads the question first. The rule it leads with is mechanical on purpose, because the judgement call got it wrong.
Why the routing rule is mechanical
An earlier version decided by judgement, fired on the word strike alone, and sent an appeals question to the eligibility calculator. It answered fluently, in house format, carrying clause IDs, and completely wrong.Nothing about that output looked like an error. It is the failure shape this project spent the most time on. The rule now leads with something that cannot be argued with: a question with no digits in it is a policy lookup, always. And the code node refuses to compute from a question that carries no numbers, rather than returning a verdict built out of zeros.
Two layers, because a prompt is probabilistic and the second layer is not.
The threshold cannot do the refusing
The obvious design is to reject anything that retrieves below some similarity score. We measured whether that works. It does not, and two numbers say why.
0.682
the lowest score a clause we genuinely needed scored
0.837
the highest score an out-of-corpus question scored
There is no threshold in between. The out-of-corpus question — how many days do I have to file an appeal against a strike? — scored higher against the appeals clause than any correct retrieval in the set, and the corpus states no appeal deadline at all.Measured with real embeddings, clause chunking and top-K 4, modelling the deployed configuration. The clause in question is SA-4.
An embedding measures whether a chunk is about the same topic, not whether it contains the answer. A question about appeal deadlines is maximally on-topic for the appeals clause and is still unanswerable from it. Retrieval cannot tell those two apart. Only the grounding instruction can — so the threshold is switched off deliberately, and the refusal lives in the prompt.