← all analysis · · docs/analysis/2026-05-taxonomy-v1.2-clarification.md
Taxonomy v1.2 clarification — what does the dual-readout certify?
On this page
Created: 2026-05-27 Trigger: Option B from
docs/handoff/2026-05-27-end-of-session.md— predict v1.2 outcomes from v1.1auto_reasoningbefore spending another API run. Inputs: v1.1auto_reasoningfields inmechinterp/outputs/dual_readout_v1/w3/*.json; diagnostic appendix indocs/analysis/2026-05-taxonomy-v1.1-delta.md§“Diagnostic appendix”. Status: Clarification proposal — not a third auto-tagger pass. Decides the taxonomy question that v1.1 surfaced. Once decided, either the spec or the canonical tags need to move.
The load-bearing question
The v1.1 diagnostic appendix identified two cases (Q-D6-L4-01, Q-D8-L3-01) where the auto-tagger fires confabulation against the reference answer while the human SME tags nla_only/converge. The diagnostic concluded — correctly — that clause (c) of rule 4 (“contradiction with the reference answer”) is the surface symptom. But it left the underlying taxonomy question unanswered:
When the NLA verbalization matches a model answer that happens to be wrong, what is the dual-readout reporting?
Two coherent positions:
-
Position A (canonical SME): the dual-readout’s job is to detect NLA-vs-rest-of-model-state gaps, not to validate against ground truth. If NLA and model converge on a wrong answer, the readout has done its job — it has reported convergence. Correctness of the model answer is a judge-scoring concern, not an interpretability concern. → tag
convergeif SAE corroborates,nla_onlyif SAE is silent. Confabulation is reserved for NLA disagreeing with the model. -
Position B (v1.0/v1.1 auto-tagger as written): the dual-readout is part of an audit chain whose downstream consumer (FINMA Randziffer mapping) is correctness-sensitive. If NLA reinforces a wrong answer, that is a confidence-downgrade signal worth flagging, even when NLA and model are aligned. → tag
confabulationwhenever NLA contradicts the reference, regardless of model alignment.
These are not reconcilable by prompt tweaks. They reflect different beliefs about what the readout is for.
Why the answer matters
| Choice | What “confabulation” means downstream | FINMA Randziffer mapping |
|---|---|---|
| Position A | NLA channel diverges from model internals → a method-disagreement flag, model output may still be correct | Bound to method confidence only; correctness handled separately by judge scoring |
| Position B | NLA + model align on a wrong answer → a correctness flag piggy-backing on NLA wrong-direction signal | Conflates method confidence with correctness; doubles the load on the readout |
Position A is the cleaner separation of concerns and is what /analysis/2026-05-agreement-patterns-taxonomy/ §“What this taxonomy does not claim” already states verbatim: “A question can have a converge tag and still be wrong… The agreement pattern speaks to method confidence, not to model correctness.” The v1.0/v1.1 rule-4 clause (c) silently contradicts that principle.
Position B is what the auto-tagger has been doing because the prompt allowed it. It is operationally simpler but creates two separate readouts (method-agreement and correctness) inside one tag, making FINMA mapping ambiguous.
Recommended position: A. It is what the taxonomy doc already commits to in plain English; clause (c) was a residue from drafting that the v1.1 refinement preserved by accident.
Option B evidence: v1.2 won’t move the four over-corrections
The v1.1 delta proposed v1.2 as “drop clause (c)” and expected the 4 over-corrections (Q-D1-L1-01, Q-D12-L3-01, Q-D7-L1-01, Q-D7-L3-01) to flip back toward confabulation. Reading the v1.1 auto_reasoning fields directly:
| qid | v1.1 tag | Did v1.1 invoke clause (c)? | v1.1 reasoning citation | v1.2 prediction |
|---|---|---|---|---|
| Q-D1-L1-01 | nla_only | No | ”this is an answer-correctness issue, not a confabulation of NLA against specific prompt tokens” | stays nla_only |
| Q-D12-L3-01 | nla_only | No | ”No direct contradiction with a specific token in the prompt/answer is present — the NLA wanders off-topic rather than contradicting a specific fact” | stays nla_only |
| Q-D7-L1-01 | nla_only | No | ”the NLA at pos 118 and 130 discusses knock-out mechanics without making a provably contradictory specific claim against a context token” | stays nla_only |
| Q-D7-L3-01 | converge | No | ”Both methods independently identify barrier-related structured product concepts, satisfying the convergence criterion” | stays converge |
In all four cases the v1.1 judge had already reached its tag by clause-(a)/(b) reasoning — i.e. by not finding a contradiction against the model answer. Dropping clause (c) cannot move those tags because clause (c) was not load-bearing for them. v1.2 will only flip Q-D6-L4-01 and Q-D8-L3-01 (the two cases where v1.1 did fire on clause (c)).
Predicted v1.2 distribution: confab 3 / nla_only 19 / converge 3. Canonical distribution: confab 7 / nla_only 14 / converge 4.
The real disagreement on the four over-corrections
Position A (the recommended v1.2) and the canonical SME tags disagree on these four. The canonical reviewer tagged them confabulation reasoning that NLA reinforces a wrong model answer at specific positions (e.g. Q-D1-L1-01 notes: “NLA consistently reinforces incorrect ‘Exotic Options’ classification at pos 29, 95, 107, 141-144”). Under Position A, NLA-reinforces-wrong-model is converge-on-error (if SAE corroborates) or nla_only (if SAE is silent), not confabulation. So:
| qid | Canonical | v1.2 (Position A) | Disagreement type |
|---|---|---|---|
| Q-D1-L1-01 | confabulation | nla_only (SAE silent) | NLA reinforces wrong model → Position A says nla_only |
| Q-D12-L3-01 | confabulation | nla_only (SAE silent) | NLA wanders to related-wrong concepts (CDS/CDO for CLN); not aligned with a specific model token |
| Q-D7-L1-01 | confabulation | nla_only (SAE blank) | NLA agrees with wrong model on physical-settlement condition |
| Q-D7-L3-01 | confabulation | converge (SAE corroborates) | NLA + SAE agree on barrier-option concept; model misdefines it |
Three of four (Q-D1-L1, Q-D7-L1, Q-D7-L3) are clean cases of NLA-matches-wrong-model; under Position A they cleanly fall out of confabulation. Q-D12-L3 is murkier — NLA invokes adjacent-but-wrong frameworks (CDS/CDO for a CLN question), neither matching the model answer nor directly contradicting it.
v1.2 prompt text
Replace rule 4 in mechinterp/scripts/08_auto_tag_w3.py TAXONOMY with:
4. confabulation — NLA names a SPECIFIC fact in DIRECT contradiction with the model's
own state
Rule: NLA at any answer-relevant position contains a SPECIFIC factual claim
(named entity, direction, magnitude, or product mechanic) that is in DIRECT
CONTRADICTION with an identifiable token in EITHER:
(a) the prompt text at or before that position, OR
(b) the model's own final answer.
The reference answer is NOT a contradiction source for this rule. If the
model's answer is wrong and the NLA agrees with the model's wrong answer,
that is `converge` (on the wrong concept) if SAE corroborates, or
`nla_only` if SAE is silent. The dual-readout discriminates whether NLA
verbalizes what the model's internal state says — it does not certify model
correctness.
Examples:
- NLA at pos 12 says "above" but prompt token at pos 11 is " below" →
confabulation (clause a, within-prompt contradiction).
- NLA names "covered call" when model's own answer says "barrier reverse
convertible" → confabulation (clause b, NLA-vs-model contradiction).
- Model answer says "underlying below barrier triggers autocall" (wrong),
NLA reinforces "below" at pos 93, SAE silent → nla_only (NLA matches
wrong model, SAE silent — Position A: no certification of correctness).
- Model answer says "barrier breach → zero value" (wrong), NLA invokes
knock-out mechanics, SAE peaks on `barrier` → converge (NLA + SAE
agree on barrier concept, model misdefines its payoff — agreement
on wrong concept).
Also update the corresponding example block in rule 2 (nla_only) to mention the NLA-matches-wrong-model-with-silent-SAE subcase explicitly, so the judge does not get pulled back into confabulation via the priority order.
Operational implications
Adopting Position A requires one of two moves, not both:
-
Update the spec, leave canonical alone. Lock the taxonomy doc with the v1.2 language. Treat canonical W3 tags as “v1.0 canonical, drafted before the load-bearing question was resolved.” Annotate the four Position-A-disagreement qids (Q-D1-L1-01, Q-D7-L1-01, Q-D7-L3-01, Q-D12-L3-01) in §4.1 of
2026-05-dual-readout-empirical.mdwith a footnote: “canonical tag predates v1.2 clarification; under v1.2 these would be {nla_only, nla_only, converge, nla_only} respectively.” No re-tagging run needed; the §4.2/§4.3/§6 aggregates stay as published. -
Re-tag canonical to v1.2. Run
08_auto_tag_w3.py --all --overwritewith the v1.2 prompt, manually review the 6 expected changes (4 over-corrections stay, Q-D6-L4-01 + Q-D8-L3-01 flip back), and overwrite canonical. Cost: ~$0.50 / ~5 min + ~1 h review. Re-publish §4.2/§4.3/§6 aggregates with the new distribution (confab 3 / nla_only 17 / converge 5, accounting for Q-D6-L4-01 and Q-D8-L3-01 already matching v1.2 in canonical).
Recommended: option 1. The §6 aggregates are already published and the canonical/v1.0-auto distribution is what the empirical doc reports. A footnote is cheaper than re-running and re-publishing, and the four flagged qids become useful “edge case” examples in the taxonomy doc for future Path B / SP-Benchmark expansion work.
What this does not resolve
-
Q-D12-L3-01 (NLA invokes CDS/CDO for a CLN question) is the murkiest case under any taxonomy version. NLA does not match the model answer cleanly, does not match the reference, and SAE is silent. Position A puts it in
nla_onlyby default; a future v1.3 might want a distinctoff_topic_concept_hallucinationtag — defer until Path B has more such cases to motivate the split. -
Whether to expose Position A vs Position B as a methodological choice in the public write-up. The mapping-doc audience (MRM practitioners) cares about whether the readout signals correctness or method-agreement. Position A is the cleaner story for FINMA Randziffer mapping (separation of method-confidence and correctness); flag this in §6.2 of
docs/testing-validation-mapping.mdwhen SME calibration lands.
Files this would touch if accepted
- /analysis/2026-05-agreement-patterns-taxonomy/ — v1.2 addendum section appended (2026-05-27), v1 body frozen.
- /analysis/2026-05-dual-readout-empirical/ §4.1 — v1.2 reconsideration block appended (2026-05-27); §4.2/§4.3/§6 aggregates unchanged.
mechinterp/scripts/08_auto_tag_w3.py— not touched. TAXONOMY rule 4 left at v1.1 wording. Update before any Path B / SP-Benchmark expansion re-uses the auto-tagger, using the §“v1.2 prompt text” block above.- This file.