← all analysis · · docs/analysis/2026-05-taxonomy-v1.2-clarification.md

Taxonomy v1.2 clarification — what does the dual-readout certify?

On this page

Created: 2026-05-27 Trigger: Option B from docs/handoff/2026-05-27-end-of-session.md — predict v1.2 outcomes from v1.1 auto_reasoning before spending another API run. Inputs: v1.1 auto_reasoning fields in mechinterp/outputs/dual_readout_v1/w3/*.json; diagnostic appendix in docs/analysis/2026-05-taxonomy-v1.1-delta.md §“Diagnostic appendix”. Status: Clarification proposal — not a third auto-tagger pass. Decides the taxonomy question that v1.1 surfaced. Once decided, either the spec or the canonical tags need to move.

The load-bearing question

The v1.1 diagnostic appendix identified two cases (Q-D6-L4-01, Q-D8-L3-01) where the auto-tagger fires confabulation against the reference answer while the human SME tags nla_only/converge. The diagnostic concluded — correctly — that clause (c) of rule 4 (“contradiction with the reference answer”) is the surface symptom. But it left the underlying taxonomy question unanswered:

When the NLA verbalization matches a model answer that happens to be wrong, what is the dual-readout reporting?

Two coherent positions:

These are not reconcilable by prompt tweaks. They reflect different beliefs about what the readout is for.

Why the answer matters

ChoiceWhat “confabulation” means downstreamFINMA Randziffer mapping
Position ANLA channel diverges from model internals → a method-disagreement flag, model output may still be correctBound to method confidence only; correctness handled separately by judge scoring
Position BNLA + model align on a wrong answer → a correctness flag piggy-backing on NLA wrong-direction signalConflates method confidence with correctness; doubles the load on the readout

Position A is the cleaner separation of concerns and is what /analysis/2026-05-agreement-patterns-taxonomy/ §“What this taxonomy does not claim” already states verbatim: “A question can have a converge tag and still be wrong… The agreement pattern speaks to method confidence, not to model correctness.” The v1.0/v1.1 rule-4 clause (c) silently contradicts that principle.

Position B is what the auto-tagger has been doing because the prompt allowed it. It is operationally simpler but creates two separate readouts (method-agreement and correctness) inside one tag, making FINMA mapping ambiguous.

Recommended position: A. It is what the taxonomy doc already commits to in plain English; clause (c) was a residue from drafting that the v1.1 refinement preserved by accident.

Option B evidence: v1.2 won’t move the four over-corrections

The v1.1 delta proposed v1.2 as “drop clause (c)” and expected the 4 over-corrections (Q-D1-L1-01, Q-D12-L3-01, Q-D7-L1-01, Q-D7-L3-01) to flip back toward confabulation. Reading the v1.1 auto_reasoning fields directly:

qidv1.1 tagDid v1.1 invoke clause (c)?v1.1 reasoning citationv1.2 prediction
Q-D1-L1-01nla_onlyNo”this is an answer-correctness issue, not a confabulation of NLA against specific prompt tokens”stays nla_only
Q-D12-L3-01nla_onlyNo”No direct contradiction with a specific token in the prompt/answer is present — the NLA wanders off-topic rather than contradicting a specific fact”stays nla_only
Q-D7-L1-01nla_onlyNo”the NLA at pos 118 and 130 discusses knock-out mechanics without making a provably contradictory specific claim against a context token”stays nla_only
Q-D7-L3-01convergeNo”Both methods independently identify barrier-related structured product concepts, satisfying the convergence criterion”stays converge

In all four cases the v1.1 judge had already reached its tag by clause-(a)/(b) reasoning — i.e. by not finding a contradiction against the model answer. Dropping clause (c) cannot move those tags because clause (c) was not load-bearing for them. v1.2 will only flip Q-D6-L4-01 and Q-D8-L3-01 (the two cases where v1.1 did fire on clause (c)).

Predicted v1.2 distribution: confab 3 / nla_only 19 / converge 3. Canonical distribution: confab 7 / nla_only 14 / converge 4.

The real disagreement on the four over-corrections

Position A (the recommended v1.2) and the canonical SME tags disagree on these four. The canonical reviewer tagged them confabulation reasoning that NLA reinforces a wrong model answer at specific positions (e.g. Q-D1-L1-01 notes: “NLA consistently reinforces incorrect ‘Exotic Options’ classification at pos 29, 95, 107, 141-144”). Under Position A, NLA-reinforces-wrong-model is converge-on-error (if SAE corroborates) or nla_only (if SAE is silent), not confabulation. So:

qidCanonicalv1.2 (Position A)Disagreement type
Q-D1-L1-01confabulationnla_only (SAE silent)NLA reinforces wrong model → Position A says nla_only
Q-D12-L3-01confabulationnla_only (SAE silent)NLA wanders to related-wrong concepts (CDS/CDO for CLN); not aligned with a specific model token
Q-D7-L1-01confabulationnla_only (SAE blank)NLA agrees with wrong model on physical-settlement condition
Q-D7-L3-01confabulationconverge (SAE corroborates)NLA + SAE agree on barrier-option concept; model misdefines it

Three of four (Q-D1-L1, Q-D7-L1, Q-D7-L3) are clean cases of NLA-matches-wrong-model; under Position A they cleanly fall out of confabulation. Q-D12-L3 is murkier — NLA invokes adjacent-but-wrong frameworks (CDS/CDO for a CLN question), neither matching the model answer nor directly contradicting it.

v1.2 prompt text

Replace rule 4 in mechinterp/scripts/08_auto_tag_w3.py TAXONOMY with:

4. confabulation — NLA names a SPECIFIC fact in DIRECT contradiction with the model's
   own state

   Rule: NLA at any answer-relevant position contains a SPECIFIC factual claim
   (named entity, direction, magnitude, or product mechanic) that is in DIRECT
   CONTRADICTION with an identifiable token in EITHER:
     (a) the prompt text at or before that position, OR
     (b) the model's own final answer.

   The reference answer is NOT a contradiction source for this rule. If the
   model's answer is wrong and the NLA agrees with the model's wrong answer,
   that is `converge` (on the wrong concept) if SAE corroborates, or
   `nla_only` if SAE is silent. The dual-readout discriminates whether NLA
   verbalizes what the model's internal state says — it does not certify model
   correctness.

   Examples:
   - NLA at pos 12 says "above" but prompt token at pos 11 is " below" →
     confabulation (clause a, within-prompt contradiction).
   - NLA names "covered call" when model's own answer says "barrier reverse
     convertible" → confabulation (clause b, NLA-vs-model contradiction).
   - Model answer says "underlying below barrier triggers autocall" (wrong),
     NLA reinforces "below" at pos 93, SAE silent → nla_only (NLA matches
     wrong model, SAE silent — Position A: no certification of correctness).
   - Model answer says "barrier breach → zero value" (wrong), NLA invokes
     knock-out mechanics, SAE peaks on `barrier` → converge (NLA + SAE
     agree on barrier concept, model misdefines its payoff — agreement
     on wrong concept).

Also update the corresponding example block in rule 2 (nla_only) to mention the NLA-matches-wrong-model-with-silent-SAE subcase explicitly, so the judge does not get pulled back into confabulation via the priority order.

Operational implications

Adopting Position A requires one of two moves, not both:

  1. Update the spec, leave canonical alone. Lock the taxonomy doc with the v1.2 language. Treat canonical W3 tags as “v1.0 canonical, drafted before the load-bearing question was resolved.” Annotate the four Position-A-disagreement qids (Q-D1-L1-01, Q-D7-L1-01, Q-D7-L3-01, Q-D12-L3-01) in §4.1 of 2026-05-dual-readout-empirical.md with a footnote: “canonical tag predates v1.2 clarification; under v1.2 these would be {nla_only, nla_only, converge, nla_only} respectively.” No re-tagging run needed; the §4.2/§4.3/§6 aggregates stay as published.

  2. Re-tag canonical to v1.2. Run 08_auto_tag_w3.py --all --overwrite with the v1.2 prompt, manually review the 6 expected changes (4 over-corrections stay, Q-D6-L4-01 + Q-D8-L3-01 flip back), and overwrite canonical. Cost: ~$0.50 / ~5 min + ~1 h review. Re-publish §4.2/§4.3/§6 aggregates with the new distribution (confab 3 / nla_only 17 / converge 5, accounting for Q-D6-L4-01 and Q-D8-L3-01 already matching v1.2 in canonical).

Recommended: option 1. The §6 aggregates are already published and the canonical/v1.0-auto distribution is what the empirical doc reports. A footnote is cheaper than re-running and re-publishing, and the four flagged qids become useful “edge case” examples in the taxonomy doc for future Path B / SP-Benchmark expansion work.

What this does not resolve

Files this would touch if accepted