Analysis
Citable research artefacts produced across the modelrisk.ch research
tracks. Each entry mirrors a working document from the
docs/analysis/ working set of the research project,
re-published here with web-friendly cross-references. In-flight,
internal-process,
and non-citable artefacts (SME calibration in progress, taxonomy
deltas, debug notes) remain internal.
- Judge hardening v0.3: the product-identity check — design, validation, re-baseline
Judge v0.3.1: auditable product-identity check closes the cousin-product blind spot (planted detection 5/10→9/10, 8/10→10/10); full 750-result re-baseline with honest costs.
- Contamination audit — perturbed-twin memorization index
Memorization-vs-reasoning audit via perturbed twins (one parameter inverted). No contamination signature for claude-sonnet-4-6; doubles as an FM3-robustness probe.
- Judge validation — planted-error set
By-construction validation of the SP-Benchmark judge: 21 planted-error cases, no human rater. Surfaces a cousin-product mis-attribution blind spot.
- Dual Readout — NLA + SAE on Gemma-3-12B-IT L32 (Path A Week 2)
Week 2 of Path A — per-token dual readout (NLA + GemmaScope-2 SAE) on the barrier triple; proof of protocol.
- Contrastive Attribution — Does the Barrier Direction Flip the Circuit?
Contrastive attribution on Gemma-2-2B: does flipping below→above the barrier flip the circuit? Negative result.
- SP-Benchmark Panel — 8 open/cloud models, dual-judge (April 2026)
8-model dual-judge benchmark run on 56 SP-Benchmark questions, scored by Claude Opus 4.6 and GPT-5.4 at temperature 0.
- Agreement-Pattern Taxonomy v1 (Path A Week 3 tagging frame)
Frozen v1 taxonomy of NLA-vs-SAE agreement patterns: converge, nla_only, sae_only, confabulation.
- Dual Readout — SP-Benchmark Empirical Run (Path A Week 3)
Week 3 of Path A — per-token dual readout on 25 SP-Benchmark questions; tag distribution, case studies, aggregates.
- First Attribution Graph — Barrier Knock-in on Gemma-2-2B
First circuit-tracer attribution graph on the barrier-knock-in prompt — Gemma-2-2B, all 26 layers, full causal-chain writeup.
- NLA Sidecar — Cross-method Comparison on Barrier Knock-in
Week 1 of Path A — natural-language autoencoder sidecar on Gemma-3-12B-IT, final-token readout, infrastructure proof.
- Taxonomy v1.2 clarification — what does the dual-readout certify?
Taxonomy v1.2 — what does the dual-readout certify? Reserves confabulation for NLA-vs-model gaps, not NLA-vs-reference.