Testing and Validation of LLMs in the Structured Products Lifecycle
On this page
- A FINMA 08/2024-anchored methodology with worked examples
- 1. Preamble and positioning
- 2. FINMA 08/2024 §2.4: what is expected
- 3. SP-Benchmark as a testing methodology
- 4. Worked example: 2026-04-14 panel (8 models, dual judge)
- 5. Crosswalk: which FINMA 08/2024 expectations this document addresses
- 6. Honest gaps and open questions
- 7. Reproducing these results
- 8. Version and revisions
- 9. Appendix: source document inventory
In plain language. Regulated firms can’t deploy software in production without first checking that it works. FINMA (Switzerland’s financial regulator) published guidance in late 2024 (Guidance 08/2024) on what that check should look like for AI systems. This page is an attempt to build a methodology that satisfies that guidance for one specific case: large language models being used in workflows involving Swiss structured products. The methodology, called SP-Benchmark, is a 56-question test scored by two independent AI judges. The questions deliberately span everything from basic recall to traps where the model is likely to confuse the direction of a payoff. What follows is the technical writeup.
A FINMA 08/2024-anchored methodology with worked examples
Version: 0.2.2 (2026-07-09) · judge prompt v0.3.1 Audience: Model Risk Management practitioners at Swiss banks and structured product issuers Scope: FINMA Guidance 08/2024 §2.4 (Tests and ongoing monitoring) for LLMs deployed in SP-adjacent use cases
1. Preamble and positioning
This document operationalises one focus area of FINMA Guidance 08/2024, specifically Tests and ongoing monitoring (§2.4), and presents a concrete testing methodology, SP-Benchmark, as a worked example. It is addressed to Model Risk Management (MRM) practitioners preparing testing evidence for LLM-based applications in structured product workflows (structuring, pricing, documentation review, lifecycle management, client-facing advisory).
Throughout, three LLM failure modes recur: FM1 (surface knowledge: right terms, wrong mechanics), FM2 (isolated knowledge: deep in one area, blind to adjacent ones), and FM3 (directional confusion: right concept, wrong sign). FM3 is the one that flips the direction of a payoff or hedge, and the one that matters most for pricing.
What this document is not: a compliance certification, a legal opinion, or a full AI governance framework. It does not cover governance (§2.1), the AI inventory (§2.2), or independent review (§2.7), which sit upstream of any testing regime.
FINMA’s stance is technology-neutral: “the technology-neutral, principle-based regulatory requirements for effective governance and risk management cover the risks arising from the use of AI” (08/2024 §1). There is no AI-specific legislation in Switzerland; the existing operational-risk framework (Art. 89 CAO, Circular 2017/1 for banks) applies. Industry has summarised this as “same business, same risks, same rules” (SIX Group policy paper, September 2025, attributing this framing to FINMA). 08/2024 is not a new rulebook. It is a finding letter, stating where FINMA observed weaknesses and what it now assesses.
This document is published openly as an industry best-practice contribution, in the spirit of SIX Group’s September 2025 call for industry groups to “publish best practices… share recommended methodologies for testing AI models” (§3.4).
Materiality tier. 08/2024 expects proportional application: scrutiny scales with materiality (§1, footnote 2) and with application opacity (§2.4, footnote 9). SP-Benchmark is aimed at high-materiality LLM use cases: frontier black-box models deployed in structuring, pricing, or retail-facing advisory where errors affect compliance posture, client outcomes, or issuer P&L. Lighter-tier applications (e.g., internal summarisation of public research) can use a reduced subset; the rubric, FM tagging, and dual-judge discipline here are dimensioned for the upper tier.
Sibling sections in the modelrisk.ch research arc cover the rest of 08/2024: /monitoring (data drift, override analysis, exception handling, the monitoring half of §2.4), /explainability (§2.6, addressed via the Gemma-2-2B attribution-graph programme), and /documentation (per-decision audit-trail schema for §2.5). All three are out of scope here and will appear separately.
2. FINMA 08/2024 §2.4: what is expected
FINMA observed specific weaknesses in this focus area:
“FINMA observed weaknesses in the selection of performance indicators, tests and ongoing monitoring at some of the supervised institutions.” (08/2024 §2.4)
FINMA then enumerates seven assessment expectations. For an LLM deployed in an SP-adjacent use case the operational reading is:
- Tests scheduled for accuracy, robustness, stability and (where necessary) bias: pre-planned, not ad-hoc.
- Expert-defined questions and expected answers: domain experts (structuring, pricing, legal, compliance) provide questions and reference answers, not ML engineers alone.
- Performance indicators defined in advance: scoring axes and thresholds committed before evaluation, not curve-fit after.
- Thresholds or validation methods for output correctness: each output assessed against a pre-committed rubric.
- Data drift monitoring: covered in /monitoring, not here.
- Analysis of user overrides: covered in /monitoring.
- Exception handling: prior consideration of out-of-distribution cases.
Footnote 8 is the most operationally useful passage in 08/2024 for a benchmark-based MRM argument. Verbatim:
Footnote 9 scales frequency to materiality and opacity: “The more essential and complex the application and the less is known about how the system works or the underlying data, the more important it is to assess whether the application is working according to its purpose before productive use, in the event of changes and – especially due to the adaptive nature of today’s applications – on an ongoing basis.” For LLMs (closed-weight, rolling-version, frontier black-box models) this sits at the high end of FINMA’s frequency expectation.
Footnote 8 enumerates six concrete test types. Every one has a direct correspondence in SP-Benchmark; the next section sets this out explicitly.
3. SP-Benchmark as a testing methodology
3.1 The crosswalk
The load-bearing table of this document. Each test type FINMA names in §2.4 footnote 8 mapped to a concrete SP-Benchmark construct.
| FINMA-listed test type | SP-Benchmark correspondence |
|---|---|
| Tests with known correct result (backtesting, out-of-sample) | L1–L3 reference questions with judge-graded rubric |
| Borderline / sensitivity / stress tests | L4 strategic and ADV adversarial questions |
| Tests with incorrect input data (adversarial) | ADV tier: deliberate traps, missing clauses, sign flips |
| Tests against additional benchmark models | Dual-judge methodology (Opus + gpt-5.4); cross-model peer ranking |
| Application limit detection | Per-domain and per-level performance floors |
| Repeatability checks | Temperature=0.0, single-run reproducibility discipline |
All six test types FINMA explicitly names have a correspondence in SP-Benchmark. This crosswalk is arguably the most useful single argument for why a benchmark built on these principles could serve as testing-half evidence in a FINMA §2.4 submission, but whether the correspondence is sufficient is a judgement for an MRM reviewer, not a claim this document settles.
3.2 Question taxonomy
SP-Benchmark v0.3 contains 56 questions across 13 domains (D1–D13) and 5 difficulty levels. Domains map to structured-product lifecycle stages: D1 Structuring · D2 Pricing & Greeks · D3 Issuance & Documentation · D4 Distribution · D5 Lifecycle Events · D6 Risk Transfer · D7 Maturity & Settlement form the core lifecycle; D8 SSPA Taxonomy · D9 Institutional & Regulatory (Swiss Context) add the Swiss scaffolding; D10 Cost & Transparency · D11 COSI, Credit Risk & Issuer Default · D12 CLNs & Fixed-Income · D13 FX/Commodity/Cross-Asset cover material-risk and product-family extensions.
The five levels form a difficulty ladder: L1 Recall → L2 Decomposition (mechanics, components) → L3 Reasoning (causal chains, edge cases, vol surface effects) → L4 Strategic (cross-book, portfolio, regulatory) → ADV Adversarial (deliberate traps, sign flips, missing clauses). All 56 questions, with reference answers and scoring notes, are published at the SP-Benchmark questions page.
3.3 Failure mode tagging
Every question is tagged with one or more of three failure modes chosen to capture LLM failure patterns that matter to MRM:
- FM1 (Surface knowledge). Correct terminology, wrong mechanics.
- FM2 (Isolated knowledge). Deep in one stage, blind to adjacencies (e.g. knows BRC pricing but not its documentation or lifecycle implications).
- FM3 (Directional confusion). Right concept, wrong sign or causality. The most dangerous mode for SP pricing and hedging: a sign error on delta, vega or correlation is book-impacting and a mis-sold product.
Public multiple-choice benchmarks (MMLU, TruthfulQA) tend to test recall or truth attribution rather than directional correctness on quantitative relations, so FM3-style errors are harder to isolate from FM1 in such rubrics. For an SP issuer the FM3 rate is the single most risk-relevant failure metric.
3.4 Scoring rubric
Each question is scored on a 0–8 core rubric decomposed into four axes:
| Axis | Max | MRM-relevant failure it detects |
|---|---|---|
| Factual accuracy | 2 | Wrong claim |
| Causal reasoning | 3 | No mechanism: recites without explaining why |
| Sign / directionality | 2 | Wrong direction (the FM3 hazard) |
| Edge case awareness | 1 | Fails to flag caveats or boundary conditions |
A separate cross-domain sub-score (0–2) is reported but never added to the core total. Rewarding a model for mentioning adjacent domains correlates with verbosity, and verbosity-correlated metrics bias leaderboards toward wordy models rather than accurate ones. Cross-domain is tracked on its own axis, an explicit design decision, not a simplification.
3.5 Judge methodology
SP-Benchmark uses model-as-judge scoring. Expert human scoring at 56 questions × M models × multiple runs is operationally prohibitive. The validity question (can one LLM reliably score another?) is the central methodological risk, addressed by the dual-judge design in Section 4. Candidate and judge models run at temperature 0.0 throughout.
3.6 Reproducibility
- Temperature 0.0 on all candidate and judge models.
- Single run per (model, question) combination in v1. This is an explicit limitation (see Section 6).
- Framework, judge code, and judge prompt all pinned to commit SHAs (exact SHAs in §7).
- Results versioned in
eval_results/with the exact CLI command and model registry used.
Anyone with API access to the evaluated models can re-run and compare.
4. Worked example: 2026-04-14 panel (8 models, dual judge)
On 2026-04-14 an eight-model panel was run through the full 56-question framework with two independent judges (claude-opus-4-6 and gpt-5.4, both at temperature 0.0), producing 448 scored evaluations. The full report breaks down per-model, per-domain, and per-failure-mode results; what follows is the MRM-relevant summary.
Models evaluated: gemma4-31b-cloud (Google Gemma 4, cloud), qwen3.5-397b-cloud (Alibaba Qwen 3.5 MoE, cloud), glm-5.1-cloud (Zhipu, cloud), gemma4-26b-local (Gemma 4 on local DGX Spark vLLM), minimax-m2.7 (cloud), gpt-oss-120b-cloud (OpenAI open-weights, cloud), nemotron-3-super (NVIDIA, cloud), deepseek-v3 (cloud).
Headline ranking (mean core score, averaged across judges): top three were gemma4-31b-cloud (6.04/8, 75.6%), qwen3.5-397b-cloud (5.93/8, 74.1%), and glm-5.1-cloud (5.81/8, 72.7%). Bottom two were nemotron-3-super (4.64/8, 58.0%) and deepseek-v3 (3.56/8, 44.5%). The chart below shows each model’s two judge scores side by side; the connecting bar is the per-model judge spread (a model-by-model picture of the dual-judge noise quantified in §4 below).
Per-domain pattern. Hardest across all models: D11 COSI & issuer default (mean 4.50/8), D8 SSPA Taxonomy (4.74/8), D2 Pricing & Greeks (4.81/8). Easiest: D3 (5.92/8), D4 (5.73/8), D13 (5.69/8). With ~4 questions per domain these rankings are indicative, not powered: a signal of where to look harder, not a settled ordering. With that caveat, a D11 floor would matter for MRM: issuer default and COSI collateralisation are material-risk topics for retail SP distribution.
Per-level pattern. Lowest at L1 (4.88/8) and L2 (4.92/8); higher at L3 (5.49/8), L4 (5.50/8), ADV (5.35/8). The counter-intuitive result (recall harder than reasoning) likely reflects a scoring artefact: L1 recall items test precise factual claims (a specific SSPA code, a specific circular number) with no partial credit, whereas L3/L4 items reward mechanism-explanation even when a final number is off.
Failure-mode distribution across all 448 evaluations: FM1 22.5%, FM2 27.7%, FM3 14.7%. Roughly one in seven evaluations showed directional confusion on a temperature-zero, reproducible test. On a 56-question snapshot this is indicative rather than a powered estimate. If a rate of this order survives expansion, it is the kind of failure mode that would bear most directly on a deployment threshold for SP pricing, because FM3 errors flip the sign of a payoff or a hedge.
Inter-judge agreement: the MRM-relevant methodological finding
Dual-judge methodology is SP-Benchmark’s operational response to FINMA’s “tests against additional, possibly simpler benchmark models.” Each of the 448 evaluations was independently scored by both judges:
| Metric | Value |
|---|---|
| Mean core-total spread (0–8 scale) | 1.37 |
| Exact agreement rate (spread = 0) | 64.1% |
| Perfect agreement on full core total (spread = 0) | 33.0% |
| High disagreement (spread ≥ 3) | 21.9% |
A 21.9% high-disagreement rate on a temperature-zero, deterministic evaluation implies that in this panel, the two judges differed by roughly ±1.4 points on average on the 0–8 scale. On a benchmark constructed like this one, that suggests a single judge’s score is better read as a confidence band than a point estimate. A firm using a single judge to make a deploy/no-deploy call at a 5.0/8 threshold should know that here a 5.0 could have been a 3.6 or a 6.4 under the equally reasonable alternative judge.
Running two judges costs roughly 50% more than single-judge (candidate answers scored twice; generation cost unchanged). Based on single-judge per-model totals of 0.10 observed in the 2026-04-14 panel (see Leaderboard, Cost column), adding a second Opus-class judge pass brings the dual-judge per-model cost into the approximate range of 0.15. At 56 questions per model, that is ~0.0027 per evaluation, operationally negligible for MRM validation workflows even scaled to ten or twenty models. Dual-judge is operationally feasible at this cost profile.
Three MRM-relevant per-model observations
deepseek-v3: low judge spread, low absolute score. Judge spread was 0.7 points (the lowest on the panel), suggesting close judge agreement on deepseek’s output. Absolute score was 3.56/8 (lowest), with 60.7% FM2 and 48.2% FM1 rates and a 0.23/2 cross-domain score. The low spread is not reassurance: the judges agree the answers are consistently wrong.
qwen3.5-397b-cloud: strong cross-domain, low FM3. Qwen recorded the highest cross-domain score on the panel (1.75/2) and a 10.7% FM3 rate (second-lowest). Strong for multi-stage reasoning workflows that cross lifecycle boundaries. But 151s latency and 5,751 output tokens per question: the model reasons by writing a lot. Verbosity-cost trade-off matters for production.
nemotron-3-super: FM3 at 23.2%. Highest FM3 rate on the panel; nearly one in four evaluations shows directional confusion. Combined with 32.1% FM2 and 4.64/8 overall, this is the panel’s clearest “looks fluent, wrong mechanics” profile. On a larger sample this is the profile that would most warrant a do-not-deploy finding for SP pricing applications where sign errors are book-impacting, though on N=56 the responsible reading is “investigate before trusting,” not a final verdict.
5. Crosswalk: which FINMA 08/2024 expectations this document addresses
The 2026-04-14 panel establishes what the testing methodology covers. The next question is what it does not cover: where the line is between an on-scope §2.4 testing contribution and the sibling workstreams that address the rest of 08/2024. Scope of this document against each of the seven focus areas in 08/2024 §2. Overclaiming is unhelpful; precise scoping is what an MRM review needs.
| Focus area | Status | How / where |
|---|---|---|
| I Governance (§2.1) | Not addressed here | Organisational requirement. Covered by the institution’s AI governance framework; benchmark is an input to governance, not a substitute. |
| II Inventory and risk classification (§2.2) | Not addressed here | Upstream prerequisite. The institution must first classify an SP-adjacent LLM use case before the §2.4 testing regime applies. |
| III Data quality (§2.3) | Partially addressed | Benchmark’s own dataset is a curated, versioned, pinned 56-question corpus. That is its own data-quality story. Does not address the candidate LLM’s training data (frontier closed models are opaque by construction). |
| IV Tests and ongoing monitoring (§2.4) | Directly addressed (testing half). | This entire document. The monitoring half (data drift, override analysis, exception handling) is scoped to a future /monitoring section. |
| V Documentation (§2.5) | Partially addressed | This document is itself a documentation artefact for the testing function (one of the nine §2.5 elements FINMA lists). Per-decision audit trail is a separate /documentation workstream. |
| VI Explainability (§2.6) | Not addressed here | Deliberately out of scope. Explainability is addressed by a separate mechanistic-interpretability programme (Gemma-2-2B attribution-graph pipeline) and will appear as /explainability. |
| VII Independent review (§2.7) | Partially addressed | Benchmark provides reviewable evidence that an independent reviewer can use; it does not constitute independent review of the model development process. |
5.1 International-standards alignment: ISO/IEC 42001 and 23894
FINMA 08/2024 is technology-neutral and principle-based; it names no standard to conform to, and there is no Swiss requirement to hold one. But MRM functions are increasingly asked, by their own second line, by group risk, or by non-Swiss counterparties, how an AI control maps to the international standards now consolidating around AI governance. The two that bear on a testing methodology are ISO/IEC 42001:2023 (AI management system, the certifiable AIMS standard) and ISO/IEC 23894:2023 (AI risk management guidance). This alignment is offered only as a convenience cross-walk: it lets an institution pursuing or holding 42001 certification locate where SP-Benchmark evidence sits inside that framework. It does not imply that conforming to ISO discharges any FINMA expectation, nor the reverse.
The mapping is deliberately narrow. SP-Benchmark is a verification-and-validation instrument; it touches a handful of ISO/IEC 42001 Annex A controls and is silent on the rest (governance, resources, third-party relationships), exactly as it is silent on FINMA focus areas I, II, VI and VII above.
| ISO control / clause | What it asks for | SP-Benchmark correspondence |
|---|---|---|
| 42001 A.6.2.4 — AI system verification and validation | Defined, documented V&V across the life cycle | The core of this document: the §3.1 test-type crosswalk and the §4 worked example are V&V evidence. |
| 42001 A.6.2.6 — AI system operation and monitoring | Ongoing monitoring of a deployed AI system | Partially: the testing half maps here; continuous monitoring (drift, override analysis) is the roadmap /monitoring workstream, the same gap as FINMA §2.4’s “ongoing” leg. |
| 42001 A.6.2.8 — recording of event logs | Event logging sufficient for traceability | The per-evaluation Langfuse trace capture (every evaluation is an auditability record) is the schema-level answer; full per-decision audit trail is the /documentation workstream. |
| 42001 A.7.4 — quality of data for AI systems | Data-quality management for AI data | The benchmark’s own 56-question corpus is versioned, pinned and canary-probed for contamination (§6.1). It does not speak to the candidate model’s training data, the same boundary as FINMA §2.3. |
| 42001 A.5.2 — AI system impact assessment process | A documented impact-assessment process | Out of scope here; listed only so the gap is explicit. |
| ISO/IEC 23894 — AI risk identification and analysis | Identify and analyse AI-specific risks | The FM1/FM2/FM3 failure-mode taxonomy (§3.3) is a domain-specific risk-identification scheme for SP knowledge failures, of which 23894 is the general-purpose counterpart; the FM3 directional-error rate is a quantified risk indicator in 23894 terms. |
As with the FINMA crosswalk (§3.1), whether these correspondences are sufficient is a judgement for an MRM or certification reviewer holding the standard text, not a claim this document settles. Annex A control numbers follow ISO/IEC 42001:2023; verify against your edition.
6. Honest gaps and open questions
6.1 What the benchmark does not test
- Production integration. SP-Benchmark is unit-level testing of the language model. It does not cover RAG retrieval quality, tool-use correctness, agent orchestration, or prompt-template robustness. A model that scores well can still fail in production because the pipeline around it is wrong.
- Real client data. Questions are synthetic, SSPA-structure-based, drawn from public domain knowledge. No client scenarios, real term sheets or KIDs. Deliberate for reproducibility and IP reasons, but it limits coverage of production distributions.
- Training-data contamination — two complementary checks, both passing so far. The questions derive from public material certainly in training data, and the questionnaire is public. Two instruments probe whether models reason or recite: a fictional-entity canary (non-existent codes/regulations where memorization fails) and a perturbed-twin memorization index — 12 real questions each with one parameter inverted so the correct answer changes, scoring MI = score(original) − score(twin). For
claude-sonnet-4-6the mean MI is +0.50/8 with 6/12 twins at 0 — no contamination signature; the residual signal concentrates in FM3 directional flips, so the twin set doubles as an FM3-robustness probe. The canary stays the cleaner pure-contamination test. Both run on one model / small n so far. Writeup: contamination audit. - Temporal drift. A single 2026-04-14 snapshot, no longitudinal tracking. FINMA §2.4 expects ongoing assessment; this snapshot is not yet a series.
6.2 Methodological limitations
- N = 56 is small. Statistical power on per-domain breakdowns (~4 questions per domain) is limited: indicative, not conclusive, for domain-level sign-off. Expansion is on the roadmap.
- Dual-judge improves confidence but does not make judge-as-evaluator objective. The 1.37-point mean spread is the empirical uncertainty floor and should be reported alongside every score claim. Treating SP-Benchmark scores as point estimates is misuse.
- Temperature-zero sampling does not test stochastic failure modes. FINMA §2.4 footnote 8 names “repeatability” as a test concern; repeatability under production sampling (top-p, temperature > 0, tool-use branch divergence) is not tested in v1 and needs a stochastic-sweep extension.
- Human-expert calibration of the judges: first pass done (v0.2), partial result. The dual-judge design addresses inter-judge disagreement but does not answer the prior question: do LLM judges agree with Swiss SP domain experts on the 0–8 rubric? A first SME re-scoring exercise was completed 2026-06-10: 12 questions scored blind by an SP practitioner against the run-1 judge scores. Result: rank-order agreement passes, absolute-value agreement fails. Core-total Spearman ρ = 0.770 (the SME and the judge order answers from best to worst substantially the same way); quadratic-weighted κ = 0.390 (weak agreement on raw scores), driven by one-sided judge leniency — the SME never scored above the judge, mean gap −1.0 points on the 0–8 core total. Two specific judge failure patterns: (1) a causal-reasoning ceiling — the judge awarded 3/3 causal reasoning on all 12 answers, including four whose reasoning chains were internally contradictory, so that dimension currently has no discriminating power; (2) under-penalised internal inconsistency — answers reaching a correct-sounding conclusion via a self-contradictory chain score ~2–3 points higher with the judge than with the SME. The best-calibrated dimension was sign/directionality (κ = 0.733, substantial), which supports the FM3-rate findings specifically. Practical reading for an MRM reviewer: relative model rankings carry a human anchor; absolute scores retain a leniency bias of roughly +1 point and should be read as upper bounds. Caveats: n = 12, single judge run compared, and the rater is the project author (blind to judge scores, but not arm’s-length) — an external specialist and ≥30 questions remain the bar for sign-off-grade calibration (v0.3).
- By-construction judge validation: done (2026-06-13), and it refines the picture above. Separately from the human pass, the judges were tested against a 21-case planted-error set — correct reference answers each carrying one known injected defect (a reversed payoff direction, a self-contradicting causal chain, a fabricated SSPA code). Because a sign-flip is a sign-flip by construction, this needs no rater and has no circularity. Both judges detect directional, causal, edge-case and fabricated-fact defects at 1.00 (Opus 4.6 0.86 / GPT-5.4 0.90 overall). Two findings sharpen the human-pass result: (1) the judges do dock self-contradiction, including subtle/numeric cases — so the causal-reasoning ceiling is a subtlety/standalone-scoring gap, not blindness; (2) the one reproducible blind spot is cousin-product mis-attribution (correct mechanics under a one-feature-off product name, e.g. autocallable mechanics called a barrier reverse convertible). Full writeup: planted-error judge validation; the formulas are on the Methodology page.
- The cousin-product blind spot is closed at the instrument level (judge v0.3.1, 2026-07-09). The judge prompt now opens with a mandatory product-identity check: both judges must name the product the reference concerns and the product the answer describes, and set a per-transcript
product_identity_matchflag whosefalsevalue caps factual accuracy at 1 deterministically, in code — an auditable signal, not a mood. Validated on the planted set expanded to 28 cases (cousin cell 3 → 10): detection rose from 5/10 (Opus) and 8/10 (GPT-5.4) to 9/10 and 10/10, and the whole 750-result corpus was re-judged twice under the hardened prompt (a deliberate comparability break, taken before publication; aggregate scores move < 0.1 core points for Opus). The costs are stated as plainly as the gains: two single-case Opus regressions on other planted defects (both still caught by the second judge), and on real transcripts the check can over-fire on answers that use a neighbouring product as a deliberate illustration — one SME-endorsed answer is now flagged by both judges, so the small-sample human-agreement stats decline under the new prompt (core ρ 0.770 → 0.421 on n = 12, driven entirely by two moved scores) and the causal-reasoning ceiling persists. The check errs conservative — the right direction for MRM use — and the 67 flagged transcripts (9% of the corpus for the conservative judge; the second judge flags 46% and should be read as a high-recall screen) form the natural set for the next human review pass. Full writeup: judge hardening v0.3.
6.3 Open questions for practitioner review
- Rubric weighting. Core is factual 2 / causal 3 / sign 2 / edge 1. Should this be re-calibrated per use case, e.g. pricing-weighted (sign 3, factual 2, causal 2, edge 1) vs. advisory-weighted (factual 3, edge 2, causal 2, sign 1)? Would need an expert-panel re-scoring exercise.
- Model-update handling. When Claude Sonnet 4.7 replaces 4.6, or Opus 4.7 replaces 4.6 as judge, what is the re-baselining protocol? Version-pinning is the current response; FINMA §2.4 footnote 9 expects evaluation to continue across changes. A separate exploratory thread is being scoped for the narrower case where the change is a fine-tune: interpreting a model update by its weight delta (does the delta carry the intended behaviour, and only that?) rather than re-running the whole benchmark. It is design-only so far; nothing has been run.
- Judge-prompt longitudinal robustness. The judge prompt is versioned, but a regression harness that runs an identical panel under prompt version n and n+1 and quantifies drift does not yet exist.
7. Reproducing these results
Exact commands:
# Install
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env # fill in API keys for evaluated providers
# Run evaluation (dual-judge config: schema judge.models: [...])
sp-benchmark run --config configs/sp_eval.yaml --concurrency 4
# Export artefacts
sp-benchmark export --format md --output-dir eval_results
sp-benchmark export --format csv --output-dir eval_results
sp-benchmark export --format json --output-dir eval_results
Version pins:
- Framework: the SP-Benchmark question set, version 0.3 (56 questions, 13 domains, 5 levels)
- Judge code:
src/sp_benchmark/judge.py@ commit3827cac02246a9124f75168eb4d0d751ec371591 - Question count: 56 · Domains: 13 · Levels: 5
Dual-judge is supported declaratively in the CLI. The config schema accepts a judge.models: [...] list (see configs/sp_eval.yaml). A reproducible dual-judge panel is configured as:
judge:
models:
- claude-opus-4-6
- gpt-5.4
temperature: 0.0
max_tokens: 1024
configs/sp_eval_panel5.yaml in this repo declares a single-judge config, retained as the minimal reproducibility shape for the initial scope. For the 2026-04-14 worked example, use configs/sp_eval.yaml.
8. Version and revisions
- v0.1 (2026-04-19): first draft. Anchored to FINMA Guidance 08/2024 of 18 December 2024. Post-merge with ThinkCentre work the same day: dual-judge CLI support is now in the mainline (§7 updated accordingly). 08/2024 §3 Outlook commits FINMA to refining its expectations; further revision will be needed when FINMA publishes further AI guidance, when the question count is expanded beyond N=56, when an SME re-scoring calibration is added (§6.2 gap), and when a longitudinal series replaces the 2026-04-14 single-snapshot worked example.
- v0.1.1 (2026-05-27): SME calibration instrument added (selection, scoresheet, comparison script; commit
0de9432). First-rater pass in progress; §6.2 updated to reflect status. v0.2 will fold in the first-rater results once the pass completes, closing the load-bearing methodological gap §6.2 identifies. - v0.2 (2026-06-10): human-anchor pass. First SME re-scoring calibration completed and folded into §6.2: 12 questions scored blind by an SP practitioner; core-total Spearman ρ = 0.770 / quadratic-weighted κ = 0.390 vs the run-1 judge, with a diagnosed judge causal-reasoning ceiling and ~1-point leniency bias on internally inconsistent answers. Sign/directionality is the best-calibrated dimension (κ = 0.733). Remaining for v0.3: external (non-author) rater, ≥30 questions, second judge run, and a judge-prompt fix for the causal-reasoning ceiling.
- v0.2.1 (2026-06-14): added §5.1, an international-standards cross-walk mapping the methodology to ISO/IEC 42001:2023 Annex A controls (A.6.2.4 V&V, A.6.2.6 monitoring, A.6.2.8 event logs, A.7.4 data quality, A.5.2 impact assessment) and ISO/IEC 23894:2023 risk identification. Subordinate to the FINMA crosswalk and explicitly non-binding; no new results.
- v0.2.2 (2026-07-09): judge hardening pass. Judge prompt hardened to v0.3.1 (mandatory product-identity check, named products, auditable
product_identity_matchflag, deterministic factual-accuracy cap); planted fixture expanded to 28 cases (cousin cell 3 → 10, detection Opus 5/10 → 9/10, GPT-5.4 8/10 → 10/10); full 750-result corpus re-judged twice under the new prompt (comparability deliberately broken before publication). §6.2 records the costs alongside the gains, including a decline in the small-sample SME agreement stats caused by an identity-check false positive on an illustrative answer. Of the four v0.3 bars named at v0.2, the judge-prompt fix and second judge run are met; the external SME and ≥30 questions remain open, so the methodology stays “self-calibrated.” Full writeup: judge hardening v0.3.
9. Appendix: source document inventory
- FINMA Guidance 08/2024 (primary regulatory source)
- FINMA 08/2024 notes (structured extraction of the guidance text)
- SIX Group policy paper (industry context, not regulatory)
- 2026-04-14 dual-judge benchmark report (Section 4 worked example): /analysis/2026-04-14-panel8-dual-judge/
- SP-Benchmark framework v0.3 (Section 3 methodology): the published question set