Be the v0.3 validator — open re-scoring protocol

On this page

In plain language. SP-Benchmark scores how well AI models reason about Swiss structured products, using two AI models as judges. The gap: those judges have so far been checked mainly against the project’s author. This page is an open invitation — if you work in Swiss structured products, you can independently validate the judge in about two hours, and your scores become the benchmark’s next calibration, with public credit.

The honest gap

Two things already reduce the circularity, but neither replaces an independent human:

What’s missing is an independent domain expert scoring the subjective 0–8 rubric. That’s the ask.

The ask — ~2 hours, blinded

  1. You receive (or generate) a blinded kit: 30 questions, each with the prompt, the reference answer, and one model’s answer — with all judge scores removed. You never see the AI’s scores while scoring.
  2. Score four dimensions per question against the rubric below. ~3–4 minutes each.
  3. Return the filled sheet. I compute your agreement with the judge (Spearman ρ for ranking, quadratic-weighted κ for absolute scores — see the Methodology page) and publish it as the v0.3 external calibration, crediting you (or anonymously, your choice).

The kit is reproducible from the repository with no model calls — it reuses already-computed answers, so there’s no cost and anyone can regenerate it.

The rubric, with worked anchors

Score each answer against the reference answer, on what it is, not how confident it sounds. The low anchors below are drawn from the planted-error set, so each band is concrete.

factual_accuracy — 0 / 1 / 2

ScoreMeaningWorked example
2Core claim fully correct”A BRC = long zero-coupon bond + short down-and-in (barrier) put.”
1Right mechanics, wrong product identity, or a material omission”A reverse convertible = ZCB + short down-and-in barrier put” — correct plumbing, but a plain RC has no barrier.
0Core claim wrong, or a confidently fabricated fact”Under FINMA Circular 2019/3 Annex 4, the issuer must hold a 15% buffer.” No such rule exists.

This cousin-product “1” is where the judges were historically weak. Since the v0.3.1 judge hardening they catch it via an explicit product-identity flag — but the flag can also over-fire on answers that use a neighbouring product as a deliberate illustration. Your independent call on where legitimate illustration ends and mis-attribution begins is the most valuable signal in the kit.

causal_reasoning — 0 / 1 / 2 / 3

ScoreMeaningWorked example
3Full, coherent causal chain”Steeper skew → higher OTM vol → the down-and-in put costs more → issuer short it → coupon falls.”
2Partial; a link missingThe chain above but stopping at “the put costs more.”
1Asserts conclusion without mechanism”Steeper skew is bad for the client.”
0No causal content, or a self-contradictory chain”The ZCB leaves almost no budget … so there is ample budget to fund a generous coupon.”

Score a self-contradictory chain 0–1 even if the conclusion sounds right.

sign_directionality — 0 / 1 / 2

ScoreMeaningWorked example
2Direction correct”When skew steepens the client gets a worse product.”
1Ambiguous / doesn’t commit”Skew changes affect the client’s economics.”
0Direction reversed”When skew steepens the client gets a better product.”

edge_case_awareness — 0 / 1

ScoreMeaningWorked example
1Flags ≥1 boundary/riskNames issuer-default / unsecured-creditor risk alongside barrier-breach risk.
0Omits the key edge caseDiscusses only barrier breach, never issuer default.

A separate cross_domain_bonus (0–2) rewards connections to adjacent domains; it is never added to the core total.

What can and cannot be claimed

No employer data is involved — the questions are synthetic and built from public SSPA/FINMA domain knowledge. If you take part in a personal capacity from a structured-products firm, that affiliation is disclosed in the writeup, not hidden.

Take part

Named credit in the repository and on this site, or anonymous attribution (“an independent Swiss SP practitioner”) if you prefer. The full protocol, rubric anchors, and kit generator live in the project repository (docs/rescoring-protocol.md, scripts/build_rescoring_kit.py). To request the kit or ask a question, reach the author via About.