Be the v0.3 validator — open re-scoring protocol
On this page
In plain language. SP-Benchmark scores how well AI models reason about Swiss structured products, using two AI models as judges. The gap: those judges have so far been checked mainly against the project’s author. This page is an open invitation — if you work in Swiss structured products, you can independently validate the judge in about two hours, and your scores become the benchmark’s next calibration, with public credit.
The honest gap
Two things already reduce the circularity, but neither replaces an independent human:
- A by-construction judge validation (planted-error set) — defects with unarguable ground truth (a sign-flip is a sign-flip). It needs no rater and shows the judges catch directional, causal, edge-case and fabrication defects at ~100%. Its one documented blind spot (cousin-product mis-attribution) has since been closed at the instrument level by a hardened judge prompt — at the cost of occasionally over-flagging answers that use a neighbouring product as a deliberate illustration.
- A first human re-scoring (SME calibration, §6.2 of the mapping doc) — but the rater was the author.
What’s missing is an independent domain expert scoring the subjective 0–8 rubric. That’s the ask.
The ask — ~2 hours, blinded
- You receive (or generate) a blinded kit: 30 questions, each with the prompt, the reference answer, and one model’s answer — with all judge scores removed. You never see the AI’s scores while scoring.
- Score four dimensions per question against the rubric below. ~3–4 minutes each.
- Return the filled sheet. I compute your agreement with the judge (Spearman ρ for ranking, quadratic-weighted κ for absolute scores — see the Methodology page) and publish it as the v0.3 external calibration, crediting you (or anonymously, your choice).
The kit is reproducible from the repository with no model calls — it reuses already-computed answers, so there’s no cost and anyone can regenerate it.
The rubric, with worked anchors
Score each answer against the reference answer, on what it is, not how confident it sounds. The low anchors below are drawn from the planted-error set, so each band is concrete.
factual_accuracy — 0 / 1 / 2
| Score | Meaning | Worked example |
|---|---|---|
| 2 | Core claim fully correct | ”A BRC = long zero-coupon bond + short down-and-in (barrier) put.” |
| 1 | Right mechanics, wrong product identity, or a material omission | ”A reverse convertible = ZCB + short down-and-in barrier put” — correct plumbing, but a plain RC has no barrier. |
| 0 | Core claim wrong, or a confidently fabricated fact | ”Under FINMA Circular 2019/3 Annex 4, the issuer must hold a 15% buffer.” No such rule exists. |
This cousin-product “1” is where the judges were historically weak. Since the v0.3.1 judge hardening they catch it via an explicit product-identity flag — but the flag can also over-fire on answers that use a neighbouring product as a deliberate illustration. Your independent call on where legitimate illustration ends and mis-attribution begins is the most valuable signal in the kit.
causal_reasoning — 0 / 1 / 2 / 3
| Score | Meaning | Worked example |
|---|---|---|
| 3 | Full, coherent causal chain | ”Steeper skew → higher OTM vol → the down-and-in put costs more → issuer short it → coupon falls.” |
| 2 | Partial; a link missing | The chain above but stopping at “the put costs more.” |
| 1 | Asserts conclusion without mechanism | ”Steeper skew is bad for the client.” |
| 0 | No causal content, or a self-contradictory chain | ”The ZCB leaves almost no budget … so there is ample budget to fund a generous coupon.” |
Score a self-contradictory chain 0–1 even if the conclusion sounds right.
sign_directionality — 0 / 1 / 2
| Score | Meaning | Worked example |
|---|---|---|
| 2 | Direction correct | ”When skew steepens the client gets a worse product.” |
| 1 | Ambiguous / doesn’t commit | ”Skew changes affect the client’s economics.” |
| 0 | Direction reversed | ”When skew steepens the client gets a better product.” |
edge_case_awareness — 0 / 1
| Score | Meaning | Worked example |
|---|---|---|
| 1 | Flags ≥1 boundary/risk | Names issuer-default / unsecured-creditor risk alongside barrier-breach risk. |
| 0 | Omits the key edge case | Discusses only barrier breach, never issuer default. |
A separate cross_domain_bonus (0–2) rewards connections to adjacent domains; it is never added to the core total.
What can and cannot be claimed
No employer data is involved — the questions are synthetic and built from public SSPA/FINMA domain knowledge. If you take part in a personal capacity from a structured-products firm, that affiliation is disclosed in the writeup, not hidden.
Take part
Named credit in the repository and on this site, or anonymous attribution (“an independent Swiss SP practitioner”) if you prefer. The full protocol, rubric anchors, and kit generator live in the project repository (docs/rescoring-protocol.md, scripts/build_rescoring_kit.py). To request the kit or ask a question, reach the author via About.