Notes on AI auditability for Swiss structured products
modelrisk.ch is my independent research project, exploring how a large language model deployed inside a Swiss bank or structured products issuer could be audited in a way that could work toward FINMA's supervisory expectations under Guidance 08/2024. The methodology is still being shaped; findings are partial.
The work is independent of any employer. It uses no employer data or IP. All structured-product examples are synthetic, based on public SSPA product structures, not real Swiss structured products in the regulatory sense.
What I'm exploring
Four research threads, at different stages of maturity. Two have worked examples; two are still ahead. The status chips say which is which.
How well do current LLMs understand Swiss structured products? SP-Benchmark is a 56-question open evaluation, scored by two independent LLM judges and mapped to FINMA Guidance 08/2024.
Browse all 56 questions →Can we produce a causal, per-decision audit trail for an LLM's output, the way a model risk manager would expect for any other quantitative model? Worked examples using attribution graphs and a per-token dual readout of model internals.
Can we detect when a deployed model starts behaving differently than it did in testing? Early experiments using sparse autoencoders.
Separately from what a model knows, how does it behave under client pressure, evaluation, or adversarial framing? A first instrument (a directional-error transcript scanner) is built and validated; the audit itself has not started.
Underlying analysis artefacts (citable research outputs across these threads) are published alongside.
Last updated 2026-07-19.