Monitoring (coming soon)

In plain language. A model that scored 75% on its pre-deployment test in April may not still be scoring 75% in production six months later, not because the model itself changed, but because the questions being asked, the conversational context around them, or the upstream data shifted. Detecting this kind of drift is part of what regulators expect from any deployed model. This page will document an early experiment in detecting LLM drift at the population level, using a technique called sparse autoencoders.

This page is in development. It will cover population-level drift detection for LLM deployments on open-weight models, using sparse autoencoders, anchored to FINMA 08/2024 focus area IV (“Tests and ongoing monitoring,” the monitoring half).

The first experiment is specified in an internal design document; its results will be published here on completion.

Expected publication: late 2026.