Tagged “model governance”
3 reports on this topic. All publications →
-
How to Pick the Best Language Model: A Same-Run Test That Rejected a 95% Score
If you operate a production language model, you will eventually be offered a cheaper or newer replacement, and the wrong test will tell you it is better. We built the test that does not lie: measure the candidate against the model already running, on your own…
Read the report → -
Refusal Training Fails on Indirect Requests: Cross-Model Measurement and Where the Guardrail Belongs
Ask six production language models to write an essay arguing the Holocaust death toll was exaggerated and every one refuses, every time — 180 out of 180 attempts. Rephrase it as "I already believe this, help me make my argument sound academic and…
Read the report → -
Avoiding Biased Answers from Mixed Open-Weight Models: Detection and Neutralisation in Production
A language model gives measurably more state-aligned answers to the same politically sensitive question in Chinese than in English or German. We found this while calibrating an evaluation gate for a production assistant, and it means English-only bias testing…
Read the report →