Tagged “production LLM systems”
3 reports on this topic. All publications →
-
How to Pick the Best Language Model: A Same-Run Test That Rejected a 95% Score
If you operate a production language model, you will eventually be offered a cheaper or newer replacement, and the wrong test will tell you it is better. We built the test that does not lie: measure the candidate against the model already running, on your own…
Read the report → -
A List of Prohibitions Is the Weakest Way to Steer a Language Model: Persona-First Prompting and Deterministic Repair
Every production chatbot has to be steered, and the reflex is to write that steering as prohibitions — a growing list of "never do X". A year of production tuning forced the opposite conclusion on us: a list of prohibitions is the weakest way to steer a…
Read the report → -
Too Many Tools Break Mid-Size Models: A Two-Stage Method for Reliable Tool Use
Model cards imply broad tool support. Operating a Mistral-based assistant with more than twenty tools, we found reliable behaviour only up to roughly four to ten tools per call, with reproducible failures past that: the model answers from memory in prose…
Read the report →