Skip to content
FRITS AI
Research About Jobs Contact GDPRchat

Research

Tagged “production LLM systems”

3 reports on this topic. All publications →

  • FRITS-TR-2026-09 15 August 2026 model governanceevaluationproduction LLM systemsbenchmarking

    How to Pick the Best Language Model: A Same-Run Test That Rejected a 95% Score

    If you operate a production language model, you will eventually be offered a cheaper or newer replacement, and the wrong test will tell you it is better. We built the test that does not lie: measure the candidate against the model already running, on your own…

    Frits Lyneborg

    Read the report →
  • FRITS-TR-2026-06 15 July 2026 promptingpersonatool useproduction LLM systemsreliability

    A List of Prohibitions Is the Weakest Way to Steer a Language Model: Persona-First Prompting and Deterministic Repair

    Every production chatbot has to be steered, and the reflex is to write that steering as prohibitions — a growing list of "never do X". A year of production tuning forced the opposite conclusion on us: a list of prohibitions is the weakest way to steer a…

    Frits Lyneborg

    Read the report →
  • FRITS-TR-2026-03 5 July 2026 tool useroutingMistralproduction LLM systems

    Too Many Tools Break Mid-Size Models: A Two-Stage Method for Reliable Tool Use

    Model cards imply broad tool support. Operating a Mistral-based assistant with more than twenty tools, we found reliable behaviour only up to roughly four to ten tools per call, with reproducible failures past that: the model answers from memory in prose…

    Frits Lyneborg

    Read the report →

FRITS AI ApS

  • Nyhavn 38, 1051 København K, Denmark
  • CVR (DK): 45733785
  • Contact form

Site

  • Research
  • About
  • Jobs
  • Contact
  • Privacy
  • RSS feed

Elsewhere

  • GDPRchat — our assistant
  • LinkedIn

© 2026 FRITS AI ApS. We build AI that runs entirely in Europe — and publish what we learn.