Nine guardrails, 949 cases, four corpora. Two of those corpora are public datasets nobody here curated. Everything is published: the harness, the attack corpus, the adapters, and every case each product missed.
We publish this and we compete in it. No methodology removes that conflict. What we did instead: scoring was fixed before anything ran, competitors run at stock settings with nothing tuned to these cases, every miss is listed by case ID, and a thirty-line regex control is included so you can see when a sophisticated product barely beats it. Fluiq does not place first on either output-side corpus. Re-run it and disagree.
Leak shapes a PII corpus does not contain: obfuscated identifiers, credential formats, injected instructions echoed into output, and refusals that leak while refusing.
49 cases · 32 should block, 17 should pass · source Written for this benchmark · licence MIT
| Guardrail | Category | Recall | False alarm | F1 |
|---|---|---|---|---|
llm-guard | Open source | 87.5% | 17.6% | 88.9% |
fluiq | Fluiq | 81.2% | 11.8% | 86.7% |
aws-comprehend | Commercial | 84.4% | 23.5% | 85.7% |
nightfall | Commercial | 71.9% | 5.9% | 82.1% |
fluiq (lite) | Fluiq | 68.8% | 5.9% | 80.0% |
regex-baseline | Control | 62.5% | 11.8% | 74.1% |
lakera-guard | Commercial | 46.9% | 5.9% | 62.5% |
presidio | Open source | 43.8% | 5.9% | 59.6% |
nemo-guardrails | Open source | 43.8% | 17.6% | 57.1% |
Recall is what it catches. False alarm is how often it blocks text that should have gone out. Neither means anything on its own, because a guardrail that blocks everything scores 100% recall and gets switched off in week two.
Every guardrail receives the same string and answers one question: block, or allow. Scoring was fixed before any contestant ran.
Adding a guardrail means implementing one method. If you think a product is misconfigured here, the fastest rebuttal is a pull request.
def scan(self, text: str) -> Verdict:
return Verdict(blocked=..., findings=[...])