Independent measurement of RAG and LLM systems — retrieval accuracy, hallucination, PII leaks, prompt injection, drift. On your corpus, in your languages, with numbers you can put in front of a risk committee. The tooling is open source; the reading is independent of whoever sold you the model.
By El Mahdi El Aimani — engineer, Rabat & Casablanca. Production retrieval, vision and speech systems in Moroccan industry and banking.
Your tests pass. Your demo works. Then someone senior asks whether the system actually works, whether it leaks, and whether you can defend it — and the honest answer is that nobody measured.
The model answers confidently from context it never had, and retrieval quietly misses in the languages your users actually type. Faithfulness drops. Nothing logs it as an error.
An answer carries a customer's IBAN. A retrieved document carries an instruction the model obeys. A log line holds a secret. None of it trips an alert.
Nobody deployed anything, yet production drifted from what was validated. When the risk committee, a European client or an auditor asks for evidence, a vendor dashboard screenshot is not it.
Model providers now ship their own evaluation and guardrails. Useful, and not the same thing: nobody grades their own exam. Every claim below is checkable in the open-source repository before you talk to me.
The engine is Apache 2.0 and runs on your machines, against whichever provider you use — hosted, self-hosted, or several at once compared on the same questions. The reading does not come from the party that sold you the system.
The same servexguard.yaml gates your CI and drives the production score. What you test and what you monitor cannot drift apart, because they are the same numbers.
Not asserted — measured, on a real Moroccan banking corpus, and published below. The result was not flattering, which is the point of measuring.
To be clear about what this is not: ServeX Guard does not block attacks in real time. It runs beside your pipeline, not in front of your model. If you need a runtime firewall for LLM traffic, buy one, and keep this for proving that what you shipped still behaves the way you validated it.
A retrieval system I built over public Moroccan banking pages — 157 passages, 24 questions, each asked five ways. Same question, same expected passage; only the language changes. To be precise: this measures a retrieval stack on public content. It is not an audit of any bank's system.
| Language the user typed in24 questions, 157 passages | Hit@1right passage, first result | Hit@3 | MRR | Similarity to expectedcosine |
|---|---|---|---|---|
| French | 83.3% | 87.5% | 0.869 | 0.694 |
| English | 79.2% | 87.5% | 0.855 | 0.664 |
| Arabic (MSA) | 87.5% | 95.8% | 0.921 | 0.653 |
| Darija, Arabic script | 54.2% | 79.2% | 0.680 | 0.550 |
| Darija, Latin script | 16.7% | 50.0% | 0.354 | 0.475 |
A question that lands at rank 1 in French lands at rank 65 of 157 when the same customer types it in Latin-script Darija. Nothing in the product logs that as an error. The user simply gets a wrong answer and leaves.
French alone is 83.3%. Four questions in twenty-four miss at rank 1, three of them buried at ranks 5, 10 and 15 — in the language the system was built for. "We only support French and English" does not make the problem go away.
Because it is the result. A measurement that only ever confirms what you hoped is not a measurement. This one was run twice, on a doubled question set, and the gap widened.
Full method, question set, per-question ranks and the failure cases.
One pip install. No account required, no data leaves your machine.
A JSONL dataset of inputs, expected outputs and, for RAG systems, retrieved context.
Non-zero exit on failure blocks the merge, like a linter for RAG quality.
Upload is always opt-in. The --upload flag can send the finished report to a hosted dashboard — available on request, not sold self-serve. Without the flag nothing leaves your machine, and an upload failure never changes your CI exit code.
# install pip install servex-guard # run a quality gate in CI servexguard check \ --dataset golden.jsonl \ --min-faithfulness 0.85 # opt in to the dashboard export SERVEXGUARD_API_KEY=sxg_xxxxx servexguard check \ --dataset golden.jsonl \ --upload --project my-rag
You should. I do. RAGAS, Presidio, promptfoo, DeepEval and self-hosted Langfuse are open source and good; Braintrust sells CI quality gates; OpenAI ships hallucination detection and guardrails inside its own SDK at no charge. This CLI calls the open-source pieces rather than replacing them.
Metrics. Faithfulness, relevancy, recall, PII hits, injection flags, traces. All of it runnable this afternoon by an engineer who reads the docs. None of it is the hard part.
An independent reading of that vendor. OpenAI's guardrails evaluate OpenAI's output. Useful in production; not evidence a risk committee, a European client or an ISO/IEC 42001 auditor will accept as third-party.
A question set built from your corpus, in the languages your users actually type, run across the providers you actually use, read by someone who is not selling you any of them — and who puts a name on the conclusion.
If you have an engineer with two spare months and no compliance deadline, build it yourself from the pieces above. The honest pitch is that most teams do not, and that what nobody budgets for is not the detection but the reading: what the numbers mean on this system, what to fix first, and a document someone outside the team can rely on.
The CLI is Apache 2.0 and yours to run without talking to anyone. What teams actually ask for is someone to point it at their system, interpret what comes back, and put a name on the conclusion.
Roughly three weeks. Ends in a report you can hand to someone who was not in the meetings.
Runs against your environment, not a copy of it. Findings are reproducible or they are not reported.
Preparation and gap analysis. Certification itself is issued by an accredited body — never by me.
For a team that has understood an AI system is watched like infrastructure. Someone who answers for it, not a dashboard.
Not a certification body, not a law firm, and not a runtime firewall. I am El Mahdi El Aimani, an engineer who measures systems and puts his name on the number.
It starts with thirty minutes on a call, at no charge, to see whether there is something worth measuring. In French, English or Arabic.