Your RAG system looks healthy.
Until someone measures it.

Independent measurement of RAG and LLM systems — retrieval accuracy, hallucination, PII leaks, prompt injection, drift. On your corpus, in your languages, with numbers you can put in front of a risk committee. The tooling is open source; the reading is independent of whoever sold you the model.

By El Mahdi El Aimani — engineer, Rabat & Casablanca. Production retrieval, vision and speech systems in Moroccan industry and banking.

Apache 2.0 CI passing Python 3.10–3.12 Runs on your machines FR · EN · AR · Darija
bench/run_bench.py — actual output, not a mock-up
$
corpus: 157 chunks · questions: 24 x 5 languages
embedding corpus .......... done
Hit@1 Hit@3 Hit@5 MRR gold sim
French 83.3% 87.5% 91.7% 0.869 0.694
English 79.2% 87.5% 91.7% 0.855 0.664
Arabic (MSA) 87.5% 95.8% 95.8% 0.921 0.653
Darija (Arabic) 54.2% 79.2% 79.2% 0.680 0.550
Darija (Latin) 16.7% 50.0% 62.5% 0.354 0.475
 
Darija (Latin) vs French: Hit@1 -66.6 pts, Hit@3 -37.5 pts, MRR -0.515
The problem

Three questions, and no number in the room

Your tests pass. Your demo works. Then someone senior asks whether the system actually works, whether it leaks, and whether you can defend it — and the honest answer is that nobody measured.

Does it work?

The model answers confidently from context it never had, and retrieval quietly misses in the languages your users actually type. Faithfulness drops. Nothing logs it as an error.

Does it leak?

An answer carries a customer's IBAN. A retrieved document carries an instruction the model obeys. A log line holds a secret. None of it trips an alert.

Can you defend it?

Nobody deployed anything, yet production drifted from what was validated. When the risk committee, a European client or an auditor asks for evidence, a vendor dashboard screenshot is not it.

Independence

A vendor cannot audit itself

Model providers now ship their own evaluation and guardrails. Useful, and not the same thing: nobody grades their own exam. Every claim below is checkable in the open-source repository before you talk to me.

Independent of the model you bought

The engine is Apache 2.0 and runs on your machines, against whichever provider you use — hosted, self-hosted, or several at once compared on the same questions. The reading does not come from the party that sold you the system.

One threshold, two places

The same servexguard.yaml gates your CI and drives the production score. What you test and what you monitor cannot drift apart, because they are the same numbers.

Measured in FR, AR and Darija

Not asserted — measured, on a real Moroccan banking corpus, and published below. The result was not flattering, which is the point of measuring.

To be clear about what this is not: ServeX Guard does not block attacks in real time. It runs beside your pipeline, not in front of your model. If you need a runtime firewall for LLM traffic, buy one, and keep this for proving that what you shipped still behaves the way you validated it.

Evidence

What a measurement actually looks like

A retrieval system I built over public Moroccan banking pages — 157 passages, 24 questions, each asked five ways. Same question, same expected passage; only the language changes. To be precise: this measures a retrieval stack on public content. It is not an audit of any bank's system.

Language the user typed in24 questions, 157 passages Hit@1right passage, first result Hit@3 MRR Similarity to expectedcosine
French83.3%87.5%0.8690.694
English79.2%87.5%0.8550.664
Arabic (MSA)87.5%95.8%0.9210.653
Darija, Arabic script54.2%79.2%0.6800.550
Darija, Latin script16.7%50.0%0.3540.475

The headline finding

A question that lands at rank 1 in French lands at rank 65 of 157 when the same customer types it in Latin-script Darija. Nothing in the product logs that as an error. The user simply gets a wrong answer and leaves.

The uncomfortable part

French alone is 83.3%. Four questions in twenty-four miss at rank 1, three of them buried at ranks 5, 10 and 15 — in the language the system was built for. "We only support French and English" does not make the problem go away.

Why publish a bad result

Because it is the result. A measurement that only ever confirms what you hoped is not a measurement. This one was run twice, on a doubled question set, and the gap widened.

Full method, question set, per-question ranks and the failure cases.

How it works

One command. Runs in CI. Fully offline by default.

1

Install

One pip install. No account required, no data leaves your machine.

2

Point it at your golden set

A JSONL dataset of inputs, expected outputs and, for RAG systems, retrieved context.

3

Gate your deploy

Non-zero exit on failure blocks the merge, like a linter for RAG quality.

4

Keep the history, if you want it

Upload is always opt-in. The --upload flag can send the finished report to a hosted dashboard — available on request, not sold self-serve. Without the flag nothing leaves your machine, and an upload failure never changes your CI exit code.

# install
pip install servex-guard

# run a quality gate in CI
servexguard check \
  --dataset golden.jsonl \
  --min-faithfulness 0.85

# opt in to the dashboard
export SERVEXGUARD_API_KEY=sxg_xxxxx
servexguard check \
  --dataset golden.jsonl \
  --upload --project my-rag
Honest comparison

Why not just use what is free?

You should. I do. RAGAS, Presidio, promptfoo, DeepEval and self-hosted Langfuse are open source and good; Braintrust sells CI quality gates; OpenAI ships hallucination detection and guardrails inside its own SDK at no charge. This CLI calls the open-source pieces rather than replacing them.

What free tooling gives you

Metrics. Faithfulness, relevancy, recall, PII hits, injection flags, traces. All of it runnable this afternoon by an engineer who reads the docs. None of it is the hard part.

What a vendor's own tooling cannot give you

An independent reading of that vendor. OpenAI's guardrails evaluate OpenAI's output. Useful in production; not evidence a risk committee, a European client or an ISO/IEC 42001 auditor will accept as third-party.

What none of them give you

A question set built from your corpus, in the languages your users actually type, run across the providers you actually use, read by someone who is not selling you any of them — and who puts a name on the conclusion.

If you have an engineer with two spare months and no compliance deadline, build it yourself from the pieces above. The honest pitch is that most teams do not, and that what nobody budgets for is not the detection but the reading: what the numbers mean on this system, what to fix first, and a document someone outside the team can rely on.

Work with me

The tooling is free. The reading is the work.

The CLI is Apache 2.0 and yours to run without talking to anyone. What teams actually ask for is someone to point it at their system, interpret what comes back, and put a name on the conclusion.

01
Measurement
About three weeks
  • Retrieval and answer quality on your own corpus, per language
  • Hallucination and refusal rates on a question set built with your team
  • Cost per request and real latency
  • A prioritised, costed correction plan

Roughly three weeks. Ends in a report you can hand to someone who was not in the meetings.

03
Governance readiness
Six to ten weeks
  • Inventory and risk classification of your AI systems
  • ISO/IEC 42001 gap analysis against Annex A
  • EU AI Act positioning and evidence you will be asked for
  • Law 09-08 / CNDP register and retention

Preparation and gap analysis. Certification itself is issued by an accredited body — never by me.

04
Retainer
Four days a month
  • Drift watched on real production samples
  • Evaluation reviewed before every release
  • On call for model incidents
  • Quarterly compliance review

For a team that has understood an AI system is watched like infrastructure. Someone who answers for it, not a dashboard.

Not a certification body, not a law firm, and not a runtime firewall. I am El Mahdi El Aimani, an engineer who measures systems and puts his name on the number.

It starts with thirty minutes on a call, at no charge, to see whether there is something worth measuring. In French, English or Arabic.