Retrieval measurement 27 August 2026 El Mahdi El Aimani

Five languages, one banking corpus

Twenty-four questions asked identically in French, English, Modern Standard Arabic and Darija, against 157 French passages from a Moroccan bank's public site. French, English and Standard Arabic hold between 79% and 88%. Darija typed in Latin script falls to 17%.

Version française

01 What was measured

The protocol fits in one sentence: same question, same expected passage, only the language changes.

A Moroccan bank publishes its documentation in French. Its customers write in French, in English, in Standard Arabic, and in Darija — which most of them type in Latin characters. The question is whether the system still finds the right document when the language of the question drifts away from the language of the documents.

Corpus
157 unique passages, extracted from 40 public pages of bankofafrica.ma (French). The content serves as a test corpus; what is measured is a retrieval stack built for the exercise, not that bank's system nor anyone else's.
Questions
24, each written in the 5 forms, tied to the same expected passage
Topics
account opening, consumer credit, mortgage, cards, limits, remote banking, mobile payment
Embedding
BAAI/bge-m3, 1024 dimensions
Retrieval
cosine similarity over the whole corpus, no reranker, no query rewriting

The content is public: no personal data, no authenticated access, no production system was touched.

02 Results

Hit@1: does the right document come first. MRR: mean reciprocal rank, 1.0 means always first.

Language of the question Hit@1 Hit@3 Hit@5 MRR Similarity vs FR
French reference 83.3% 87.5%91.7%0.8690.694 —
English 79.2% 87.5%91.7%0.8550.664 −4.2
Standard Arabic 87.5% 95.8%95.8%0.9210.653 +4.2
Darija Arabic script 54.2% 79.2%79.2%0.6800.550 −29.2
Darija Latin script 16.7% 50.0%62.5%0.3540.475 −66.7

Language is not a problem until Darija. French, English and Standard Arabic sit within a hair of each other: over 24 questions, 4.2 points is a single question. An assistant advertised as "French and English" has no significant multilingual gap.

The gap appears with Darija, and doubles when it is written in Latin characters — wach ne9der nhel 7sab, with digits standing in for Arabic sounds that have no Latin equivalent. That is the form customers actually type into a chat window.

03 What a rank of 65 means

The corpus holds 157 passages. The position of the right document decides what the model receives.

French — "Est-ce qu'il y a des frais pour ouvrir un compte ?" rank 1 / 157
Latin-script Darija — "wach kaynin chi masarif bach nhel hsab?" rank 65 / 157
Same question, same expected passage. No production system passes 65 passages to the model: at that depth the document is not ranked badly, it is out of reach.

The other gaps follow the same pattern: the consumer-credit rate goes from rank 1 to rank 45, the repayment deferral from rank 5 to rank 47, cardless withdrawal from rank 15 to rank 77.

None of these failures surface anywhere. The system does not go down, latency does not move, no error is logged. The model receives an irrelevant document, retrieved with confidence, and writes an answer on top of it.

04 French alone is not at 100%

The number that concerns everyone, whatever languages are supported.

On French questions against French documents, 4 out of 24 do not put the right passage first, and three bury it beyond rank 5: cardless counter withdrawal at rank 15, credit response time at rank 10, repayment deferral at rank 5.

And this is the favourable case: short, keyword-rich marketing content, with questions written to be answerable. A base of internal procedures, longer and more specialised, behaves worse.

05 What can be fixed, and what cannot

The two gaps do not have the same cause, so they do not have the same remedy.

Darija in Arabic script: tuning

The gap at Hit@3 is only 8.3 points, against 29.2 at Hit@1. The right document is often present, just not first. Passing the top 3 to 5 passages to the model instead of the first alone recovers most of the loss.

Darija in Latin script: normalisation

Here the same tuning is not enough: the gap stays at 37.5 points at Hit@3 and 29.2 at Hit@5. Transliteration is not an approximate translation. masarif is neither a French word nor an Arabic string: the tokeniser cuts it into fragments that mean nothing in any language seen at training time. The lead is to normalise the script before embedding, which is development work, not configuration.

06 Why these numbers are a floor

Three reasons to expect a real system to do worse, not better.

  1. bge-m3 is among the best multilingual models available, and it was chosen for exactly that property. A system built on text-embedding-3 or on a French-specialised model should do worse.
  2. The corpus is public marketing content: short, repetitive, keyword-rich. Real internal procedures widen the gap rather than narrow it.
  3. No reranker, no query rewriting, no hybrid BM25. A production system may have them, which would improve the result. It may also not.

07 Limits

What this measurement does not show, and does not claim to show.

08 Reproduce

The corpus is rebuilt from the public pages; the measurement then replays offline.

# rebuild the corpus from the public pages
python build_corpus.py     → 40 pages, 157 passages

# embed and score the 24 questions x 5 forms
python run_bench.py        → Hit@1, Hit@3, Hit@5, MRR, similarity

Vectors are cached locally: the first run costs about 180 embedding calls, later runs are close to free. The measurement script raises when an embedding fails, instead of returning a zero vector that would yield a clean, false score.

The question set and scripts are available on request.