Twenty-four questions asked identically in French, English, Modern Standard Arabic and Darija, against 157 French passages from a Moroccan bank's public site. French, English and Standard Arabic hold between 79% and 88%. Darija typed in Latin script falls to 17%.
The protocol fits in one sentence: same question, same expected passage, only the language changes.
A Moroccan bank publishes its documentation in French. Its customers write in French, in English, in Standard Arabic, and in Darija — which most of them type in Latin characters. The question is whether the system still finds the right document when the language of the question drifts away from the language of the documents.
The content is public: no personal data, no authenticated access, no production system was touched.
Hit@1: does the right document come first. MRR: mean reciprocal rank, 1.0 means always first.
| Language of the question | Hit@1 | Hit@3 | Hit@5 | MRR | Similarity | vs FR |
|---|---|---|---|---|---|---|
| French reference | 83.3% | 87.5% | 91.7% | 0.869 | 0.694 | — |
| English | 79.2% | 87.5% | 91.7% | 0.855 | 0.664 | −4.2 |
| Standard Arabic | 87.5% | 95.8% | 95.8% | 0.921 | 0.653 | +4.2 |
| Darija Arabic script | 54.2% | 79.2% | 79.2% | 0.680 | 0.550 | −29.2 |
| Darija Latin script | 16.7% | 50.0% | 62.5% | 0.354 | 0.475 | −66.7 |
Language is not a problem until Darija. French, English and Standard Arabic sit within a hair of each other: over 24 questions, 4.2 points is a single question. An assistant advertised as "French and English" has no significant multilingual gap.
The gap appears with Darija, and doubles when it is written in Latin characters — wach ne9der nhel 7sab, with digits standing in for Arabic sounds that have no Latin equivalent. That is the form customers actually type into a chat window.
The corpus holds 157 passages. The position of the right document decides what the model receives.
The other gaps follow the same pattern: the consumer-credit rate goes from rank 1 to rank 45, the repayment deferral from rank 5 to rank 47, cardless withdrawal from rank 15 to rank 77.
None of these failures surface anywhere. The system does not go down, latency does not move, no error is logged. The model receives an irrelevant document, retrieved with confidence, and writes an answer on top of it.
The number that concerns everyone, whatever languages are supported.
On French questions against French documents, 4 out of 24 do not put the right passage first, and three bury it beyond rank 5: cardless counter withdrawal at rank 15, credit response time at rank 10, repayment deferral at rank 5.
And this is the favourable case: short, keyword-rich marketing content, with questions written to be answerable. A base of internal procedures, longer and more specialised, behaves worse.
The two gaps do not have the same cause, so they do not have the same remedy.
The gap at Hit@3 is only 8.3 points, against 29.2 at Hit@1. The right document is often present, just not first. Passing the top 3 to 5 passages to the model instead of the first alone recovers most of the loss.
Here the same tuning is not enough: the gap stays at 37.5 points at Hit@3 and 29.2 at Hit@5. Transliteration is not an approximate translation. masarif is neither a French word nor an Arabic string: the tokeniser cuts it into fragments that mean nothing in any language seen at training time. The lead is to normalise the script before embedding, which is development work, not configuration.
Three reasons to expect a real system to do worse, not better.
What this measurement does not show, and does not claim to show.
The corpus is rebuilt from the public pages; the measurement then replays offline.
# rebuild the corpus from the public pages python build_corpus.py → 40 pages, 157 passages # embed and score the 24 questions x 5 forms python run_bench.py → Hit@1, Hit@3, Hit@5, MRR, similarity
Vectors are cached locally: the first run costs about 180 embedding calls, later runs are close to free. The measurement script raises when an embedding fails, instead of returning a zero vector that would yield a clean, false score.
The question set and scripts are available on request.