Topic review
Arabic-English retrieval evaluation: evidence across query, source, and document direction
A regional release guide that evaluates Arabic, English, both cross-language directions, and mixed-language retrieval as separate evidence slices with explicit blocking decisions.
- Published by
- LangData Editorial, Editorial and architecture team
- Review owner
- Sourav Chandra, Co-Founder, LangData · Approved
- Published
- Updated
Swipe or scroll horizontally to inspect the full diagram.
“Arabic-English retrieval” is often reported as one capability and one score. That summary can conceal several different systems. An Arabic query retrieving Arabic sources depends on Arabic extraction, representation, terminology and ranking. An English query retrieving Arabic material adds cross-language matching and meaning transfer. A bilingual query over mixed pages creates another path, with script boundaries, names, numerals and document structure all interacting. A large English test set can make the headline result look stable while a required Arabic or cross-language route fails.
The release thesis is simple: Arabic, English, each cross-language direction, and mixed-language retrieval require separate test slices, separate error analysis, separate reviewers, and separate dispositions. A total score may be shown for exploration, but it cannot release a direction. The decision must identify the exact query-source path, corpus snapshot, configuration, expected sources or refusal, observed evidence, and blocking cases.
Regional official material makes AI, digital government, data, inclusion and human-centric governance visible public topics. The UAE AI Strategy, UAE Digital Government Strategy 2025, SDAIA About page, and Oman MTCIT AI sector page provide bounded public context in different jurisdictions. None supplies a retrieval test set, language threshold, model endorsement, or regional technical requirement. The evaluation below is an original method for customer-approved workflows.
Freeze the evaluation object before selecting metrics
Give the review a journey or use-case ID, release ID, corpus snapshot, test-set version, and evidence-pack version. State the task: policy lookup, case assistance, document discovery, form support, contact-centre knowledge, or another bounded action. Identify intended users, permissions, document families, source authority, required languages, and expected response form. State excluded scripts, dialect contexts, layouts, channels, and actions.
Record each component and configuration that can change retrieval: parser and OCR, layout reconstruction, language identification, normalization, metadata, chunking, embedding, lexical index, query expansion, translation if used, reranker, filters, prompt or orchestration, model, citation renderer, and policy integration. Preserve versions with every case. “Same model” is insufficient when the corpus or ranking configuration changed.
Use customer-approved evaluation material. Production records should not be copied into a test set merely because they are available. Record the source, permission, allowed use, reviewer, and retention instruction for test items. Synthetic or redacted cases can be useful, but label them so reviewers understand where they differ from operational material.
Validate extraction before retrieval
A retrieval failure can begin before the query. Arabic and bilingual documents may include right-to-left text, embedded left-to-right numbers, Arabic-Indic and European digits, tables, footnotes, headers, stamps, scanned pages, multi-column layouts, diacritics, ligatures, and mixed scripts. English documents can also fail through scans, tables, headings and version extraction. A text index cannot recover structure that the parser lost.
Build an extraction set by document family and layout. Compare approved pages with extracted text, reading order, tables, fields, page locators, language tags and metadata. Record missing regions, reordered blocks, merged columns, detached headings, uncertain fields and page-reference errors. Quarantine failed or ambiguous pages rather than allowing convenient fragments into the retrieval set.
Keep extraction status in the source register. Every expected retrieval case should point to an eligible source version and extraction version. If a source was not successfully prepared, classify the retrieval case as blocked or outside scope; do not count it as a ranking miss and then adjust a search threshold to compensate.
Slice S1: Arabic query to Arabic source
This slice tests Arabic retrieval without crossing the source language. Include queries that vary morphology, spelling, diacritics, hamza forms, Arabic-Indic and European digits, abbreviations, named entities, dates, domain terminology, and the approved dialect or register contexts relevant to the journey. Include short and long queries, ambiguous terms, unsupported questions, and source versions that should no longer be current.
Expected evidence should identify eligible Arabic source IDs and locators, not only a paraphrased answer. Native-language and domain reviewers should judge whether the retrieved passage addresses the intended meaning and whether a refusal case truly lacks approved support. Record error classes such as extraction, normalization, terminology, query ambiguity, ranking, authorization, citation, or source gap.
A passing S1 decision applies only to the reviewed material and cases. It does not establish general Arabic support, dialect coverage, or quality for another corpus.
Slice S2: English query to English source
English-to-English is a separate baseline, not the denominator that can outweigh other paths. Use the same task categories and consequence levels intended for the service. Test terminology variants, abbreviations, names, numerals, conflicting sources, withdrawn material, permission boundaries and “not found” cases.
The slice helps isolate architecture faults. If S2 and S1 both fail on the same document family, extraction or source metadata may be the cause. If S2 passes while S1 fails, language-specific representation, normalization, terminology or reviewer coverage becomes more likely. This comparison supports diagnosis, but each slice retains its own release result.
Slice S3: Arabic query to English source
Arabic-to-English retrieval adds a directional mapping. The system may use multilingual representations, query translation, terminology expansion, transliteration, or a combination. Record which path was executed for each case. A hidden translation step must have a version and evidence because changed negation, entity, date, unit or technical term can alter retrieval intent.
Cases should include Arabic questions whose authoritative support exists only in approved English material, bilingual terminology lists, names with multiple transliterations, acronyms, and questions that become ambiguous when translated. The expected record should state the intended English source and the meaning that must survive. Review both the Arabic query intent and the selected English passage; a high rank is not useful if the mapping changed the question.
Slice S4: English query to Arabic source
English-to-Arabic is not interchangeable with the reverse direction. English users may use transliterated Arabic names, translated organization terms, Romanized locations, English acronyms, or terminology that has several Arabic equivalents. The query expansion or representation can favor a different meaning than the customer expects.
Create independent cases and expected Arabic source locators. Native reviewers should inspect selected passages in context, while domain reviewers confirm that the English query was understood correctly. When the service returns an English answer from Arabic evidence, evaluate meaning transfer separately from retrieval. A correct Arabic passage can still be summarized incorrectly; do not assign that generation fault to ranking.
Slice S5: mixed-language query and mixed corpus
Mixed-language behavior deserves its own slice. Queries may code-switch mid-sentence, combine an English product name with Arabic instructions, use an acronym in one script and a name in another, or contain numbers with mixed directionality. Documents may present Arabic and English side by side, switch language by section, or repeat content with unequal version dates.
Record the actual query language spans and the language of each eligible source segment. Test whether chunking separates a heading from its corresponding text, whether duplicate bilingual sections compete, and whether metadata points to the authoritative representation. Include cases where only one language version is current and cases where users must see both versions to understand a term.
Do not collapse mixed behavior into S1 or S2 based on a dominant-language detector. The detector output is part of the observed evidence, not an unquestioned ground truth.
Separate retrieval, citation, and answer behavior
Evaluate stages independently. Extraction asks whether source text and structure were preserved. Retrieval asks whether eligible evidence was selected for the query and identity. Citation asks whether the user can resolve a claim to the correct authorized source version and locator. Answer review asks whether generated text remains within that evidence and preserves meaning.
Use case-level expected source IDs, acceptable alternatives, prohibited sources, and expected refusals. Record observed ranking and locators, but choose customer-reviewed measures and cutoffs based on task consequence. Possible evidence includes whether an expected source appears in the reviewed window, whether prohibited material appears, whether a citation resolves, and whether the system refuses when no eligible support exists. No metric alone decides release.
Authorization precedes language matching. A cross-language representation must retain document entitlement, source status and user identity. Test contrasting identities in every required slice. Query translation, multilingual embeddings, cached results and generated summaries must not widen the eligible corpus.
Diagnose by error class and release by slice
Review failed cases with a taxonomy that points to action: source authority, extraction, layout, language ID, normalization, terminology, transliteration, query mapping, translation, chunking, metadata, embedding, lexical matching, ranking, authorization, citation, generation, or test-data gap. One visible miss can have several contributing causes; link the correction and rerun the affected cases.
Do not hide blocked cases. A test may lack native review, approved material, source authority, a valid identity, an inspectable citation, or reproducible configuration. Mark it blocked and keep it in the slice summary. A required slice with a blocking authorization fault, unsupported material claim, unreviewed meaning transfer, or expired evidence cannot be released.
Exceptions are directional. If S3 is withheld, that does not automatically stop S1 or S2 when the journey can truthfully remove the cross-language path. The service UI, routing and monitoring must enforce the narrower scope. Every exception needs an owner, permitted cases, prohibited cases, evidence, expiry and retest trigger.
Complete text equivalent: directional evaluation board
These tables fully reproduce the SVG’s operational content. Complete the cover, one record per case, and an independent summary for all five slices.
| Evaluation cover | Required entry |
|---|---|
| Review identity | Journey or use-case ID, release ID, corpus snapshot, test-set version, evidence-pack version, environment and date |
| Scope | Users, tasks, permissions, document families, languages, layouts, channels and required slices; explicit exclusions |
| Executed configuration | Parser/OCR, language ID, normalization, metadata, chunking, lexical and vector retrieval, translation or expansion, reranker, filters, prompt/model, citation and policy versions |
| Test-material governance | Approved material references, source owners, allowed use, synthetic/redacted labels, access, retention and native/domain reviewers |
| Decision roles | Evaluation owner, Arabic reviewer, English/domain reviewer, security/authorization reviewer, service owner, exception owner and approver |
| Slice | Case prefix and required coverage | Independent blocking condition |
|---|---|---|
| S1 Arabic → Arabic | AR-AR; morphology, spelling, diacritics, numerals, names, terms, layouts, permissions, current/withdrawn sources and refusals | Required Arabic meaning or authorized evidence is not established by approved native/domain review |
| S2 English → English | EN-EN; task terms, names, numerals, conflicts, permissions, current/withdrawn sources and unsupported questions | Required English evidence, refusal, authorization or citation case fails or remains blocked |
| S3 Arabic → English | AR-EN; directional mapping, translation/expansion version, transliteration, entities, acronyms, ambiguity and meaning preservation | Arabic intent does not map to eligible English support or the directional transfer lacks reviewable evidence |
| S4 English → Arabic | EN-AR; translated terms, Romanized names, acronyms, alternative Arabic terms, context and any answer-language transfer | Eligible Arabic support or English-query intent is not established, or returned meaning is unreviewed |
| S5 Mixed → mixed | MX-MX; code-switching, script spans, names, numerals, bilingual pages, duplicate sections, unequal versions and permissions | Dominant-language routing hides a required span, wrong-version text wins, or mixed meaning lacks approved review |
| Per-case record | Required entry |
|---|---|
| Identity | Stable case ID and slice, task, approved query reference, expected language direction, permission identity and document-layout class |
| Source expectation | Eligible source IDs and versions, expected locators or refusal, acceptable alternatives, prohibited sources and authority state |
| Executed setup | Corpus snapshot; extraction, chunking, embedding, lexical, expansion/translation, ranking, policy, prompt and model versions |
| Observed evidence | Retrieved source IDs/ranks/locators, authorization result, citation result, answer-stage result if in scope, evidence URI, executor and date |
| Review and error | Arabic/native reviewer, English/domain reviewer, error class, observed meaning, reviewer date and unresolved question |
| Decision rule | Case threshold or blocking rule, pass / fail / blocked, affected slice and release consequence |
| Exception and recovery | Exception ID, permitted scope, owner, expiry, slice disable or rollback, corrected configuration and retest trigger |
| Slice decision summary | Required entry |
|---|---|
| Disposition | Separate release, release-with-conditions or stop result for S1, S2, S3, S4 and S5 |
| Reconciliation | Pass, fail and blocked counts per slice; every blocking case and accepted exception ID |
| Enforcement | UI/routing restrictions for withheld slices, source and permission enforcement, monitoring, slice-disable action and incident owner |
| Approval | Release ID, accepted corpus/configuration, decision owner, approver, decision date and named language/domain reviewers |
| Revalidation | Evidence expiry and triggers for source, extraction, terminology, translation, retrieval, model, policy, interface or reviewer change |
Publish a directional decision, not a multilingual claim
The final release note should say which slices, document families, tasks, users, corpus snapshot and configuration were accepted, with failures and conditions visible. It may release Arabic-to-Arabic while withholding one cross-language direction. It may allow mixed queries for one named journey but not another. That precision is a feature, not an embarrassment.
No official strategy proves retrieval quality. No public benchmark proves behavior on customer documents. The defensible outcome is a set of versioned slice decisions supported by approved material, native and domain review, authorization tests, resolvable evidence, explicit exceptions, and a mechanism to disable or retest a direction when its sources or components change.
Annotated primary sources
What the public material supports—and what it does not
These official pages provide bounded public context. They do not prescribe this design, decide a customer’s obligations, or establish any LangData affiliation or conformity.
-
1.UAE Government — UAE Strategy for Artificial Intelligence
Supports: Official UAE public context for artificial intelligence in government and multiple sectors, including an integrated digital-system framing.
Boundary: The strategy does not define Arabic-English retrieval evaluation, certify a model, prove language quality, provide customer test material, prescribe this method, or establish LangData participation.
-
2.UAE Government — Digital Government Strategy 2025
Supports: Official public context for inclusive, resilient, user-driven, digital-by-design, data-driven, open, and proactive digital-government dimensions.
Boundary: The page does not establish that a retrieval workflow satisfies those dimensions, is accessible, covers Arabic or English, meets a current requirement, or has official approval.
-
3.Saudi Data & AI Authority — About SDAIA
Supports: Official Saudi institutional context for national data and AI direction, data governance, capabilities, operation, research, and innovation.
Boundary: It provides no benchmark, corpus, language-direction threshold, customer permission, service result, regional mandate, model endorsement, or evidence of LangData affiliation.
-
4.Oman MTCIT — Artificial Intelligence sector
Supports: Official Oman context for an AI and advanced digital technologies program, adoption, capability, infrastructure, research, and human-centric governance themes.
Boundary: Program context does not establish retrieval behavior, Arabic or English quality, implementation access, evaluation acceptance, customer outcomes, or an architecture required across the region.
Review the operating boundary before choosing the stack.
Bring one workflow, its approved source inventory, and the people who own authorization, review, and incidents.
Discuss the architecture review