Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertili — Danlu Chen, Ka Sing He, Jiahe Tian
View PDF
Abstract:The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertility Ratio (F), Retrieval Proxy (R), Pre-training Exposure (E), and Corpus Diversity (D) and serve as dataset-intrinsic metrics to contextualize reported scores. These metrics reveal that a significant portion of result variability is explained by train-test overlap and pre-training exposure rather than model capability. Additionally, we identify that some languages -- particularly extinct and non-Latin indigenous languages -- suffer from poor tokenization coverage (high token fertility), highlighting a fundamental limitation of transferring models from high-resource languages that lack a shared vocabulary. By providing these indices alongside performance scores, we enable more transparent evaluation of cross-lingual transfer and provide a more reliable foundation for the XLR MT community.
Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2603.25222 [cs.CL]
(or arXiv:2603.25222v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2603.25222
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Danlu Chen [view email] [v1] Thu, 26 Mar 2026 09:20:17 UTC (6,934 KB)
Sign in to highlight and annotate this article

Conversation starters
Daily AI Digest
Get the top 5 AI stories delivered to your inbox every morning.
More about
researchpaperarxiv7 million euros available for research into AI applications in the agriculture, horticulture, water, and food sectors - nwo.nl
<a href="https://news.google.com/rss/articles/CBMi2AFBVV95cUxOZkhpc1NkQkZacGs4RmI0a0RRa2RsdmdVbzF6Z0RLS2xZS3J5eWtVR282LUFFcURWM1ltdFl0UTZ6RVNSM2h4bnVlbVdsREVpX3YzdHR4X2VoRkxKQUp2R1ZXQXJNbU1Cblh5QlJDc2RIbUVBSlhJdnhQRlB5QUtoVG5oaDc1T0l3TFZyQnhEYS0xXzdWZVZkV1E2Sk0tYzVLeUc4bUV2Z25pX1R3am1kWHctaFhnOFlaR2tZQzRPMmxWS2tKY2djZFEyNHFxbHZZM1lpQ1o3aS0?oc=5" target="_blank">7 million euros available for research into AI applications in the agriculture, horticulture, water, and food sectors</a> <font color="#6f6f6f">nwo.nl</font>

Europe urged to ‘learn to fight for itself’ in case US-China truce collapses
European governments breathed a sigh of relief in October when the US and China sealed a fragile trade truce that paused more sweeping Chinese rare earth restrictions and papered over a Sino-Dutch row over chipmaker Nexperia. Now, however, the European Union is being urged to come up with a battle plan should the ceasefire fail or expire. A spike in superpower tensions could expose the EU to Chinese export controls, potentially pulverising its military support for Ukraine, its own efforts to...
[D] TurboQuant author replies on OpenReview
<!-- SC_OFF --><div class="md"><p>I wanted to follow up to <a href="https://www.reddit.com/r/MachineLearning/comments/1s7m7rn/comment/odaect4/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button">yesterday's thread</a> and see if anyone wanted to weigh in on it. This work is far outside of my niche, but it strikes me as an attempt to reframe the issue instead of addressing concerns head on. </p> <p>OpenReview link for reference: <a href="https://openreview.net/forum?id=tO3ASKZlok">https://openreview.net/forum?id=tO3ASKZlok</a></p> <blockquote> <p>In response to recent commentary regarding our paper, "TurboQuant," we provide the following technical clarifications to correct the record.</p> <p>TurboQuant did not derive its core method from
Knowledge Map
Connected Articles — Knowledge Graph
This article is connected to other articles through shared AI topics and tags.
More in Research Papers
[D] TurboQuant author replies on OpenReview
<!-- SC_OFF --><div class="md"><p>I wanted to follow up to <a href="https://www.reddit.com/r/MachineLearning/comments/1s7m7rn/comment/odaect4/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button">yesterday's thread</a> and see if anyone wanted to weigh in on it. This work is far outside of my niche, but it strikes me as an attempt to reframe the issue instead of addressing concerns head on. </p> <p>OpenReview link for reference: <a href="https://openreview.net/forum?id=tO3ASKZlok">https://openreview.net/forum?id=tO3ASKZlok</a></p> <blockquote> <p>In response to recent commentary regarding our paper, "TurboQuant," we provide the following technical clarifications to correct the record.</p> <p>TurboQuant did not derive its core method from

Carnegie Mellon Researchers Rethink Chronic Pain
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/260123D_WTM_Yttri_Lab_RD002.jpg.webp?itok=DozEPS-g" width="900" height="508" alt="Eric Yttri"> </p> Across neuroscience, biomedical engineering and artificial intelligence, researchers from Carnegie Mellon University are exploring how pain is measured, understood and treated to support safer, more effective care.

Five CMU Faculty Members Named 2026 Sloan Research Fellows
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/sloan-collage-2000%20copy.jpg.webp?itok=Ih-KH3Na" width="900" height="508" alt="2026 Sloan Awardees"> </p> Five Carnegie Mellon University faculty members are among the 126 recipients of 2026 Sloan Research Fellowships, which honor early career scholars whose achievements put them among the best scientific minds working today.
Discussion
Sign in to join the discussion
No comments yet — be the first to share your thoughts!