Research Papers research paper arxiv nlp language-models

A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations

arXivMarch 26, 202610 min read0 views

Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic and comprehensive compilation of the dialectal data currently available in Basque. Two types of data sources have been distinguished: online data originally written in some dialect, and standard-to-dialect adapted data. The former includes all dialectal data that can be found online, such as news and radio sites, informal tweets, as well as online resources such as dictionaries — Jaione Bengoetxea, Itziar Gonzalez-Dios, Rodrigo Agerri

View PDF HTML (experimental)

Abstract:Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic and comprehensive compilation of the dialectal data currently available in Basque. Two types of data sources have been distinguished: online data originally written in some dialect, and standard-to-dialect adapted data. The former includes all dialectal data that can be found online, such as news and radio sites, informal tweets, as well as online resources such as dictionaries, atlases, grammar rules, or videos. The latter consists of data that has been adapted from the standard variety to dialectal varieties, either manually or automatically. Regarding the manual adaptation, the test split of the XNLI Natural Language Inference dataset was manually adapted into three Basque dialects: Western, Central, and Navarrese-Lapurdian, yielding a high-quality parallel gold standard evaluation dataset. With respect to the automatic dialectal adaptation, the automatically adapted physical commonsense dataset (BasPhyCowest) underwent additional manual evaluation by native speakers to assess its quality and determine whether it could serve as a viable substitute for full manual adaptation (i.e., silver data creation).

Subjects:

Computation and Language (cs.CL)

Cite as: arXiv:2603.25189 [cs.CL]

(or arXiv:2603.25189v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2603.25189

arXiv-issued DOI via DataCite

Submission history

From: Jaione Bengoetxea [view email] [v1] Thu, 26 Mar 2026 08:55:23 UTC (2,348 KB)

Original source

arXiv

https://arxiv.org/abs/2603.25189v1

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Research Papers

Humboldt Fellow from the US conducts research in robotics to one day harvest energy from ocean waves

is.mpg.de

1m8 months ago

Models

Howard University and Google Research Enhance A.I. Speech Recognition of African American English - The Dig at Howard University

<a href="https://news.google.com/rss/articles/CBMiygFBVV95cUxQRTh4T2h6cVRsdEF2cjlkWGQyT2tWZnVTTmh4czBJV3ZpSmd1T1Z2eG5Ld1dvQWhNckpjRDItVEtiZ2hMdjBVLWJ0b0xTY0pieG82U0VibXFBLWVUN0tlQ3J1dzBFa2ZBekF1YXJPZlpHNGtkOWZjdWFCSlVTQTctcTNvcURtOER4MnhnYk1BQUt4WllmekE4WkVERTA4Wi1VcnFCY2xYSml6ak9GM1o1NmI0VWtXb2xERlVZVFNBTTQyQ1FBWThESk53?oc=5" target="_blank">Howard University and Google Research Enhance A.I. Speech Recognition of African American English</a> The Dig at Howard University

GNews AI voice

1m9 months ago

Products

Speech-to-Retrieval (S2R): A new approach to voice search - research.google

<a href="https://news.google.com/rss/articles/CBMijAFBVV95cUxQekN0T0VkREpJVGk0U25zMVcyX0VYV0V4eVRJY2ozVW02ampCVXFMRDJybk56blpMdWVhdkRsWWI2S19JemlYM3dHd2dBSkx0SWxtNnNfN18zcjBKLWVXN3JZUnVFdndndTBnSVlVSGhVdWwyS1V3TkRCSUJ5SnRkYXJBV1NfZWUwa3ByWA?oc=5" target="_blank">Speech-to-Retrieval (S2R): A new approach to voice search</a> research.google

GNews AI voice

1m6 months ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 97 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research Papers

Humboldt Fellow from the US conducts research in robotics to one day harvest energy from ocean waves

is.mpg.de

1m8 months ago

Research Papers

AI-driven digital manipulation ‘tested’ Dutch election integrity, researchers warn - EUobserver

<a href="https://news.google.com/rss/articles/CBMirwFBVV95cUxQcERTcUc5ZndxZ054endXTXNwTlhtYjRyLXBHWVJmRXloNV9JUUpFZnBrLUdDeUpSNklZRFJuUXl0bThIT2ZzbFd6ZU02TW9yaXBPbHducUlHaXVUbWprS0pla0JENkxpSkZfWW9vdTRvcjIzc2ZzWGF6ZmJPMXRVRkFnNmp5NWpLZTBIRk9LamF2RUtkdnQ2bFJXRVZMdVkxZWNHVUl1SzZZeE1JT3R3?oc=5" target="_blank">AI-driven digital manipulation ‘tested’ Dutch election integrity, researchers warn</a> EUobserver

GNews AI Netherlands

1m2 months ago

Research PapersLive

Why Drug Toxicity Can’t Be Predicted in Isolation — Building EIRION with Graph Neural Networks

How we built a graph neural network that finally sees the whole play — not just the audition Every year, drugs that passed early safety tests go on to harm people in ways nobody predicted. Not because the chemistry was wrong. Not because the researchers were careless. But because we kept evaluating drugs the way a talent agent judges an actor from a solo audition tape. Isolated. Out of context. No script. No co-stars. No stage. In real theatre, a performance is never just about one actor. It depends on who they share the stage with, which scene they appear in, what the story demands at that moment. A brilliant performer in the wrong play, surrounded by the wrong cast, in the wrong context — can still wreck the whole production. That is exactly how drug toxicity works. And that is exactly t

Towards AI

17mabout 1 hour ago

Research PapersLive

It's Not Smarter Models — It's Cheaper Memory: TurboQuant's Real Impact, Wall Street Panic & Academic Storm

<blockquote> One-line summary: TurboQuant is a genuinely important engineering breakthrough — but Google's marketing, academic ethics controversy, and Wall Street's overreaction made the story far more dramatic than the technology itself. </blockquote> <h2> 0. What This Article Answers </h2> Google Research published TurboQuant at ICLR 2026 (<a href="https://arxiv.org/abs/2504.19874" rel="noopener noreferrer">arXiv 2504.19874</a>), claiming 6x memory compression, 8x speedup, and zero accuracy loss for LLM KV caches. Then, in the same week: <ol> <li>Global memory stocks lost over $90 billion in market cap</li> <li>An ETH Zürich researcher publicly accused the paper of academic plagiarism and experimental fraud </li> <li

DEV Community

10mabout 1 hour ago

A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations

Submission history

Daily AI Digest

More about

Humboldt Fellow from the US conducts research in robotics to one day harvest energy from ocean waves

Howard University and Google Research Enhance A.I. Speech Recognition of African American English - The Dig at Howard University

​​Speech-to-Retrieval (S2R): A new approach to voice search - research.google

Knowledge Map

Connected Articles — Knowledge Graph

Discussion

More in Research Papers

Humboldt Fellow from the US conducts research in robotics to one day harvest energy from ocean waves

AI-driven digital manipulation ‘tested’ Dutch election integrity, researchers warn - EUobserver

Why Drug Toxicity Can’t Be Predicted in Isolation — Building EIRION with Graph Neural Networks

It's Not Smarter Models — It's Cheaper Memory: TurboQuant's Real Impact, Wall Street Panic & Academic Storm

Speech-to-Retrieval (S2R): A new approach to voice search - research.google