Research Papers research paper arxiv ai artificial-intelligence

AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation

arXivMarch 31, 202610 min read0 views

arXiv:2509.16952v2 Announce Type: replace-cross Abstract: The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are capable of automating question answering (QA) workflows for scientific papers, there still lacks a comprehensive and realistic benchmark to evaluate their capabilities. Moreover, training an interactive agent for this specific task is hindered by the shortage of high-quality interaction trajectories. In this work, we propose AirQA, a human-annotated comprehen — Tiancheng Huang, Ruisheng Cao, Yuxin Zhang, Zhangyi Kang, Zijian Wang, Chenrun Wang, Yijie Luo, Hang Zheng, Lirong Qian, Lu Chen, Kai Yu

Authors:Tiancheng Huang, Ruisheng Cao, Yuxin Zhang, Zhangyi Kang, Zijian Wang, Chenrun Wang, Yijie Luo, Hang Zheng, Lirong Qian, Lu Chen, Kai Yu

View PDF HTML (experimental)

Abstract:The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are capable of automating question answering (QA) workflows for scientific papers, there still lacks a comprehensive and realistic benchmark to evaluate their capabilities. Moreover, training an interactive agent for this specific task is hindered by the shortage of high-quality interaction trajectories. In this work, we propose AirQA, a human-annotated comprehensive paper QA dataset in the field of artificial intelligence (AI), with 13,956 papers and 1,246 questions, that encompasses multi-task, multi-modal and instance-level evaluation. Furthermore, we propose ExTrActor, an automated framework for instruction data synthesis. With three LLM-based agents, ExTrActor can perform example generation and trajectory collection without human intervention. Evaluations of multiple open-source and proprietary models show that most models underperform on AirQA, demonstrating the quality of our dataset. Extensive experiments confirm that ExTrActor consistently improves the multi-turn tool-use capability of small models, enabling them to achieve performance comparable to larger ones.

Comments: 29 page, 6 figures, 17 tables, accepted to ICLR 2026

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as: arXiv:2509.16952 [cs.CL]

(or arXiv:2509.16952v2 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2509.16952

arXiv-issued DOI via DataCite

Submission history

From: Tiancheng Huang [view email] [v1] Sun, 21 Sep 2025 07:24:17 UTC (2,855 KB) [v2] Mon, 30 Mar 2026 06:47:00 UTC (3,139 KB)

Original source

arXiv

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Research Papers

How a Nonprofit Transforms Data with Cloudera and AI

The organization developed data pipelines that extract and structure information from various scientific sources, significantly accelerating the research process.

AI Business

1m13 days ago

Research Papers

Scientists should use AI as a tool, not an oracle

How AI hype leads to flawed research that fuels more hype

AI Snake Oil

1malmost 2 years ago

Models

New paper: AI agents that matter

Rethinking AI agent benchmarking and evaluation

AI Snake Oil

1mover 1 year ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 163 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research Papers

How a Nonprofit Transforms Data with Cloudera and AI

The organization developed data pipelines that extract and structure information from various scientific sources, significantly accelerating the research process.

AI Business

1m13 days ago

Research Papers

Scientists should use AI as a tool, not an oracle

How AI hype leads to flawed research that fuels more hype

AI Snake Oil

1malmost 2 years ago

Research Papers

Start reading the AI Snake Oil book online

The book was published September 2024

AI Snake Oil

1mover 1 year ago

Research Papers

Alibaba Poaches Google DeepMind Research Scientist For Qwen AI Push - Yahoo Finance

<a href="https://news.google.com/rss/articles/CBMijwFBVV95cUxOYTZwZk0walRzazJQampab1FCM2k4Uy1SYk12UWZraENkUXYzZU9kbnlGTGZJS0pFaTZIUFlKZFkwVnJkRzhKbXhNV3lNdUZpdF8tSU1LMklqcTZlUDZERDZ3VzdWbjNQYUN4T2d2ZkRQT1R1MUc0LXdYNndPQTNzbXBXMXJhb3ZEZE00ZFMtaw?oc=5" target="_blank">Alibaba Poaches Google DeepMind Research Scientist For Qwen AI Push</a> <font color="#6f6f6f">Yahoo Finance</font>

Google News: DeepMind

1m25 days ago