Research Papers research paper arxiv nlp language-models

FinTruthQA: A Benchmark for AI-Driven Financial Disclosure Quality Assessment in Investor -- Firm Interactions

arXivMarch 30, 202610 min read0 views

arXiv:2406.12009v4 Announce Type: replace Abstract: Accurate and transparent financial information disclosure is essential for market efficiency, investor decision-making, and corporate governance. Chinese stock exchanges' investor interactive platforms provide a widely used channel through which listed firms respond to investor concerns, yet these responses are often limited or non-substantive, making disclosure quality difficult to assess at scale. To address this challenge, we introduce FinTruthQA, to our knowledge the first benchmark for AI-driven assessment of financial disclosure quality — Peilin Zhou, Ziyue Xu, Xinyu Shi, Jiageng Wu, Yikang Jiang, Dading Chong, Bin Ke, Jie Yang

View PDF HTML (experimental)

Abstract:Accurate and transparent financial information disclosure is essential for market efficiency, investor decision-making, and corporate governance. Chinese stock exchanges' investor interactive platforms provide a widely used channel through which listed firms respond to investor concerns, yet these responses are often limited or non-substantive, making disclosure quality difficult to assess at scale. To address this challenge, we introduce FinTruthQA, to our knowledge the first benchmark for AI-driven assessment of financial disclosure quality in investor-firm interactions. FinTruthQA comprises 6,000 real-world financial Q&A entries, each manually annotated based on four key evaluation criteria: question identification, question relevance, answer readability, and answer relevance. We benchmark statistical machine learning models, pre-trained language models and their fine-tuned variants, as well as large language models (LLMs), on FinTruthQA. Experiments show that existing models achieve strong performance on question identification and question relevance (F1 > 95%), but remain substantially weaker on answer readability (Micro F1 approximately 88%) and especially answer relevance (Micro F1 approximately 80%), highlighting the nontrivial difficulty of fine-grained disclosure quality assessment. Domain- and task-adapted pre-trained language models consistently outperform general-purpose models and LLM-based prompting on the most challenging settings. These findings position FinTruthQA as a practical foundation for AI-driven disclosure monitoring in capital markets, with value for regulatory oversight, investor protection, and disclosure governance in real-world financial settings.

Subjects:

Computation and Language (cs.CL)

Cite as: arXiv:2406.12009 [cs.CL]

(or arXiv:2406.12009v4 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2406.12009

arXiv-issued DOI via DataCite

Submission history

From: Ziyue Xu [view email] [v1] Mon, 17 Jun 2024 18:25:02 UTC (1,216 KB) [v2] Sat, 7 Dec 2024 15:47:26 UTC (1,273 KB) [v3] Tue, 11 Feb 2025 16:49:17 UTC (1,042 KB) [v4] Fri, 27 Mar 2026 08:49:53 UTC (778 KB)

Original source

arXiv

https://arxiv.org/abs/2406.12009

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - wsj.com

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxPV0Z4UERPZU5fUFY4QXBHeXRud2w2ZWN1WTRaUVFSQ081QnM0YXdKLVFtclZTb2l5SE5QNUxzQTE4eV9HMjhqYkE0RE1HN1hITFVOMFU1c1FSZjcxR3F0Y2w3NHVrVndjUERpS1FTX3JQS0Y2MjF6U1dpY1J0elBFUVN2VEZULXVqUUxGeHVjYUFNUERVVUdlb1F4cGxQQmRKY3poVVlVTVphbDV5SlU0X0ZZVHlUTmFlUWRFNHFfbm9nODVBV3pWZjVGOUZyOFlSZlBlLVNnS2o2eXJxWEtTUEJOaUlnTUlIN29tWTJoWXcwWGxXOXlzU1BLWkNQUzBURWxqRXk5RWVOQ0JDT1FoMVMtUVJfdk9LYlpLb3RSeGpRb2poUlYyZUZRNk54cFBKYTVqcXI1SkdKQlBGRGZKLWcwNDNjTHRaSGtVV1o0dHItTnpvNzJzZ3Qteks2NVFKSndoOVZhdFFHbXdCQXlkbFlSLUJfeFphTmlKN1FOcXVUQVNYcFBaMEk2WU5CSjM3M01lbUZROFlWbHZvaFpkR1I5ZGVlazdxSVJuNmVoN3hQai1GTDBkZw?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> wsj.com

Google News: LLM

1mabout 16 hours ago

Research PapersLive

AI Inspires New Research Topics In Materials Science - miragenews.com

<a href="https://news.google.com/rss/articles/CBMihwFBVV95cUxQRlVFdkRBaHRvYkJJdFRlMTZmajEzeFRPU0hGWWdfbi02V1FnTUdVQ2pmY2VZLUV2NlB4V3BFdEVlSVZkUlhRSTZaNWFKMmcyWXJYbnNqbUhMTmp0NnFtMEppOXlPZkJSNHJfck5VSEVYcmUtX1k2QkJlR1BvUEdTTkp3UmlYRkk?oc=5" target="_blank">AI Inspires New Research Topics In Materials Science</a> miragenews.com

Google News: Machine Learning

1m44 minutes ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNNWh0OTV4cnNDLVdHdHVUdE02cWRiaE03VENfdWJFbFlyaXZmbWtJNm9OdU05TXVsRjd4dVFwUzl0WkRfLVJoVzlaNkhKRVl4S0Y0Um5jN2QzZzhsb0twMElFOEpSZjdjX1pZZzNacXIxU2U4Ulloam5nR1hQeXg3TWhoMEE1ZzFzQmdiSjktRG1rUEs2YVVhZ0VMMk0wS3J6SWNJdTdJZlAtTEE0SUdaaFl5QWFUWS05NGFDN1FudnNRN2ZpcnFmM0N1bGVpSjNYZmZ6MUJKSkpMWk5tRWFSN2s4V0tEdi1EVVBuUTdnZm92Sjk0MEVYZWRieTkxNWMwRzRiQmxWVHpvaGEwUnpEZGJ1UVFhQmoydGxSTW93XzFVR1ZHeG5mMTZOLWthOVVKZTZMeGdsS0dDaUROelpWc1l4QmJLNWkzRkhGUGdua3hnOHFWYUpXQWp3RktyemZiN0VBTFhfNGFZUHpNaV9jX2U0Sk9Fb2k1dXhOZHdENWpPc2dRU2ZQeHZoMnBZNEN6RHJnNU1YYk9SSzRYNzZrbXRnQ3VOdE0ydGFCM3ZBVTdHeFJLY29Feg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 16 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 159 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersLive

AI Inspires New Research Topics In Materials Science - miragenews.com

Google News: Machine Learning

1m44 minutes ago

Research Papers

From brain scans to alloys: Teaching AI to make sense of complex research data - Penn State University

<a href="https://news.google.com/rss/articles/CBMiwAFBVV95cUxPZDFHdkptQ2VUM2hmWjhqQkxoRnBiTWoxMXRRR21MUG5TamdUMlFRWmhvYVNHaFVNREVKU3VmSnVOdDVZYnNLb2ppYXRVRTZmVFVMV1pLTlVhUm9ybTNZbGtvZTdIMnIyMHNpOEk5aU9TSmxxS2Y4V2MwazYwY3JlX1Axbk1nd3pfcWhFdUJaaDJWRXJaMFIyTTROcmFHeXI3ZzFudXJ2M1h6UHI1LW1Ca1dta2RkM3BiYndocGk3Yjg?oc=5" target="_blank">From brain scans to alloys: Teaching AI to make sense of complex research data</a> Penn State University

GNews AI materials

1m3 months ago

Research PapersFresh

Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work

arXiv:2505.24246v4 Announce Type: replace Abstract: As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection process

arXiv cs.HC

2mabout 6 hours ago

Research PapersFresh

Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability

arXiv:2505.01000v5 Announce Type: replace Abstract: Scheduling is a perennial-and often challenging-problem for many groups. Existing tools are mostly static, showing an identical set of choices to everyone, regardless of the current status of attendees' inputs and preferences. In this paper, we propose Togedule, an adaptive scheduling tool that uses large language models to dynamically adjust the pool of choices and their presentation format. With the initial prototype, we conducted a formative study (N=10) and identified the potential benefits and risks of such an adaptive scheduling tool. Then, after enhancing the system, we conducted two controlled experiments, one each for attendees and organizers (total N=66). For each experiment, we compared scheduling with verbal messages, shared c

arXiv cs.HC

1mabout 6 hours ago