Research Papers research paper arxiv ai artificial-intelligence

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

arXivMarch 31, 202610 min read0 views

arXiv:2603.28737v1 Announce Type: cross Abstract: We introduce ParaSpeechCLAP, a dual-encoder contrastive model that maps speech and text style captions into a common embedding space, supporting a wide range of intrinsic (speaker-level) and situational (utterance-level) descriptors (such as pitch, texture and emotion) far beyond the narrow set handled by existing models. We train specialized ParaSpeechCLAP-Intrinsic and ParaSpeechCLAP-Situational models alongside a unified ParaSpeechCLAP-Combined model, finding that specialization yields stronger performance on individual style dimensions whil — Anuj Diwan, Eunsol Choi, David Harwath

View PDF HTML (experimental)

Abstract:We introduce ParaSpeechCLAP, a dual-encoder contrastive model that maps speech and text style captions into a common embedding space, supporting a wide range of intrinsic (speaker-level) and situational (utterance-level) descriptors (such as pitch, texture and emotion) far beyond the narrow set handled by existing models. We train specialized ParaSpeechCLAP-Intrinsic and ParaSpeechCLAP-Situational models alongside a unified ParaSpeechCLAP-Combined model, finding that specialization yields stronger performance on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate our models' performance on style caption retrieval, speech attribute classification and as an inference-time reward model that improves style-prompted TTS without additional training. ParaSpeechCLAP outperforms baselines on most metrics across all three applications. Our models and code are released at this https URL .

Comments: Under review

Subjects:

Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)

Cite as: arXiv:2603.28737 [eess.AS]

(or arXiv:2603.28737v1 [eess.AS] for this version)

https://doi.org/10.48550/arXiv.2603.28737

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anuj Diwan [view email] [v1] Mon, 30 Mar 2026 17:50:07 UTC (132 KB)

Original source

arXiv

https://arxiv.org/abs/2603.28737

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Products

7 million euros available for research into AI applications in the agriculture, horticulture, water, and food sectors - nwo.nl

<a href="https://news.google.com/rss/articles/CBMi2AFBVV95cUxOZkhpc1NkQkZacGs4RmI0a0RRa2RsdmdVbzF6Z0RLS2xZS3J5eWtVR282LUFFcURWM1ltdFl0UTZ6RVNSM2h4bnVlbVdsREVpX3YzdHR4X2VoRkxKQUp2R1ZXQXJNbU1Cblh5QlJDc2RIbUVBSlhJdnhQRlB5QUtoVG5oaDc1T0l3TFZyQnhEYS0xXzdWZVZkV1E2Sk0tYzVLeUc4bUV2Z25pX1R3am1kWHctaFhnOFlaR2tZQzRPMmxWS2tKY2djZFEyNHFxbHZZM1lpQ1o3aS0?oc=5" target="_blank">7 million euros available for research into AI applications in the agriculture, horticulture, water, and food sectors</a> nwo.nl

GNews AI agriculture

1m22 days ago

CountriesRecent

Europe urged to ‘learn to fight for itself’ in case US-China truce collapses

European governments breathed a sigh of relief in October when the US and China sealed a fragile trade truce that paused more sweeping Chinese rare earth restrictions and papered over a Sino-Dutch row over chipmaker Nexperia. Now, however, the European Union is being urged to come up with a battle plan should the ceasefire fail or expire. A spike in superpower tensions could expose the EU to Chinese export controls, potentially pulverising its military support for Ukraine, its own efforts to...

SCMP Tech (Asia AI)

2mabout 14 hours ago

Research PapersLive

[D] TurboQuant author replies on OpenReview

<div class="md">I wanted to follow up to <a href="https://www.reddit.com/r/MachineLearning/comments/1s7m7rn/comment/odaect4/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button">yesterday's thread</a> and see if anyone wanted to weigh in on it. This work is far outside of my niche, but it strikes me as an attempt to reframe the issue instead of addressing concerns head on. OpenReview link for reference: <a href="https://openreview.net/forum?id=tO3ASKZlok">https://openreview.net/forum?id=tO3ASKZlok</a> <blockquote> In response to recent commentary regarding our paper, "TurboQuant," we provide the following technical clarifications to correct the record. TurboQuant did not derive its core method from

Reddit r/MachineLearning

2m37 minutes ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 128 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersLive

[D] TurboQuant author replies on OpenReview

Reddit r/MachineLearning

2m37 minutes ago

Research Papers

Carnegie Mellon Researchers Rethink Chronic Pain

<img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/260123D_WTM_Yttri_Lab_RD002.jpg.webp?itok=DozEPS-g" width="900" height="508" alt="Eric Yttri"> Across neuroscience, biomedical engineering and artificial intelligence, researchers from Carnegie Mellon University are exploring how pain is measured, understood and treated to support safer, more effective care.

Carnegie Mellon News

1mabout 2 months ago

Research Papers

Five CMU Faculty Members Named 2026 Sloan Research Fellows

<img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/sloan-collage-2000%20copy.jpg.webp?itok=Ih-KH3Na" width="900" height="508" alt="2026 Sloan Awardees"> Five Carnegie Mellon University faculty members are among the 126 recipients of 2026 Sloan Research Fellowships, which honor early career scholars whose achievements put them among the best scientific minds working today.

Carnegie Mellon News

1mabout 1 month ago

Research Papers

NIST Researchers Develop More Accurate Formula for Measuring Particle Concentration

The new method will be useful in various fields, including nanomedicine, food science, environmental science and advanced manufacturing.

nist.gov

1m7 months ago