Live
Black Hat USAAI BusinessBlack Hat AsiaAI BusinessMassachusetts Sen. Ed Markey is putting AV firms on blast for using human staffersFast Company TechRTX 60 series leaks are everywhere, but Nvidia hasn't finalized the GPUs yetTechSpotWhen Your LLM Becomes Your Twin (and Starts Judging Your Code) 🤖👀DEV CommunityUnderstanding Data Modelling in Power BI: Joins, Relationships and Schemes ExplainedDEV CommunityUnderstanding Attention Mechanisms – Part 4: Turning Similarity Scores into Attention WeightsDEV CommunityQ/A: How engineers must design AVs to drive safelyFierce ElectronicsBosch’s pressure sensor is part of Qualcomm’s new wearables chipFierce ElectronicsQ/A: Lumotive CTO talks software-defined optical sensingFierce ElectronicsST’s smart IMU bolsters Qualcomm’s monster AI chip for wearablesFierce ElectronicsRound three: More Rising Stars 2026Fierce ElectronicsMy Obsidian Tab-to-Vault Workflow (with a Free Chrome Extension)DEV CommunityOpenAI contract with U.S. Cyber Command went unnoticed amid degradation of transparency and veracity of U.S. procurement database - All-Source Intelligence | Jack PoulsonGoogle News: OpenAIBlack Hat USAAI BusinessBlack Hat AsiaAI BusinessMassachusetts Sen. Ed Markey is putting AV firms on blast for using human staffersFast Company TechRTX 60 series leaks are everywhere, but Nvidia hasn't finalized the GPUs yetTechSpotWhen Your LLM Becomes Your Twin (and Starts Judging Your Code) 🤖👀DEV CommunityUnderstanding Data Modelling in Power BI: Joins, Relationships and Schemes ExplainedDEV CommunityUnderstanding Attention Mechanisms – Part 4: Turning Similarity Scores into Attention WeightsDEV CommunityQ/A: How engineers must design AVs to drive safelyFierce ElectronicsBosch’s pressure sensor is part of Qualcomm’s new wearables chipFierce ElectronicsQ/A: Lumotive CTO talks software-defined optical sensingFierce ElectronicsST’s smart IMU bolsters Qualcomm’s monster AI chip for wearablesFierce ElectronicsRound three: More Rising Stars 2026Fierce ElectronicsMy Obsidian Tab-to-Vault Workflow (with a Free Chrome Extension)DEV CommunityOpenAI contract with U.S. Cyber Command went unnoticed amid degradation of transparency and veracity of U.S. procurement database - All-Source Intelligence | Jack PoulsonGoogle News: OpenAI

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

arXivMarch 31, 20262 min read0 views
Source Quiz

arXiv:2512.14177v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) often produce plausible but unreliable outputs, making robust uncertainty estimation essential. Recent work on semantic uncertainty estimates relies on external models to cluster multiple sampled responses and measure their semantic consistency. However, these clustering methods are often fragile, highly sensitive to minor phrasing variations, and can incorrectly group or separate semantically similar answers, leading to unreliable uncertainty estimates. We propose Semantic Gaussian Process Uncertainty (SG — Joseph Hoche, Andrei Bursuc, David Brellmann, Gilles Louppe, Pavel Izmailov, Angela Yao, Gianni Franchi

View PDF HTML (experimental)

Abstract:Large Vision-Language Models (LVLMs) often produce plausible but unreliable outputs, making robust uncertainty estimation essential. Recent work on semantic uncertainty estimates relies on external models to cluster multiple sampled responses and measure their semantic consistency. However, these clustering methods are often fragile, highly sensitive to minor phrasing variations, and can incorrectly group or separate semantically similar answers, leading to unreliable uncertainty estimates. We propose Semantic Gaussian Process Uncertainty (SGPU), a Bayesian framework that quantifies semantic uncertainty by analyzing the geometric structure of answer embeddings, avoiding brittle clustering. SGPU maps generated answers into a dense semantic space, computes the Gram matrix of their embeddings, and summarizes their semantic configuration via the eigenspectrum. This spectral representation is then fed into a Gaussian Process Classifier that learns to map patterns of semantic consistency to predictive uncertainty, and that can be applied in both black-box and white-box settings. Across six LLMs and LVLMs on eight datasets spanning VQA, image classification, and textual QA, SGPU consistently achieves state-of-the-art calibration (ECE) and discriminative (AUROC, AUARC) performance. We further show that SGPU transfers across models and modalities, indicating that its spectral representation captures general patterns of semantic uncertainty.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2512.14177 [cs.CV]

(or arXiv:2512.14177v2 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2512.14177

arXiv-issued DOI via DataCite

Submission history

From: Joseph Hoche [view email] [v1] Tue, 16 Dec 2025 08:15:24 UTC (2,200 KB) [v2] Mon, 30 Mar 2026 12:43:20 UTC (1,325 KB)

Original source

arXiv

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by AI News Hub · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
Improving S…researchpaperarxivcomputer-vi…image-recog…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 207 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers