Research Papers research paper arxiv machine-learning deep-learning

Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models

arXivMarch 26, 202610 min read0 views

Out-of-distribution (OOD) detection aims to identify samples that deviate from in-distribution (ID). One popular pipeline addresses this by introducing negative labels distant from ID classes and detecting OOD based on their distance to these labels. However, such labels may present poor activation on OOD samples, failing to capture the OOD characteristics. To address this, we propose \underline{T}est-time \underline{A}ctivated \underline{N}egative \underline{L}abels (TANL) by dynamically evaluating activation levels across the corpus dataset and mining candidate labels with high activation re — Yabin Zhang, Maya Varma, Yunhe Gao

View PDF HTML (experimental)

Abstract:Out-of-distribution (OOD) detection aims to identify samples that deviate from in-distribution (ID). One popular pipeline addresses this by introducing negative labels distant from ID classes and detecting OOD based on their distance to these labels. However, such labels may present poor activation on OOD samples, failing to capture the OOD characteristics. To address this, we propose \underline{T}est-time \underline{A}ctivated \underline{N}egative \underline{L}abels (TANL) by dynamically evaluating activation levels across the corpus dataset and mining candidate labels with high activation responses during the testing process. Specifically, TANL identifies high-confidence test images online and accumulates their assignment probabilities over the corpus to construct a label activation metric. Such a metric leverages historical test samples to adaptively align with the test distribution, enabling the selection of distribution-adaptive activated negative labels. By further exploring the activation information within the current testing batch, we introduce a more fine-grained, batch-adaptive variant. To fully utilize label activation knowledge, we propose an activation-aware score function that emphasizes negative labels with stronger activations, boosting performance and enhancing its robustness to the label number. Our TANL is training-free, test-efficient, and grounded in theoretical justification. Experiments on diverse backbones and wide task settings validate its effectiveness. Notably, on the large-scale ImageNet benchmark, TANL significantly reduces the FPR95 from 17.5% to 9.8%. Codes are available at \href{this https URL}{YBZh/OpenOOD-VLM}.

Comments: CVPR 2026 main track, Codes are available at this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Cite as: arXiv:2603.25250 [cs.CV]

(or arXiv:2603.25250v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.25250

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yabin Zhang [view email] [v1] Thu, 26 Mar 2026 09:53:04 UTC (3,661 KB)

Original source

arXiv

https://arxiv.org/abs/2603.25250v1

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Research PapersLive

I ran by instinct for years. Then I built an AI running coach.

How a 50 km trail race, a broken ChatGPT workflow, and 60+ research papers led me to create Coach Leo. Continue reading on Medium »

Medium AI

1mabout 1 hour ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models WSJ

Google News: LLM

1m1 day ago

ModelsFresh

How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models

arXiv:2511.06676v2 Announce Type: replace-cross Abstract: Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online post flagged as "inappropriate" was not simply the victim of a biased algorithm? This paper investigates this problem using a dual approach. First, I conduct a quantitative benchmark of a widely used toxicity model (unitary/toxic-bert) to measure performance disparity between text in African-American English (AAE) and Standard American English (SAE). The benchmark reveals a clear, systematic bias: on average, the model scores AAE text as 1.8 times more toxic and 8.8 times higher for "identity hate"

arXiv cs.HC

1mabout 7 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 172 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models

Submission history

Daily AI Digest

More about

I ran by instinct for years. Then I built an AI running coach.

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models

Knowledge Map

Connected Articles — Knowledge Graph

Discussion

More in Research Papers

I ran by instinct for years. Then I built an AI running coach.

“It's not about gatekeeping."

Adversaries have under-protected APIs in their sights

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI - WSJ