Research Papers research paper arxiv ai artificial-intelligence

Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs

arXivMarch 31, 202610 min read0 views

arXiv:2503.05371v3 Announce Type: replace-cross Abstract: We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We compute 8 steering vectors, each corresponding to a different social bias axis, such as age, gender, or race, on a training subset of the BBQ dataset and compare the effectiveness of these to 3 additional bias mitigation methods across 4 datasets. When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.8% on BBQ, 8.3% on CLEAR-B — Zara Siddique, Irtaza Khalid, Liam D. Turner, Luis Espinosa-Anke

View PDF HTML (experimental)

Abstract:We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We compute 8 steering vectors, each corresponding to a different social bias axis, such as age, gender, or race, on a training subset of the BBQ dataset and compare the effectiveness of these to 3 additional bias mitigation methods across 4 datasets. When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.8% on BBQ, 8.3% on CLEAR-Bias, and 1% on StereoSet, and show improvements over prompting and Self-Debias in all cases, and improvements over fine-tuning in 12 out of 17 evaluations. In addition, steering vectors showed the lowest impact on MMLU scores of the four bias mitigation methods tested. The work presents the first systematic investigation of steering vectors for bias mitigation, and we demonstrate that they are a powerful and computationally efficient strategy for reducing bias in LLMs, with broader implications for enhancing AI safety.

Comments: Published to EACL Findings 2026

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as: arXiv:2503.05371 [cs.LG]

(or arXiv:2503.05371v3 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2503.05371

arXiv-issued DOI via DataCite

Submission history

From: Zara Siddique [view email] [v1] Fri, 7 Mar 2025 12:25:29 UTC (947 KB) [v2] Wed, 13 Aug 2025 12:45:25 UTC (677 KB) [v3] Sat, 28 Mar 2026 13:41:57 UTC (674 KB)

Original source

arXiv

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Research Papers

How a Nonprofit Transforms Data with Cloudera and AI

The organization developed data pipelines that extract and structure information from various scientific sources, significantly accelerating the research process.

AI Business

1m13 days ago

Research Papers

Scientists should use AI as a tool, not an oracle

How AI hype leads to flawed research that fuels more hype

AI Snake Oil

1malmost 2 years ago

Models

New paper: AI agents that matter

Rethinking AI agent benchmarking and evaluation

AI Snake Oil

1mover 1 year ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 158 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research Papers

How a Nonprofit Transforms Data with Cloudera and AI

The organization developed data pipelines that extract and structure information from various scientific sources, significantly accelerating the research process.

AI Business

1m13 days ago

Research Papers

Scientists should use AI as a tool, not an oracle

How AI hype leads to flawed research that fuels more hype

AI Snake Oil

1malmost 2 years ago

Research Papers

Start reading the AI Snake Oil book online

The book was published September 2024

AI Snake Oil

1mover 1 year ago

Research Papers

Alibaba Poaches Google DeepMind Research Scientist For Qwen AI Push - Yahoo Finance

<a href="https://news.google.com/rss/articles/CBMijwFBVV95cUxOYTZwZk0walRzazJQampab1FCM2k4Uy1SYk12UWZraENkUXYzZU9kbnlGTGZJS0pFaTZIUFlKZFkwVnJkRzhKbXhNV3lNdUZpdF8tSU1LMklqcTZlUDZERDZ3VzdWbjNQYUN4T2d2ZkRQT1R1MUc0LXdYNndPQTNzbXBXMXJhb3ZEZE00ZFMtaw?oc=5" target="_blank">Alibaba Poaches Google DeepMind Research Scientist For Qwen AI Push</a> <font color="#6f6f6f">Yahoo Finance</font>

Google News: DeepMind

1m25 days ago