Live
Black Hat USADark ReadingBlack Hat AsiaAI BusinessDante-2B: I'm training a 2.1B bilingual fully open Italian/English LLM from scratch on 2×H200. Phase 1 done — here's what I've built.Reddit r/LocalLLaMAMastering AI Careers in 90 Days: Transformative OpportunitiesMedium AIPSSU: The Minimal Architecture for Persistent AIDev.to AIComplete Guide to MCP (Model Context Protocol) in 2026 — Architecture, Implementation, and Enterprise RoadmapDev.to AIFrom Answers to ProcessesMedium AIUnlocking Document Intelligence: A Comprehensive Guide to Multimodal ExtractionMedium AII Studied 40 Viral AI Reels to Find What Actually Works (With Real Numbers)Dev.to AIFive Questions Every AI Investor Should Ask About Intelligence ArchitectureDev.to AIОдин промпт заменил мне 2 часа работы в деньDev.to AIThe 12 AI Tools Actually Worth Using in ClassroomsDev.to AICode Ignition: How AI Sparks Innovation in Software DevelopmentDev.to AIThe Silent Freeze: When Your Model Runs Out of Credits Mid-ConversationDev.to AIBlack Hat USADark ReadingBlack Hat AsiaAI BusinessDante-2B: I'm training a 2.1B bilingual fully open Italian/English LLM from scratch on 2×H200. Phase 1 done — here's what I've built.Reddit r/LocalLLaMAMastering AI Careers in 90 Days: Transformative OpportunitiesMedium AIPSSU: The Minimal Architecture for Persistent AIDev.to AIComplete Guide to MCP (Model Context Protocol) in 2026 — Architecture, Implementation, and Enterprise RoadmapDev.to AIFrom Answers to ProcessesMedium AIUnlocking Document Intelligence: A Comprehensive Guide to Multimodal ExtractionMedium AII Studied 40 Viral AI Reels to Find What Actually Works (With Real Numbers)Dev.to AIFive Questions Every AI Investor Should Ask About Intelligence ArchitectureDev.to AIОдин промпт заменил мне 2 часа работы в деньDev.to AIThe 12 AI Tools Actually Worth Using in ClassroomsDev.to AICode Ignition: How AI Sparks Innovation in Software DevelopmentDev.to AIThe Silent Freeze: When Your Model Runs Out of Credits Mid-ConversationDev.to AI
AI NEWS HUBbyEIGENVECTOREigenvector

SkillRouter: Skill Routing for LLM Agents at Scale

arXivby [Submitted on 23 Mar 2026 (v1), last revised 31 Mar 2026 (this version, v3)]March 31, 20262 min read1 views
Source Quiz

arXiv:2603.22455v2 Announce Type: replace Abstract: Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference time becomes infeasible. This creates a skill-routing problem: given a user task, the system must identify relevant skills before downstream planning or execution. Existing agent stacks often rely on progressive disclosure, exposing only skill names and descriptions while hiding the full implementation body. We examine — YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu

Authors:YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu

View PDF HTML (experimental)

Abstract:Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference time becomes infeasible. This creates a skill-routing problem: given a user task, the system must identify relevant skills before downstream planning or execution. Existing agent stacks often rely on progressive disclosure, exposing only skill names and descriptions while hiding the full implementation body. We examine this design choice on a SkillsBench-derived benchmark with approximately 80K candidate skills, targeting the practically important setting of large skill registries with heavy overlap. Across representative sparse, dense, and reranking baselines on this setting, hiding the skill body causes a 31--44 percentage point drop in routing accuracy, showing that full skill text is a critical routing signal in this setting rather than a minor metadata refinement. Motivated by this finding, we present SkillRouter, a compact 1.2B full-text retrieve-and-rerank pipeline. SkillRouter achieves 74.0% Hit@1 on our benchmark -- the strongest average top-1 routing performance among the baselines we evaluate -- while using 13$\times$ fewer parameters and running 5.8$\times$ faster than the strongest base pipeline. The ranking gains further generalize to a supplementary benchmark independently constructed from three skill sources. In a complementary end-to-end study across four coding agents, routing gains transfer to improved task success, with larger gains for more capable agents.

Subjects:

Machine Learning (cs.LG)

Cite as: arXiv:2603.22455 [cs.LG]

(or arXiv:2603.22455v3 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2603.22455

arXiv-issued DOI via DataCite

Submission history

From: Yanzhao Zheng [view email] [v1] Mon, 23 Mar 2026 18:23:59 UTC (545 KB) [v2] Mon, 30 Mar 2026 09:19:32 UTC (438 KB) [v3] Tue, 31 Mar 2026 16:28:22 UTC (439 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by Eigenvector · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
SkillRouter…researchpaperarxivmachine-lea…deep-learni…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 155 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!