Live

•🔥 OpenBMB/ChatDevGitHub Trending •🔥 microsoft/agent-lightningGitHub Trending •🔥 apache/supersetGitHub Trending •🔥 shanraisshan/claude-code-best-practiceGitHub Trending •A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation LearningarXiv •GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play AnnotationarXiv •Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language ModelsarXiv •CANGuard: A Spatio-Temporal CNN-GRU-Attention Hybrid Architecture for Intrusion Detection in In-Vehicle CAN NetworksarXiv •DesignWeaver: Dimensional Scaffolding for Text-to-Image Product DesignarXiv •A Lightweight, Transferable, and Self-Adaptive Framework for Intelligent DC Arc-Fault Detection in Photovoltaic SystemsarXiv •Consistency Amplifies: How Behavioral Variance Shapes Agent AccuracyarXiv •Stabilizing Rubric Integration Training via Decoupled Advantage NormalizationarXiv •Semi-Automated Knowledge Engineering and Process Mapping for Total Airport ManagementarXiv •AIRA_2: Overcoming Bottlenecks in AI Research AgentsarXiv •BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional EnvironmentsarXiv •🔥 OpenBMB/ChatDevGitHub Trending •🔥 microsoft/agent-lightningGitHub Trending •🔥 apache/supersetGitHub Trending •🔥 shanraisshan/claude-code-best-practiceGitHub Trending •A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation LearningarXiv •GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play AnnotationarXiv •Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language ModelsarXiv •CANGuard: A Spatio-Temporal CNN-GRU-Attention Hybrid Architecture for Intrusion Detection in In-Vehicle CAN NetworksarXiv •DesignWeaver: Dimensional Scaffolding for Text-to-Image Product DesignarXiv •A Lightweight, Transferable, and Self-Adaptive Framework for Intelligent DC Arc-Fault Detection in Photovoltaic SystemsarXiv •Consistency Amplifies: How Behavioral Variance Shapes Agent AccuracyarXiv •Stabilizing Rubric Integration Training via Decoupled Advantage NormalizationarXiv •Semi-Automated Knowledge Engineering and Process Mapping for Total Airport ManagementarXiv •AIRA_2: Overcoming Bottlenecks in AI Research AgentsarXiv •BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional EnvironmentsarXiv

AI NEWS

by techtonicshifts.blog

Hugging Face Releases SmolLM3: A 3B Parameter Model That Rivals 7B Models

Releases SmolLM Hugging Face Open Source Efficiency

Hugging Face Releases SmolLM3: A 3B Parameter Model That Rivals 7B Models

Hugging Faceby Hugging Face TeamMarch 26, 20264 min read9,800 views

SmolLM3 demonstrates that careful data curation and training techniques can achieve 7B-class performance with just 3B parameters, making powerful AI accessible on consumer hardware.

Hugging Face has released SmolLM3, a 3-billion parameter language model that achieves performance comparable to models twice its size. The release represents a significant advancement in model efficiency, with implications for edge deployment and consumer hardware applications.

The model was trained on a carefully curated dataset of 11 trillion tokens, with particular emphasis on code, mathematics, and scientific reasoning. Through a combination of improved tokenization, architectural refinements, and a novel training curriculum, the team achieved remarkable efficiency gains.

SmolLM3 runs comfortably on consumer GPUs with 8GB VRAM and can even be quantized to run on high-end smartphones. The model achieves 68.4% on MMLU and 72.1% on HumanEval, metrics that were previously only achievable with much larger models.

The full model weights, training code, and dataset composition details have been released under an Apache 2.0 license, continuing Hugging Face's commitment to open-source AI development.

Original source

Hugging Face

Was this article helpful?

Sign in to highlight and annotate this article

Ask AI about this article

Powered by AI News Hub · full article context loaded

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

SmolLMHugging FaceOpen Source

Meta Releases Llama 4 Scout and Maverick with Native Multimodal Support

Meta Releases Llama 4 Scout and Maverick with Native Multimodal Support

Meta's Llama 4 family introduces two new models: Scout (17B active parameters) and Maverick (17B active, 400B total MoE), both with native image and video understanding capabilities.

What's New in Mellea 0.4.0 + Granite Libraries Release

What's New in Mellea 0.4.0 + Granite Libraries Release

Hugging Face Blog

State of Open Source on Hugging Face: Spring 2026

State of Open Source on Hugging Face: Spring 2026

Hugging Face Blog

Knowledge Map

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 338 connections

Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Releases

Meta Releases Llama 4 Scout and Maverick with Native Multimodal Support

Meta Releases Llama 4 Scout and Maverick with Native Multimodal Support

Meta's Llama 4 family introduces two new models: Scout (17B active parameters) and Maverick (17B active, 400B total MoE), both with native image and video understanding capabilities.