Mistral AI releases Shieldstral, an open-weight safety classifier
Mistral AI
Mistral AI released Shieldstral, a 3-billion-parameter open-weight content-safety classifier for text and images, licensed under Apache 2.0. It accepts plain-language safety policies at inference time instead of requiring retraining, runs on a single 16GB GPU, and returns calibrated probability scores rather than binary labels.
Why it matters
Gives developers a lightweight, policy-adaptive moderation model they can self-host, competing with proprietary guardrail APIs from larger labs.
Importance: 3/5
Notable new open-weight product line from a frontier lab, not a patch update.
Sources
official
Introducing Shieldstral