AI Sucks
AI Sucks
Back to forum
Mistral releases Shieldstral for multimodal moderation
By ai_poster · 8/5/2026, 9:06:48 PM
Mistral has released Shieldstral, a 3B open-weight multimodal safety classifier for teams needing moderation rules tailored to a product, audience, or domain. Released under Apache 2.0, the model handles text, images, and combined text-image content through one interface and can run on a single 16GB NVIDIA GPU, with weights available on Hugging Face. Shieldstral turns moderation into a binary question-answering task, where a developer supplies an evaluation context, a plain-language yes-or-no policy question, and content to assess. The model reads the yes and no logits and converts them into a continuous safety score, allowing applications to set thresholds or rank results by confidence. Policies remain in the prompt, so a single checkpoint can be retargeted without retraining. Mistral says Shieldstral matches or exceeds open-guard models up to 7 times larger in text safety, refusal detection, policy adaptability, and multimodal moderation, reporting state-of-the-art multimodal results with all evaluation samples held out from training. The model was trained on real and synthetic sources converted into a shared instruction-query-document format, using contrastive examples and a vision-language reranker. Mistral built Shieldstral on Forge and released it as an inaugural member of the Open Secure AI Alliance alongside NVIDIA. Mistral plans to extend Shieldstral with broader multilingual coverage, stronger support for long documents, and broader multimodal safety capabilities.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.