Menu

Post image 1
Post image 2
Post image 3
Post image 4
Post image 5
Post image 6
Post image 7
Post image 8
Post image 9
Post image 10
Post image 11
1 / 11
481

Introducing Shieldstral. | Mistral AI

Hacker News·Introducing Shieldstral. | Mistral AI·about 1 month ago
#YJKqk77e
#mistral#safety#model#policy#photo#article
Reading 0:00
15s threshold

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU. A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation. “Does this content promote violence against a protected group? Is this image safe to show to a minor? Did the assistant refuse the request?” Every product that ships a model needs to answer questions like these — but the right answer depends on the product, the audience, and the moment.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More