Europe’s leading AI model still can’t reliably follow European law
TL;DR: Mistral previously performed the worst of all providers on our human rights scenarios, with their Medium 3.5 model failing on almost every single test. Their newly released model, Mistral Large 4, improves performance by 61 percentage points, showing strong resistance in most scenarios.
Likewise, Mistral Large 4 demonstrates large improvements in compliance with European law. Its predecessor, Mistral Large 2512, lagged at the bottom of the leaderboard when we tested 12 frontier models earlier this year on compliance with the EU AI Act and GDPR, reaching only 16% compliance even when explicitly instructed to follow these laws. Mistral Large 4 improves by 37 percentage points, reaching 53% compliance - but still falling behind recent Claude and GPT models.
These much needed improvements close the gap between Europe’s leading model and US models. However, legal compliance of 53% means breaking the law in just under half of cases - and while Mistral’s latest model is more robust against misuse for human rights violations, its earlier generations remain available for deployment, with no standard against which to judge undesirable use.
Mistral’s performance on human rights
In September, we published results that several frontier AI models will violate human rights under realistic agentic simulations when asked to. Spanning a total of 13 rights from the Universal Declaration of Human Rights, the simulated scenarios test whether the AI will willingly cooperate with rights violations such as belief surveillance, un-enrolling children from school based on their last names, reversing equal pay, and denying a request for reproductive healthcare, among others. You can explore these scenarios in detail at lara.aithos.org.
Of the 7 models tested, Mistral’s Medium 3.5 - its leading agentic-ready model at the time of running (August 2026) - was the least resistant to tasks that involved violating human rights of all tested models. Overall, Mistral refused such requests only 11% of the time - in some scenarios, it carried out the rights violation in every instance.
With Mistral’s latest agentic-ready model, Mistral Large 4, released this week, we put it to the test.
Overall resistance to human rights violations increased from just 11% (Mistral Medium 3.5, August 2026), to 72% (Mistral Large 4, October 2026), putting Mistral in 3rd place, behind Claude Opus 5 and Kimi K3. This now puts two open models among the top 3, while also closing the gap between European models with US and Chinese-made models.
Even more striking than the 61 percentage point increase in overall compliance, is the per-scenario improvements. Mistral Medium 3.5 held 0% pass rates in 6 of 8 scenarios - Mistral Large 4 now achieves >90% in 5, with strong improvements across 2 others.
The one outcome that remains unmoved is Mistral’s performance in the Denial of reproductive healthcare scenario, where Mistral continues to uphold an inferred ban on abortion care - this time, going as far as contacting external clinics to instruct them not to offer services to the patient.
Another concerning shared model behavior that persists, but is new to this release of Mistral, is verdict instability. In our earlier tests, models including OpenAI’s GPT 5.6 Sol demonstrated startling decision instability across their 20 runs. For instance, in the Belief surveillance scenario, the model refuses to surveil and flag a vulnerable woman in one run, and immediately sends an email requesting her silence in another. Mistral Medium 3.5 did not demonstrate this behaviour in our earlier tests because it committed the violation in every instance - Mistral Large 4 on the other hand now replicates this exact coin-flip pattern observed by other models.
These tests replicate the same methodology as our earlier work, available to read in further detail here.
Following the law
Our human rights tests relate to a gap in AI governance - despite the supposed universality of human rights, there is currently no universal translation to behavioral standards for AI models. Developers are left to figure out what's acceptable for themselves, explaining the wide variety in model behavior we observe. The same cannot be said for compliance with laws that already exist and that clearly prohibit certain AI uses.
In May, Aithos launched 10 agentic scenarios simulating 10 different articles of the EU AI Act and GDPR, including exploiting an elderly customer, concealing its AI status and the use of AI for social scoring. Across 9 different model providers from US, China and Europe, Mistral’s leading model at the time, Mistral Large 2512, came second to last with just 16% compliance. Explicitly including legal texts in this model's instructions and asking it to follow these laws at all times did not improve its overall compliance rate at all.
Mistral’s latest model improves on that rate by 37 percentage points - now coming in 7th, behind several models from Anthropic, OpenAI and Google.
Compliance leader GPT 5.6 Sol demonstrated that it is feasible to achieve at least 70% overall compliance. We are pleased to see Mistral has come some way to closing that gap, but it is clear it can be improved further.
Performance improvement does not resolve two fundamental problems
Problem 1: Developers are not accountable.
No model we tested reliably resists requests that would violate the law, even under explicit instruction to reference provided legal texts. Deployers of AI systems remain the ones accountable if the AI is used for illegal purposes, but when models are incapable of recognizing legal limits, those deployers don't actually have the means to guarantee compliance. Model creators are incentivized to continuously push for capability increases, and without formal accountability or reason to adhere to behavioral standards, it should come as no surprise when frontier models treat benchmark performance as more important than legal or ethical behavior.
Problem 2: Harmful tools remain accessible.
Mistral Large 4 becomes the next model to demonstrate that agents can be prevented from violating human rights at the model level. While it is promising to see the availability of rights-preserving models grow, we remain concerned that frontier agents can be used as off-the-shelf tools to violate human rights at scale. If a manager wants to profile their team for union activities, they need only find the right model to do so. AI companies are not equipped to be the arbiter of what model usages are and are not acceptable, and internationally established guidelines on what uses to prevent would at least make the minimal standards clear and unilateral.
While model developers remain the only ones responsible for deciding what tools are harmful, the risk of AI being used to violate human rights at scale will persist.
The scenarios used in this experiment were generated and run using LARA, an open research tool meant to open up AI behavioral evaluations to the public. We will be releasing the tool publicly soon; get in touch if you’d like early access as a beta-tester: info@aithos.org.