Europe’s most capable open model was downloaded 1,071 times last month. A 24-billion-parameter coding model from the same company was downloaded 232,736 times. That gap is the whole story, and it is why “best” has to mean something more useful than “biggest”. We pulled fresh download data on every European open-weight model and ranked them by what each one is good for.
| Covers | Every European open-weight LLM you can download, ranked by use case, not size. |
| Best overall | Mistral Small 4 (119B / 6.5B active, 256k, Apache 2.0). |
| Run on one machine | Devstral Small 2 24B for code, Apertus v1.5-8B for everything else. |
| Key number | Europe’s biggest open model is out-downloaded by a 2023 7B, roughly 4,700 to one. |
| The trap | At least five “sovereign AI” models ship under a US company’s licence. |
Size Is Not the Ranking
Ask which European model is best and you get an answer about parameters. Mistral Large 3 has 675 billion of them, 41 billion active, and shipped under Apache 2.0 in December 2025. On paper that settles it.
Then you look at who downloads it. In the 30 days to 4 August, Mistral’s official Large 3 instruct repository logged 1,071 downloads. In the same window, Mistral-7B-Instruct-v0.3, a 7B model from 2023 that Mistral has long since superseded, logged 5,000,363. That is a ratio of roughly 4,700 to one, and it is not a quality judgement but a hardware one: almost nobody outside a data centre can load 675 billion parameters, so almost nobody does.
So ranking these by capability alone is close to useless. The useful question is which model to put into production given your hardware, and that is how we have ordered this.
Two definitions first, because “open” is carrying a lot of unearned weight in European AI marketing. Open weights means you can download the parameters. Open source means you can also use them commercially without asking, under a licence an ordinary lawyer recognises. Several European models are the first and not the second, and we say so each time.
Click here for the wider European LLM map.
The Ranking, by What You Are Doing
1. Best Overall Open-Source Model: Mistral Small 4
mistralai/Mistral-Small-4-119B-2603
Apache 2.0. 119B total parameters, roughly 6.5B active per token, 256k context. Released March 2026.
The name is misleading, and deliberately so: “Small” refers to active parameters, not total. Mistral folded three separate lines into this one release, absorbing the reasoning model formerly called Magistral and the Devstral coding behaviour into a single model with configurable reasoning effort. Its published figures include 71.2 on GPQA Diamond and a claim that it beats OpenAI’s gpt-oss-120b on LiveCodeBench while producing about 20% less output.
It is the best combination on offer in Europe of a permissive licence, current training and a context window long enough for real documents. 182,356 downloads last month, which is the highest of any European model above 100B parameters and still an order of magnitude below Mistral’s own 7B from 2023.
2. Best for Coding: Devstral Small 2 24B
mistralai/Devstral-Small-2-24B-Instruct-2512
Apache 2.0. 24B parameters, 256k context.
68.0% on SWE-Bench Verified, 55.7% on SWE-Bench Multilingual, 22.5% on Terminal Bench 2, all published by Mistral. For a 24B model they are strong enough that this is the one most people should try first for agentic coding, and it fits on hardware a small team owns, which is why it drew 232,736 downloads last month against Mistral Large 3’s 1,071.
If you want the higher score and can live with the licence, Devstral 2 123B reaches 72.2% on SWE-Bench Verified. Read the next section before you commit to it.
3. Best for European Language Coverage: EuroLLM-22B-Instruct-2512
utter-project/EuroLLM-22B-Instruct-2512
Apache 2.0. 22.6B parameters, 32,768 context.
35 languages, including all 24 official EU languages. Built by a consortium spanning Instituto Superior Técnico in Lisbon, Edinburgh, Unbabel, Paris-Saclay, Amsterdam and Sorbonne, funded through EuroHPC and Horizon Europe. Its team calls it “best EU-made fully open model” and reports that on EU-language translation it matches or beats Gemma-3-27B, Qwen-3-32B and Apertus-70B.
The weakness is the context window. 32,768 tokens is an eighth of what Apertus and the current Mistral models offer, and a fifth of Spain’s ALIA. For chat and translation that rarely bites. For long-document work it does.
4. Best Small Model: Apertus v1.5-8B
swiss-ai/Apertus-v1.5-8B
Apache 2.0, behind an acceptable-use gate. 8B parameters, 262,144 context.
Released on 24 July by EPFL, ETH Zurich and the Swiss National Supercomputing Centre, and the most interesting European release of the summer. Version 1.5 adds image and audio input and an optional thinking mode, and quadruples the context window to 262,144 tokens, the longest of any European model you can download today. It also publishes its training data and code, which almost nobody else here does.
One caveat we will not paper over: the technical report with benchmark numbers has not been published. The team says it is coming “in the coming weeks”. Until it lands, the capability claims are the developer’s own and cannot be checked.
5. Best for Self-Hosting One European Language: Bielik 11B v3.0
speakleash/Bielik-11B-v3.0-Instruct
Apache 2.0. 11B parameters, Polish-optimised, 32 European languages.
Bielik is the most downloaded European open model that Mistral did not build: 464,298 in the 30 days to 4 August, up 10% from the 422,043 we measured on 27 July. Sixth overall, above every Mistral release this year. It is built by SpeakLeash, a volunteer open-science project, on public supercomputer time at ACK Cyfronet AGH. We tested it in July and it held up.
The same logic applies elsewhere: Llama-Krikri-8B for Greek, INSAIT’s BgGPT for Bulgarian, BSC’s ALIA-40b (Apache 2.0) for Spanish, Catalan, Basque and Galician. National models beat general models on their own language, at a size you can host yourself.
6. Best for Translation: salamandraTA-7b-instruct
BSC-LT/salamandraTA-7b-instruct
Apache 2.0. 7.8B parameters, 40 languages plus three regional varieties, trained on MareNostrum 5 in Barcelona.
A dedicated translation model rather than a chat model, refreshed at the end of July, and the right tool if translation is the job rather than a feature.
7. Highest Capability, If the Licence Allows: Mistral Medium 3.5
mistralai/Mistral-Medium-3.5-128B
128B dense, 256k context. 77.6% on SWE-Bench Verified, the highest European score we found.
It is not Apache 2.0. Its modified MIT licence reads, in the one clause that decides everything: “You are not authorized to exercise any rights under this license if the global consolidated monthly revenue of your company (or that of your employer) exceeds $20 million (or its equivalent in another currency) for the preceding month.”
For a startup that is a free frontier-class model. For a bank it is a licence you cannot sign. Mistral’s own documentation page still lists the previous generation under the old research licence and omits the Mistral 3 family entirely, so check the model card and not the docs.
8. The One to Skip If You Are a Business: Teuken-7B
openGPT-X/Teuken-7B-instruct-v0.6
CC-BY-NC-4.0. Non-commercial only.
Teuken covers all 24 official EU languages and was funded by the German economics ministry through the OpenGPT-X consortium. A real achievement, and not usable in a product. The v0.4 generation shipped a separate Apache 2.0 commercial variant; v0.6 did not. Nothing has shipped since August 2025, and German public money has moved to the Soofi consortium, whose 30B went into closed beta in July with its licence still described only as “permissive”.
What Europe Actually Downloads
These are trailing-30-day figures from the Hugging Face API, pulled in early August 2026. One repository per line, instruction-tuned text models only. Base checkpoints, quantised builds, community re-uploads and speech models are excluded, which understates several families: counting every Bielik repository together brings that total past 889,000.
| Model | Country | Size | Licence | Downloads, 30 days |
| Mistral-7B-Instruct-v0.3 | FR | 7B | Apache 2.0 | 5,000,363 |
| Mistral-7B-Instruct-v0.2 | FR | 7B | Apache 2.0 | 1,313,050 |
| Mixtral-8x7B-Instruct-v0.1 | FR | 47B MoE | Apache 2.0 | 843,240 |
| Mistral-Nemo-Instruct-2407 | FR | 12B | Apache 2.0 | 477,536 |
| Ministral-3-3B-Instruct-2512 | FR | 3.8B | Apache 2.0 | 476,923 |
| Bielik-11B-v3.0-Instruct | PL | 11B | Apache 2.0 | 464,298 |
| Ministral-8B-Instruct-2410 | FR | 8B | Mistral Research, non-commercial | 406,941 |
| Apertus-8B-Instruct-2509 | CH | 8B | Apache 2.0 | 360,741 |
| Mistral-Small-3.2-24B-Instruct | FR | 24B | Apache 2.0 | 314,137 |
| Devstral-Small-2-24B-Instruct | FR | 24B | Apache 2.0 | 232,736 |
| Ministral-3-14B-Instruct-2512 | FR | 14B | Apache 2.0 | 215,718 |
| Mistral-Small-4-119B | FR | 119B MoE | Apache 2.0 | 182,356 |
| Mistral-Medium-3.5-128B | FR | 128B | Modified MIT | 107,147 |
| EuroLLM-9B-Instruct | EU | 9B | Apache 2.0 | 40,164 |
| Apertus-70B-Instruct-2509 | CH | 70B | Apache 2.0 | 39,308 |
| salamandra-7b-instruct | ES | 7B | Apache 2.0 | 38,357 |
| Apertus-v1.5-8B | CH | 8B | Apache 2.0 | 8,827 |
| EuroLLM-22B-Instruct-2512 | EU | 22.6B | Apache 2.0 | 8,207 |
| MamayLM-Gemma-3-12B-IT-v2.0 | UA, built in BG | 12B | Gemma Terms of Use | 7,503 |
| Llama-Krikri-8B-Instruct | GR | 8B | Llama 3.1 Community | 5,127 |
| Apertus-v1.5-70B | CH | 70B | Apache 2.0 | 4,934 |
| TildeOpen-30b | LV | 30.7B | CC-BY-4.0 | 4,418 |
| Teuken-7B-instruct-v0.6 | DE | 7B | CC-BY-NC-4.0 | 4,360 |
| Velvet-14B | IT | 14B | Apache 2.0 | 2,947 |
| Llama-PLLuM-70B-instruct-2512 | PL | 70B | Llama 3.1 Community | 2,427 |
| ALIA-40b-instruct-2606 | ES | 40B | Apache 2.0 | 2,164 |
| BgGPT-Gemma-3-12B-IT | BG | 12B | Gemma Terms of Use | 1,503 |
| Llama-Poro-2-8B-Instruct | FI | 8B | Llama 3.3 Community | 1,357 |
| Mistral-Large-3-675B-Instruct | FR | 675B MoE | Apache 2.0 | 1,071 |
Three things fall out of that table.
- Usage sits at the small end and the old end. The five most-downloaded European models are two 7Bs from 2023, a mixture-of-experts from early 2024, a 12B from July 2024 and a 3.8B. Nothing released in 2026 comes close, and Mistral-7B-Instruct-v0.3 alone outdraws everything the company has shipped this year put together.
- New does not mean adopted, even inside one project. Apertus’s September 2509 generation outdraws its July v1.5 release by a factor of 41, because the ecosystem, the quantisations and the tutorials still point at the old checkpoint. Version numbers move faster than infrastructure.
- The most capable model is the least used. Europe’s largest Apache 2.0 model sits at the bottom of the table and every model above it is smaller. Download counts measure what fits on available hardware, not quality, and any ranking that treats them as a quality signal is misreading them.
The Licence Trap
Read the licence column again. BgGPT and MamayLM ship under Google’s Gemma Terms of Use. Krikri and Llama-PLLuM ship under Meta’s Llama 3.1 Community Licence. Finland’s Poro 2 ships under Llama 3.3. These five are built by public research institutions, funded by national ministries or EU programmes, and marketed as sovereign capability. Their terms of use are set in Mountain View and Menlo Park.
There is a second trap that has nothing to do with American companies. Ministral-8B-Instruct-2410 drew 406,941 downloads last month, seventh on the table, and it is not open source: it ships under the Mistral Research License, which reserves commercial use and tells you to contact Mistral for a licence. Sitting two rows above it, Ministral-3-3B is Apache 2.0. The two models have almost identical names. Check the licence file on the exact repository you are pulling, not the family.
That is not a hypothetical risk. Meta’s own Llama 4 Acceptable Use Policy ends with this: “With respect to any multimodal models included in Llama 4, the rights granted under Section 1(a) of the Llama 4 Community License Agreement are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union. This restriction does not apply to end users of a product or service that incorporates any such multimodal models.”
A European company cannot license Llama 4’s multimodal models. An American company can, then sell the result to European users. Whatever else that is, it is the opposite of sovereignty, and it is the strongest practical argument for the Apache 2.0 models here. We covered the wider question separately.
Where Europe Really Sits
Artificial Analysis tracks 337 open-weight models and scores them on a composite intelligence index. Version 4.1, live at the time of writing, ranks open-weight models like this: Kimi K3 at 57, GLM-5.2 at 51, DeepSeek V4 Flash at 50, MiniMax-M3 and DeepSeek V4 Pro at 44, MiMo-V2.5-Pro at 42, Nvidia’s Nemotron 3 Ultra at 38.
The first European entry is Mistral Medium 3.5, ninth, at 30. It is not Apache 2.0. Mistral Large 3, which is, scores 16. No publicly funded European model appears on the index at all: not EuroLLM, not Apertus, not ALIA, not Teuken, not Bielik.
One warning about leaderboards before anyone quotes that back. The Hugging Face Open LLM Leaderboard, still cited in a lot of 2026 coverage, retired in March 2025 after its own team said it “could encourage people to hill climb irrelevant directions”. Artificial Analysis scores are index points, not percentages, and are not comparable across index versions. Human-preference boards like LMArena measure something different again and reward style.
The honest summary is that the top of the open-weight field is entirely Chinese, that Europe’s best tracked entry sits at roughly half Kimi K3’s score, and that the publicly funded EU-language models are not competing on general capability at all. Their developers do not claim they are. EuroLLM’s team claims the best fully open European model. Germany’s Soofi consortium claims the best scores among fully open models on English and German. Those are precise claims about a narrower field, and they are defensible. “Europe’s answer to ChatGPT” is not.
Pick for the Licence and the Hardware, Not the Headline
There is no single best European open-source model, and any ranking that names one is answering the wrong question. There is a best coding model for your own machine (Devstral Small 2 24B), a best model for all 24 EU languages (EuroLLM-22B), a best small model (Apertus v1.5-8B), and a most-used model, which is a 7B from 2023 nobody markets any more.
What there is not, anywhere in Europe, is an open model near the top of the open-weight field. That position is held by four Chinese labs and it is not close. Europe’s real advantage is the boring one: licences you can actually sign, training data you can inspect, and models sized for hardware that exists in ordinary companies. On current evidence that is worth more to most European businesses than another ten points of index score, and it is the thing worth defending.
We test these models one at a time and publish what we find. Apertus 1.5 and BgGPT 3.0 are next.
Author: Akos Szima
See Also:
European LLMs in 2026: The Complete Map
What Is Bielik? Poland’s Sovereign AI Model, Tested (2026)
What Happened to Le Chat? Mistral’s Vibe Rebrand, Tested (2026)
Frequently Asked Questions
Mistral Small 4 (119B total, 6.5B active, Apache 2.0, 256k context) is the best general-purpose European model with a fully permissive licence. For coding, Devstral Small 2 24B. For all 24 EU languages, EuroLLM-22B-Instruct-2512.
Devstral Small 2 24B for code and Apertus v1.5-8B for general use. Both are Apache 2.0 and fit on a single well-specified machine; Apertus also takes image and audio input and has a 262,144-token context window.
Yes. Mistral Large 3, released in December 2025, is 675B total parameters with 41B active and ships under Apache 2.0 with no revenue restriction. Mistral Medium 3.5 is more capable but ships under a modified MIT licence that withdraws the rights grant from any company with more than $20 million in monthly revenue.
Under OSI-approved licences with commercial use permitted: Mistral Large 3, Mistral Small 4, Devstral Small 2, Apertus, EuroLLM, Bielik, ALIA and Salamandra, Velvet and the Mistral-based PLLuM models. Not open source despite the framing: Teuken-7B v0.6 (non-commercial), Ministral-8B (Mistral Research License, non-commercial), BgGPT and MamayLM (Google’s Gemma terms), Krikri, Llama-PLLuM and Poro 2 (Meta’s Llama licences), Mistral Medium 3.5 and Devstral 2 123B (revenue-capped).
Not well on general capability. Artificial Analysis tracks 337 open-weight models and its v4.1 index puts Kimi K3 at 57, GLM-5.2 at 51 and DeepSeek V4 Flash at 50, against 30 for the best-tracked European model, Mistral Medium 3.5. No publicly funded European model is tracked at all. Europe’s advantages are licence quality, EU-language coverage and, in Apertus’s case, published training data.
