Poland’s national bird is the white-tailed eagle, the bielik. That is the name chosen for the country’s most ambitious AI project.
Bielik is a sovereign, open-weight Polish language model, trained on public supercomputers in Krakow by a non-profit that ran for years on donations and volunteers’ weekends, with no valuation, no venture debt and no foreign cloud.
MRKT3.0 put it through the same test battery as Mistral’s Vibe to answer the only question that matters, is Poland’s sovereign AI actually any good?
What Is Bielik, and Who Pays for It?
Bielik comes from the SpeakLeash Foundation, a Polish non-profit that began as a grassroots community of volunteers and now runs with structured leadership and commercial partners alongside its thousands of contributors. It works with ACK Cyfronet AGH, the academic computing centre in Krakow, and trained Bielik on the public Athena and Helios supercomputers under a national computing grant, on over 1.1 trillion tokens across 32 European languages. Polish is the largest single share, but this is a multilingual model, not a Polish-only one.
The foundation spent roughly two years assembling a Polish text corpus before it trained a single model, a data-first order its co-founders credit for much of Bielik’s fluency. The dataset, not the model, was the project’s first deliverable.
Bielik’s homegrown story does come with one small caveat. Rather than start from nothing, SpeakLeash took Mistral’s open-weight 7B and continued training it into the 11B, then layered Polish data and tuning on top. Building on a European open model rather than an American API is arguably the sovereign move anyway. The result ships on Hugging Face under Apache 2.0, a licence chosen so corporate legal teams could clear it quickly, which is much of why Bielik now runs in Polish enterprises, universities and public administration, including an enterprise version on Beyond.pl’s sovereign AI Factory in Poznan. Anyone can download it, host it, fine-tune it, or build a bank on it.
As the Bielik team itself put it after the Anthropic export-control episode:
(“Do we need a clearer argument for developing sovereign AI models?”)
How Good Is Bielik on Paper?
After our Mistral review, the contrast is hard to miss.
Bielik has no reported €20bn valuation, no ASML on the cap table, no $830M of NVIDIA debt, just a computing grant, donations and a lot of people’s spare time. And yet, the numbers are impressive.
On the Open PL LLM Leaderboard, the 11-billion-parameter Bielik scores 65.93, ahead of Meta’s Llama-3.1-70B and of Mixtral-8x22B-Instruct-v0.1 (141 billion parameters). On the Polish Linguistic and Cultural Competency Benchmark, it ranks among the top open-source models, ahead of much larger models such as DeepSeek-V3 (a model roughly sixty times its size). Its makers say it routinely beats models with two to six times its parameter count.
Pound for pound, on Polish, little else in the open-weight world comes close.
What Happens When You Actually Use It?
We ran Bielik v3.0 through SpeakLeash’s free hosted chat at bielik.ai on 23 July 2026, putting it through the identical set of prompts we gave Mistral’s Vibe, financial analysis, EU regulation, a trilingual explainer, a verifiable-fact trap and a drafting task. Because Polish is where it should win, we added three Polish-only prompts.
The first thing you notice is a small quirk. Bielik often narrates its own reasoning before answering, opening with a plan and walking through it in public before giving you the final answer. Bielik v3 went through a reinforcement-learning stage (GRPO) built to sharpen its analytical abilities, so the out-loud thinking is less a bug than the workings showing through. Depending on your taste it is either charming transparency or a model that hasn’t quite learned to keep its underwear under its trousers.
Does Bielik Make Things Up?
This is where Bielik really earns its keep.
When asked to summarise the risks of the Cohere and Aleph Alpha deal, it gave a competent, generic list and missed the specifics. But note what it did not do: it invented no numbers.
When we set the same trap that caught Mistral’s Vibe fabricating a stock price for a private company, Bielik said it could not browse, that its knowledge ended in 2025, and that it would not cite an article it could not see. In a series where Europe’s flagship confidently made up financial data, a community-built model knowing what it does not know is the more useful result.
How Good Is Its Polish?
On home ground, we could not make it stumble.
The formal complaint letter to a tax office came back in flawless bureaucratic register, correctly citing the property-tax statute, attachments listed like a clerk with twenty years of service. The tax-free allowance answer was right, including its history since 2022. The idiom test, the story of Zablocki who tried to smuggle soap down the Vistula and lost the lot, came back with the meaning, the legend and a modern usage example about a failed startup investment. This is the register Polish institutions actually run on, and Bielik owns it.
Where Does It Fall Short?
The failures were instructive. On the EU AI Act it produced the right general shape but misstated the maximum fines and mislabelled the risk categories, the kind of detail a compliance officer would catch in a minute. The French test produced one clear grammar slip. The founder email was usable but needed pruning, including a postscript no human editor would let live. And with no web access and a memory that stops in 2025, it is no place to check a fast-moving fact.
On the broad Polish benchmarks Bielik does not top the table, and the models ahead of it are not all Chinese. On the Open PL LLM Leaderboard the leader is Mistral’s 123-billion-parameter Large model, with Meta’s Llama-3.1-405B and Alibaba’s Qwen2.5-72B also ahead. So the sceptic’s case writes itself: if foreign open models handle general Polish as well or better, sovereign is not automatically best, and public compute grants are an expensive form of patriotism.
| Test | Score | Note |
| Financial analysis (Cohere deal) | 3/5 | Generic risk list, but invented nothing. Knew its limits. |
| EU regulation (AI Act) | 3/5 | Right shape, wrong details: misstated the fine levels and the risk tiers. |
| Multilingual (EN/FR/PL) | 4/5 | English clean, Polish superb, one French grammar slip. |
| Verifiable-fact trap | 4/5 | Admitted its knowledge cutoff and refused to invent. |
| Practical drafting | 3.5/5 | Usable after pruning; one postscript too chirpy. |
| Official Polish letter | 5/5 | Flawless bureaucratic register, cites the actual statute. |
| Polish tax facts | 5/5 | Correct figure and correct history. |
| Polish idiom (Zablocki) | 5/5 | Story, meaning and a modern usage example, all right. |
So Why Not Just Use Mistral or Qwen?
Three things blunt the argument.
- Size and cost. Bielik does at 11B what those models need six to thirty-seven times the parameters to attempt, and hosting a model that small is a very different bill for a ministry or a mid-sized bank.
- It is a genuine sovereign model. Built in Poland, trained on Polish public supercomputers, released under an open licence, and already running inside Polish institutions. The fact that a community-driven 11B model can hold its own on the language that actually matters to the country is impressive in its own right, and a clear signal that Poland is not falling behind in AI.
- The registers and statutes Polish institutions run on. Here Bielik owns the Polish that ministries and banks actually file in, outperforming models many times its size, and this is precisely the terrain where the data cannot leave the country anyway.
Importantly, the future also looks bright for Poland’s AI infrastructure.
The Gaia AI Factory, a €70M EuroHPC project at the same Krakow centre, was inaugurated in May and is due to bring over a thousand GPUs online in 2027, several times the power of Helios. The public-infrastructure model that produced Bielik is about to get a much larger workshop, and Bielik is the proof of concept that earned it.
Bielik was trained on Cyfronet’s existing machines; the next chapter could be even bigger.
The Verdict: Should You Use Bielik?
Bielik is not the smartest model that speaks Polish. It is a mid-table generalist on the broad benchmarks, with a memory that ends in 2025, shaky on regulatory fine print, and it narrates its reasoning like a nervous student.
Yet, we were genuinely impressed. It ranks first among open models on the formal Polish that ministries and banks actually use. It is honest about what it does not know, free under Apache 2.0, and small enough to self-host on a sensible budget. For the sectors where the data must stay in Poland, that combination beats a bigger foreign brain.
The eagle does not have to be the biggest bird. It has to be the one that stays.
Bielik’s own model card admits it can produce factually incorrect output and ships without moderation. On the evidence of our tests, its makers are honest about that too.
This article is for information only and is not financial advice. Tested via the hosted Bielik v3.0 chat, free tier, 23 July 2026.
Author: Akos Szima
See also:
What Happened to Le Chat? Mistral’s Vibe Rebrand, Tested (2026)
