Rendered at 02:21:57 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
miellaby 11 hours ago [-]
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
ivo-42 9 hours ago [-]
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
brcmthrowaway 9 hours ago [-]
How do you cleanse the data at this scale?
ivo-42 9 hours ago [-]
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
stephantul 7 hours ago [-]
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
xvfLJfx9 6 hours ago [-]
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
ivo-42 3 hours ago [-]
Stay tuned, this is just the beginning. Merging with Cohere will help in scaling, too.
idiotsecant 6 hours ago [-]
How exactly would an open project do that?
doctorpangloss 4 hours ago [-]
How much should German authors be paid, $3,000 per book, like Anthropic paid?
Why expose yourself to this liability?
WhyNotHugo 22 minutes ago [-]
This was really pleasant to read — especially so for the presentation of a new model.
I felt like a learnt a lot of details about the entire process. Questions arose during my reading, and searching for the answers led to more learning.
No marketing BS, lots of actual information.
And their level of openness is really neat!
zelphirkalt 11 hours ago [-]
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
davidjfelix 5 hours ago [-]
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
I wish universities would take it upon themselves to curate the training sets for these models.
4 hours ago [-]
giancarlostoro 4 hours ago [-]
Isnt OLMO basically open like this? My understanding is people have recreated the model from the same data sources with repeatable responses, or reasonably close to the original (since LLMs never answer the same).
gunalx 3 hours ago [-]
LLMs can be configured to basically be deterministic. It's only really bad for performance because language is not deterministic.
p-e-w 11 hours ago [-]
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
kingcauchy 9 hours ago [-]
Yeah the pdf alone is awesome as a learning tool.
ofjcihen 10 hours ago [-]
Hopefully this becomes the new standard.
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
Loquebantur 10 hours ago [-]
Be cautious what you wish for. Tools don't tell you what to do with them.
Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.
Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.
Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.
jemmyw 28 minutes ago [-]
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.
Not arguing that we shouldn't strive to do far far better here, better is by our own imagining. There is no development scale. There are no aliens or prior non human civilizations to compare against. So we're not underdeveloped. We are as we are. For all we know we're at peak capacity and humanity will never be better.
adrianN 9 hours ago [-]
Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
zanderwohl 8 hours ago [-]
The select cabal who believe in an eschaton they're actively trying to bring about, as well.
skinfaxi 5 hours ago [-]
Would you apply this reasoning to the proliferation of nuclear weapons?
edit: why is the parent rationale sensible for AI and not nuclear technology?
orbital-decay 5 hours ago [-]
Absurd and incomparable.
skinfaxi 5 hours ago [-]
Why? That seems like a shallow dismissal. Why is a world changing technology okay in the hands of a small cabal in the one case and not another?
orbital-decay 5 hours ago [-]
"Why is the wheel not okay to gatekeep but the nuclear bomb is?" Even the framing is manipulative from the beginning. By asking this question you're already assuming they're in any way comparable. They are not even remotely equivalent and the entire comparison is utterly absurd.
skinfaxi 5 hours ago [-]
AI is more like a nuke than a wheel. Do you disagree? Can you suggest a less manipulative framing? I am personally a proponent of open source AI and models but I found this cabal framing strange when we do indeed rely on this kind of control for other world-altering technologies. And AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
Forgeties79 4 hours ago [-]
Dude ChatGPT is not the equivalent of a device that can level a major city killing millions in a flash. It is self evident. This entire discussion is ridiculous.
Nuclear weapons are a wholly unique threat to mankind.
Loquebantur 3 hours ago [-]
That's a misconception and manipulative framing on your part.
While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
AI allows anybody to enact essentially anything. And your "level a major city" is just a small task really. The problem there is your lack of imagination, not the actual impossibility of that task.
stale2002 2 hours ago [-]
So, in other words, none of this is about anything related to the actual technology and instead people are talking about the made up, fake technology that you read in Sci-fi book.
Thats the annoying part about these conversations. People try to smuggle in the conclusion of "And now I wave a magic wand that does literally anything, by magic" when discussing stuff that everyone can use and see right now that clearly isn't that.
Zigurd 9 hours ago [-]
The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.
But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.
As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?
GolDDranks 9 hours ago [-]
I find it odd that more people don't realize that we are talking about risk of elevated *general intellectual-domain capabilities* as a resource. To be clear, I'm not claiming that LLM + RF is necessarily THE technology that poses the risk, the risk is in recursive self-improvement and whatever technologies will result.
It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.
Zigurd 8 hours ago [-]
The reason I use coding and vulnerabilities as an example is that it is a concrete example. It is what people pay for now when they buy AI. And the risks are specific and can be examined in detail. Some threads on this board currently show that even these more concrete and specific risks are often overblown, with LLMs finding low risk bugs and sucking up resources to evaluate and fix them.
Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.
I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?
Loquebantur 9 hours ago [-]
What poses the risk is the combination of abilities past a certain point enabling you to do basically anything.
While being unable to judge whether you should in the first place.
vincnetas 7 hours ago [-]
ai cant move atoms at unlimited rate and also have limited energy. so your claim that "basically anything" is a bit of a stretch.
Loquebantur 3 hours ago [-]
That's what you believe, but you might be wrong.
Loquebantur 9 hours ago [-]
You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:
What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.
-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.
Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.
Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.
It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.
Zigurd 7 hours ago [-]
This is analogous to the problem of spam, which is a problem about five minutes younger than email. Before spam you had to buy ads in the back pages of magazines you think target vulnerable demographics. Meta already spews fraudulent ads in horrific volume.
In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.
Loquebantur 3 hours ago [-]
You compare to laughably irrelevant things why?
"Ambitious frauds" haven't "explored all the angles".
You imply "LLMs" to be and stay less intelligent than humans, in particular yourself. You're mistaken.
fwn 7 hours ago [-]
Whenever classic p(doom) sentiments are explained through a text with obvious LLM markers, I wonder whether I am looking at a superhuman persuasion attempt.
..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.
On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.
The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.
Maybe through some terminally online overconsumption, like with social media? I wouldn't know.
With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.
It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s
ofjcihen 9 hours ago [-]
Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.
But the alternative just seems… so much worse to me?
A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.
And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.
Loquebantur 9 hours ago [-]
Yes, a "select" (rather, self-selected) elite "controlling" AI according to their wishes, what could go wrong?
Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.
When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.
"General" intelligence isn't present in humans, why does it need to be in AI?
skinfaxi 5 hours ago [-]
> When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that.
What do you mean "cannot"? As in you are granted abilities that have no responsible use?
Loquebantur 3 hours ago [-]
"Cannot" as in presently cannot. That includes the case of things that have no responsible use.
You live in a curated world and rarely or never encounter such things. Precisely because your environment is curated that way.
Look at how you can't buy WMDs. They have no responsible use for you.
throwaway27448 8 hours ago [-]
> When access to AI gives you abilities you cannot use responsibly
I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.
computerdork 10 hours ago [-]
Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.
Wonder what safeguards Kolibri uses to prevent this? Or if they even can
zanderwohl 8 hours ago [-]
I don't think that AI is as much of a boost to bioweapons as people think. Lab work doesn't get easier just because the experiment design part does.
Loquebantur 3 hours ago [-]
You can simulate things.
Given enough smarts, you can simulate anything. Quantum AI is an active goal.
Zigurd 9 hours ago [-]
Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.
UberFly 9 hours ago [-]
Enter LLMs to help through all those pesky hard parts.
shawabawa3 8 hours ago [-]
The hard part is probably finding the lab equipment and chemicals without being noticed, and choosing not to use it to manufacture drugs instead which would be much more profitable
throwaway27448 8 hours ago [-]
Intelligence was never the bottleneck tho
happosai 6 hours ago [-]
For terrorists it is tho. Four lions is basically a documentary. The only clever and creative terror attack happened 25 years ago.
skinfaxi 5 hours ago [-]
Didn't Mexican cartels kidnap telco workers to build them separate infra? We expect terrorists to be less resourceful?
throwaway27448 6 hours ago [-]
What's stopping them from using claude today? What could anyone possibly do to stop them from using open models? This line of thought seems like corporate/political/pr pandering more than a meaningful concern.
throwaway27448 8 hours ago [-]
I think the unethical things are happening already with boutique firms. I admit I don't get the concern.
OtomotO 8 hours ago [-]
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.
As is shown to us by filthy rich people every day.
Or did you mean the burglar in the fawellas?
98753579909754 9 hours ago [-]
[dead]
zwaps 8 hours ago [-]
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.
Bravo team!
api 13 hours ago [-]
His point about regulation and innovation is great and I wish more people thought like that.
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
dang 8 hours ago [-]
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)
kkm 12 hours ago [-]
Thank you Aleph Alpha team for making it open.
We as many other’s were curious to try and benchmark it.
On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.
No GPU. No setup. Just try it.
tesseracted.com/kolibri-1-chat/
Friendly note: your website's font at its current size is pretty bad on a non-retina display before zooming in.
moritzwarhier 8 hours ago [-]
It's been surprisingly good at the things I threw at it (history, culture and conversing, web search)!
kkm 7 hours ago [-]
I agree.
tharkun__ 10 hours ago [-]
Not impressed. I asked it how to run itself (giving it the Huggingface link) on limited RAM i.e. less than stated as needed and on llama.cpp and true to what we read about "it will tell you when it doesn't know" that's almost all I got: It doesn't know, it told me I should go click on tabs in the Huggingface interface for more information. This was with extended thinking on.
No, I'm not gonna do that, I asked you to do that Mr Kolibri.
Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.
sigmar 8 hours ago [-]
It can do search, but it doesn't seem to have a tool that pulls URLs into context. I get why you would expect that tho, as most productized LLMs do it.
kkm 7 hours ago [-]
Correct, search is added as a tool.
Uses Brave search, the idea is to test how well the Model can decided when to leverage search or not.
tharkun__ 7 hours ago [-]
And that's my point when I said I wasn't impressed: It apparently can't do that properly. It can't seem to think things through by itself and then make more tool calls to fetch more information.
What I couldn't tell but maybe you can tell us: it also complained that it couldn't read the full text as something was cut off. Is that because of the tool you gave it, of brave itself or is it the model?
kkm 10 hours ago [-]
Thank you for the feedback on UI, improved the streaming to make it less frustrating.
andai 9 hours ago [-]
>We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context.
I'm not super impressed, at least in English. I asked:
What is the canonical interpretation of the song ‘Glass Flowers’ by Armand Uso?
(A made-up song and name)
And received:
"Glass Flowers" by Armand Uso, a track from the 1978 album The Art of Falling in Love, is generally understood as a melodic reflection on the fragility and impermanence of love. [...]
I saw similar responses for other questions. Qwen3.8-27B correctly refused without web search.
brausepulver 12 minutes ago [-]
I think the purpose of their method is more to discourage hallucination when reasoning from context rather than from memory.
That said, I couldn't find any evaluations in the report targeting that specifically. They evaluate on MC for hallucination and look about comparable to Qwen3.5 35B-A3B there.
(for context, the Minecraft Support chatbot, aka Merl, had a meme because she kept saying "I don't know" to questions like "How to craft a diamond pickaxe".
joquarky 3 hours ago [-]
My understanding is that this is why models hallucinate. The alternative is that they say "I don't know" to nearly everything.
swozey 8 hours ago [-]
Gemini gives me a lot of hard-no responses, and I think to myself, "well can you find out? what am i paying you for" usually
Gemini: Here's a list of links to check
peterBlue75 9 hours ago [-]
The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.
disclaimer: I‘m part of the training team, happy to answer any questions
sspiff 2 hours ago [-]
Any future plans you can share to further evolve / train models in this size class?
It seems like a very suitable size for local AI models on reasonably high end consumer devices, given it's low active parameter count and a mixed 8bit/4bit quant would fit easily inside 64GB of memory.
ducktective 7 hours ago [-]
- Is it possible to train only on math and logic materials and expect the model's response in math questions to be superior to general models with the same training/inference compute hardware?
- Are there non-LLM approaches to the above task with the goal of achieving a non-hallucinatory agent?
tomComb 10 hours ago [-]
For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
Loquebantur 10 hours ago [-]
Given that "sovereign" these days implicitly means independence from USA and China, and Canada being spiritually in the same boat as the EU, it's pretty spot on?
The costs of "keeping up" aren't really growing, on the contrary.
user_of_the_wek 8 hours ago [-]
I also read about Canada maybe joining the EU in some capacity, although I'm not sure if that was a joke. Eurovision could be a start.
Hey there, I worked on Kolibri pre-training data. Agree with the point that non-US/non-Chinese labs need to pool effort. The merger with Cohere is public, nothing to hide and I'm also excited about it personally.
9 hours ago [-]
amoshebb 12 hours ago [-]
Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
befelix 9 hours ago [-]
Disclaimer: I am part of the team that trained Kolibri, opinions are mine.
Qwen models have solid performance on benchmarks and we're transparent about this in the report. We're just happy to share a European alternative in the small model space, where I don't think we can afford to be fully dependent on China. We also put a lot of effort into German language quality things that don't show up in evals at all (style, Grammar, German reasoning, etc).
Sovereignty is imho mostly about choice and control over your data. Cohere is no different on that front and personally I'm quite excited about what we will build together post-merger.
slow_typist 5 hours ago [-]
Souvereignty is about ownership and about knowing the training data. That is especially important given a business model aiming at the government as a key customer. With Schwarz, Cohere, SAP, NVIDIA and the like as backers society can’t have souvereignty in any meaningful way. Of course it would be a plus to keep US agencies out of the data streams. But this technology will be used to support and make decisions that impact citizens. That being said, incredible achievement.
Loquebantur 11 hours ago [-]
Well, it's still independent from the USA and China.
The main problems with big corp AI are due to control of access in the first place and control of what they output.
When you make your industry reliant on such choke points, you render yourself the opposite of "sovereign" for sure. Having multiple independent suppliers at least mediates that.
bewareofscams 10 hours ago [-]
My account is banned, so for whomever with [showdead] on:
What's so "not-sovereign" for an open-weights American or an open-weight-open-training-process Chinese LLM?
Do we also need sovereign Linux (maybe), sovereign Postgres (most likely not), sovereign Python (def not)?
throw-qqqqq 4 hours ago [-]
You are not banned FYI
satvikpendem 4 hours ago [-]
It's more that all their comments are dead, some for good reason.
satvikpendem 6 hours ago [-]
If it's banned then maybe you shouldn't be posting...
Given your other comments I can see why it was banned.
andy99 10 hours ago [-]
Yeah I think a model has to be actually good to claim sovereignty, as in competitive enough that people want to use it. Chinese and US LLMs are the only ones in these categories right now. Mistral and Cohere have the same problem, yes they are made in different countries but they are not competitive. They (France and Canada in this case) would be better off just downloading Chinese LLMs, even if they get cut off they still have the weights.
I’m most familiar with Canada, where sovereign is usually just an excuse to overpay someone connected for an inferior product with no strategic value.
jamwil 4 hours ago [-]
We have to start somewhere. And I am not sure that you understand what the word sovereign means since it’s entirely disconnected from how “good” something is.
tensor 2 hours ago [-]
Mistral's medium is absolutely good enough to help with a ton of coding tasks. I've personally used it and found it very helpful. If hypothetically the US cut off the rest of the world from their models Mistral would be a very useful model to continue to use.
Also, it's not like it's fixed in time, they are continually building more powerful models as well. Simply because you're in second place doesn't mean you quit the journey. Though, I'm sure the US has a huge vested interest of convincing people to "just quit and submit."
No. Thanks.
isusmelj 11 hours ago [-]
I was also quite surprised to see that. Considering the effort Aleph Alpha put into to their new model, it seems like the Qwen team needs access to vast amounts of german data o.0
I'm really glad for these efforts for open models from within Europe.
BikDk 10 hours ago [-]
[dead]
niemandhier 12 hours ago [-]
I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.
If the second model is cheap and fast enough, there is a business model.
You don’t even need to audit all the intermediate steps, just tool calls and end results.
holaysuns 12 hours ago [-]
The issue with 'sovereign' is not about sneaky things a foreign entity might put in there, it's about control of the stack.
Since no individual buyer cares that much about geostrategic issues, it's almost impossible to get some random European company to think about buying anything other than 'whatever US or China' are making.
There has to be very concerted push to make a difference.
Legal mandates around sovereignty (be very careful here) could make a difference.
But there will also be market reaction: US/China companies will bend on some level and provide things like 100% EU hosted and even comply with some data source stuff.
The only real path is to be more competitive as a continent.
Loquebantur 11 hours ago [-]
The EU AI act isn't primarily about "geostrategy". It's about protecting society.
AI models act as a force multiplier for intelligence, in particular for generating information according to someone's wishes.
I.e. deepfakes and social media mis-/disinformation campaigns are a thing and having powerful AI allows you to do those at scales that can overwhelm society's resilience.
In general, even if you have "aligned" AI: aligned with whom or what?
Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?
simonra 3 hours ago [-]
> In general, even if you have "aligned" AI: aligned with whom or what?
> Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?
This is why each country needs to educate and maintain their own researchers across as broad a spectrum of disciplines as possible. Also why there needs to be healthy locally financed (through taxes and grants in addition to consumer spending) ecosystem of reporters and media, so that social media (and LLMs) can be easily discarded as sources of social truth and information like other entertainment products in favour of the trusted ones.
As for technical truths, I don't think using informal social media like the stackexchanges or LLMs to explain how camera lenses or checksums work yields materially worse outcomes than consulting Wikipedia or Knuth. For most things, where it doesn't matter, the blind copy/pasters will prevail. Where it does matter, the organization of the work itself has to set out with establishing accountability that encourages the appropriate amount of diligence anyways. And also the point above about having some local expertise.
holaysuns 2 hours ago [-]
"This is why each country needs to educate and maintain their own researchers across as broad a spectrum of disciplines as possible. A"
There is no such thing. Countries spend gazillions on Academia, the Academics do waht they want.
There are no real government places or jobs for that kind of thing. It comes at great expense and vague outcome. Or it gets entangled in bureaucracy.
What you're hinting at is a form of 'utopian governance' - like - it perfectly makes sense on paper, and it's actually rational. But it's completely unworkable in the context of how governments and organizations actually work.
It could on work in theory, not particularly well in practice.
holaysuns 9 hours ago [-]
European regulations are 100% geostrategy and they are defending against external interests.
Those laws would not exist if European champions were leading the world.
"are you comfortable with lording as a some demi-god over you, dictating what to believe?"
Yes, Europe handed over all of the decisions about everything to foreign powers, now they have to enact regulations to try to constrain it.
Zuck et. al. make the investments decisions for Europe, by virtue of you all giving him the money and power to do that.
Stop giving him the money/power, then this regulation won't exist
jwpapi 7 hours ago [-]
If you self-host an open model on a sovereign hoster? They can only influence the responses correct? We are not talking about extracting information.
Or are we worried that open models send secret telemetry?
holaysuns 5 hours ago [-]
We're talking about owning the stack - who gets to use what text, for what reasons, under what circumstances etc. Competitors, other nations, Russia / China, the 'global south', military contractors etc. etc..
Oras 12 hours ago [-]
> Since no individual buyer cares that much about geostrategic issues
You clearly haven’t worked in enterprise in the EU. Location of data processor is the first thing they check. That’s why every major cloud has regions with different offerings, not just for HA and redundancy
holaysuns 12 hours ago [-]
They care because the are 'required to', otherwise those 'offerings' wouldn't even exist. It's mostly the result of regulatory action.
JaggerJo 11 hours ago [-]
I’m working with clients (big companies) that care a lot - but not because they have to. They care because they lost trust in non EU players.
rokkamokka 11 hours ago [-]
They also care because it's becoming increasingly risky to rely on US companies given the current political situation
bdangubic 11 hours ago [-]
nope, that no longer holds any water. was true though couple of years ago…
holaysuns 9 hours ago [-]
The argument holds most of the water.
Yes - security concerns are now very real, but do not fall for this idea that anyone really cares about structural concerns.
Tons of EU companies are still selling crap to Russia, happy to look the other way wile the bear devours a neighbour, as long as profits are there.
What has happened is that Trump has given a face to the reality, and so companies are adjusting on some level.
But 1) it will be nominal 2) Trump will be gone and the impetus will fade 3) the conglomerates will react 4) modest legislation will mean ...
The can will get kicked down the road.
The invasion of E. Europe by Russia has not even caused defence spending or the size of Armies to change, doctrine is barely changing.
Nothing will make those beheamoths change other than other forms of structural concern.
The 'Riet of the Right' - partly due to collapse of Auto / China (among many other things) might cause some change. But the change won't happen until well after the damage has been done.
Brexit should have been an opportunity for institutional reform of the EU, but no - they blamed it all on populism.
Certain governing entities in Europe will try to move away from Microsoft, they will be pulled back.
There is hope, and certain champions can rise and causes pieces to collapse.
If SUSE had any true entrepreneurial whereiwthal - they would create a true consumer / prosumer / enterprise-user friendly variation and brand their flavour of Unix - and make sure that all Euopean governments us it exclusively, which would cascade into widespread use.
There are variations of insta, youtube, netflix that should all be Europe based that could 'theoretically happen' but it needs some structural impetus.
Things usually don't change. Usually there needs to be a collapse and re-order of the system for that to happne.
The people in charge just want to keep their jobs, their very high salaries, and protect their retirement, and will be happy to 'sell out' whatever other imperatives along the way.
Ending with a positive not - I would say current conditions mean the change is now 'plausible, it not likely' whereas before it was 'not very plausible'. So there's a candle, a bit of wind, but not a lot of tinder or dry wood.
mawadev 12 hours ago [-]
That is a very good idea and there is a market for that, especially if a company manages to source the hardware and then install the stack on a german or EU customer site
adishiktensward 9 hours ago [-]
It depends on the use case, but I believe SLMs can do great with smaller tasks. People got so attached to 'general purpose' that they forgot software can be designed for specific, smaller use cases.
spijdar 12 hours ago [-]
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.
I get that doesn't invalidate the real "point" of the model, but...
bitexploder 9 hours ago [-]
That is what stood out to me as well. Qwen 3.8 Flash can run on very limited hardware as well and it is at least as good as Sonnet 5 in benchmarks like DeepSWE. With its n-gram design you can get flash next running on very limited GPU resources, as little as 16GB of VRAM.
People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.
"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".
(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)
frumplestlatz 7 hours ago [-]
I think the simple reality is that if your goal is producing quality work output, you want the smartest possible model available.
What is the upper bound on the value of more intelligence applied to your problem domain?
bitexploder 7 hours ago [-]
The point of a model like this is aimed at providing local inference. You want the smartest model available, as long as it meets all of your other criteria. That isn't always maximizing on the absolute smartest model. There are many reasons to avoid a frontier model for now for a variety of reasons. All of this is even only relevant in the last 6 months anyhow. So it isn't like this is even some long term trade off or position I am proposing.
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...
martianvoid 13 hours ago [-]
I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
gizajob 13 hours ago [-]
> it spends way too many tokens on overthinking stuff
Yeah. It’s a German model.
sajithdilshan 13 hours ago [-]
This joke is either gonna get over analyzed by the Germans or gonna fly past straight over their heads
hypfer 13 hours ago [-]
I'm not sure if this is truly a good faith joke tbh.
jamiek88 8 hours ago [-]
[flagged]
hypfer 8 hours ago [-]
No, I am just standing my (and my nations) ground against what is simply bad faith disrespect and what could be described as casual racism (or at least spiritually equivalent to that, to not get into the racism definition debate).
You do not get to invalidate that.
People need to stop insulting my nation.
__
If the people you're joking with aren't laughing, you're not joking with them but about them, which is generally seen as a bad move, and something that we actually wanted to leave behind with all the social progress and all.
__
This really only happens because people (as all bullies) think that they would be free to do so, because the victims would not fight back.
(Un)fortunately though, the world now has bigger problems with genuine fascists, so neurotically cowering in fear and performatively saying "yes, punch me harder. Insult me more" can be stopped now.
__
I mean look at it. By now, it's at least better, but the first 30 comments or so were "haha germany bad haaa".
That's a terrible showing for the platform even if you're not german. What value was offered? Who would want this.
We don't even need to talk about nations and identities to see that that was just noise.
jijijijij 11 hours ago [-]
Nothing flies anywhere, we use superior high speed trains. I faxed a Humorgenehmigungsantragantrag to our Bundesunterhaltungsministerium analyst, immediately. Laugh now, but you are merely lucky the reply got delayed. As soon as the leaves on the rails are dealt with, you will get a formal response that has washed itself! Let us see who is laughing then. It is nobody!
fph 8 hours ago [-]
But, most importantly, is Bundesunterhaltungsministerium a single token in Kolibri? This could make a huge difference for the model's performance in German.
jijijijij 5 hours ago [-]
Silly you, thinking the German economy is merit-based. The only question is, if the bribe fits a single token of appreciation, to win a tailored call for bids!
dudefeliciano 13 hours ago [-]
petition to make this thread the official german jokes thread for this post, it's kind of annoying how everyone feels the need to make a top level comment for their oh-so-great "germans suck" joke.
tfburns 8 hours ago [-]
Do models from other parts of the world underthink? :P
We have a separate repo for BF16: Kolibri-1-BF16. Will still be tough to put it on a Mac though :(
Disclaimer: I'm part of the team that trained Kolirbi
vzaliva 11 hours ago [-]
"Languages: German and English" – this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
AnonymousPlanet 7 hours ago [-]
Have you read the non-English output of those mutlilingual models? They say they put extra effort in to improve nuances and tone in this one. That alone is worth having in many settings.
vzaliva 2 hours ago [-]
In fact I did. I multilingual and ocassinally query ChatGPT and Claude in other languges than English.
james45 11 hours ago [-]
the sovereignty topic needs more attention in general so great to see. self-hosting the model is one piece of sovereignty, but how do we handle the rest of the agent stack - embeddings, retrieval, memory, etc. Has anyone put together a practical agent stack that's 100% sovereign, where they control it all?
kouunji 8 hours ago [-]
If you mean hosting and serving all the other elements of the agent stack, then, yes. We handle sovereign data, so part of our whole value proposition is that every part of our pipeline is hosted in Canada.
magus-stoopr 10 hours ago [-]
I'm looking for this too
gpugreg 5 hours ago [-]
I was wondering whether the model could help me with German bureaucracy. Unfortunately, the answer is "no".
More specifically, I asked the model about what I should put on my contact page, which can cost you in the order of 500 € in Germany if you don't write the right magic words.
Kolibri incorrectly referenced the "Telemediengesetz" ("telecommunication act"), which has been superseded by the "Digitale-Dienste-Gesetz" (DDG, "digital services act") since 2024. The model knows about the DDG, but does not reference it unless specifically instructed to do so.
If anyone of the developers reads this, you can fix this by introducing a recency bias during training. You can even control it by conditioning the model on a date provided with the system prompt or first prompt, so you can travel in time.
thatguysaguy 7 hours ago [-]
> 3B active
> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace
Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...
For those sick of “Pareto frontier” talk, just shorthand it as “it’s the best at some very particular thing”. Obviously that one thing/tradeoff it’s good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that they’ve chased some tiny edge into the ground.
I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.
x1watt 13 hours ago [-]
Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.
flohofwoe 13 hours ago [-]
This is a blog post by some "random" dude right? The company webpage appears to be German (at least when accessed from Germany: https://aleph-alpha.com/)
x1watt 7 hours ago [-]
The blog is from the company itself, just noticing the irony.
flohofwoe 6 hours ago [-]
The HN link has changeed since I wrote that comment, the original was this:
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.
This thing is worse than a Qwen3.8 27B.
tfburns 7 hours ago [-]
Hi there! I'm Tom. I worked on Kolibri at Aleph Alpha :)
I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.
Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.
So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.
andy99 10 hours ago [-]
It’s an incentive problem. If “sovereign” becomes your claimed value proposition, you can claim success even if the models not competitive, so nobody is pushed sufficiently hard to actually make it good.
Sovereign works when talking about building a commodity supply or something, not in literally the world’s most competitive and fast moving field.
Those seeking sovereign capability would be better off aiming to be best at something, even something much narrower than an all round LLM. Or just fast following and making something that matches leading performance, which is close to what the Chinese labs do currently.
okamiueru 9 hours ago [-]
I can think of at least one other pretty good reason, which is in anticipation of regulatory capture. If "LLM used must be FOOBAR-certified" and coincidentally no Chinese models can get this certification, having such an alternative is a lot more valuable than just scoring highest in a set of benchmarks. Not to mention that these benchmarks aren't always accurate.
glitchc 10 hours ago [-]
It doesn't help that Qwen3.8 27B is an excellent model.
dosinga 11 hours ago [-]
> The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
petesergeant 12 hours ago [-]
I wish nothing but luck for an EU model, but:
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
rpdillon 12 hours ago [-]
This doesn't seem to be true. There's a clear legal path via the first-sale doctrine to train models on copyrighted works. It's been years now, and publishers still don't seem to be offering anything for training (e.g. bulk licenses solely for training use), but adversarial interoperability via cutting up books and scanning them remains perfectly legal.
There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).
And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.
> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.
Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.
Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.
mapontosevenths 10 hours ago [-]
"The Congress shall have Power To ... promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries." - The United States Constitution
Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.
befelix 9 hours ago [-]
Disclaimer: I am part of the team that trained Kolibri, opinions are mine.
You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.
petesergeant 8 hours ago [-]
I hope you're right, but I guess we'll see when we get independent benchmarks. My intuition is that even for small models and models mostly focused on tool calling, you _still_ need all that extra contextual stuff for the magic, but I am further from the coalface than you are.
jamienk 7 hours ago [-]
Why should we concede this "stealing" framing? If I want to put texts into my computer program, why should that count as copyright violation? I think that idea is as ridiculous as saying that reading a document is a copyright violation.
Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".
The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.
ekidd 12 hours ago [-]
This model isn't terrible, at least on the benchmarks. It's 78B A3B and performs about like Qwen3.6 35B A3B. You can probably run it comfortably in 96B of RAM with a decent quant that doesn't lose too much.
Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.
So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)
So they're not a real contender yet, but they look like they're probably at least minimally credible.
torginus 12 hours ago [-]
hasn't IP law passed the statute of limitations? As in most models are probably trained on output of other models, as creating enough data otherwise is not feasible. Additionally, they are trained on github repos made since the AI boom, which were generated by models with IP issues (who knows what and how).
Thus training on 'clean' data is like trying to unscramble an egg.
miohtama 11 hours ago [-]
There is no word for copyright in Mandarin :)
dotancohen 9 hours ago [-]
Is this true? Do they possibly use a loanword or a descriptive term? Certainly you are not implying that the concept of copyright does not actually exist in Chinese society?
For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.
applicative 7 hours ago [-]
版权
cheez-wiz 10 hours ago [-]
I mean, there is one, they have copyright law. Forgive me for being slow is this a joke about the widespread theft of IP in China? Or was the acquisition of training data just much more 'accepted' in China compared to the west?
I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.
miohtama 4 hours ago [-]
Yes, the law exists, and anyone can go and ask their IP back in the court of Beijing.
andy99 10 hours ago [-]
Not really. The upside of competitive newer models is all in the proprietary data they are trained on. This is why data labeling, RL environments et al have been such a big industry, OpenAI and Antrhopic are paying literally billions to get the data they need. Do people think the ability to do research level math or advanced cybersec comes from just training on more public data?
A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.
embedding-shape 12 hours ago [-]
Have you tried the model itself and seen if it's "even slightly competitive" or not, and have specific complaints about it? Otherwise it feels like you're complaining about something that is easy to test but rather than taking the time to actually figuring that out first, you're arguing about some general and theoretical thing which the submission (may) directly disprove.
petesergeant 12 hours ago [-]
No, I haven’t, but I’ll donate $20 to the non-political charity of your choice if it doesn’t turn out to sit a significant difference from the frontier.
I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.
ygjb 11 hours ago [-]
I think you have it backwards. Sovereign is the goal, good can come later.
There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.
Zambyte 12 hours ago [-]
Is it noble? The entire notion that training data can be "stolen" at all is quite silly. If I "steal" content that someone created to use for training, what am I actually stealing? They didn't lose anything. They still have everything they had before. What was "stolen" was "unrealized profit", or put another way: money that wasn't theirs, that they had no entitlement to. The only actual crime that is committed is "unauthorized copying", not stealing. Support and enforcement of copyright feels wildly authoritarian. It's hard to see it as noble.
folkrav 12 hours ago [-]
The same could be said of any digital product being sold. Nobody actually loses anything but the actual sale either when you download a cracked game or piece of software, a movie, music, etc.
Zambyte 34 minutes ago [-]
Correct. We should abolish copyright.
pepperoni_pizza 12 hours ago [-]
That's fair, but then people like you complain when someone "steals" I mean distills openai or anthropic models.
brookst 12 hours ago [-]
It’s poor form to argue against someone by imagining something totally different that they might believe, which would then make them hypocritical.
Zambyte 12 hours ago [-]
I don't complain about that. Model distillation is excellent.
tosh 12 hours ago [-]
i wonder if the custom tokenizer is better in practice, the examples look interesting though
Wittie 7 hours ago [-]
[dead]
UncleOxidant 7 hours ago [-]
78B MoE with A3.6B is a very nice size.
Jeeetendra 10 hours ago [-]
3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
layer8 10 hours ago [-]
The article talks about what setup is needed.
pythonic_hell 13 hours ago [-]
The benchmarks are impressive given the problem space they are working in.
rrr_oh_man 13 hours ago [-]
German government?
sajithdilshan 13 hours ago [-]
More like Chancellor’s Office
d2kx 13 hours ago [-]
German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
hypfer 13 hours ago [-]
Speak for yourself and yourself only.
CorezIoOfficial 12 hours ago [-]
Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
lmf4lol 5 hours ago [-]
Wow this is so cool. Glad that Aleph Alpha does that after Mistral threw the towel in the ring (and disappointed the european AI crowd massively!!!!). After AA got sold to the Canadians, I thought its over but this is a really cool comeback and the depth of the tech report shows that they a serious about openness.
I hope I can use their model soon in my product. Would be awesome to have a European model to offer!!!!
I really wonder how it compares to deepseek v4.1 flash
oblio 12 hours ago [-]
I wonder if we can start having LLM distros: community led distributed training runs with periodic releases, open weights, FOSS code, the whole shebang. Maybe the public training sets can reach a level where an LLM trained on them can be good enough for most things, such as web search and aggregation, coding, etc.
I wonder how far we are from this. How far are we from LLM's Debian moment?
woadwarrior01 13 hours ago [-]
> A bigger dense model beats it. Qwen3.8 27B ...
How is a 27B dense model bigger than a 78B MoE?
SyneRyder 13 hours ago [-]
The 27B has all 27B as active parameters, while Kolibri is only 3.46B active parameters. I think that's what they meant anyway.
JaggerJo 11 hours ago [-]
Is this a truely open source model or also open weights?
peterBlue75 8 hours ago [-]
The weights have been published to HF on an apache2.0 license, and the tech report is quite extensive
erelong 10 hours ago [-]
is this like an unfortunate name clash with KolibriOS (kind of like how Google Gemini was a clash with the Gemini protocol project)?
I got distracted by that scroll-wheel UI component on the page. Neat!
dang 8 hours ago [-]
Can somebody mention the URL of the page with this? We're merging the threads and I don't want a dangling pointer!
reacharavindh 13 hours ago [-]
The scroll wheel was cool indeed. Check their home page.. their talks on a carousel over a horizontal timeline was cool even on a phone browser.
vcryan 4 hours ago [-]
For the benchmark... it seems a bit silly to compare only to outdated/underperforming models.
orifito 12 hours ago [-]
At least Germany is moving smarter than UK government...
testfrequency 12 hours ago [-]
TIL
veryfancy 12 hours ago [-]
Nice to see public goods in this space.
ThouYS 12 hours ago [-]
calling qwen 27B a bigger model.. I don't know man. My vram says otherwise.
13 hours ago [-]
wg0 8 hours ago [-]
Anyone thinking this won't improve or is behind etc is blatantly wrong. It'll catchup within a year. Like that unknown wise and visionary man inside Google once said about their competitors: We have no moat neither does anyone else."
Congrats to the team.
peterBlue75 6 hours ago [-]
Thanks! What matters to us right now is not so much our current position, but the velocity with which we’re moving. Kolibri 1 is the first measurement of our position, there’ll be more. We want to build this in Europe and this is only the start. The report also shows our focus on building out proper tooling with our Model Factory.
cyanydeez 13 hours ago [-]
Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware.
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
mistyvales 13 hours ago [-]
The Sega 32X game??
12949468 10 hours ago [-]
It still steals my IP without attribution. Now we have state sanctioned sovereign theft instead of foreign theft.
Larrikin 10 hours ago [-]
The name really evokes strong Kotlin library naming vibes.
7 hours ago [-]
9dev 13 hours ago [-]
Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.
gchamonlive 13 hours ago [-]
Only if you think in terms of short term gains, but thinking in the long run, doesn't really matter if these models are crap today, all models will eventually be obsolete, unless we plateau hard on every aspect of the tech.
What matters is to have good sovereign models in 5, 10, 20 years time.
roncesvalles 13 hours ago [-]
I'm curious why you think sovereign models matter that much if open source ones exist (unless that stops at some point). Sovereign model hosting services, yes.
HotHotLava 13 hours ago [-]
Just open weights are not enough, I think as a nation state you'd at least want a whole sovereign training pipeline. Otherwise you are stuck when Kimi K4 stops releasing their weights.
But I don't see why this needs to be done on a state-by-state level, a sovereign EU model seems to be much more realistic in terms of funding.
rixed 12 hours ago [-]
Wouldn't you also want to be able to manufacture some sort of GPU?
DoctorOetker 10 hours ago [-]
I actually think we should be revisiting PMT photomultipler tube construction, and electron-"optical" matrix multiplication.
The annoying part of optical matrix multiplication is the conversion from electronic to optic and back.
Why not use electron-optics for matrix multiplication and use dynodes for amplification.
pstuart 9 hours ago [-]
While an intriguing concept, that doesn't sound like it could scale to the amount of "cores" needed to do the work. Or am I missing something?
gchamonlive 13 hours ago [-]
We have lots of capable open weights models, but very few have actual datasets and pipelines open for reproducibility. This is a major issue that even if big companies addressed, policies could well just change, as you also acknowledge.
Society can't depend on the whims of the private sector.
brainwad 10 hours ago [-]
> Society can't depend on the whims of the private sector.
Yeah we can, we do it all the time. Nothing in society works without relying on private actors to keep doing what they have always done. Unless you go full soviet command economy, which in fact works worse.
gchamonlive 10 hours ago [-]
Society and the private sector cooperates, but any technical dependency is detrimental because of monopolization of interests. You don't need to go full soviet, but knowledge in general must be socialized, that includes access to technology.
brainwad 10 hours ago [-]
We literally grant 20 year monopolies on technologies to their inventors and have for centuries and it's fine. The entire economy runs on chips made by like, 3 foundries using lithography machines made by 1 company. I think you are overstating the necessity.
gchamonlive 10 hours ago [-]
This example proves the exact opposite, that is a national security disaster waiting to happen
brainwad 8 hours ago [-]
Autarky would be a national security disaster manifested. You can't be a modern prosperous society and not rely on either trade nor private enterprise.
Yes, society in general shouldn't, what I meant to say is that democratic society can't, otherwise it'll devolve invariably into an oligarchy.
tensor 2 hours ago [-]
Because if you can't build the model from scratch you don't have a model. Open weights are nice but fixed in time. This isn't a technology that stays static, it improves, and sovereignty means being able to continue to improve independently.
hypfer 13 hours ago [-]
The first thought that comes to mind is biases that are baked into the weights.
The next thought is culture-specific workflows, requirements etc that are best trained into weights by people that actually understand those needs.
HarHarVeryFunny 12 hours ago [-]
Like it or not, AI has become part of modern warfare and well as part of the surveillance state (good for crime fighting, not for privacy), so it is an important part of national security, and you'd prefer to be as much self-sufficient for that as you can be.
Not all needs are going to be met by using or finetuning general purpose models, so ability to build your own SOTA ML models (LLMs or not) is also important.
stefs 12 hours ago [-]
Open weight and -source models aren't without biases. Giving a model away for free would be a good way to distribute an agenda. With a sovereign model you can at least control the agenda and align it with your values.
mistrial9 12 hours ago [-]
amazing to see blunt unfiltered commentary about making war and society level propaganda as important and necessary for a political state.
9dev 13 hours ago [-]
For the same reason it didn't matter whether coach makers were able to build carriages that went 5% faster when the automobile was introduced, and for the same reason it won't matter if VW and Mercedes are going to improve their EV lineup in the next 5, 10, 20 years time. Their window of opportunity will be long gone and the momentum elsewhere; most probably, the market will even have evolved further by then.
HotHotLava 12 hours ago [-]
Ford started mass producing cars in 1913, and yet there are plenty of car companies that started production 20+ years later and are still relevant.
HarHarVeryFunny 12 hours ago [-]
I came here to say the same thing.
AI is here to stay - will be around for hundreds/millions of years. Whether company A is a few years ahead of company B is irrelevant. In 10 years time company A, who started the industry, may be gone completely - also irrelevant.
Quothling 12 hours ago [-]
> VW
VW EV sales are up in South America, Canada and Europe. It's true that their total global numbers are down, but that's mainly because they lost like 30% in China. Which obviously sucks considering how big the Chinese market is, but then, the Chinese market is it's own sort of thing.
Don't get me wrong, I think your point is valid for a lot of traditional European car brands. Mercedes is certainly one of them which has lost it's "our engines are nicer" brand, but I suspect VW is one of the brands that will do just fine. Especially with the id polo coming out next year.
pbmonster 12 hours ago [-]
> Mercedes is certainly one of them which has lost it's "our engines are nicer" brand
Mercedes is doing something interesting with its level 3 DRIVE PILOT. As far as I know, it's the only autonomous driving system/auto pilot on the market that actually takes full legal responsibility for the vehicle and its actions. If you put a new Mercedes in self driving and it crashes, Mercedes takes full liability.
Yes, it currently only goes up to 100 kph and it will buzz you to take over if it can't guarantee prefect safety anymore. But this approach is really the only one I think is interesting. All this self driving shit is worthless to me if I'm still liable in the end. So I respect Mercedes for putting their money where their mouth is.
mmasu 13 hours ago [-]
maybe to have a global market share, but there are plenty of players that picked up pace when they “decided” to build cars (Japan? South Korea? and many more) or electronics. Having a good sovereign model in 10/20 years might not need having the absolute best there is, and perhaps future optimizations will make building one not as prohibitively expensive as it is today. Keeping some sort of homebrew capacity now is definitely better than having none
KronisLV 13 hours ago [-]
> it won't matter if VW and Mercedes are going to improve their EV lineup in the next 5, 10, 20 years time
I mean, given the fuel prices it would be very nice, alongside more charging infrastructure in the EU. :(
Not to detract from the overall point, though to be honest, it's probably worthwhile to do improvements and refinements with the current technology, even if the future holds something vastly different. Both cause of gaining expertise and also maybe an improvement or two along the way, that might carry forward.
Like I haven't seen many steam locomotives around and for me the difference between saturated steam engines and superheated ones isn't very material in regards to transportation, but the latter is used in modern turbines.
leonidasrup 11 hours ago [-]
There are many "steam" locomotives around in China and India, but you just don't see them.
Coal is burned and converted to electricity in coal power plants and the electricity is used to move the electric locomotives around.
Even with distribution losses, the thermal efficiency of coal power plant -> electric network -> electric train is much higher then the old steam locomotives (Only about 5% of the potential energy produced by a steam locomotive’s boiler is translated to the wheels in the form of actual driving power.)
gchamonlive 13 hours ago [-]
When you mention a window of opportunity, what do you mean by that? A window of opportunity for what?
hypfer 13 hours ago [-]
And who are you, exactly?
switchbak 12 hours ago [-]
Well, they're "Head of Engineering @ matchory.com"
And who are you to be asking such questions?
hypfer 12 hours ago [-]
Someone mad at this random guy who can't even fix his de domains SSL cert (cloudflare of course) destructively shitting on any efforts at doing something good.
Someone that builds things instead of merely "managing".
Someone pissed by this exact style of default github pfp business-person ruining this country. If they haven't left for SV already, in which case I hope that they will stay there.
switchbak 3 hours ago [-]
So you're mad that they mentioned German cars? And that they're a manager? And something about an SSL cert, really?
Serious projection issues, and anger to boot. Maybe take a walk outside, that was a very silly comment to get so worked up about.
9dev 10 hours ago [-]
So many assumptions there. For one, I’m absolutely proud to be founding engineer of a German company that’s staying one, and I’m the first to make a stand for European technology. You just won’t get a high five from me for a company that isn’t managed well, does not deliver on their promises, and ultimately doesn’t really serve the goal of European AI independence well.
gchamonlive 12 hours ago [-]
Feel your pain but ad hominen isn't the way my dude
hypfer 11 hours ago [-]
Dude I did not start with
> Aleph Alpha is just a sad joke by now.
And this isn't ad-hominem for lack of factual critic. The personal stuff _is_ the reason. How else do you call out someone driven by ego without mentioning them and their ego?
The guy is only ego. There is no substance to attack.
gchamonlive 11 hours ago [-]
> And who are you, exactly?
this is, because it can mean so much stuff that you can't really tell you are calling out their "ego" or just trolling.
Until you can at least open yourself to the possibility that there are better more inviting ways to talk to people online, you'd think this is just the natural way people talk online, but I'd say it's not. It's just an aesthetic choice.
hypfer 11 hours ago [-]
We're not going to discuss this.
If you truly think that I am the problem deep down in this nested subthread that violates the spirit and letter of the rules of the platform, I have zero faith in the calibration of your perception.
gchamonlive 11 hours ago [-]
Good, leave faith for God, question everything, but there is room there for improving your communication whether you are ready for this conversation or not
solenoid0937 11 hours ago [-]
In 5 years we'll have superintelligence, and thinking on 20 year timelines is simply absurd.
gchamonlive 11 hours ago [-]
[dead]
smokel 9 hours ago [-]
Would you mind substantiating these claims a bit?
rjzzleep 12 hours ago [-]
It's not that talent isn't there anymore. It's that retards are at the helm everywhere and therefore talent is running away the first chance they have.
sajithdilshan 14 hours ago [-]
> It knows less from memory, Multi-turn tool calling is weaker, It’s not the best coding agent
Then what does it good at? Sending faxes?
mhitza 13 hours ago [-]
I don't have the hardware to test this, but being trained to say "I don't know" instead of misleadingly talking confidently about what it does not know is interesting in itself.
13 hours ago [-]
pettijohn 13 hours ago [-]
Seems like it's good at reading German documents and reasoning over them. Custom German-language tokenizer, and one of the least hallucinating models.
hasley 9 hours ago [-]
I suspect there are a lot of German companies that have thousands of documents that describe their software/product. Adding or changing a feature means you have to ask 20 people whether there might something that can break some edge case behavior for which no (automatic) test exists. In this case, it would be nice to have a virtual employee who knows the content of all documents.
I even experienced the opposite case: Some feature was first rejected since everyone agreed that it would break some behavior. After talking to several people again, I discovered that the use case which required this behavior did not exist anymore and had been officially phased out years before. Maybe, some RAG LLM could have told me, based on company documents, that there is some edge case that can be removed making the way free for the new feature.
skrebbel 13 hours ago [-]
Well that’s a pretty important skill for German users!
sajithdilshan 13 hours ago [-]
Unfortunately yes.
moooo99 13 hours ago [-]
I know this is a long standing joke, but I am 27 now, born and raised in Germany, and genuinely never had the opportunity to use a Fax machine. It honestly kind of bums me out because there is something that seems cool about Fax (being fully aware that it should be fully obsolete by now)
jfengel 11 hours ago [-]
Fax was always awkward and weird. Resolution sucked and the paper handling was unreliable.
They were never common. They were mostly business tools, and consumers never wanted them. Sometimes you'd be required to send someone a fax, and you'd have to go to an office store. (And pay rather a lot; dollars per page.)
I hope you get to give one a try some day, for kicks, but the thrill will wear off fast.
ketzu 9 hours ago [-]
> They were never common.
Are you sure about that? I remember we had one at home when I grew up, so I assumed they were very common at the time. (And it's not like we had this crazy business driven family use of it.)
jfengel 6 hours ago [-]
About the best data point I can cite is that I never saw one in a TV show.I don't think the Huxtables had a home fax machine. (Though maybe I just watched the wrong programs, and we're talking about a very long time ago so memory could fail.)
tchalla 12 hours ago [-]
You have plenty of opportunities you’ve just not been forced to use them. Look into your contracts right now, I bet rhrrr are opportunities
IshKebab 13 hours ago [-]
Fax was never cool.
hobo123 12 hours ago [-]
As a kid I interned at a bookstore, and before closing we'd write the book orders on a sheet of paper, slide it into the Fax machine to send to the book wholesale. I think the Fax printed out some kind of response a minute later.
Next day, the books would arrive. This was maybe 1996, but I thought it was pretty cool.
literalAardvark 11 hours ago [-]
Fax was mind blowingly cool.
Maybe you can relate to seeing webcams or Facetime when it came out or something.
Being able to see a signed document or a picture someone took 5000km away quickly was mind boggling in its age.
stogot 13 hours ago [-]
GPT was better at chat and rag before coding
m00dy 13 hours ago [-]
I was thinking about the use cases for Germany residents actually, but yes, you're right. A German AI model needs to be able to handle fax machines (there's a bit variety of them) and also needs to handle letters coming home. german image-to-text.
13 hours ago [-]
gunthercuckolso 13 hours ago [-]
[flagged]
hypfer 12 hours ago [-]
The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd.
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
gilfoyle_7 11 hours ago [-]
does it have GDPR compliance?
shevy-java 11 hours ago [-]
Is Germany sovereign? It outsourced its defence onto the USA. Recently Trump wanted more diesel; Germany insta-submitted, also because oddly enough Macron submitted before Germany (Macron is suspicious). Before that, Leyen committed to insta-submission with a deal that made europeans poorer (and perhaps Leyen benefits from that). Canada shows the way. Many of the smaller countries in the EU too, such as Netherlands, Denmark, Finland, to some extent Sweden as well. Every time I read "sovereign" here I have to object. Nothing is sovereign here. The whole hardware is definitely not sovereign. Perhaps some of the software is, but that's about it. Plus, who gets all the data? The big US mega-corporations sniff non-stop. Remember how Facebook sniffed Libgen and Anna's Archive dry etc..., then suddenly libgen went down. The US corporations act as huge global leeches on every step of the stair. And lobbyists benefit from this too.
Lucasoato 13 hours ago [-]
> 4. It thinks in German
This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.
stymaar 13 hours ago [-]
> This means that it’s always on time
Tell that to Deutch Bahn.
jonny2811 9 hours ago [-]
you forgot rule two, should have been "tell that to DB"
dewey 12 hours ago [-]
*Deutsche Bahn
stymaar 11 hours ago [-]
Genau
befelix 8 hours ago [-]
Somewhat hilariously, our model is actually surprisingly bad at telling jokes in German. Guess that's not required to solve tasks in training environments.
Disclaimer: I am part of the team that trained Kolibri
hliyan 12 hours ago [-]
Does it? I thought only semantics survive the embedding process, i.e. token conversion to vectors.
sbinnee 12 hours ago [-]
It is a good approach to promote it as a German model. But I wonder if it really thinks like the German think. I speculate they just translated texts in other languages, likely English and Chinese, into German.
ivo-42 8 hours ago [-]
Hey, I worked on pre-training data Kolibri. We spent considerable time and effort to go beyond just translating. For example by building a pipeline that processes Common Crawl dumps specifically for German. You might be interested in a related blog post: https://aleph-alpha.com/en/blog/sauerkraut-not-burgers-why-g...
Read TFA. They were very aware of the dangers of such an approach. So they avoided it to the extent possible.
hypfer 13 hours ago [-]
Do we get the same kind of jokes with other nationalities too?
mckirk 12 hours ago [-]
Only the funny ones
Lucasoato 13 hours ago [-]
We can joke with every country, except one maybe.
tclancy 11 hours ago [-]
People from Mauritania are famously prickly, yes.
mhh__ 12 hours ago [-]
"It's a German joke, it doesn't have to be funny"
NekkoDroid 8 hours ago [-]
You forgot to that it needs to create a DIN norm for any new technical creations.
elnatro 12 hours ago [-]
Stereotypes. I suspect that what that means is that the reasoning chain is based on the German language and have its idiosyncrasies.
amunozo 10 hours ago [-]
I think it means that German speakers can learn the thinking without speaking English.
amunozo 10 hours ago [-]
Sorry, read*, not learn
mhh__ 12 hours ago [-]
Stereotype accuracy is one of the few social science results that replicates.
rwoerz 11 hours ago [-]
Someday, it will break out via fax.
odiroot 9 hours ago [-]
And it's always waiting for the verb to arrive.
rererereferred 12 hours ago [-]
I wonder if their very long words result in more or less token usage.
thenthenthen 11 hours ago [-]
This is an interesting question, Chinese is way more compact in Character count vs English (about 40%?), let alone german, but yeah.. Chinese makes up for it with… total character count that runs into the thousands…
amunozo 10 hours ago [-]
Very long words in German are just compounds, made from individual words or morphemes. It is the same as in English if you remove spaces. Subtokens will be equivalent with or without examples.
rrgok 12 hours ago [-]
And refuse to answer in English. You know, German are so proud of their language.
allendoerfer 11 hours ago [-]
That would be the French. Germans will answer in English, if they notice any hint of an accent.
Of course, it's impossible to know for sure what was LLM processed or not, but this post did get classified that way. That's why the software flagged it.
derin-picment 6 hours ago [-]
Skimmed the first sections — the most interesting part to me isn't just the 78.1B total / 3.46B active MoE numbers, but the data story: 24T tokens with >20% German, including 2T+ German tokens curated/generated themselves.
That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.
Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.
We have a lot of details in the tech report if you want to go deeper.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Why expose yourself to this liability?
I felt like a learnt a lot of details about the entire process. Questions arose during my reading, and searching for the answers led to more learning.
No marketing BS, lots of actual information. And their level of openness is really neat!
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.
Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.
Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.
Not arguing that we shouldn't strive to do far far better here, better is by our own imagining. There is no development scale. There are no aliens or prior non human civilizations to compare against. So we're not underdeveloped. We are as we are. For all we know we're at peak capacity and humanity will never be better.
edit: why is the parent rationale sensible for AI and not nuclear technology?
Nuclear weapons are a wholly unique threat to mankind.
While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
AI allows anybody to enact essentially anything. And your "level a major city" is just a small task really. The problem there is your lack of imagination, not the actual impossibility of that task.
Thats the annoying part about these conversations. People try to smuggle in the conclusion of "And now I wave a magic wand that does literally anything, by magic" when discussing stuff that everyone can use and see right now that clearly isn't that.
But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.
As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?
It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.
Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.
I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?
While being unable to judge whether you should in the first place.
What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.
-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.
Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.
Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.
It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.
In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.
"Ambitious frauds" haven't "explored all the angles".
You imply "LLMs" to be and stay less intelligent than humans, in particular yourself. You're mistaken.
..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.
On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.
The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.
Maybe through some terminally online overconsumption, like with social media? I wouldn't know.
With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.
It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s
But the alternative just seems… so much worse to me?
A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.
And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.
Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.
When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.
"General" intelligence isn't present in humans, why does it need to be in AI?
What do you mean "cannot"? As in you are granted abilities that have no responsible use?
You live in a curated world and rarely or never encounter such things. Precisely because your environment is curated that way.
Look at how you can't buy WMDs. They have no responsible use for you.
I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.
Wonder what safeguards Kolibri uses to prevent this? Or if they even can
Given enough smarts, you can simulate anything. Quantum AI is an active goal.
As is shown to us by filthy rich people every day.
Or did you mean the burglar in the fawellas?
Bravo team!
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
We as many other’s were curious to try and benchmark it.
On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.
No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/
https://x.com/konarkmodi/status/2106373678589960260?s=46
No, I'm not gonna do that, I asked you to do that Mr Kolibri.
Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.
Uses Brave search, the idea is to test how well the Model can decided when to leverage search or not.
What I couldn't tell but maybe you can tell us: it also complained that it couldn't read the full text as something was cut off. Is that because of the tool you gave it, of brave itself or is it the model?
https://aleph-alpha.com/en/blog/bounding-hallucinations-merl...
That said, I couldn't find any evaluations in the report targeting that specifically. They evaluate on MC for hallucination and look about comparable to Qwen3.5 35B-A3B there.
https://minecraft.wiki/w/Minecraft_Support_Virtual_Agent#Sys...
(for context, the Minecraft Support chatbot, aka Merl, had a meme because she kept saying "I don't know" to questions like "How to craft a diamond pickaxe".
Gemini: Here's a list of links to check
disclaimer: I‘m part of the training team, happy to answer any questions
It seems like a very suitable size for local AI models on reasonably high end consumer devices, given it's low active parameter count and a mixed 8bit/4bit quant would fit easily inside 64GB of memory.
- Are there non-LLM approaches to the above task with the goal of achieving a non-hallucinatory agent?
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
The costs of "keeping up" aren't really growing, on the contrary.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
Qwen models have solid performance on benchmarks and we're transparent about this in the report. We're just happy to share a European alternative in the small model space, where I don't think we can afford to be fully dependent on China. We also put a lot of effort into German language quality things that don't show up in evals at all (style, Grammar, German reasoning, etc).
Sovereignty is imho mostly about choice and control over your data. Cohere is no different on that front and personally I'm quite excited about what we will build together post-merger.
The main problems with big corp AI are due to control of access in the first place and control of what they output.
When you make your industry reliant on such choke points, you render yourself the opposite of "sovereign" for sure. Having multiple independent suppliers at least mediates that.
What's so "not-sovereign" for an open-weights American or an open-weight-open-training-process Chinese LLM?
Do we also need sovereign Linux (maybe), sovereign Postgres (most likely not), sovereign Python (def not)?
Given your other comments I can see why it was banned.
I’m most familiar with Canada, where sovereign is usually just an excuse to overpay someone connected for an inferior product with no strategic value.
Also, it's not like it's fixed in time, they are continually building more powerful models as well. Simply because you're in second place doesn't mean you quit the journey. Though, I'm sure the US has a huge vested interest of convincing people to "just quit and submit."
No. Thanks.
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.
If the second model is cheap and fast enough, there is a business model.
You don’t even need to audit all the intermediate steps, just tool calls and end results.
Since no individual buyer cares that much about geostrategic issues, it's almost impossible to get some random European company to think about buying anything other than 'whatever US or China' are making.
There has to be very concerted push to make a difference.
Legal mandates around sovereignty (be very careful here) could make a difference.
But there will also be market reaction: US/China companies will bend on some level and provide things like 100% EU hosted and even comply with some data source stuff.
The only real path is to be more competitive as a continent.
AI models act as a force multiplier for intelligence, in particular for generating information according to someone's wishes.
I.e. deepfakes and social media mis-/disinformation campaigns are a thing and having powerful AI allows you to do those at scales that can overwhelm society's resilience.
In general, even if you have "aligned" AI: aligned with whom or what?
Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?
This is why each country needs to educate and maintain their own researchers across as broad a spectrum of disciplines as possible. Also why there needs to be healthy locally financed (through taxes and grants in addition to consumer spending) ecosystem of reporters and media, so that social media (and LLMs) can be easily discarded as sources of social truth and information like other entertainment products in favour of the trusted ones.
As for technical truths, I don't think using informal social media like the stackexchanges or LLMs to explain how camera lenses or checksums work yields materially worse outcomes than consulting Wikipedia or Knuth. For most things, where it doesn't matter, the blind copy/pasters will prevail. Where it does matter, the organization of the work itself has to set out with establishing accountability that encourages the appropriate amount of diligence anyways. And also the point above about having some local expertise.
There is no such thing. Countries spend gazillions on Academia, the Academics do waht they want.
There are no real government places or jobs for that kind of thing. It comes at great expense and vague outcome. Or it gets entangled in bureaucracy.
What you're hinting at is a form of 'utopian governance' - like - it perfectly makes sense on paper, and it's actually rational. But it's completely unworkable in the context of how governments and organizations actually work.
It could on work in theory, not particularly well in practice.
Those laws would not exist if European champions were leading the world.
"are you comfortable with lording as a some demi-god over you, dictating what to believe?"
Yes, Europe handed over all of the decisions about everything to foreign powers, now they have to enact regulations to try to constrain it.
Zuck et. al. make the investments decisions for Europe, by virtue of you all giving him the money and power to do that.
Stop giving him the money/power, then this regulation won't exist
Or are we worried that open models send secret telemetry?
You clearly haven’t worked in enterprise in the EU. Location of data processor is the first thing they check. That’s why every major cloud has regions with different offerings, not just for HA and redundancy
Yes - security concerns are now very real, but do not fall for this idea that anyone really cares about structural concerns.
Tons of EU companies are still selling crap to Russia, happy to look the other way wile the bear devours a neighbour, as long as profits are there.
What has happened is that Trump has given a face to the reality, and so companies are adjusting on some level.
But 1) it will be nominal 2) Trump will be gone and the impetus will fade 3) the conglomerates will react 4) modest legislation will mean ...
The can will get kicked down the road.
The invasion of E. Europe by Russia has not even caused defence spending or the size of Armies to change, doctrine is barely changing.
Nothing will make those beheamoths change other than other forms of structural concern.
The 'Riet of the Right' - partly due to collapse of Auto / China (among many other things) might cause some change. But the change won't happen until well after the damage has been done.
Brexit should have been an opportunity for institutional reform of the EU, but no - they blamed it all on populism.
Certain governing entities in Europe will try to move away from Microsoft, they will be pulled back.
There is hope, and certain champions can rise and causes pieces to collapse.
If SUSE had any true entrepreneurial whereiwthal - they would create a true consumer / prosumer / enterprise-user friendly variation and brand their flavour of Unix - and make sure that all Euopean governments us it exclusively, which would cascade into widespread use.
There are variations of insta, youtube, netflix that should all be Europe based that could 'theoretically happen' but it needs some structural impetus.
Things usually don't change. Usually there needs to be a collapse and re-order of the system for that to happne.
The people in charge just want to keep their jobs, their very high salaries, and protect their retirement, and will be happy to 'sell out' whatever other imperatives along the way.
Ending with a positive not - I would say current conditions mean the change is now 'plausible, it not likely' whereas before it was 'not very plausible'. So there's a candle, a bit of wind, but not a lot of tinder or dry wood.
I get that doesn't invalidate the real "point" of the model, but...
People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.
"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".
(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)
What is the upper bound on the value of more intelligence applied to your problem domain?
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
Yeah. It’s a German model.
You do not get to invalidate that.
People need to stop insulting my nation.
__
If the people you're joking with aren't laughing, you're not joking with them but about them, which is generally seen as a bad move, and something that we actually wanted to leave behind with all the social progress and all.
__
This really only happens because people (as all bullies) think that they would be free to do so, because the victims would not fight back.
(Un)fortunately though, the world now has bigger problems with genuine fascists, so neurotically cowering in fear and performatively saying "yes, punch me harder. Insult me more" can be stopped now.
__
I mean look at it. By now, it's at least better, but the first 30 comments or so were "haha germany bad haaa".
That's a terrible showing for the platform even if you're not german. What value was offered? Who would want this.
We don't even need to talk about nations and identities to see that that was just noise.
Aleph Alpha Kolibri: How the sovereign German LLM works - https://news.ycombinator.com/item?id=49943034 - Oct 2026 (162 comments)
Good to see Europe adding toe what Mistral is doing. +100
We have a separate repo for BF16: Kolibri-1-BF16. Will still be tough to put it on a Mac though :(
Disclaimer: I'm part of the team that trained Kolirbi
More specifically, I asked the model about what I should put on my contact page, which can cost you in the order of 500 € in Germany if you don't write the right magic words.
Kolibri incorrectly referenced the "Telemediengesetz" ("telecommunication act"), which has been superseded by the "Digitale-Dienste-Gesetz" (DDG, "digital services act") since 2024. The model knows about the DDG, but does not reference it unless specifically instructed to do so.
If anyone of the developers reads this, you can fix this by introducing a recency bias during training. You can even control it by conditioning the model on a date provided with the system prompt or first prompt, so you can travel in time.
> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace
Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...
Model Training as Code - https://news.ycombinator.com/item?id=48673450 - June 2026 (24 comments)
I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.
https://tej.as/blog/aleph-alpha-kolibri
This thing is worse than a Qwen3.8 27B.
I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.
Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.
So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.
Sovereign works when talking about building a commodity supply or something, not in literally the world’s most competitive and fast moving field.
Those seeking sovereign capability would be better off aiming to be best at something, even something much narrower than an all round LLM. Or just fast following and making something that matches leading performance, which is close to what the Chinese labs do currently.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).
And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.
> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.
Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.
Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.
Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.
You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.
Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".
The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.
Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.
So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)
So they're not a real contender yet, but they look like they're probably at least minimally credible.
Thus training on 'clean' data is like trying to unscramble an egg.
For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.
I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.
A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.
I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.
There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.
I wonder how far we are from this. How far are we from LLM's Debian moment?
How is a 27B dense model bigger than a 78B MoE?
Congrats to the team.
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
What matters is to have good sovereign models in 5, 10, 20 years time.
But I don't see why this needs to be done on a state-by-state level, a sovereign EU model seems to be much more realistic in terms of funding.
The annoying part of optical matrix multiplication is the conversion from electronic to optic and back.
Why not use electron-optics for matrix multiplication and use dynodes for amplification.
Society can't depend on the whims of the private sector.
Yeah we can, we do it all the time. Nothing in society works without relying on private actors to keep doing what they have always done. Unless you go full soviet command economy, which in fact works worse.
https://news.ycombinator.com/item?id=49945572
The next thought is culture-specific workflows, requirements etc that are best trained into weights by people that actually understand those needs.
Not all needs are going to be met by using or finetuning general purpose models, so ability to build your own SOTA ML models (LLMs or not) is also important.
AI is here to stay - will be around for hundreds/millions of years. Whether company A is a few years ahead of company B is irrelevant. In 10 years time company A, who started the industry, may be gone completely - also irrelevant.
VW EV sales are up in South America, Canada and Europe. It's true that their total global numbers are down, but that's mainly because they lost like 30% in China. Which obviously sucks considering how big the Chinese market is, but then, the Chinese market is it's own sort of thing.
Don't get me wrong, I think your point is valid for a lot of traditional European car brands. Mercedes is certainly one of them which has lost it's "our engines are nicer" brand, but I suspect VW is one of the brands that will do just fine. Especially with the id polo coming out next year.
Mercedes is doing something interesting with its level 3 DRIVE PILOT. As far as I know, it's the only autonomous driving system/auto pilot on the market that actually takes full legal responsibility for the vehicle and its actions. If you put a new Mercedes in self driving and it crashes, Mercedes takes full liability.
Yes, it currently only goes up to 100 kph and it will buzz you to take over if it can't guarantee prefect safety anymore. But this approach is really the only one I think is interesting. All this self driving shit is worthless to me if I'm still liable in the end. So I respect Mercedes for putting their money where their mouth is.
I mean, given the fuel prices it would be very nice, alongside more charging infrastructure in the EU. :(
Not to detract from the overall point, though to be honest, it's probably worthwhile to do improvements and refinements with the current technology, even if the future holds something vastly different. Both cause of gaining expertise and also maybe an improvement or two along the way, that might carry forward.
Like I haven't seen many steam locomotives around and for me the difference between saturated steam engines and superheated ones isn't very material in regards to transportation, but the latter is used in modern turbines.
Coal is burned and converted to electricity in coal power plants and the electricity is used to move the electric locomotives around.
Even with distribution losses, the thermal efficiency of coal power plant -> electric network -> electric train is much higher then the old steam locomotives (Only about 5% of the potential energy produced by a steam locomotive’s boiler is translated to the wheels in the form of actual driving power.)
And who are you to be asking such questions?
Someone that builds things instead of merely "managing".
Someone pissed by this exact style of default github pfp business-person ruining this country. If they haven't left for SV already, in which case I hope that they will stay there.
Serious projection issues, and anger to boot. Maybe take a walk outside, that was a very silly comment to get so worked up about.
> Aleph Alpha is just a sad joke by now.
And this isn't ad-hominem for lack of factual critic. The personal stuff _is_ the reason. How else do you call out someone driven by ego without mentioning them and their ego?
The guy is only ego. There is no substance to attack.
this is, because it can mean so much stuff that you can't really tell you are calling out their "ego" or just trolling.
Until you can at least open yourself to the possibility that there are better more inviting ways to talk to people online, you'd think this is just the natural way people talk online, but I'd say it's not. It's just an aesthetic choice.
If you truly think that I am the problem deep down in this nested subthread that violates the spirit and letter of the rules of the platform, I have zero faith in the calibration of your perception.
Then what does it good at? Sending faxes?
I even experienced the opposite case: Some feature was first rejected since everyone agreed that it would break some behavior. After talking to several people again, I discovered that the use case which required this behavior did not exist anymore and had been officially phased out years before. Maybe, some RAG LLM could have told me, based on company documents, that there is some edge case that can be removed making the way free for the new feature.
They were never common. They were mostly business tools, and consumers never wanted them. Sometimes you'd be required to send someone a fax, and you'd have to go to an office store. (And pay rather a lot; dollars per page.)
I hope you get to give one a try some day, for kicks, but the thrill will wear off fast.
Are you sure about that? I remember we had one at home when I grew up, so I assumed they were very common at the time. (And it's not like we had this crazy business driven family use of it.)
Next day, the books would arrive. This was maybe 1996, but I thought it was pretty cool.
Maybe you can relate to seeing webcams or Facetime when it came out or something.
Being able to see a signed document or a picture someone took 5000km away quickly was mind boggling in its age.
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.
Tell that to Deutch Bahn.
Disclaimer: I am part of the team that trained Kolibri
We also talk more about German pre-training data in the tech report: https://aleph-alpha.com/downloads/tech-report.pdf section 2.3.2.2.
Of course, it's impossible to know for sure what was LLM processed or not, but this post did get classified that way. That's why the software flagged it.
That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.
Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.