posted in Selfhosted

[AIT] The Codeberg ban on LLM content

Right - I see the other topic has been locked / author hasn’t returned to follow community rules.

Here’s the original topic news.ycombinator.com/item?id=49003386

and a measured (IMHO) response to it. (BTW, do your self a favour and change tabs with that site open :)

マリウス.com/i-regret-migrating-to-codeberg/

It’s a tough spot, Codeberg has found themselves in and I wish them luck. But beyond that, this is (yet) another reminder that in 2026, if you don’t self host it, the cloud is just someone else computer

news.ycombinator.comCodeberg bans vibe coded projects | Hacker News

Replying to an earlier post

From what I understand from their blogpost, it was mostly about conserving their own resources, because a lot of the vibecoded projects they got were uploading insane amounts of binary releases and wasting resources on CI/CD while having no users or other collaborators.

They wanna reserve more server space for projects, that actually productively use it, like bigger FOSS projects with actual users and contributors.

That is perfectly understandable for me, even without taking my huge distaste for AI into consideration. Anyone that thinks, that this small, community funded project is obligated to host their huge slop repos semms pretty entitled to me.

Replying to @⁨DeckPacker@piefed.social⁩

before reading the blog post I was thinking the same. now I don’t.

the worst of the LLM projects have no place on codeberg that’s for sure, but there would have been better ways than a blanket ban to limit the resource consumption of LLM and crypto projects. codeberg already has a storage quota system, they could be giving a lower quota for LLM projects, maybe also disable free CI for them, which I am a bit surprised they have given. and solve the reputation problem with banners. but no, total blanket ban based on feelings, it is.

Replying to an earlier post

There is simply no way that is true. First, the legal arguments are dodgy:

  1. There is a good chance that LLMs are sufficiently transformative that courts will decide they don’t infringe the copyright of their sources.
  2. Even if not, hosts will get DMCA-style safe harbour protection and the most they’ll be liable for is takedown requests, which is a major thing for any large host.

Second, there is no fucking way a western court is going to tell every big tech company they have to delete 99% of the code that was written since the start of the year, even if the law as written literally said verbatim, “use of any LLM output for any purpose is breach of copyright” because it doesn’t take a conspiracy theorist to realise that it’s politically impossible.

You may not like that, but it means that, again, Codeberg is doing this based on feels.

Replying to @⁨FishFace@piefed.social⁩

You are missing the point entirely. LLMs regularly generate code that is a near verbatim copy of existing copyrighted code, but with almost no way for the LLM using person to notice that. LLMs being sufficiently transformative might be an argument about the use of training material, making the resulting model not a copyright violation itself, and thus might protect the companies that produce and offer these models, but it says nothing about the actual output of a model.

It is only a question of time before some enterprising law firm decides to mass scan open-source projects and weaponize their findings similar to patent trolls or file-sharing legal threats. This has a long history in Germany where Codeberg is located, and even if a court rules that the host itself is only responsible for removing such copyright violating code, it will require significant effort to do so with constant legal fights as the attacking lawyers will try to figure out the identity of the person responsible so that they can blackmail them with cease and desist legal fees.

Politics will not care about some hobbyist open-source projects and large companies will spend a lot of effort to obfuscate their code to prevent this legal trolling to affect them.

Replying to an earlier post

You’re right I wasn’t really thinking about the possibility of verbatim reproduction. I guess we’ll see about that.

Codeberg will need efficient procedures to handle IP trolls anyway, getting a bunch more takedown requests ought not to make a big difference. And if somehow it does start to be a problem… that would be the time to blanket ban or take similar drastic action. If it’s possible to detect now, it will be possible to detect then.

Politics will absolutely care about having a precedent which leaves big tech exposed. Any case against little fish will be jumped on by Zuck and Bezos before you can say “datacenter”.

en