posted in Selfhosted

[AIT] The Codeberg ban on LLM content

Right - I see the other topic has been locked / author hasn’t returned to follow community rules.

Here’s the original topic news.ycombinator.com/item?id=49003386

and a measured (IMHO) response to it. (BTW, do your self a favour and change tabs with that site open :)

マリウス.com/i-regret-migrating-to-codeberg/

It’s a tough spot, Codeberg has found themselves in and I wish them luck. But beyond that, this is (yet) another reminder that in 2026, if you don’t self host it, the cloud is just someone else computer

news.ycombinator.comCodeberg bans vibe coded projects | Hacker News

Replying to an earlier post

From what I understand from their blogpost, it was mostly about conserving their own resources, because a lot of the vibecoded projects they got were uploading insane amounts of binary releases and wasting resources on CI/CD while having no users or other collaborators.

They wanna reserve more server space for projects, that actually productively use it, like bigger FOSS projects with actual users and contributors.

That is perfectly understandable for me, even without taking my huge distaste for AI into consideration. Anyone that thinks, that this small, community funded project is obligated to host their huge slop repos semms pretty entitled to me.

Replying to @⁨DeckPacker@piefed.social⁩

before reading the blog post I was thinking the same. now I don’t.

the worst of the LLM projects have no place on codeberg that’s for sure, but there would have been better ways than a blanket ban to limit the resource consumption of LLM and crypto projects. codeberg already has a storage quota system, they could be giving a lower quota for LLM projects, maybe also disable free CI for them, which I am a bit surprised they have given. and solve the reputation problem with banners. but no, total blanket ban based on feelings, it is.

Replying to an earlier post

There is simply no way that is true. First, the legal arguments are dodgy:

  1. There is a good chance that LLMs are sufficiently transformative that courts will decide they don’t infringe the copyright of their sources.
  2. Even if not, hosts will get DMCA-style safe harbour protection and the most they’ll be liable for is takedown requests, which is a major thing for any large host.

Second, there is no fucking way a western court is going to tell every big tech company they have to delete 99% of the code that was written since the start of the year, even if the law as written literally said verbatim, “use of any LLM output for any purpose is breach of copyright” because it doesn’t take a conspiracy theorist to realise that it’s politically impossible.

You may not like that, but it means that, again, Codeberg is doing this based on feels.

en

Replying to @⁨FishFace@piefed.social⁩

You are missing the point entirely. LLMs regularly generate code that is a near verbatim copy of existing copyrighted code, but with almost no way for the LLM using person to notice that. LLMs being sufficiently transformative might be an argument about the use of training material, making the resulting model not a copyright violation itself, and thus might protect the companies that produce and offer these models, but it says nothing about the actual output of a model.

It is only a question of time before some enterprising law firm decides to mass scan open-source projects and weaponize their findings similar to patent trolls or file-sharing legal threats. This has a long history in Germany where Codeberg is located, and even if a court rules that the host itself is only responsible for removing such copyright violating code, it will require significant effort to do so with constant legal fights as the attacking lawyers will try to figure out the identity of the person responsible so that they can blackmail them with cease and desist legal fees.

Politics will not care about some hobbyist open-source projects and large companies will spend a lot of effort to obfuscate their code to prevent this legal trolling to affect them.

Replying to @⁨poVoq@slrpnk.net⁩

People can upload code that violates copyright too. No llm required.

So if that is the concern have a rule about not violating copyright (which they may already have). And there is likely already a process for handling that.

This is the problem I have with the situation: the solution chosen (banning llm code) doesn’t really address the concerns they’ve raised directly. It’s just apologetics to make it sound like it’s reason-based rather than “we don’t like it”.

It’s their platform, they can do what they want, but let’s not pretend it’s not code-puritanism.

Replying to @⁨atzanteol@sh.itjust.works⁩

Yes, they can upload copyrighted stuff, but people typically don’t do so on large scale and when they knowingly do it these days they typically try to hide their tracks well enough that law firms know it will not be a lucrative business to try and blackmail them.

And one of Codeberg’s main points is the unknown copyright status, you just failed to understand it.

Replying to @⁨poVoq@slrpnk.net⁩

Yeah - it’s well known that scripts to automate doing bad things did not exist before llms.

People have tried using forges to host explicitly illegal material all the time. Not just “this may be a copyright issue perhaps maybe”.

The copyright complaint is a fig leaf to hide their shame.

They also outright banned cryptocurrency code with no “copyright fig leaf” apologia.

It’s code puritanism plain and simple. “We don’t like these things so we’re banning them.”

Replying to @⁨atzanteol@sh.itjust.works⁩

You are still not understanding the issue.

LLMs cause a lot of people to unknowingly violate copyright, and those people then become easy targets for malicious copyright litigation. And Codeberg is caught in the middle of that and doesn’t want to be involved in the resulting mass legal cases because they are just a small volunteer organisation without a legal team.

And Codeberg’s explanation is very clear that is isn’t a blanket ban on LLM generated code for “purity” reasons. It is a risk mitigation strategy against projects that are mostly LLM generated.

Replying to @⁨poVoq@slrpnk.net⁩

P2P torrent users are nothing like a code hosting platform. 🤣

As has been said - their liability for llm code is no different from hosting any other code. They already have the same risk. And that is typically that they must respond to take-down notices which they are already doing.

They are already dealing with all of the problems they would be dealing with with llm code.

It’s just code puritanism wrapped in pseudo legal justification. Own it.

Replying to @⁨atzanteol@sh.itjust.works⁩

It is not. There are plenty of studies showing that LLMs spit out code that is near verbatim to existing code and LLM companies even go so far as to instruct their models to not also add the corresponding license/copyright headers with that code.

There are probably law firms analysing common code patterns LLMs often use right now and are approaching copyright holders of similar enough code to buy up the rights. It might not all stand up in court, but it will be enough to scare some people into settling for fee that guarantees a profit for these law firms. This is a tried and true method for an entire industry of law firms.

Replying to @⁨poVoq@slrpnk.net⁩

You know what’s funny? Codeberg has said nothing about being concerned about liability.

Their terms service change simply mentions that they require certain licenses and that they have concern over the licensing of LLM meeting that standard.

Event their follow up communication says absolutely nothing of liability.

blog.codeberg.orgProtecting our FLOSS commons from LLMs — Codeberg News.codeberg-design ul { padding-left: revert !important; } In Brief: Two...

Replying to @⁨poVoq@slrpnk.net⁩

LLMs cause a lot of people to unknowingly violate copyright, and those people then become easy targets for malicious copyright litigation

I’d buy that. As a mediocre musician, one can even subconsciously ‘sample’ a copyrighted piece. I mean, there are only so many chords and combinations, melodies, etc, so duplicates, and even exact duplicates are bound to happen whether intentional or not. If you upload your track to say SoundCloud on the professional plan with which you intend to generate revenue, the track is heavily scrutinized, and often rejected for this very reason. You actually have to prove it’s OC by various means. They don’t want any pieces parts of any litigation that may even mildly involve them. So, I can see that.

Replying to @⁨poVoq@slrpnk.net⁩

malus.sh

AIs already pillaged all the code in the world. This is fighting a lost fight. We won’t manage to force big AI firms that literally props a country to stay afloat to make their AI “forget” or “untrained”.

Regulators let the pillaging happen so now AI does know how to code. They dont imitate they truly code.

But IMO the real fight is about things like malus that are the real dangerous and malevolent actor here. They will take your open source code and resell it to big businesses.

malus.shMALUS - Clean Room as a Service | Liberation from Open Source Attribution

Replying to @⁨Tetsuo@jlai.lu⁩

You are linking to a parody project 🤦

And you are also misunderstanding my point. Yes the cat is out of the bag and companies will absolutely copyright wash their codebases like that, but they have big legal teams to defend against copyright trolls and generally do not publish most of their code base for anyone to see.

Small hobbyist open-source projects that make up near 100% of the projects hosted on Codeberg on the other hand are easy marks for malicious litigation, just like home users torrenting movies were in the 1990ties and early 2000.

Replying to an earlier post

You’re right I wasn’t really thinking about the possibility of verbatim reproduction. I guess we’ll see about that.

Codeberg will need efficient procedures to handle IP trolls anyway, getting a bunch more takedown requests ought not to make a big difference. And if somehow it does start to be a problem… that would be the time to blanket ban or take similar drastic action. If it’s possible to detect now, it will be possible to detect then.

Politics will absolutely care about having a precedent which leaves big tech exposed. Any case against little fish will be jumped on by Zuck and Bezos before you can say “datacenter”.

Replying to @⁨FishFace@piefed.social⁩

@FishFace Nothing is "politically impossible," and AI is not inevitable.

When rich people get out of hand, there is always the same correction that takes place, sooner or later. Right up until the very moment it happens, people with varying degrees of comfort in the existing regime are squawking that the world will end if we mess with the Load-Bearing Pedophiles.

They have less chance of making their tyranny permanent this time than at any previous time it has been tried.

Prepare your mind for the idea that you and everyone you know might have to personally pick up a shovel and dig, before the mess they are making right now gets cleaned up.

@poVoq