Replying to @⁨eicker@lemmy.world⁩

The main challenge I see is that Aaron Schwartz and countless others have been prosecuted for access to information but AI companies have been rewarded. There is precedence that access like this is not legal. Had they gone through a library system or used a mechanic like that it may have worked but from what I understand, they used torrents and other mechanics to access the data. So you have companies that go after individual infringement but pursue their own mass infringement. Whether AI generated materials is infringement is above my pay grade but their consumption of the materials seems pretty straight forward as infringement.

en

Replying to @⁨assembly@lemmy.world⁩

But paying and asking for permission first would have been costly at the start and slow. They made a decision, maybe at the beginning, maybe as they realized ethical would mean lost time and placement in the race, and they said screw it, let's go. They also sidelined any AI safety research they were doing (some were making an effort on a difficult problem, but again, $$$ wins).

Replying to @⁨assembly@lemmy.world⁩

The NY Times and Pearson, two of the most valuable US publishers, each have market caps of about $10 billion.

Let’s say Pearson went after OpenAI. They devote an unlimited legal budget to the fight. OpenAI is hoping to IPO as a trillion dollar company. If Pearson went after OpenAI, rather than fight them in court, a deal could be reached first. If that didn’t work, if a deal couldn’t be reached, OpenAI could bypass the problem completely:

  1. Spend $5 billion to buy a controlling share of Pearson.
  2. Fire the entire leadership team and install OpenAI minions in their place.
  3. Once they control Pearson, sign a long-term licensing deal with OpenAI with very generous licensing terms and huge early cancellation fees.
  4. Sell the shares back on the market at a (likely slightly reduced) value.

OpenAI would likely have to spend some money on net. The value of Pearson stock would likely be a bit lower after effectively giving away the rights to their works as training data. If signing a durable rights contract the new owners can’t escape isn’t practical, buying the company and simply holding it indefinitely would also be an option.

Replying to @⁨assembly@lemmy.world⁩

Schwartz was saving and distributing copies against the terms of the agreement by which he was able to access journals. What happened to him was heinous but it was pretty dissimilar to how models train on data. And the tormented material was Anthropic, which resulted in the largest copyright settlement in history. Because it was piracy. They briefly tried an argument that their intended use made it fair use, but…that’s never how literally any of that worked.