posted in Technology

World's largest open library calls for volunteers to scan and preserve physical books as AI companies buy, scan, and destroy them — Anna's Archive says ‘time is running out’ as ‘knowledge is permanent

cross-posted from: kbin.earth/m/piracy@lemmy.dbzer0.com/t/3120649

As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.

www.tomshardware.com/tech-industry/artificial-intelligence/worlds-largest-open-library-calls-for-volunteers-to-scan-and-preserve-physical-books-as-ai-companies-buy-scan-and-destroy-them-annas-archive-says-time-is-running-out-as-knowledge-is-permanently-monopolized-on-private-servers
enPage

Replying to an earlier post

No I know AI companies are doing that, I just don’t think they would be making gifs about it.

But there’s a Tom Hanks movie that came out a few years ago where in the beginning of the movie he lives in a bunker because the UV index outside is so high the sunlight is basically radioactive, and he’s building a robot and he uses a machine to chop the spines off of books to feed them into a scanner to train the robot brain.

That all takes place in the very beginning and then he leaves the bunker with the robot and some shit happens. But that’s what I thought this gif was from.

Replying to an earlier post

Ages ago I had mimicked the “overhead scanner” with a mounted digital camera, but now I’d go with a mounted smartphone + light bars + bluetooth photo trigger. Scrolling through Bezos’ was pretty depressing in terms of cost/quality and even a $150 device seemed like a repurposed webcam, with a sole non-free confirmed buyer saying it died after 3 months while freebies praise it lol. As long as you don’t scan drawings/designs/comic books - just set lights evenly and that’d be enough, and even moreso if you OCR it and turn a scan into an e-book where visual fidelity stops to matter.

Replying to @⁨themachinestops@lemmy.dbzer0.com⁩

Anna’s Archive says ‘time is running out’ as ‘knowledge is permanently monopolized on private servers’

Not only that, of course, but easily edited. It’s a “mandala effect” generator. AKA modern gaslighting. They name the thing, point to legitimate examples, then abuse our new “understanding” to cover their crimes. 1984 indeed.

Oceania was at war with Eurasia; therefore Oceania had always been at war with Eurasia

Replying to @⁨AlecSadler@lemmy.dbzer0.com⁩

Perhaps there is a public library in your area that accepts donations? The one I live by is STARVING for support, they accept books, old computers, anything.

They do good work for the community too, like free tech literacy classes, chess for kids, etc.

But if you are dead set on an online NGO, then maybe archive.org? You have probably heard about their Wayback Machine, but they preserve books and other media too.

Replying to an earlier post

They just cut and bin the books after they’ve been scanned, the bins are taken to the local disposal, most municipalities have a garbage dump, some burn the trash.

An ai trained on unique data is more valuable if you don’t share the data it was trained on to your competition, your LLM gains an edge.

If I found a unique handwritten book from the 18th century, than I proceed to read it and memorise then burn it, I’ll be the only one to posses the knowledge of it.

Replying to an earlier post

Eh, perfect is the enemy of good. I’ll take the wins we can get. It’s not like having to buy the books is gonna stop the biggest offenders from burning investor money (along with the books).

The defeat of LLMs isn’t gonna come on that front. Let the cultists believe that more words will somehow become more than words. Their silicon gods run on power and water. That’s where they’re most vulnerable.

Replying to an earlier post

I was aiming for the poetic parallel between burning money and destroying books, but also, isn’t the difference between destroying the primary source and preventing access to it mostly academic? Either way, as you say, the AI grifters are looking to control that information. Once they do, the original text and knowledge might as well be lost, since we couldn’t verify anything without the original to compare to. Whether through deliberate human manipulation or random word soup mutation, they can’t ever be trusted to deliver a genuine, faithful reproduction of the contents.

And the conclusion is the same: Better to have it freely available for everyone (including the grifters) than the grifters only.

Replying to an earlier post

If you read the blog post, they don’t provide help or guidance for regular bookworms on how to scan. No information about what they expect.

Also, flatbed scanning a whole book takes many hours. It’s one page at a time, line up each page, press the book flat and scan, for hundreds of pages.

You’d likely need to buy, or have access to a book scanner. Which is fine, there are cheap models, but who knows if that would work for them? They don’t say. It doesn’t seem like they’re expecting help from you or me.

Replying to an earlier post

With modern phones having a document scan option with our phone cameras, is it possible we can scan and combine into a PDF file?

I am going to test it myself, see how easy it is, and find a way to ensure anyone can follow with low cognitive load, but hopefully it’s as easy as I think to do.

The sad part, is to do so would require a slow tedious task for books with hundreds of pages. It would be nice to see a large group tackle chunks at a time to the point we eventually have groups dedicated to genres and scouring to find them to preserve them.

Replying to @⁨themachinestops@lemmy.dbzer0.com⁩

AI training is distroying books

Every time I hear that claim it sounds weird to me. You can’t say that without an explanation.

Sure, scanning a print copy can mean distroying it but a book as a work still exist. We are very far from having scanned every books which physical copy are numerous. If AI trains on every romance novel published in the 80’s, it is far from distroying books by distroying physical copy.

Now if AI company are searching for rare books to scan and distroy, it is much more worrying but finding rare editions or print copies of books is a trade and the normal citizen can’t jumps in to save the day.

Replying to an earlier post

They’re destroying things that don’t have to be destroyed and compiling the information into privately owned megastructures they’ll then charge us to access while we own nothing and they can then alter the books.

Nah I’m good. Everything these AI companies stand for is fucking evil and malicious.

For every AI that researches a legit medical issue, they start literally burning fucking books, replacing jobs, screwing basically everyone, and making the world overall much worse.

If you can’t see why these companies destroying any amount of books is bad, idk how to help you. Maybe look up a list of who has ever mass destroyed books. There’s not a lot of good people there.

Replying to an earlier post

All books created today exist electronically before (a copy) gets physically printed of a book that remains existing.

Not really. There is a lot of full manuscrit books made to this day. By scholar in region where electricity is not as much available. By student compile hand-written notes from oral teaching. By people composing small things for family and friends or just for themselves. By activist and locally involed people making just on or a few copy to be share in physicial Space around them.

That’s a huge part of what exist outside the publishing industry.

Replying to an earlier post

And we used copiers and voice recorders for that since before computers. That has been used since the 70s. Even earlier.

scanning that and adding it to the common knowledge and making it accessible is something we’ve been doing since computers were a thing decades before today.

Students now still use recordings that can be parsed into text.

Suddenly it’s an issue now because someone needs so badly to be outraged by literally everything that can be distorted into a rage bait title like this one.

Replying to @⁨KairuByte@lemmy.dbzer0.com⁩

They bought each one as a copy. They aren’t taking all copies of all books.

Not sure why I have to explain basic math to you but here we are. Learn math bro.

And get if they are recycling THEIR ONE COPY OF EACH which they no longer use that isn’t illegal nor worth an outrage. That’s how recycling and reducing waste works. Better than hoarding.

Do feel free to go buy a copy and support the author if you’re this passionate to get enraged over what someone does with their own one single copy.

Replying to @⁨KairuByte@lemmy.dbzer0.com⁩

This literally isnt book burning though. it is recording and recycling of (one) copy of (a) or even each book. Not even the original. One copy of many of each. One that you didnt even care about it until now.

So you calling it book burning or comparing it to actual book burning is misinformation.

And it is you willfully wasting your time here trying to convince somoeone of a lie you’re desperately trying to push here to control the emotion of others. Manipulation. A shit rage bait piece for attention and pandering to unnecessary anxiety and trying to depress other’s no doubt.

So if I am annoying you by calling out this shitty faćade:

Good.

Replying to an earlier post

Independent bookstores all across Europe are seeing a trend like this, where they receive random, relatively large orders that are shipped to local addresses. While orders like these still come through in the digital age, especially from institutional buyers looking to fill out new libraries, they say that most “legitimate” orders often come with coordination and negotiation, not just a straight purchase order.

Replying to @⁨themachinestops@lemmy.dbzer0.com⁩

destroying (A) physical copy. in order to copy the (1) copy of a book(which most new already exist as pdf). kinda like how people do it when they scan the book. which has been happening for decades now to make knowledge permanent.

far cry from “destroying all books” which these rage bait titles suggest. we arent standing around a pit of fire.

bigger tragedy is how humans seem shittier at communication as AI learns. probably a more accurate title.

Replying to an earlier post

There are many, many non destructive scanning methods. My understanding is that most archival projects are using the non destructive options.

The easy option is to just slice off the binding and run it through a scanner.

The easy option is what is being talked about here, and those books are then fed into AI as training material. The digital copies aren’t “making knowledge permanent” as they are tucked away into a digital corner virtually no one has access to. They are as permanent as the source code for Fortnite.

If a PDF already exists, as you claim, then why go to the trouble of obtaining, destroying and scanning the books again?

Replying to @⁨KairuByte@lemmy.dbzer0.com⁩

They bought a copy of a book. They gonna do with it what they want. It’s their’s. Not your’s. Their’s.

You didn’t buy it. And you didn’t care about it till now. Over something you didn’t buy. Someone else did.

So now you are outraged they are making copies? This has been happening since copy machines were invented.

You’re outraged they destroyed THE ONE COPY they bought? (NOT ALL. JUST ONE)

They recycled rather than hoarded. Good for them.

This happens every day since books and other things were invented to manage unwanted junk.

If you want it go buy a copy and support the author, go for it. You want to recycle? Go for it. Both things are not illegal. Nor should they be.

I don’t get your problem here. Seems like you’re too stupid to know what to outrage about.

It’s like you discovered that dirt exists for the first time and how dare someone shovels dirt in some unmarked area you never even knew or cared existed to make a trench.

Replying to @⁨themachinestops@lemmy.dbzer0.com⁩

Yall keep overreacting to this bs, what’s even a rare book with no other copies, every article I saw says the books were scheduled to be destroyed or taken to a dump, and those are the books they are swooping up. I have a feeling rare means like 1st edition or one with typos or some sht. These are books ppl already weren’t buying that they are grabving in mass. All knowledge and all books aren’t important, plenty of below average intelligence ppl dropping slop long before ai became a thing. Any book that you value obviously already has tons of physical copies and digital backups.

This isn’t them looting some ancient libraries/museums taking some hidden/secret knowledge from humanity like yall are making up in your heads.

Replying to an earlier post

every article I saw says the books were scheduled to be destroyed or taken to a dump

Do you have a source for this? Because I’ve seen literally nothing about this being the way books are being sourced. And no offense but this reads way too close to weasel words. For example: Every article I’ve read on the subject says my pants are sentient and enjoys my legs being in them. Also, I’ve read no articles on the subject so that isn’t a lie.

Replying to an earlier post

Fair point. Beside the point, is it technically correct to make a positive assertion about a negative thing — only because any positive assertion about that negative thing is inherently part of an empty set?

Meaning like… our logic technically rules based on “no available contradictions, even if there is no evidence” as opposed to “at least one evidence is required?”

Replying to an earlier post

In certain situations I’d agree that the absence of contradiction is evidence enough, but those situations are pretty specific. For example, we don’t have absolute proof that orange juice doesn’t cause cancer, our “proof” comes from the lack of reliable evidence showing that it does, along with relevant research. But that doesn’t mean every claim can be treated the same way. But that doesn’t mean every claim can be treated the same way, especially a claim based only on “every article I saw,” without a source being provided.

Replying to an earlier post

For an AI to troubleshoot problems with your computer (something it’s really good at today), it needs content like reddit’s. Books don’t and won’t detail weird software issues and their workarounds. One the current training set becomes irrelevant to whatever software people are running, this is a specific area where Reddit’s data is way more valuable than books.

Of course Reddit locking down access to that information leads to fewer people contributing. Reddit is dead, it’s just struggling to go down.

Replying to an earlier post

there’s gold in there, but it’s absolutely buried in shite

yep, there is a lot of nuanced depth, but measured against the overall sea of chuds lol… and yep, the moment they started putting up walls and controlling the community it was doomed. it’s gonna thrash for a while then, when the only ones left are advertisers and bots, it’ll shit itself and stop twitching.

Replying to an earlier post

im guessing reddit has been banning so aggressively lately, that AI scraping isnt getting as much “useful data” as much anymore (chatgtp/google dropped "references for reddit in thier slop summaries). so they are trying to force logins now to keep up with the user generated content so they can squeeze every last penny they can from reddit, i dont think ads were ever a big presence in reddit? its the data. people try to post content, like posts, subreddits, those can instantly get removed by redit filters. i think lost alot of traffic recently to, which was likely AI scraping.