World's largest open library calls for volunteers to scan and preserve physical books as AI companies buy, scan, and destroy them — Anna's Archive says ‘time is running out’ as ‘knowledge is permanent
As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.
No I know AI companies are doing that, I just don’t think they would be making gifs about it.
But there’s a Tom Hanks movie that came out a few years ago where in the beginning of the movie he lives in a bunker because the UV index outside is so high the sunlight is basically radioactive, and he’s building a robot and he uses a machine to chop the spines off of books to feed them into a scanner to train the robot brain.
That all takes place in the very beginning and then he leaves the bunker with the robot and some shit happens. But that’s what I thought this gif was from.
Ages ago I had mimicked the “overhead scanner” with a mounted digital camera, but now I’d go with a mounted smartphone + light bars + bluetooth photo trigger. Scrolling through Bezos’ was pretty depressing in terms of cost/quality and even a $150 device seemed like a repurposed webcam, with a sole non-free confirmed buyer saying it died after 3 months while freebies praise it lol. As long as you don’t scan drawings/designs/comic books - just set lights evenly and that’d be enough, and even moreso if you OCR it and turn a scan into an e-book where visual fidelity stops to matter.
I don’t know about buying, but the ones I have seen basically hold the book open at a 45 degreeish angle, to not damage the spine. Then you press the pages with a 45 degree set of glass plates. The actual “scan” is a camera of some kind that slodes back.and forth between the two halves on a mount.
texas was supppose to be filled to brim with AI datacenters, until hot wheels so his election chances diminishing, at leas until he wins again,or another R.
Anna’s Archive says ‘time is running out’ as ‘knowledge is permanently monopolized on private servers’
Not only that, of course, but easily edited. It’s a “mandala effect” generator. AKA modern gaslighting. They name the thing, point to legitimate examples, then abuse our new “understanding” to cover their crimes. 1984 indeed.
Oceania was at war with Eurasia; therefore Oceania had always been at war with Eurasia
Any reputable orgs I can donate to for assisting the cause? I just don’t feel I have the means otherwise to make a dent in the buying and preservation.
Perhaps there is a public library in your area that accepts donations? The one I live by is STARVING for support, they accept books, old computers, anything.
They do good work for the community too, like free tech literacy classes, chess for kids, etc.
But if you are dead set on an online NGO, then maybe archive.org? You have probably heard about their Wayback Machine, but they preserve books and other media too.
So, my question is, of all these books being scanned and then destroyed, aren’t there many more copies of these books? And BTW, knowledge isn’t permanent. It can be lost.
If you had clicked on the article before commenting instead of just reading the headline, you’d have seen that it was truncated on the lemmy post
And if you had read the small number of other comments before commenting instead of just reading the headline, you’d have seen this correction already be posted: lemmy.world/comment/25456917
They just cut and bin the books after they’ve been scanned, the bins are taken to the local disposal, most municipalities have a garbage dump, some burn the trash.
An ai trained on unique data is more valuable if you don’t share the data it was trained on to your competition, your LLM gains an edge.
If I found a unique handwritten book from the 18th century, than I proceed to read it and memorise then burn it, I’ll be the only one to posses the knowledge of it.
There are very few, if any, rare, one of a kind books that would go through this process. If one of these books were scanned for an LLM it would be done delicately in a more time consuming manner and the book would be returned.
I am all up for starving machine learning of new knowledge but I view this as the least friction path. Otherwise, if the machines try to brute force their entry, what harm can kt cause?
Eh, perfect is the enemy of good. I’ll take the wins we can get. It’s not like having to buy the books is gonna stop the biggest offenders from burning investor money (along with the books).
The defeat of LLMs isn’t gonna come on that front. Let the cultists believe that more words will somehow become more than words. Their silicon gods run on power and water. That’s where they’re most vulnerable.
It’s not necessarily trying to stop them burning the books, just stop their ability to monopolise the knowledge, that’s their aim, lock the knowledge behind a subscription, Anna’s archive will always be available for all, so why pay for their subscription if its available for free.
I was aiming for the poetic parallel between burning money and destroying books, but also, isn’t the difference between destroying the primary source and preventing access to it mostly academic? Either way, as you say, the AI grifters are looking to control that information. Once they do, the original text and knowledge might as well be lost, since we couldn’t verify anything without the original to compare to. Whether through deliberate human manipulation or random word soup mutation, they can’t ever be trusted to deliver a genuine, faithful reproduction of the contents.
And the conclusion is the same: Better to have it freely available for everyone (including the grifters) than the grifters only.
Those who want to conduct large-scale scans and upload many titles can reach out to the archive for support, as the archive says that it “can help pay for the scanning fees and other rewards.”
My local library has some large format scanners. I’ve considered using them to do my photo albums. I have the individual photos scanned and backed up, but the layout and margin notes are worth preserving also.
If your local library doesn’t own a book scanner, any colleges near you may have one. You usually don’t have to be a student to walk into the library, and the book scanners I’ve used don’t require a login. YMMV on whether this is how it’s set up where you are.
I might try this. Im concerned that some of the books may still be under copyright. Not all of my rare books are that old. I collect indie poetry books and comics as well as old religious texts/pamphlets.
Also concerned that some might be literally some of the only remaining copies and I really don’t want to risk damaging them.
It’s worse than that. I’m sure they’re getting wet dreams about eradicating as many books as possible so that they can finally sell AI slop books at a premium price, claiming that will be “the last way to get something in [author] style!”.
I’m rather thinking that hey want to make their slop machines the single source of truth. Much easier to get control over information, when all information is located in your own house.
If you read the blog post, they don’t provide help or guidance for regular bookworms on how to scan. No information about what they expect.
Also, flatbed scanning a whole book takes many hours. It’s one page at a time, line up each page, press the book flat and scan, for hundreds of pages.
You’d likely need to buy, or have access to a book scanner. Which is fine, there are cheap models, but who knows if that would work for them? They don’t say. It doesn’t seem like they’re expecting help from you or me.
With modern phones having a document scan option with our phone cameras, is it possible we can scan and combine into a PDF file?
I am going to test it myself, see how easy it is, and find a way to ensure anyone can follow with low cognitive load, but hopefully it’s as easy as I think to do.
The sad part, is to do so would require a slow tedious task for books with hundreds of pages. It would be nice to see a large group tackle chunks at a time to the point we eventually have groups dedicated to genres and scouring to find them to preserve them.
Which app is good for this task? Last time I checked the result was garbage, as in not really better than simply the picture itself. But that was years ago…
Looking similar does mean that two things are similar. You could have tried to make your point in a variety of ways, but you chose the way that can be countered by just, understanding English.
Okay I’ll give you a hint. The bad thing about the nazis burning books wasn’t that they were destroying physical copies of books. I’ll give you a few keywords intent, selection, political control
You already fucked up the point you were making. By all means, insult people’s intelligence at the start of the conversation, before you fuck up. Doing it at the same time as fucking up looks bad… but insults after you fuck up looks even worse.
Every time I hear that claim it sounds weird to me. You can’t say that without an explanation.
Sure, scanning a print copy can mean distroying it but a book as a work still exist. We are very far from having scanned every books which physical copy are numerous. If AI trains on every romance novel published in the 80’s, it is far from distroying books by distroying physical copy.
Now if AI company are searching for rare books to scan and distroy, it is much more worrying but finding rare editions or print copies of books is a trade and the normal citizen can’t jumps in to save the day.
Many should probably be distroy without training a LLM or we will get models that are aggressivly mysoginist and who verbally abuse female users, on top of every other problem they cause.
They’re destroying things that don’t have to be destroyed and compiling the information into privately owned megastructures they’ll then charge us to access while we own nothing and they can then alter the books.
Nah I’m good. Everything these AI companies stand for is fucking evil and malicious.
For every AI that researches a legit medical issue, they start literally burning fucking books, replacing jobs, screwing basically everyone, and making the world overall much worse.
If you can’t see why these companies destroying any amount of books is bad, idk how to help you. Maybe look up a list of who has ever mass destroyed books. There’s not a lot of good people there.
Humans have been scanning books for decades. All books created today exist electronically before (a copy) gets physically printed of a book that remains existing.
Type writers are a museum piece. not a utility anymore. Calm down.
hurry, you better take up reading books to learn the meaning of these words you’re using before the AI ‘destroys’ the useless copy you never wanted nor thought about until now.
All books created today exist electronically before (a copy) gets physically printed of a book that remains existing.
Not really. There is a lot of full manuscrit books made to this day. By scholar in region where electricity is not as much available. By student compile hand-written notes from oral teaching. By people composing small things for family and friends or just for themselves. By activist and locally involed people making just on or a few copy to be share in physicial Space around them.
That’s a huge part of what exist outside the publishing industry.
And we used copiers and voice recorders for that since before computers. That has been used since the 70s. Even earlier.
scanning that and adding it to the common knowledge and making it accessible is something we’ve been doing since computers were a thing decades before today.
Students now still use recordings that can be parsed into text.
Suddenly it’s an issue now because someone needs so badly to be outraged by literally everything that can be distorted into a rage bait title like this one.
And how about you? Do you have the hard originals of every book? Why do you suddenly care only now that books are being copied?? They’ve been getting scanned for half a century now. You only care now to be outraged about it something for today
I don’t think that’s the argument you think it is.
The outrage is over the books being destructively scanned, en mass, and held in private after the fact. Very, very, very different than the majority of scanning which is non destructive, and usually shared.
If they buy 100k books and destroy them all, even if every one is unique, that is en mass. I don’t know why I have to explain that to you, but there you go.
They bought each one as a copy. They aren’t taking all copies of all books.
Not sure why I have to explain basic math to you but here we are.
Learn math bro.
And get if they are recycling THEIR ONE COPY OF EACH which they no longer use that isn’t illegal nor worth an outrage. That’s how recycling and reducing waste works. Better than hoarding.
Do feel free to go buy a copy and support the author if you’re this passionate to get enraged over what someone does with their own one single copy.
Lmao, I don’t see much point in talking with you, you’d be happily skipping around while books were being burned telling people it’s fine, there are plenty of other books, these ones aren’t theirs.
This literally isnt book burning though. it is recording and recycling of (one) copy of (a) or even each book. Not even the original. One copy of many of each. One that you didnt even care about it until now.
So you calling it book burning or comparing it to actual book burning is misinformation.
And it is you willfully wasting your time here trying to convince somoeone of a lie you’re desperately trying to push here to control the emotion of others. Manipulation. A shit rage bait piece for attention and pandering to unnecessary anxiety and trying to depress other’s no doubt.
So if I am annoying you by calling out this shitty faćade:
They are literally searching for rare books, paying premium costs for them, then cutting off their spines because it is faster to scan them that way.
Not only do they not need to destroy the books to scan them, but Google already did this 2 decades ago, and they did it non-destructively.
I just went on and read more article of the topics people shared in the comment. That’s sounds worrying indeed but my point is more that Annans Archive appeal is very abstract in a way that it présent something that doesn’t have to be a problem in a very sensationnalising manner instead of being clear about what is whappening and what we can do.
Independent bookstores all across Europe are seeing a trend like this, where they receive random, relatively large orders that are shipped to local addresses. While orders like these still come through in the digital age, especially from institutional buyers looking to fill out new libraries, they say that most “legitimate” orders often come with coordination and negotiation, not just a straight purchase order.
I guess the next step is picking specific people to memorize specific books and pass that knowledge down to an assistant/acolyte who will keep the memories alive. Rinse and repeat. Fahrenheit 451 anybody?
like im a bit lost here how it is burning all books or that it rob the idea making the book worthless? why cant this part be explained plainly when you bring this topic to the table here?
and if AI is destroying stuff, why are we being so shitty at communication?did AI destroy your ability to communicate proper meaning to others?
destroying (A) physical copy. in order to copy the (1) copy of a book(which most new already exist as pdf). kinda like how people do it when they scan the book. which has been happening for decades now to make knowledge permanent.
far cry from “destroying all books” which these rage bait titles suggest. we arent standing around a pit of fire.
bigger tragedy is how humans seem shittier at communication as AI learns. probably a more accurate title.
There are many, many non destructive scanning methods. My understanding is that most archival projects are using the non destructive options.
The easy option is to just slice off the binding and run it through a scanner.
The easy option is what is being talked about here, and those books are then fed into AI as training material. The digital copies aren’t “making knowledge permanent” as they are tucked away into a digital corner virtually no one has access to. They are as permanent as the source code for Fortnite.
If a PDF already exists, as you claim, then why go to the trouble of obtaining, destroying and scanning the books again?
They bought a copy of a book. They gonna do with it what they want. It’s their’s. Not your’s. Their’s.
You didn’t buy it. And you didn’t care about it till now. Over something you didn’t buy. Someone else did.
So now you are outraged they are making copies? This has been happening since copy machines were invented.
You’re outraged they destroyed THE ONE COPY they bought? (NOT ALL. JUST ONE)
They recycled rather than hoarded. Good for them.
This happens every day since books and other things were invented to manage unwanted junk.
If you want it go buy a copy and support the author, go for it. You want to recycle? Go for it. Both things are not illegal. Nor should they be.
I don’t get your problem here. Seems like you’re too stupid to know what to outrage about.
It’s like you discovered that dirt exists for the first time and how dare someone shovels dirt in some unmarked area you never even knew or cared existed to make a trench.
Yall keep overreacting to this bs, what’s even a rare book with no other copies, every article I saw says the books were scheduled to be destroyed or taken to a dump, and those are the books they are swooping up. I have a feeling rare means like 1st edition or one with typos or some sht. These are books ppl already weren’t buying that they are grabving in mass. All knowledge and all books aren’t important, plenty of below average intelligence ppl dropping slop long before ai became a thing. Any book that you value obviously already has tons of physical copies and digital backups.
This isn’t them looting some ancient libraries/museums taking some hidden/secret knowledge from humanity like yall are making up in your heads.
every article I saw says the books were scheduled to be destroyed or taken to a dump
Do you have a source for this? Because I’ve seen literally nothing about this being the way books are being sourced. And no offense but this reads way too close to weasel words. For example: Every article I’ve read on the subject says my pants are sentient and enjoys my legs being in them. Also, I’ve read no articles on the subject so that isn’t a lie.
Fair point. Beside the point, is it technically correct to make a positive assertion about a negative thing — only because any positive assertion about that negative thing is inherently part of an empty set?
Meaning like… our logic technically rules based on “no available contradictions, even if there is no evidence” as opposed to “at least one evidence is required?”
In certain situations I’d agree that the absence of contradiction is evidence enough, but those situations are pretty specific. For example, we don’t have absolute proof that orange juice doesn’t cause cancer, our “proof” comes from the lack of reliable evidence showing that it does, along with relevant research. But that doesn’t mean every claim can be treated the same way. But that doesn’t mean every claim can be treated the same way, especially a claim based only on “every article I saw,” without a source being provided.
the AI ran out of Reddit slop material to scan, because reddit has been purging alot of content/accounts even suspected of spamming, so they go for physical books now.
that’s what I’m saying. I’ve never understood the hoopla of reddit selling it’s corpus, aside from being fucking gross (spot on for spez) it’s a shitton of garbage. there’s gold in there, but it’s absolutely buried in shite
For an AI to troubleshoot problems with your computer (something it’s really good at today), it needs content like reddit’s. Books don’t and won’t detail weird software issues and their workarounds. One the current training set becomes irrelevant to whatever software people are running, this is a specific area where Reddit’s data is way more valuable than books.
Of course Reddit locking down access to that information leads to fewer people contributing. Reddit is dead, it’s just struggling to go down.
there’s gold in there, but it’s absolutely buried in shite
yep, there is a lot of nuanced depth, but measured against the overall sea of chuds lol… and yep, the moment they started putting up walls and controlling the community it was doomed. it’s gonna thrash for a while then, when the only ones left are advertisers and bots, it’ll shit itself and stop twitching.
im guessing reddit has been banning so aggressively lately, that AI scraping isnt getting as much “useful data” as much anymore (chatgtp/google dropped "references for reddit in thier slop summaries). so they are trying to force logins now to keep up with the user generated content so they can squeeze every last penny they can from reddit, i dont think ads were ever a big presence in reddit? its the data. people try to post content, like posts, subreddits, those can instantly get removed by redit filters. i think lost alot of traffic recently to, which was likely AI scraping.