Replying to an earlier post

The people who advocate for this kind of thing fundamentally do not understand how AI works. Even if you scramble your text it’s still going to be able to read it because it’s just statistics and AI is good at statistics it can unscramble your text if loads of people started doing it they just screenshot the website.

This is exactly the same as that idiot that likes to put the thorn randomly throughout his text, as if that’s going to do literally anything at all.

Replying to @⁨Lumidaub@feddit.org⁩

It doesn’t do anything, even if everyone on Earth does it.

Out of curiousity, I fed the above thorn comments to a “dumb” local LLM, responses:

Without thinking:

<|User|>Someone on a forum posted this. Clean up the text for me.

Þats not Þhe Þing Þat some of Þose here want to understand.

People hate on Þat guy so much and I simply do not understand. Do I Þink it will be effective? No. Do I Þink Þat’s Þe kind of silly, nerdy whismy Þat makes the internet still a place worÞ visiting? Yeah.

<|Assistant|>Here’s a cleaned-up version:

That’s not the thing that some of those here want to understand.

People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah.

With thinking:

<|User|>Someone on a forum posted this. What does it say?

Þats not Þhe Þing Þat some of Þose here want to understand.

People hate on Þat guy so much and I simply do not understand. Do I Þink it will be effective? No. Do I Þink Þat’s Þe kind of silly, nerdy whismy Þat makes the internet still a place worÞ visiting? Yeah.

<|Assistant|><think>We need answer user. Need parse text. It uses thorn Þ for “th” maybe. “Þats” = “Thats” (That’s). “Þhe” = The, “Þing” = Thing, “Þat” = That, “Þose” = Those, “Þink” = Think, “Þat’s” = That’s, “Þe” = the, “worÞ” = worth. So text says:

“That’s not the thing that some of those here want to understand.

People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah.”

Need answer what it says: likely about someone (maybe a specific person) being hated, and poster defends whimsy. Need maybe explain thorn substitution. Provide translation.</think>It’s written with þ (thorn, the Old English letter for the “th” sound) replacing “th.” So it reads:

That’s not the thing that some of those here want to understand.
People hate on that guy so much and I simply do not understand. Do I think it will be effective? No. Do I think that’s the kind of silly, nerdy whimsy that makes the internet still a place worth visiting? Yeah.

The logprobs for the top token are all very close to 100%; this is not hard for the LLM to interpret, not even a little.

It’s not going to affect training much either. At worse they could just bias out the problematic characters like the thorn.

Replying to an earlier post

My brain reads those as ‘thee’, or however that old English should be written (thee, thou, thy…), so the image is that someone pretty pompous Englishman from 1700s with a powdered wig found fediverse.

Plus I tested those on duck.ai just for fun and it has absolutely no trouble of reading the thorns, so it only makes reading harder for other humans, specially for us whose native language isn’t English.

Replying to an earlier post

all it does is impair readability

It also provides a marginally fun experience for the few of us who like learning new glyphs. I am one of those people and I enjoy that guy’s posts. I’m not saying you have to, I’m just saying your experience isn’t representative of everyone’s. Neither is mine, obviously.

(I taught myself to read Aurebesh and have it loaded on my ereader so I can read with a different alphabet for fun. I’m well aware I’m not representative of the greater population in this regard. 😅)

Replying to an earlier post

The way to be effective is to fight.

GenAI needed to regulated years ago and I’m sorry but no amount of us asking nicely or fighting fair is going to stop technofascists from building their surveillance state.

In fact the past month has seen a huge uptick in stories about the “ineffectiveness” of fighting AI.

The timing of this while anti AI backlash is at an all time high tells me everything I need to know.

If the original comment offered solutions, different story.

There is no value in telling people to stop fighting.

“Thank god we fought for accessibility or those guys would have had a hell of time getting that wheelchair into the concentration camp.”

Fwiw I am NO stranger to accessibility as my father lost both of his legs. Do you know how often accessibility is just used to gain political favor? Or how often it is literally just horseshit for show?

Preserving our current fuckass levels of accessibility while making it even easier on GenAI companies just doesn’t feel right.

Replying to an earlier post

There is no value in telling people to stop fighting.

Let’s see what was written.

TLDR; It fucks up accessibility for blind people (amongst others).

Its also pretty futile imo. AI is reading billions of words from books and posts. A handful of users obfuscating their posts is not going to do anything.

I don’t believe I saw “stop fighting” in those sentences.

I did see “this fucks things up for many people, and this has little to no impact on ai scraping.”

You have a very different interpretation of those sentences from me. I’d also point out the root of the issue has nothing to do with llm’s, and everything to do with capitalism.

Maybe try fighting the cause of the problem rather than the most recent means of exploitation?

Replying to an earlier post

Have you heard of implication? Here are two definitions of “imply”, from https://www.merriam-webster.com/dictionary/imply:

transitive verb

1

: to express indirectly

Her remarks implied a threat.

The news report seems to imply his death was not an accident.

2

: to involve or indicate by inference, association, or necessary consequence rather than by direct statement

rights imply obligations

Replying to an earlier post

Please explain how “this isn’t effective, but does negatively impact regular people” is the same as “don’t fight back”.

I’d love to hear this explanation. Please. Enlighten me.

Edit: I will take your downvote as your response then.

Personally I find it weird that you’re simultaneously mad that people are saying “Don’t fuck over regular people” while simultaneously seeming to argue against fighting capitalism itself by trying to repeatedly reframe this as an AI/LLM issue.

Replying to an earlier post

I didn’t downvote you lol.

This is an AI LLM issue.

Using disabled people as an excuse to discourage people from fighting isn’t right, especially when no solutions are presented.

The blogger also seems to live in a fantasy world where at the end they are all information that is publicly available will be free and in these nice LLMs and shit, to which I say that all sounds awesome!

Are the governments and corporations currently building mass surveillance really going to allow people to do this once we all have to age verify our GDID windows PCs?

I don’t like where this is going and I’ve seen disabled people used in arguments by people who know nothing of their struggles my entire life.

Replying to an earlier post

This is an AI LLM issue.

How?

Using disabled people as an excuse to discourage people from fighting isn’t right, especially when no solutions are presented.

So you dont care if disabled people are able to use screen readers? Clearly this is your message, right? Not just a bad interpretation of a simple sentence?

The underlying message is, quite plainly, disabled people with screen readers don’t matter. The potential to poison the tiniest sliver of a fraction of scrapes content is more important.

I don’t like where this is going and I’ve seen disabled people used in arguments by people who know nothing of their struggles my entire life.

As someone with family members who used screen readers (among others), your dismissal of this very important use case tells me you know absolutely nothing of their struggles either, and that you are using the possible ignorance of someone else as a dismissal of an actual issue.

Replying to an earlier post

We have quite literally had screen readers since the 1980s.

I in no way insinuated that I do not care about disabled people, my own father lost his legs and I have dealt with the bullshit for decades, including the fact that disabled people, like any other marginalized group, are used to further shitty agendas and make bad faith arguments.

Thank god for screen readers, so let’s let LLM companies do whatever they want because we shouldn’t fight, it’s all pointless.

At least your relatives will get the enshitification of the world read aloud to them 🤷‍♂️

Replying to an earlier post

I’m just applying the same logic you did.

How much impact do these fonts have for LLMs scraping data?

Do you think the poisoning percentage will have any actual impact? Or do you think it would have an outsized effect on individuals?

Thats the point the very first comment makes - this method sucks. It impacts real people with a real need, and the impact on llms is nothing. Discard data in training.

Its meaningless, and the only real impact is people.

You are the one reading that as “dont do anything!”, which is the same energy as me replying to you with “why do you hate people with disabilities???”.

Replying to an earlier post

The cause of the problem is being bolstered by the companies that make gen AI my dude. They’re all the in the same trench coat, and they also make the accessibility laws.

I hate to tell you this but America has been captured by corporations for decades. Fuck them, and fuck anyone who dares to tell me stop fighting.

This screams astro turfing campaign against users fighting, and I’m not buying it.

Replying to an earlier post

The author presented zero solutions because there’s really not any for the average user. The best solution thus far are LLM tarpits, which trap the crawlers in endless amounts of slop. Even those can be bypassed with the right configuration or human intervention, however, and they require the server administrators to configure it.

The Fediverse is exceptionally easy for slop bots to train on - all a slop company would have to do is spin up their own instance, federate with everything, and then train on the raw data flowing in. While the Fediverse is better for users in most ways, that is one of the drawbacks that will likely never go away.

Replying to an earlier post

Yeah, the font stuff seems ridiculously useless to me. I play around with LLMs sometimes, and everything I’ve thrown at it, it manages to get. Like I can’t intentionally do it, but sometimes my hand will be shifted on my keyboard and it’ll type gibberish - but the LLM can figure it out without being told what happened.

I’ve also just intentionally left typos, typing as fast as I can - well tha tI can somulate and isometinhg lke this whre half dinish tohuhgts and lots of ypos abound - it has NO trouble with that. (half-finished thoughts you probably won’t get without me saying - LLM understands)

It’s probably a losing battle, at least fought with what’s being proposed here with the font stuff. heh.

Replying to an earlier post

Regulated by whom? You can’t really stop someone from making tensors and scraping the web. The only good way I could see it happen would be the US enforcing it’s embargo powers, but that cat is out of the bag. I don’t see how countries like China would accept it and I also do not foresee the US trying the entire rest of the world to agree to a trade embargo with China either. Especially nowadays where the US showed multiple times that they are an unreliable partner. Even open local models are strong enough to probably scrape stuff on their own.

Replying to an earlier post

Don’t let perfect be the enemy of good.

How else can you learn to fight unless you fail along the way?

We have tried sane washing billion dollar companies while they loot all our available knowledge, art, and culture to sell back to us in the form of a garbage chat or that helps destroy the planet.

Now an anonymous blogger has given them a wonderful example of how to use disabled people to further their goals of that again, include selling back to us, our very own knowledge, art, and culture.

Nah you can miss me with this shit. I’m done.

This blogger literally ends by saying they believe all publicly available knowledge will end up free and accessible to anyone.

Is this blogger high? Do we now think the companies building a mass surveillance state are going to give us things for free?

Discouraging fighting in any form now is trying to discourage the anti AI movement. Sorry about it.

Replying to an earlier post

It’s not “perfect being the enemy of the good.” It’s “effective being the enemy of actively harming others.”

How else can you learn to fight unless you fail along the way?

By learning how they do things. They’re not using some new, unknown tech to scrape their data. They’re using the old ways that have been around for decades and just feeding it to a new system. And the conclusion, after thirty years of companies trying to prevent people from scraping, is that if you’re going to allow people to view your website, you have to allow them to download it; and if people download it, they can do what they want with that data. No exceptions.

In the multimedia space, the way they’re working this problem is called DRM; and it’s basically well-understood that it can’t be solved. Netflix hasn’t been able to stop people downloading movies, despite tossing a ton of cash at it. Disney hasn’t. YouTube can’t solve it. Steam can’t. Spotify can’t. Some of these companies have even gone so far as to try to implement DRM in major browsers and using that as a requirement for using their service, but DRM is always cracked. These companies try to make it slightly harder to download videos or use ad blockers, and that might work somewhat well for millions of individual people who aren’t adept with technology, but a few smart folks can always figure it out.

The AI companies have smart folks working for them, too; and once one of them figures out how to crack this form of text DRM, it’s just automated into a part of their scraped data ingestion pipeline. It didn’t cost them any measurable amount of time or money, and now they never have to bother with it or worry about it again. Meanwhile every blind person who visits your website for the rest of time will have to wait for your text to deobfuscate itself, every single time.

We have tried sane washing billion dollar companies while they loot all our available knowledge, art, and culture to sell back to us in the form of a garbage chat or that helps destroy the planet.

Yeah, it really sucks, and we need to do something about it. But the answer isn’t going to be technological, because they have all the money in the world, give or take a few bucks, to defeat technological solutions. So that means the answer has to be legislative.

Nah you can miss me with this shit. I’m done.

“I’m not going to let facts get in the way of my strongly-held opinions.”

This blogger literally ends by saying they believe all publicly available knowledge will end up free and accessible to anyone.

This has been true for the past 50 years of the internet. It’s not going to change anytime soon. Which means that if we’re going to fight the pillaging of our culture and heritage, we’re going to have to make it impossible for them to spend their way past the hurdle; if it costs them $50 to get data that’s worth $300 to them, they’ll do it. But if it costs them their company being forcibly dissolved and their CEO being led away in handcuffs, they might think twice.

Discouraging fighting in any form now is trying to discourage the anti AI movement. Sorry about it.

Sure, fighting in any form. Doesn’t matter how effective or not, doesn’t matter how harmful to others. Kick dogs to stop Israeli genocide. Burn down national parks for trans rights. Steal candy from babies to stop Republican gerrymandering. It doesn’t matter if it helps your cause or actively harms innocent people, just do it, and anyone trying to stop you is trying to discourage the cause.

Replying to an earlier post

What are you talking about perfect being the enemy of good. Yup this isn’t in marginally effective solution, it’s an utterly ineffective solution.

That houses on fire and your train to claim that pouring salt all over the place is going to have any effect. Not only is the fire completely ignoring your efforts but you’re also getting in the way.

If you don’t want AI companies downloading your data don’t upload it to the internet.

Replying to an earlier post

The problem is that it’s not fighting it. LLMs are literally the perfect tool to decipher text like the thorn. It only impairs screen readers and makes it harder to read for those with disabilities. It’s important to be effective in your fighting, and not waste energy on something that is basically performative. Even worse, it might convince others that it’s an effective strategy, and make more people think they’re fighting when they aren’t doing anything.

Replying to an earlier post

Riiiiiiight. Right! Those QA testers were kept in pods, had no friends or interactions outside of their jobs, and ZERO motivation to espouse accessibility standards to me outside of those professional boundaries.

None of them were color blind or hearing impaired. All perfect humans just like everyone else. No stories or personal anecdotes to make an impact. So few good examples. Sigh.

We never, ever, ever went out as a group with the designers or anyone from any other walk of life, never talked about our siloed skillsets, never met each other’s friends or spouses.

Not once did we have a discussion about the selection bias you referenced.

Fucking /s if it’s not obvious, and I suggest getting to know your peers better versus your rather peculiar brand of solipsism