Replying to @⁨lemmydividebyzero@reddthat.com⁩

The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!

Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.

en

Replying to @⁨artyom@piefed.social⁩

The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.

Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.

That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.

Replying to @⁨artyom@piefed.social⁩

Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?

(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)

Replying to an earlier post

I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.

Replying to @⁨grue@lemmy.world⁩

The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API

Reddit does provide RSS feeds, e.g.: www.reddit.com/r/SonicTheHedgehog/.rss

Frankly, I’m surprised they still offer RSS feeds. They’ve been slowly but surely killing off all ways of accessing their content for years. One day they’ll disable them, but for now they still work.