posted in Technology
So Reddit has decided that plain HTML is unsafe
www.cole-k.com/2026/07/21/reddit/posted in Technology
So Reddit has decided that plain HTML is unsafe
www.cole-k.com/2026/07/21/reddit/Replying to @lemmydividebyzero@reddthat.com
The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves.
Replying to @artyom@piefed.social
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
Replying to @grue@lemmy.world
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
Replying to @Toga77@lemmy.world
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
Replying to @grue@lemmy.world
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Replying to @rudyharrelson@lemmy.radio
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.