posted in Technology
So Reddit has decided that plain HTML is unsafe
www.cole-k.com/2026/07/21/reddit/posted in Technology
So Reddit has decided that plain HTML is unsafe
www.cole-k.com/2026/07/21/reddit/Replying to @lemmydividebyzero@reddthat.com
The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
Replying to @grue@lemmy.world
Hmffh, anti-copyright. After all, every view is a copy to your machine. Just information being free.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves.
Replying to @artyom@piefed.social
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API
That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.
Replying to @artyom@piefed.social
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Replying to @grue@lemmy.world
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
Replying to @Toga77@lemmy.world
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
Replying to @grue@lemmy.world
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Replying to @rudyharrelson@lemmy.radio
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.
Replying to @grue@lemmy.world
Reddit does have RSS feeds
Replying to @grue@lemmy.world
Well they had an API, but...
Replying to @grue@lemmy.world
so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.
I blocked the fuckers, no qualms at all.
Even if they were using the API they were not being nice about their shit.
Replying to @grue@lemmy.world
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API
Reddit does provide RSS feeds, e.g.: www.reddit.com/r/SonicTheHedgehog/.rss
Frankly, I’m surprised they still offer RSS feeds. They’ve been slowly but surely killing off all ways of accessing their content for years. One day they’ll disable them, but for now they still work.
Its amazing to me how consistently the shitty behaviours of these billionaire techbro oligarchs impact disabled or marginalised people… even when the point isnt to directly shit on them. Its fucking vile.
I honestly think many (too many, but certainly not all! I am one) programmers are some of the immoral, ethically spurious people around in the 21st century.
Replying to @PurpleFanatic@quokk.au
That is why they want AI to replace programmers. AI morals are programmed, so they can be designed to do shitty things that a normal person would refuse.
Replying to @grue@lemmy.world
Semantic web was a separate initiative by Tim Berners-Lee when the web already used un-semantic HTML. And it never went anywhere.
Replying to @lemmydividebyzero@reddthat.com
afaik supposedly it was because a large chunk of the bot network went through the old.reddit portal over the standard reddit portal.
Replying to @Dudewitbow@lemmy.zip
Only because the bots were already set up to do it that way. It’s quicker to use what exists when it works.
Replying to @Dudewitbow@lemmy.zip
That’s just an excuse by reddit. Moving to the new style allows them to choke down and control the way that posts and replies are displayed and nested. This is good for them, because it allows them to offer white glove PR services to paying customers. It also obfuscates useful user supplied content so that it can be sold wholesale to anyone who has the money to buy it. That’s more important to them than offering a good user experience and useful website to the proles.
Anyone still posting on reddit (who isn’t a bot) is working for free for an unscrupulous company.
Replying to @voluble@lemmy.ca
Anyone still posting on reddit (who isn’t a bot) is working for free for an unscrupulous company.
Reddit was founded during the “crowdsourcing” craze. People realized that they could launch websites where all the content was created by the users, and it would snowball into daily views.
My point is that posting and commenting on Reddit has always been doing free work for an unscrupulous company. You could argue that it wasn’t unscrupulous before it went public, but even that is debatable.
Replying to @lemmydividebyzero@reddthat.com
even logged in I can’t access old reddit anymore.
Replying to @clanker_victim_555@lemmy.world
if you use the reddit enhanced suite extension you can get redirected
Replying to @GottaHaveFaith@fedia.io
If I wasn’t permabanned I might bother trying work arounds.
Replying to @clanker_victim_555@lemmy.world
I just checked my old.reddit login on the desktop… still working just fine today. I don’t think you need work-arounds, you just need to access it without any work-arounds.
Replying to @Shdwdrgn@mander.xyz
You think that if I do the thing that’s not working that it will somehow work.
Replying to @clanker_victim_555@lemmy.world
working for me
Replying to @lemmydividebyzero@reddthat.com
How dare you question King Steven the Turd, Greediest of Pigboys? If he proclaims HTML to be unsafe, it must be so. That’s a King’s job, to tell the Landed Gentry how to behave…
Replying to @dhork@lemmy.world
I’m not forgetting or forgiving that shit either.
Replying to @lemmydividebyzero@reddthat.com
I get blocked and asked to login to reddit no matter if it’s old or new. safereddit still works though so I’m using that for now.
I am so glad I’m off reddit. Even though I could use their help on a great number of subjects. Neverheless, because Isreal I can’t use reddit apparently. Not permabanned yet but they are on my shit with fake violations like right away now. I abandon them after a 2nd violation, go through one every 3 to 6 months when using it.
Every single company that goes public gets worse, reddit will be no exception. Those craven amoral cockscum are not to be trusted, nor patronized with our words they can use for their own benefit.
Replying to @lemmydividebyzero@reddthat.com
I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once.
For anyone who doesn’t know, you can download Wikipedia and host it yourself! I got the top 50k version (~7G) on my RPI3 and now no matter what fuckery they pull or the government pulls, I’ve got a pretty decent source of general information.
Replying to @HAL_9_TRILLION@lemmy.world
Yeah, I did that about 20 months ago (wonder why?) and the problem is, there it sits but I forget where on my LAN I stored it, how to access it - I suppose, were I absolutely desperate, I could launch a several hours research project and find it, maybe even figure out how to access (I believe I stored README with it about how to do all that), but… why? And, even worse, one why? might be to see how articles have drifted over time, which I guarantee they have and continue to do, Wikipedia is actually quite fluid, but does knowing how current articles compare to the ones you archived in 2024 do anything of greater value for you than the negative value of being pissed off at how the world is lying to itself (same as it ever has…)?
Replying to @MangoCats@feddit.it
I have my Pi-Hole doing DNS for me so I just set up an easy to remember redirect, wikipedia.local. I did it more for independence. I got Home Assistant and Voice PE so I could have a local smart speaker to get big tech out of my house. I don’t do anything on the cloud, so it stands to reason that if I want to make sure I always have access to rudimentary information, having my own version of WP is nice. I also just think it’s kind of cool and it never begs me for money.
Replying to @HAL_9_TRILLION@lemmy.world
It’s definitely cool - just not one of my top 100 cool projects on the list so after having made it work for 5 minutes I got the old ADHD onset and haven’t looked back at it for a couple of years.
Replying to @HAL_9_TRILLION@lemmy.world
Could you share what you did to achieve this? I’ve been planning on doing just that, and have it auto-update every week or so (keeping the previous versions archived, of course) by using kiwix-serve for a static ‘.zim’ file and maybe a cron job for the auto-update. But if you have a better solution, I’d love to know. The deployment I am planning is kind of convoluted to be honest.
Replying to @jjlinux@lemmy.zip
No, that’s exactly what I did, I’m running kiwix-serve, but I’m not going to bother updating it because I’m really worried about information degrading now that fascists are basically calling the shots on everything (and WP’s jackboot co-founder has a hard on for it). If I feel enough time has gone by to warrant an update I’ll just do it manually.
Replying to @HAL_9_TRILLION@lemmy.world
What year was your cutoff?
Replying to @Truscape@lemmy.blahaj.zone
About 6 months ago.
Replying to @HAL_9_TRILLION@lemmy.world
I ended up using mediawiki instead of kiwix-serve. Left my proxmox grabbing the data and its at around 170,000 pages right now. I’m still going to be versioning to make sure I keep the most up to date data, but still have access to the previous versions, as well as getting new articles, and keeping any removed ones if it happens.
Replying to @HAL_9_TRILLION@lemmy.world
I am currently pretty tapped out on all storage and backup drives but if I had space this post would have motivated me. Just sayin
Replying to @HAL_9_TRILLION@lemmy.world
How do you do that?
Replying to @MrOtingocni@lemmy.world
You run a server on a machine inside your house, it can be any computer on your local LAN/wifi, but it’s obviously best if it’s a machine that’s always on. I use a Raspberry Pi 3B+ (these can be had for about $50) that I have plugged into my wifi router and it’s running a little program called Kiwix-Server (free and open source). You download the WP file (it’s a huge single file with a .zim extension) and point the server to it and boom.
Replying to @HAL_9_TRILLION@lemmy.world
Wow, neat! Thanks for the explanation
Replying to @lemmydividebyzero@reddthat.com
Whats reddit?
Replying to @Steve@startrek.website
typo, it’s spelled “read it!”
It’s shitty moves like this that have me feeling deeply grateful about the fediverse. It’s NOT without its myriad of problems, but how lucky are we? We’re insulated from all this bullshit.
Mastodon is every bit as good (and better) as it was when I started using it in 2018. Can the same be said for Reddit, Instagram, Facebook or YouTube? Absolutely the fuck not.
Replying to @PurpleFanatic@quokk.au
I’ll admit I’m having trouble moving from yt to peertube :/ and still use the old gmail accounts
Replying to @Gsus4@mander.xyz
PT, nebula, loops, odysee, and floatplane together can’t fill the yt content gap.
You can start with moving to newpipe/grayjay and curate your own content which will lower your surface area.
At some point YT will manage to widevine and we’ll be torrenting the best of that shit.
Replying to @rumba@lemmy.zip
Ahead of the curve, already downloaded all of the greats onto my selfhosted storage.
Replying to @Gsus4@mander.xyz
You can use something like GrayJay to watch YT and other sources (and have offline playlists and subscriptions) without a google account whatsoever. That helped me make the jump to fully degoogle.
Replying to @Truscape@lemmy.blahaj.zone
The lack of a “cross-platform” (for lack of a better term) account is why I don’t use things like GrayJay because I watch on different devices and want to ensure they all have similar recommendations and a watch history.
I use SmartTube next on tv like 80% of the time I watch YouTube.
Replying to @MrScottyTay@sh.itjust.works
Grayjay allows you to sync between devices if desired, including platforms. I have a desktop, laptop, and a phone, and I can sync everything locally by just pairing the devices and having them at least 2 active for the transfer.
Replying to @Truscape@lemmy.blahaj.zone
So another has to be active at the same time as accessing another? I can’t always guarantee that so that’s a bit too much of a faff around and having to do that multiple times a day would be annoying. I’m glad it’s there for those if works for, if it works that way though still think it’s not right for me sadly.
Replying to @MrScottyTay@sh.itjust.works
You can use a “routing server” from FUTO (the guys behind it) to make it happen as well, although obviously that just means shifting traffic through a benevolent third party rather than only the devices you own and manage.
Replying to @MrScottyTay@sh.itjust.works
Recommendations are their method of control. Eshew the algorithm, curate your own choices of who you watch. also drastically reduces your exposure to slop.
Replying to @rumba@lemmy.zip
My recommendations have been mostly fine and just keep the channels i regularly watch at the forefront. I’ve had my account for decades now so my subscribed feed is too much of a mess to wrangle now.
That saying, I do really miss the custom folders you could once make on YouTube. Back then I would categorise certain favourite YouTubers together and mostly use that. Using something like that again would be nice.
Replying to @MrScottyTay@sh.itjust.works
They also have this feature, you have subscription groups!
As for recommendations, it does one big sync once then anytime you open the app it seemed to sync everything (ime at least)
That alongside being able to sync a video to my PC at the exact timestamp is very nice for living room experiences. Bonus you can have peer tube sources show up next to your yt videos
Replying to @DanWolfstone@leminal.space
I might have to properly spend some time to meeting about with it then. Is there an android tv app for it yet? (If you know off the top of your head, if you don’t i don’t expect you to do my research for me haha)
Replying to @MrScottyTay@sh.itjust.works
Exclusively? Not that I know of but I think the layout is dynamic so I don’t assume the normal app would be that different to operate… But then again I’ve never used an android TV with my own apps,
if you find out though lmk!
Replying to @Gsus4@mander.xyz
Yeah that’s the killer. Just not enough content yet. I think it’s cause video hosting is expensive. My hope is that individual youtubers start hosting their own peertube instances. But they’d never do that because they’d lose money.
If peertube could get functionality which would allow creators to hide videos behind a subcription which could be paid in fiat/crypto that would be a game changer. Easy to donate to your favourite creators while keeping federation.
Maybe one day.
Defo would recommend changing your email though. Takes a while to do all your accounts but once it’s done you’re free! Feels good.
Replying to @mildseason@sopuli.xyz
Defo would recommend changing your email though. Takes a while to do all your accounts but once it’s done you’re free! Feels good.
And if you move to provider that will let you use your own domain, you won’t need to update your accounts next time you switch.
Replying to @lemmydividebyzero@reddthat.com
If you consider scraping a threat then yes plain HTML might as well be giving up. The advantage of new reddit for that is quite clear: they can collect a bunch of data about your browser before deciding if you are a bot and if the rest of the page should load. The embedded recaptcha call in the screenshots is a pretty good hint. I suspect blocking trackers on new reddit will break as soon as the scrapers move over.
As for why they don’t just kill old reddit: a significant chunk of their active posters use it and are attached to it. So if they kill it entirely they will lose content. Posters are of course logged in so this change is less likely to affect them.
Replying to @carpelbridgesyndrome@sh.itjust.works
if you are a bot
Or they’re bouncing banned people.
funny, that comes just on the heels of my deciding that reddit is unsafe.
Replying to @turdburglar@piefed.social
You’re just now realizing it?
not really but it make for good joke pacing.
i’ve been here for a while. and not there for a while.
Replying to @lemmydividebyzero@reddthat.com
reddit stinks super bad
Replying to @redditStinksSuperBad@lemmy.world
a corpse of a platform
Replying to @lemmydividebyzero@reddthat.com
Whether or not it requires a login for old.reddit.com depends on client IP. I see this on a few networks, not on others.
Replying to @lemmydividebyzero@reddthat.com
Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them.
I call bs. It’s been completely destroyed for a while.
Replying to @PattyMcB@lemmy.world
I think they’re more upset AI companies scraped 'em and they didn’t get paid.
Replying to @JackbyDev@programming.dev
They told Google to pound sand and their stock took a hit. Fuck em
Replying to @PattyMcB@lemmy.world
You mean Reddit told Google? They actually have an agreement with Google where the latter get a direct feed of new posts and comments, and index them pretty much immediately. So not sure where you got the ‘told Google to pound sand’ idea.
Replying to @SlurpingPus@lemmy.world
cnbc.com/…/reddit-stock-google-ai-content-deal.ht…
They threatened to cut Google off from training AI and their stock dropped.
Replying to @PattyMcB@lemmy.world
I agree with you. But what’s important is if they can sell anything. They destroyed their own product but they’re still pretending it has value, and maybe they can fool some investors into paying for script and AI slop.
Replying to @PattyMcB@lemmy.world
While I agree, More celebs than ever have been asking for comments and ideas on their reddit pages.
Replying to @PattyMcB@lemmy.world
They are in the business of attracting new users into their new algorithmic engagement hellhole now. Show any propensity of interest, and you will get sidetracked to the most godawful side-communities that seem to have emerged to engage as many victims as possible. They do not want to focus on their old users as anything less than the content they already made that makes reddit show up as free advertisement to their new base in search engines. It’s all a game of “it’s the algorithm’s fault so you can’t blame us” now.
Replying to @lemmydividebyzero@reddthat.com
Fuck you spez.
Replying to @lemmydividebyzero@reddthat.com
Apart from all that, what would prevent an AI knowledge thieving tool from using an actual account on Reddit?
Replying to @Treczoks@lemmy.world
One would need more than 1 account…
Replying to @lemmydividebyzero@reddthat.com
As if that would stop an AI agent.
Replying to @lemmydividebyzero@reddthat.com
Gone
Replying to @DeadSquirrel@lemmy.blahaj.zone
tag users […] it’s super weird seeing people from one country pretending to be from another one.
I see issues.
Implementation A: automatic by IP)
Implementation B: field in profile settings)
Replying to @MonkderVierte@lemmy.zip
Gone
Replying to @DeadSquirrel@lemmy.blahaj.zone
Ahh, so jt’s you doing the tagging of others, only visible to you.
Replying to @MonkderVierte@lemmy.zip
Gone
Replying to @lemmydividebyzero@reddthat.com
This is because appending
site: reddit.comto a search query is basically a surefire way to find results written by genuine humans.
The article is from 2026, not 2016? That bot-ridden Reddit? Am i in the wrong film?
Replying to @MonkderVierte@lemmy.zip
wrong film
Second reality on your left, past the one where Santa Clause rules the world
Replying to @MonkderVierte@lemmy.zip
Or that site:url doesn’t return results it used too anymore.
Replying to @MonkderVierte@lemmy.zip
100% bot ridden. However site:reddit.com is used for the types of questions where otherwise, you just get page after page of SEO sites, which are even wose.
Replying to @MonkderVierte@lemmy.zip
Meanwhile, I use -site:reddit.com more and more.
We are not the same. ;)
Replying to @nibbs@lemmy.zip
I use Google hit hider by Domain ((also works on most major search engines), which also let’s you remove AI barf sites.
Replying to @MonkderVierte@lemmy.zip
Hey, those are real human made bots. Not shitty ai generated bots.
Replying to @lemmydividebyzero@reddthat.com
They’ve already disabled old.reddit for me. That makes it unusable, and thank you, Reddit. I actually used a domain blocker to block reddit, but generally would still be tempted to peek. Now, that is no longer the case. I, for one, am wholly in support of Reddit’s new anti-advertising stance!
Replying to @lemmydividebyzero@reddthat.com
Replying to @lemmydividebyzero@reddthat.com
I don’t even understand how there are still so many comments on that site. How is anyone even accessing it anymore? I just assume it’s 100% bots
Replying to @chunes@lemmy.world
The site is most hostile to visitors who aren’t logged in, and the users who comment probably mainly visit the site while logged in.
Nah I have been banned loads for insanely minor stuff. I got a sitewide ban having appealed a ban for some star trek opinion I cant remember.
My account was eight years old with a decrnt history and contributions. They give mods too luch sway because mods are losers doing work for free.
Replying to @chunes@lemmy.world
Inertia, just like Facebook.
Replying to @lemmydividebyzero@reddthat.com
“reddit has made a bad decision” is not news. it’s entirely what is expectable of them.
Replying to @lemmydividebyzero@reddthat.com
Yesterday, I visited new Reddit. No VPN, Chrome in incognito mode, so no extensions, clickedon a link directly from Google search. Got a message that my access was blocked for security reasons. Copied the link to Firefox (with uBO and a few privacy-centric extensions), changed it to old reddit (where I was already logged in), and it worked just fine. I found out that when Reddit kills old reddit, I won’t even have the choice to switch to the new one (not that I ever would) because I’d be blocked anyway.
Replying to @Bruncvik@lemmy.world
It works through Tor, by the way. I had the same experience can’t look at it but I can through Tor. They even have a .onion.
Replying to @Bruncvik@lemmy.world
seems like Old reddit is account-walled as of today maybe
Replying to @lemmydividebyzero@reddthat.com
Wait. How does that work? Wouldn’t plain html be safer because it’s JUST html?? Or are there vulnerabilities in just html that I’m not aware of?
Replying to @SCmSTR@lemmy.blahaj.zone
It’s unsafe for their profits.
Replying to @SCmSTR@lemmy.blahaj.zone
They could use a WAF to block AI scrapers, and both Cap and Shieldstral together to block most AI posters but they won’t bother.
Replying to @lemmydividebyzero@reddthat.com
I haven’t maintained my personal website in ages. That, of course, means it loads instantaneously, in comparison to all these Library of Congress websites.
Replying to @lemmydividebyzero@reddthat.com
Quit using reddit. I don’t give a shit what your excuse is, you’re propping up fascism.
Replying to @Jaysyn@lemmy.world
My cousins say they cannot find any discourse communities for their hobbies outside Reddit. I would respect the decision more if they only stuck to those subreddits but they don’t.
Replying to @OldChicoAle@lemmy.world
Again, I don’t fucking care about your excuses.
Replying to @OldChicoAle@lemmy.world
Skill issue
Replying to @lemmydividebyzero@reddthat.com
this is like when in a community i am looked at as a criminal because i don’t drink alcohol.
Replying to @lemmydividebyzero@reddthat.com
SMHTML