Replying to @⁨SorryQuick@lemmy.ca⁩

That was hardly conspiracy worthy. It’s a bit hyperbolic, but it’s a fairly legit concern. Big tech is not the friend of any consumer. And big tech buying literally anything ends up a net negative for the consumer and freedom of choice. There is evidence that tech oligarchs (openAI/Sam Altman) want to gatekeep even knowledge and charge a subscription to access it.

Huggingface is now absolutely going to become pay to play. Maybe not day one, but it will happen. Just like all the other services out there that are now subscription based that were once free and offer little to justify the move. It’s also likely to drop AMD development or at least deprioritize it so hard they may as well.

There are also movements in the industry that suggest these big tech companies want to push a move to thinclients connected to cloud services instead of local machines. This idea isn’t helped by companies like Sony stripping away physical media. Nor by the RAM manufacturing cartel colluding for overpricing and abandonment of consumer markets.

Conspiracies are only conspiracies until they come to pass, and not all conspiracies are unfounded tripe like flat earthers etc.

Replying to @⁨CosmoNova@lemmy.world⁩

Why do we need a centralized hub for AI models anyway? I’ve never figured this out. We flock to these big “friendly” fucking companies with obviously unsustainable business models trying to own and sell things that should be shared in a decentralized, democratized mesh anyway. AI models should all be distributed magnet links, not hosted files. Why do we do this to ourselves?

The way things are going at Huggingface now, I imagine we probably will have to start building the infrastructure we need to handle AI models as distributed links sooner rather than later. Let the enshittification begin, we’ll move on to a different tactic for sharing AI models while cheerfully they squeeze cash out of people and businesses too lazy to adapt. Everything working as it should, I guess.

Replying to @⁨cecilkorik@lemmy.ca⁩

The way things are going at Huggingface now, I imagine we probably will have to start building the infrastructure we need to handle AI models as distributed links sooner rather than later.

Then “we” better get started. Mega corpos have fully embraced LLMs, to the point that they’ve distrusted human input too much, which is still a vital component of successful adoption. Eventually, more and more companies will wise up and figure out how to find the right balance.

Where are we at? Oh, right… we’re too busy arguing about how all AI is bad and sticking our fucking heads in the sand until the Big Bad AI Problem goes away.

We have to fix our attitudes if we have any hope of surviving this mess and not ending up as a Cyberpunk-wannabe dystopia by the time 2077 hits. Use the fucking weapons given to us!

Replying to @⁨DarkCloud@lemmy.world⁩

Nvidia also sells super expensive graphics cards. They probably want you to buy a couple of 5090s or a Spark.

They know they can’t shut down the entire self hosted LLM market but they can direct it towards using Nvidia GPUs by making sure llama.cpp devs don’t spend much time on ROCm support, etc. And they probably know that ClosedAI and Anthropic’s days on the frontier are limited.

Replying to @⁨FlexibleToast@lemmy.world⁩

seconded, they charge $21/month/seat if you’re using the pro features; that’s $50k a year even for a smallish team of 20, just to use their website. Once you get to a bigger company with 800 seats, that’s $1M over 5 years, for something that was mostly built on free OSS software to begin with.

Not to mention the biggest bait-and-switch in history that they pulled off on Github Copilot subscriptions, that I’m guessing a lot of big orgs won’t have cancelled yet. They really are laughing all the way to the bank.

Replying to @⁨obsidian@discuss.online⁩

So they are just slapping random price tags on things now. It’s a database of AI questions. They have paid 13 billion dollars for a database, and everyone’s acting like that’s a perfectly rational sensible thing to do. No one in the financial industry has any sense anymore.

A company that’s only product is something that people either don’t want or actively hate has paid an eye-watering amount of money for a database to train their product on, this training will have no effect whatsoever on whether people want it.

I have this nice bridge with lots of training examples if anybody’s interested, 100 trillion dollars please

Replying to @⁨Alcoholicorn@mander.xyz⁩

thats completely false when you only ask the llm a very basic question like “what color is the sun?” you need to talk to the model like its smarter than a google search engine. “what color is the sun? pull the data from scientific articles and other sites that study the sun”

if you google the first question, you will have list of different pages, where if you use AI to search on your query, it will model an answer from thoese sources, with said sources linked, similar to a Wikipedia page.

Replying to @⁨1985MustangCobra@lemmy.ca⁩

There are thousands of papers on the color of the sun, its trivial to google a paper, if you need an LLM to do this, that’s on you. But if you ask it valve clearances for an indonesian motorbike in english, it will make it up, even though the service manual is available in Indonesian. If you ask it how to synthesize a chemical that nobody bothers to write about because its uninteresting or impossible, it will straight up lie.

Replying to @⁨1985MustangCobra@lemmy.ca⁩

OK, instead of “lie”, I should have said it produces incorrect information any time the answer isn’t obvious from a cursory google search, and does so in a way that deceives anyone who isn’t already familiar enough with the subject that they didn’t need to ask in the first place.

It’s fine for asking it to generate answers you could have typed out yourself, such as boilerplate code and “what color is the sun”.

Replying to @⁨1985MustangCobra@lemmy.ca⁩

Yeah because it was impossible to Google things in the past. Every time anybody comes up with a use for AI it’s basically just automating something that isn’t even that hard for you to do. If you want to automate simple tasks that’s absolutely fine but it’s not the second coming as people keep insisting.

Come back to me when an AI can actually add value to something, when it can do something that is not just automating a simple task, but it’s capable of doing things that humans are not. I keep being told that we’re only a few years away from the technological singularity, so wake me up when we actually get there.

Replying to @⁨echodot@feddit.uk⁩

when it can do something that is not just automating a simple task, but it’s capable of doing things that humans are no

No, credit when credit’s due, LLMs have arrived (by themselves) or helped to arrive at some mathematical conclusions, didn’t they? A few of Erdős, some others. Yes, it’s, a very small subset (duh), out of god knows how many they tried, and achieving the solution is made via technique called “infinite monkeys typing out Hamlet”, but credit when credit’s due.

Replying to @⁨obsidian@discuss.online⁩

The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.

Local LLMs have gotten to the point where they are a serious threat to Anthropic and OpenAI, and Nvidia has a lot of skin in the game. If Nvidia wanted to do some serious damage to local LLMs, they are now in a position to do so.

I’m also imagining they may try to squeeze out support for other GPU vendors. I’m using an AMD 7900 XTX to run Qwen 3.8 27B that I downloaded from HF to run on llama.cpp, which currently works like a dream. The 7900 XTX is the only sanely priced 24GB GPU left in 2026 (under $1k vs. $2k, $3k, $4k for Nvidia 24-32GB cards). Combined with OpenCode or Pi, a setup like this basically eliminates the need to use Anthropic or OpenAI products in the same way Jellyfin eliminates the need to use streaming services.

I’m sure Nvidia and their buddies don’t like one thing I’ve said in this comment and may very well be plotting to put a stop to it, so the community may need to step up our game and get our eggs out of the big tech basket.

Replying to @⁨melfie@lemmy.zip⁩

Well the good news is they can’t take away from you what you already have. It being an open source project, I’m assuming if they do anything to deliberately gut AMD performance, it’ll get forked.

Also

The 7900 XTX is the only sanely priced 24GB GPU left in 2026 (under $1k vs. $2k, $3k, $4k for Nvidia 24-32GB cards).

Not on sale anymore, at least not at any vendor in my country, I searched an aggregate pricing website. Amazon has a few used ones left of some models, but that’s probably a 2 or 3 digit figure across SKUs. Hold on to yours with an iron grip.

What kind of tok/s are you getting with it on Qwen 3.8 27B and how’s the output quality? I may consider getting one if I can find one used or import from abroad.

Replying to @⁨boonhet@sopuli.xyz⁩

Agreed, I was referring more to future updates. Obviously we’re good with what is available now.

I can’t speak to pricing and availability outside the U.S., but it looks like the one I got went up $100:

https://www.newegg.com/asrock-radeon-rx7900xtx-24g-radeon-rx-7900-xtx-24gb-graphics-card-triple-fans/p/N82E16814930084. 

I traded in my 3070 and my final price was in the 700s. Last I looked, used ones were going for $800 on eBay vs. $1200 for a used 3090.

I run 3.8 27B at q4 with q4 context up to 200k. Decode is generally in the 30s and pp starts in the 700s and drops to the 400s as context approaches 200k. I use mostly Sonnet 5 at work and I would rate this setup with the OpenCode desktop app as pretty comparable overall for coding at least. Let’s just say I have no reason to use any cloud models, not that I would do that voluntarily outside of being compelled to at work.

Replying to @⁨melfie@lemmy.zip⁩

The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.

I guess that explains why features like quantized KV caches are lagging behind. Maintainers are purposely dragging their feet.

the community may need to step up our game and get our eggs out of the big tech basket.

The community has chosen to not fight at all, which is worse. Anti-AI sentiment is at an all-time high.

Publicly. Privately, these hypocrites still whisper in ChatGPT’s ear when they get lazy enough. Or use some feature in Photoshop or some other software that they didn’t even understand was AI-driven.

Replying to @⁨p03locke@lemmy.dbzer0.com⁩

guess that explains why features like quantized KV caches are lagging behind

Yeah, I was using the TheTom fork for a while and not sure why TQ KV cache hasn’t merged yet.

Anti-AI sentiment is at an all-time high

I suppose the tech bros have understandably soured a lot of people on LLMs with all of the negative societal costs LLMs are created due to their greedy and irresponsible bejavior. On the other hand, the concepts of the perceptron and artificial neural networks from the 40s and 50s are finally coming to fruition and we have these things now that are legitimately artificially intelligent that we can run on our gaming PCs. They are overhyped and used in ways they make no sense, but they’re also useful as long as their limitations are kept in mind. From a technology perspective, they’re cool as hell.