Replying to @⁨obsidian@discuss.online⁩

The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.

Local LLMs have gotten to the point where they are a serious threat to Anthropic and OpenAI, and Nvidia has a lot of skin in the game. If Nvidia wanted to do some serious damage to local LLMs, they are now in a position to do so.

I’m also imagining they may try to squeeze out support for other GPU vendors. I’m using an AMD 7900 XTX to run Qwen 3.8 27B that I downloaded from HF to run on llama.cpp, which currently works like a dream. The 7900 XTX is the only sanely priced 24GB GPU left in 2026 (under $1k vs. $2k, $3k, $4k for Nvidia 24-32GB cards). Combined with OpenCode or Pi, a setup like this basically eliminates the need to use Anthropic or OpenAI products in the same way Jellyfin eliminates the need to use streaming services.

I’m sure Nvidia and their buddies don’t like one thing I’ve said in this comment and may very well be plotting to put a stop to it, so the community may need to step up our game and get our eggs out of the big tech basket.

Replying to @⁨melfie@lemmy.zip⁩

The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.

I guess that explains why features like quantized KV caches are lagging behind. Maintainers are purposely dragging their feet.

the community may need to step up our game and get our eggs out of the big tech basket.

The community has chosen to not fight at all, which is worse. Anti-AI sentiment is at an all-time high.

Publicly. Privately, these hypocrites still whisper in ChatGPT’s ear when they get lazy enough. Or use some feature in Photoshop or some other software that they didn’t even understand was AI-driven.

Replying to @⁨p03locke@lemmy.dbzer0.com⁩

guess that explains why features like quantized KV caches are lagging behind

Yeah, I was using the TheTom fork for a while and not sure why TQ KV cache hasn’t merged yet.

Anti-AI sentiment is at an all-time high

I suppose the tech bros have understandably soured a lot of people on LLMs with all of the negative societal costs LLMs are created due to their greedy and irresponsible bejavior. On the other hand, the concepts of the perceptron and artificial neural networks from the 40s and 50s are finally coming to fruition and we have these things now that are legitimately artificially intelligent that we can run on our gaming PCs. They are overhyped and used in ways they make no sense, but they’re also useful as long as their limitations are kept in mind. From a technology perspective, they’re cool as hell.

en