posted in Technology
NVIDIA reportedly buys HuggingFace for 13 billion
Damn, I hope this is not goodbye to uncensored models 😢
gizmodo.com/nvidia-reportedly-stops-flirting-with-hugging-face-and-just-buys-it-2000803681posted in Technology
NVIDIA reportedly buys HuggingFace for 13 billion
Damn, I hope this is not goodbye to uncensored models 😢
gizmodo.com/nvidia-reportedly-stops-flirting-with-hugging-face-and-just-buys-it-2000803681Replying to @obsidian@discuss.online
The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.
Local LLMs have gotten to the point where they are a serious threat to Anthropic and OpenAI, and Nvidia has a lot of skin in the game. If Nvidia wanted to do some serious damage to local LLMs, they are now in a position to do so.
I’m also imagining they may try to squeeze out support for other GPU vendors. I’m using an AMD 7900 XTX to run Qwen 3.8 27B that I downloaded from HF to run on llama.cpp, which currently works like a dream. The 7900 XTX is the only sanely priced 24GB GPU left in 2026 (under $1k vs. $2k, $3k, $4k for Nvidia 24-32GB cards). Combined with OpenCode or Pi, a setup like this basically eliminates the need to use Anthropic or OpenAI products in the same way Jellyfin eliminates the need to use streaming services.
I’m sure Nvidia and their buddies don’t like one thing I’ve said in this comment and may very well be plotting to put a stop to it, so the community may need to step up our game and get our eggs out of the big tech basket.
Replying to @melfie@lemmy.zip
Well the good news is they can’t take away from you what you already have. It being an open source project, I’m assuming if they do anything to deliberately gut AMD performance, it’ll get forked.
Also
The 7900 XTX is the only sanely priced 24GB GPU left in 2026 (under $1k vs. $2k, $3k, $4k for Nvidia 24-32GB cards).
Not on sale anymore, at least not at any vendor in my country, I searched an aggregate pricing website. Amazon has a few used ones left of some models, but that’s probably a 2 or 3 digit figure across SKUs. Hold on to yours with an iron grip.
What kind of tok/s are you getting with it on Qwen 3.8 27B and how’s the output quality? I may consider getting one if I can find one used or import from abroad.
Replying to @boonhet@sopuli.xyz
Agreed, I was referring more to future updates. Obviously we’re good with what is available now.
I can’t speak to pricing and availability outside the U.S., but it looks like the one I got went up $100:
https://www.newegg.com/asrock-radeon-rx7900xtx-24g-radeon-rx-7900-xtx-24gb-graphics-card-triple-fans/p/N82E16814930084.
I traded in my 3070 and my final price was in the 700s. Last I looked, used ones were going for $800 on eBay vs. $1200 for a used 3090.
I run 3.8 27B at q4 with q4 context up to 200k. Decode is generally in the 30s and pp starts in the 700s and drops to the 400s as context approaches 200k. I use mostly Sonnet 5 at work and I would rate this setup with the OpenCode desktop app as pretty comparable overall for coding at least. Let’s just say I have no reason to use any cloud models, not that I would do that voluntarily outside of being compelled to at work.
Replying to @boonhet@sopuli.xyz
I got mine used, around 600 imperial credits. Look for ads that provide proof of working and benchmarks (like FurMark)
Replying to @boonhet@sopuli.xyz
I get around 40 token/s and it frequently has become reliable enough to drop sonnet for me. So take that as you will
Replying to @boonhet@sopuli.xyz
Intel B50 and B60 pros are at microcenter right now perfect for this.