posted in Technology

If open weight models are the future, U.S. AI companies are going to have a hard time

www.fastcompany.com/91577359/why-u-s-ai-companies-cant-match-chinas-open-weight-frontier-models
Fast CompanyIf open weight models are the future, U.S. AI companies are going to have a hard timeOpenAI, Anthropic, and other Western labs have spent billions training the model weights at the heart of their most advanced systems.

Replying to @⁨sanitation@lemmy.today⁩

I ran DeepSeek and Llama and Mistral at home on my consumer grade gaming PC.

With a little tweaking of the system prompts and configuring web search, I was running a local LLM that felt pretty darn close to the commercial LLMs.

With this technology out in the open internet where you can download the models in a few hours I don’t see how the commercial AI companies are going to last. If selling “Artificial Intelligence” subscriptions is all your company does for revenue, you’re screwed.

I downloaded and ran an LLM that I could have a conversation with and feed basic coding problems to for basically zero dollars and ran it on my puny gaming machine…puny compared to enterprise-class hardware. It would be trivial for a company with a very moderate budget to buy some servers and start running their own LLMs that they can use to feed all the PII and HIPPA data they want.

Replying to @⁨DJKJuicy@sh.itjust.works⁩

sleepingrobots.com/dreams/stop-using-ollama/

And this is just the tip of the iceberg for ollama. They’re the same kind of scammy tech bros as OpenAI.

The best setup depends on your hardware. There is no “easy button” unfortunately, quantized LLMs are just too intense and finicky to run without making some informed choices.

It also depends on what you want to do with the LLM. For example, some are too slow or bad at long context for agenic use, some quantizations are great at scripts but terrible outside that, or vice versa.

But LM Studio and Qwen 3.5 35B Q4 is probably the “easiest” flat recommendation I can make.

Or… honestly, just pay $40 for basically unlimited usage for a year from an API, then roll your own frontend.

Friends Don't Let Friends Use OllamaSleeping RobotsFriends Don't Let Friends Use Ollama | Sleeping RobotsOllama gained traction by being the first easy llama.cpp wrapper, then spent years dodging attribution, misleading users, and pivoting to cloud, all while riding VC money earned on someone else's engine. Here's the full history, and why the alternatives are better.

Replying to @⁨naught101@lemmy.world⁩

I just meant that you have to be cognizant of what went into the quantization.

As an example, a “Q4_K_M” could be too much quantization to be usable on one model, and an inefficient waste of space on the other. Two Q4_K_Ms of the exact same model could be completely different, one totally borked. Or one particular Q4_K_M could excel in one task, but be totally useless for another, even with the exact same settings, when a slightly different sized or type of quantization would excel.

It’s a deep rabbit hole. It’s not random either; there are distinct technical reasons behind every case mentioned above.

And that’s not even at the cutting edge quantization anymore, though what’s “cutting edge” completely depends on your particular hardware and use case.

I’m trying to make this sound daunting on purpose.

Many people have really horrible experience with a default “ollama run” for this exact reason, because the defaults are terrible and the customization is critical to getting coherent, performant output.

Unquantized LLMs, on the other hand, are basically always run the same way: vllm docker image on a big server, official weights. There’s less to “go wrong” trying to squeeze it on hardware with unofficial runtimes and compressors.

Edited ⁨⁨Jul⁩ ⁨27⁩, ⁨2026⁩, ⁨02:13⁩⁩en