posted in Technology
Nobody Will Say Why Every Major AI Chatbot Suddenly Went Down Yesterday
futurism.com/artificial-intelligence/nobody-saying-why-major-chatbot-outageposted in Technology
Nobody Will Say Why Every Major AI Chatbot Suddenly Went Down Yesterday
futurism.com/artificial-intelligence/nobody-saying-why-major-chatbot-outageReplying to @GolfFoxtrotLima@sh.itjust.works
Because they all use the illegal Colossus 2 data centre from SpaceX/XAi/fascism central and that data centre went down.
Google, Anthropic, and OpenAI all have contracts with them.
When that illegally running environmental disaster of data centre goes down, all those services all go over capacity and you get 502 rate limit errors.
Replying to @panda_abyss@lemmy.ca
It’s kinda surprising,
I know specifically where one of the big ones hosts its models and its not there, but I guess they could have infrastructure in there.
Replying to @BarbecueCowboy@lemmy.dbzer0.com
They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
Replying to @panda_abyss@lemmy.ca
The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.
(It’s not super complex code, just some scripts that I would not have taken to time to write manually.)
I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬
Replying to @percent@infosec.pub
Show me your set up!
Replying to @isVeryLoud@lemmy.ca
There’s not really anything interesting to show. It’s just a home server in a 13 year old desktop ATX case.
There’s no desk, monitor, keyboard, or mouse… But also no cool server rack.
Function over form, and it sits in a spare bedroom out of sight.
EDIT: I found the receipt for the case. It’s a Cougar Volant Black Steel mid tower, purchased in 2013. So my server just looks like this:
Replying to @percent@infosec.pub
I meant your LLM stack lol. I just have an RX 6800 XT in my main Linux PC for inference, but it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.
What are you using? vLLM? llama.cpp? Which params? How much CPU offloading? Do you use draft models? Is it a MoE model? Have you tried llama-swap? Which agentic front-end are you using? I presume you set it up to access it without SSH’ing into the machine, did you do anything special or is it just a raw unsecured open port on the machine to the LAN?
Replying to @isVeryLoud@lemmy.ca
it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.
eh, just systemctl isolate multi-user.target
Replying to @Damage@feddit.it
Correct, that’s how I would do it, but then I need another machine to act as a head.
Replying to @isVeryLoud@lemmy.ca
If your MB has onboard graphics, maybe you could mask the GPU and just pass it off to a container running the LLMs I guess
Replying to @Damage@feddit.it
No onboard graphics unfortunately, that would have been the easy way out.
Replying to @isVeryLoud@lemmy.ca
Well it works anyway even with a bit of occupied vram, but you could also buy a cheap videocard to use as an output. I have small intel card like that in my server for jellyfin transcoding, I think I paid 60€ for it, it hardly uses any power
Replying to @Damage@feddit.it
I actually did exactly that previously! I had both an RX 6800 XT and an RX 6600 in my system and I used the 6600 for video output. Unfortunately, this cuts my RX 6800 XT from PCIe 4 16x to PCIe 4 8x and severely slows down model loading for llama-swap. Joys of the X570!
And yes, I do have it running right now with a bit of occupied VRAM, but I need to limit my model to 14 GB to leave 2 GB free for GNOME Shell. I really want one of those 64 GB UMA Mac Mini, I heard they work really well because the GPU has direct access to system RAM.
Replying to @isVeryLoud@lemmy.ca
So I have a framework laptop with ryzen ai cpu that uses 48gb of shared ram, and it does run Q4 llms fine enough, but I’m not sure it compares to a real GPU.
On my desktop I have an RX 7900 XTX but I’ve only dabbled in image generation so far, so right now I couldn’t really tell you the difference.
Replying to @Damage@feddit.it
24 GB VRAM. Damn, jealous! My 16 GB seems pitiful in comparison 😅
I do wonder if the Ryzen AI CPUs compare with Apple’s UMA. I’m mostly interested in LLM inference for code generation and automation.
Replying to @isVeryLoud@lemmy.ca
Yeah I was lucky to buy a 6900XT when it was near the lowest price, so I sold that and added a couple hundred for the 7900, seemed like a good future-proofing move, given the times we’re living in.
I think Apple silicon is faster than my generation of Ryzen, but the newest (Strix something?) with the LPDDRGGFASEWARGH5 memory should be faster. Of course buying all that memory right now would be quite painful.