Replying to an earlier post

Chinese open-source models are becoming the norm. You don’t need the latest model from ChatGPT (e.g., a Ferrari) to complete your task. A simple Honda will suffice (open-source models), and it uses a lot less energy since it runs on your home or company’s infrastructure.

Because of that, I predict that we’re going to see a lot of financial hardship among AI companies.

en

Replying to @⁨mereo@piefed.ca⁩

it uses a lot less energy since it runs on your home or company’s infrastructure.

A CPU-cycle requires energy no matter where it is. The open source models may be more efficient, or you may distribute the power usage over a greater geographical area. But you’re not using less energy because you’re running the model at home.

Because of that, I predict that we’re going to see a lot of financial hardship among AI companies.

I do like the idea that open source will kill tech giants.

Replying to @⁨eestileib@sh.itjust.works⁩

Current marginal cost is certainly negative.

The expectation is that development costs will get cheaper (model design solidifies), Training costs will get cheaper (no need to retrain the whole model) and running costs will get cheaper (datacenter economies of scale).

All this is still true.

OpenAI, anthropic, Google etc. will generate excessive by being able to charge more than cost because their models are vastly superior.

Is likely to be false.

Replying to @⁨Knock_Knock_Lemmy_In@lemmy.world⁩

Datacenter economies of scale do not seem to be panning out, 'cause they’re buying into a RAM cartel who knows they’re desperate.

Coreweave’s reports say that the GPUs are 85% of the cost of a data center, and that they’ll last 6 usable years. This isn’t like building a facility like TSMC, datacenters feel more like COGS than capital to me, and those aren’t businesses I would want anything to do with even if I didn’t think they were all run by conmen who are rapists or racists or both.

Replying to @⁨boonhet@sopuli.xyz⁩

I don’t know how much they’re actually doing to surpass the West tbh. A lot of the success relies on distilling closed models, and running the models cheaper with almost comparable effectiveness. If the Chinese are showing that such profit/investment strategies in closed models are weak and can be easily decimated by a competitor, what motivation would they have to adopt one?

Replying to @⁨Guilvareux@feddit.uk⁩

If their models truly are reaching parity with new western models as many claim, then it can’t be from distillation alone, they must be gathering their own datasets too. It takes months to train a new model.

Also distillation only really saves you the data collection and preparation (categorization). Training is still expensive, as is inference. It’s likely architectural changes that are making their inference cheaper (MoE vs dense models for one), not sure if they’ve gotten any good methods for making training cheaper.

Replying to @⁨boonhet@sopuli.xyz⁩

There is zero proof of distillation. Minimax 2.7 development was surrounded by moderate use of Claude. M3 is their latest generation, and pretty solid, but its performance cannot be attributed solely (or even 5%) to distillation, and that is only lab that has been accused of significant API use. These claims are all 3-4 months old by now, and Anthropic blocked China access after publishing the accusations. Repeated BS is BS from losers trying to lobby for support.