The hardware savings at the scale Anthropic or Google use would me immense, it makes me wonder why no big player has done more optimization already. When DeepSeek showed how inneficient were the models of its time, I'd expected each to create a permanent optimization team with all the talent they have hired.
I guess hardware is not that expensive to them in the grand scheme of things, at least not at this stage. OTOH, their propietary models might be thoughly optimized and we can't know, because they're still bound by supply contracts to buy the same amount of hardware nevertheless.
It makes sense when you think AI is seen by the government as a strategic asset. They'll want it to progress unhampered, but also don't want other nations/actors to catch up.
If you mean retraining, it's not even needed anymore. If you want the guardrails off, these days you just install LMStudio and download an abridged model. It's all GUI. The abridged models might have weird behavior in edge cases after the pruning, though.
The comparison is about how many tokens you buy vs how much hardware you could buy with the same money. It's as saying "if you have rib eyes at Applebee's every day, how long until cooking your own rib eyes pays for itself".
If you don't consume many of tokens, it will likely never pay for itself. If you do, though, it will have trade-offs, but you'll probably save money in the end.
> If I am offloading some of my thought processes to a machine
"Offloading thought" sounds a lot better than "outsourcing thought", but the latter is what we're really doing. Offloading implies you thought it first and then gave it to the LLM, but we're only giving it the minimun so it can do most of the work in our place,
I'm worried that I'm starting to find those glitches endearing.
reply