Hacker Newsnew | past | comments | ask | show | jobs | submit | kristofferR's commentslogin

Cloud models will always have massive benefits of scale.

Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.


KV cache is per conversation, I'm getting 100% hit rate on my single tenant local set up.

The benefits of scale are on the token generation side, you can batch rounds and generate tokens for multiple conversations per pass instead of just one token per pass.


That's not true, the journalists are hands on with it right now.

Here's a live stream of the popular obnoxious streamer, he held it a few minutes ago (also talked with MKBHD).

https://www.youtube.com/watch?v=5CPAtEmHAio


Yeah, but it's not funny since it doesn't match reality.

That makes no sense, they should be able to host chatgpt.com regardless and just say "Sorry, we're at capacity" or something instead of hard 404

Distributed systems are hard. For all we know, this could be the error code returned by some DNS service that is down because of load and a proxy cannot find the records, so 404 makes sense, or whatever. Systems like these don't always have a easy point in the infrastructure that can ALWAYS respond correctly no matter what.

> Distributed systems are hard

... when vibe-coded.


Vibe coded or not, building and running distributed systems is hard :P I think only people who never built and/or ran them would say it isn't hard, regardless of what tools you have available. Not saying OpenAI won't find it extra hard due to their vibe coding, still hard when you're a group of knowledgeable people with/without AI.

Yeah who wrote that website? Oh right

Agreed, it makes no sense to send out 404. Unless some agent finally went rogue with absolute permission and did `rm -rf`

At the very least, should be a 429

exactly - this should be trivial.

This ain't just slow/at capacity, this is dead down.

When you're dealing with millions of users and response times go up above timeouts, there isn't much difference between "at capacity" and "down" if most users can't reliably use the service.

exactly, I think it will be an interesting story in retrospective, what happened here

Wooohoo, was in bad need of a reset!

So? There's still a connection to products not provided by them.

Remember that this is a TOS, so you gotta presume that they will act as malicious as possible within the text.


That line of reasoning has no end. If you use Antigravity on anything other than a Google Chromebook or Pixel, the hardware is a 'product not provided by them'. Is that a TOS violation?

Yes. I'd avoid using services like this.

" Your Google Cloud billing account is being processed. Processing time varies from a few moments to a few weeks. "

What kind of bullshit is this? I want to test out Gemini-3.5-Transcribe in my app, but have to wait a completely indeterminate time, have already waited over 24 hours. OpenRouter took seconds to set up


> Creating a useful product that people want -- this is the hard part.

When you do that you'll just a bunch of people clowning on the project due to the politics of the founder.


Just hype, it's Kagi all over again. People said the same thing there, we're just too cynical, not used to people getting excited about really cool stuff anymore.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: