Hacker Newsnew | past | comments | ask | show | jobs | submit | AgentMasterRace's commentslogin

his 128gb Ram laptop is quite extreme


It should just about be usable in 32GB.


On a consumer hardware it would be nicer. With no GPU/iGPU or a 6-8GB VRAM.


It would be somewhat slow on a CPU only machine, but it still works.

Besides, Macbooks with 32GB RAM is consumer hardware, just maybe on the higher end.


Well maybe, but a non-macos laptop is a bit more common.


It is. I am running it on R9700


RAM is never the issue, it's always the compute power


It's absolutely not for these models. There are plenty of consumer GPUs out there with 8 or 12GB VRAM - they are comparatively very fast at inference but just aren't big enough to run lots of the models you want. Also context management is a massive pain.


I run qwen3.5-9B on an RTX 3080 with 10GB of vram. It runs at ~77tk/s with around 50k context size.

As soon as I switch to a model that doesn't fully fit into vram it tanks to <10tk/s which makes it unusable for me for most tasks.


RAM bandwidth is the main issue for running LLMs on consumer hardware...


Quite the opposite, RAM is always the issue. More specifically, high bandwidth RAM.


RAM is not “never” the issue. My iPhone and MacBook Air could both run larger and more capable models if they had more RAM.


what??? not true!

for inference the compute is the last thing we need more of.

memory bandwidth is the numebr one blocker, after that the inefficiencies that where introduced with MoE models (and all new large models are made that way)

Here is a quick read: https://news.ycombinator.com/item?id=49324600


and memory bandwidth


Give it 6 months, the capabilities will increase even further.


hell yeha bro, I still rock qwen 0.1


or use herdr, or the many other options that don't force you to use Claude code.


Yeah, herdr is great too. Will be adding other agents to this soon. Just started with Claude Code mainly because it's the one I use atm!


if you know what you're doing your results will actually be good. what a dumb take.


because they're stealing from the frontier models. they're gaming the benchmarks. look how bad glm 5.2 is on cursors evals. gmhit garbage , but it gets glazed as God tier.


They're stealing, eh?


so the answer is use grok ?


I'd say answer , the opus is no longer undisputed. grok + gpt models are very competitive + glm if you are ok to wait 3-4 times longer, unless you have some unique access to GPU


I never have have the issues most people talk about ... I feel like most were never Devs before ai and don't know what they actually need done when prompting. that on top of not utilizing good tools such as a codebase indexer, lsp and a project scaffold.


did you use Claude design, their tool meant for Web design? because if not then you're the problem .


I did. It was still Claude that’s the problem


If plenty of other people are having success with the same tool, perhaps it's not the tool.


Plenty of other people aren’t having success though. Maybe it’s the tool.


Perhaps you could explain why they are the problem.


he's writing novels


better, use oh my pi.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: