Hacker Newsnew | past | comments | ask | show | jobs | submit | sponaugle's commentslogin

The idea that someone would reach out to you and actually expect that you would take something down because students were using it is absurd. It shows a lack of understanding of both how the internet works, and more so the idea that an action like this is really improving the outcome of the students. If the student outcome depends on you taking something OFF the internet, that is the problem.

The pricing differential is just absurd now. I expanded the memory on my math clusters about 2 years ago so I would have 4 machines with 1.5TB of DDR4 memory each. Back then a 128GB DDR4 DIMM could be found for $150 used. Now those same DIMMs are $600-$800 used! I think the RAM in my homelab is worth more than all of the other hardware combined.


The best part of this workflow - which I see often - is that by having someone build custom software to automate some process they often step back away from the process being their job. That eventually translates into them understanding that some (or sometimes most or all) of that process is not needed. There are so many corporate processes that were implemented and then become the way... and if there are people who identify that process as being their job those people resist attempts to optimize that process.

I have seem several people use AI to write apps to automate a process and along they way finally ask the question 'do we even need this process?'.

Regrettably this does not happen everywhere.


"But there’s another challenge: local LLMs. It’s already possible to run LLMs on local hardware, and that’s only going to get easier in the future. Apple’s M-series chips are extremely good at doing this today. Open weight (read: free) models are widely available and good enough that most people probably couldn’t tell the difference. They also have the benefits of running on hardware that’s sipping power most of the time, rather than slurping it down in massive data centres."

This is such an odd and illogical conclusion. If a smaller model can be sufficient (which is not something I would have said), that smaller model can be ran in a datacenter. The idea that a small model running at home is 'sipping' while that same small model in a datacenter is 'slurping' is absurd. The datacenter will have much greater overall efficiency in both power usage and total cost to implement. Of course if you compare a small home model to a DC frontier model the power usage is different, but so is the output.


I’m beginning to challenge the assumption that datacenters are more efficient. I can get the same computing power out of a single Mac Mini 32 GB that I get from from an AWS virtual machine that costs hundreds of dollars per month. Even compared to cheap baremetal providers like Hetzner, the Mac Mini pays for itself in a few months of cloud costs. How exactly are datacenters more efficient? I don’t see it in the price. It may be the costs of centralizing large amounts of compute actually make it more expensive, not less, when accounting for profit margins, and considering the fact that base infrastructure (power, internet) is a given in every home anyway.

There are huge hidden costs in datacenter prices that are simply unnecessary for most casual users of compute. Salaries of staff to maintain datacenters, redundancy and high availability of nine 9s that are simply not required by most customers, as well as real estate costs are all non-existent costs in a homelab setup because those are living costs you pay for anyway, with or without a home server.


>I can get the same computing power out of a single Mac Mini 32 GB that I get from from an AWS virtual machine that costs hundreds of dollars per month.

This quickly breaks down when you're talking about large models that needs terabytes of memory to run[1]. There's no way that you're going to be able to amortize that for a single person.

[1] https://apxml.com/models/glm-51


The comment is about smaller models


Right, but what are you going to do with small models? If your time is worth anything at all you'd pay for the $100 claude code/codex pro subscription, rather than fumbling around with the models quantized enough to fit on your mac.


If you're building agentic processes (harnesses) for business processes local models are a great way to do that, while keeping your data, and any personal data, private.

If you're vibe coding a codex/claude subscription makes more sense as a more polished experience.

I don't vibe code, but I use self hosted models with codex for code review and snippet generation.


If small models keep improving for specific purposes and larger models have diminishing returns, then what?

E.g. I can see a world where you have a local model that is specialised just for producing code.


$100 isn't going to buy you much access to claude code when they start charging a profitable fee for using it.


Author here. The reason I wrote that local hardware is "sipping power most of the time" is because most of the time it's not doing LLM-related work. If you're just using your local machine (or eventually maybe even your phone) to do local LLM tasks, you're not doing that all day.

I agree that data centres will be set up to be more efficient, but we're also going to need fewer of them if local LLMs take off. If that's true, overbuilding data centres is more revenue pressure for AI companies.


Electricity is more expensive at home than where data centers are built, batch inference is more efficient at GPU/TPU inference per watt, power supplies in data centers are more efficient than in average consumer devices, entire racks can be fully powered off when not in use vs. standby power consumption, and of course the investment in hardware is amortized across many users in data centers. It allows more people to have access to larger models than everyone buying an M3 Ultra.

The economy of scale that data centers have is actually a good thing economically and environmentally for many kinds of demand.

I think that the most capable models will continue to be in high demand across the market until at least "a datacenter of PhDs" level of capability. At that point I can see a transition to more local model use if affordable consumer hardware is available (for the median human on Earth). If that turns out to be true then the hyperscaling will plateau at the level allowing sustained commercial/industrial "PhD"-level demand which we aren't at yet (all providers are still struggling to meet current demands).


What I was commenting on was the concept that a small model at home is somehow more efficient. To make a reasonable and fair comparison you would compare many people running a small model at home vs those same people using what would likely be a shared resource in a datacenter.

The core concept is that tokens/watt is tokens/watt ( for a given model of course ). A computer at home is actually less efficient overall because most of the time it is not doing tokens but still using a small footprint of power.

The revenue pressure is an interesting problem , but I suspect the actual demand math will be much more complicated.

I find local models interesting for sure, and run several on my own personal DGX cluster. I am however most certainly not power efficient!


Fully agree with you, smaller model are great for some tasks but the security concern on injection prompts etc is what really makes it for me. Great to run offline tasks etc, but whenever interacting outside the local network I still run Claude or ChatGPT depending on the task


It's technically odd and illogical, but practically probably correct and on-the-money, as the companies try to artificially create demand?


George was really into video stuff - he had stacks of 8mm video tapes in his office, and of course stacks of exabyte drives. He had many different cameras and was always trying new ones out. He was also a really early adopter of laserdiscs, and I have a few discs he gave me when I graduated.


Any chance all of that will be sent to the Internet Archive or Archive Team?


I was holding the camera for some of these videos. Such a great time!


GOOD GOD! you know that 1 oz of LoX + 1 charcoal briquette = 1 stick of dynomite. I am so glad that only grills were hurt, and a few camera lenses as the sky went dark on the video, because the light was SO BRIGHT.

I am glad to have been 100s of miles away.

Thank you for your work, and of corse for the many many laughs.


Sad to hear! I worked for George for all of my undergraduate time at Purdue. He was an amazing boss with such a passion for all things unix. For a while he had the UNIX license plate on his minivan.


I worked for him as well, from 1988 through 1990. He mentored me as I helped sysadmin various BSD machines the university was beta-testing (CCI Tahoe and Gould NP-1), and supervised my work fixing bugs in the Berkeley Pascal compiler. It was fun watching him put his early-model Motorola cell phone into service mode and tweak register values... while he was driving. And of course I enjoyed finding him in his office at all sorts of weird hours and listening to him rant about various technical topics.


That is awesome. The NP-1s were great. I spent lots of time working on en.ecn.purdue.edu - Some tape drivers, some maintenance, and lots of software projects - it was really cool that in those days everyone was on the same machine, working from terminals. Good times!


RIP —ghg

I worked for a George as an undergrad too, between Housel and Longshot. Went to many a lunch at Pizza Hut, Pepe’s, and Burger King (“run it through the broiler twice”), usually riding shotgun in the Pontiac Transport.

I didn’t do much on the NP1s for George. By that time Gould had been folded into Encore and development on the NPL architecture had been halted. While the NP1s were still key resources at ECN, a lot of our attention shifted to the Ardent Titan by then, and trying to shake the bugs out of that machine and associated software, including a port of an early version of Matlab.

George was a true renaissance engineer who set an example for me to follow my curiosity and not worry about sticking to a single, narrow field. As a result I’ve had a wonderful career that has included supercomputing, computer networks, land mobile radio, software defined radios, telecommunications, electric power, and rail transportation. Ironically, I have never been all that great of a programmer, but I feel the year or so I worked for George really opened my eyes to what being an engineer could be. I’ve tried to pass that on to the engineers I’ve mentored over the years.

—zawada


That brings back memories! Riding in the Transport to pizza hut with tanks of refrigerant banging around in the back. I remember when Ardent appeared! It was such a cool machine, and to this day I occasionally scan e-bay to see if one comes up.

Indeed what made George special was that he was a broad 'engineer' first, and a specialist later. Everything was an engineering problem that could be solved. Perhaps most of all he believed in the students, hiring as many as he could get approved.

I found a 'bug' in the debugger on the NP1s that gave me a root escalation, and when he found out he said I should come work for him and that was that! About a year in we scrapped another older system and he had a spare disk that we put in the EE NP1 (en.ecn.purdue.edu) and mounted it as '/hogs' because I was always hogging disk space with stuff I downloaded. Funny enough I still have an exabyte tape backup of that drive.

Great times for sure. RIP.


/hogs reminds me of the earlier (mid Eighties) /nightowl and nightowl group. Davie Curry and others were taking up precious disk space on utilities of dubious necessity. One example was dog, a version of cat that displayed the file to screen using curses to randomly paint the letters in the correct location. Students found them, and began hogging too much cpu cycles using them, so George migrated them to /nightowl, and only members of nightowl group could execute.


I am similar in that all of my interactions are with my real name and it is unique enough that just putting it into google will instantly identify me. There is one other 'jeff sponaugle' but I think he is far more annoyed with my presence than I would be with him.

On the plus side, someone will sometimes say while talking to me - oh your are that Subaru guy, or that youtube guy, or whatever and that is fun connection.


A fantastic read, and really interesting to see the use of Forth. I remember Forth having a bit of popularity in the 80s. This was such an amazing game, especially in that you felt like the world was huge with the encouragement to just explore.

The other game this reminds me of is a game for the TI99/4a called Tunnels of Doom. It was a cartridge game that also had a floppy or cassette data load. It had a dynamic dungeon creation so every time you played the game you got a new unique experience. That would be an equally challenging one to reverse engineer due to the oddity of the GROM/GPL architecture in the TI99/4a.


The study looks at a wide range of different tests spanning many different areas of expertise and output types. Some of the tests, like the web vis tasks used Sonnet not Opus (which was not out at the time). It is similar to testing a car to do many different things, but only one of the tests is the actual driving somehwhere and many of the others are based of the fabric used in the interior. This gives a very broad "96% failure" while missing the observation of the successes. Of course AI can't do everything, and nor can I.

One of the most interesting observations about AI is the timescale at which the favorite model and favorite task changes. Before November I found Sonnet to be interesting, but not moving that much of the needle. Once Opus came out it was clear the needle was not only moving, but moving fast.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: