I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS.
I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details).
It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automatically releases the models from RAM after 5 minutes of inactivity so it is pretty seamless to leave running in the background.
All the details are stored in a simple webapp with a SQLite db that I can search through.
> I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS.
Doesn't Apple do this already within it's OS all locally? It certainly does it for OCR and categorization.
EDIT: Also, no reason to use a generic LLM for this. This functionality exists in something like Immich (both OCR and 'context categorization'), and doesn't tie you into the Apple ecosystem either.
I personally use Apple Photos for this. It stores the original, plus makes them nicely searchable, so I have a Hazel action that takes screenshots from the desktop (and from my NAS where mobile devices back them up) and imports them.
Neither openai or xai. Anthropic maybe but not likely. Mistral is the most likely one because they're under the EU laws but I think they focus more on commercial these days.
having a 64GB mac mini m4 pro the last few years with some increasingly capable usefulness has kept me interested in this stuff in a way that using a paid platform wouldn't have. Similar to running K8s in a homelab, something about interacting with the hardware makes it more engaging/interesting, for me at least.
In general, I'm a big believer in doing more with fewer resources, within reason, and think having local setups really helps me be mindful with what's happening under the hood with these systems and managing context efficiently to get high quality results.
I guess it depends on what you're trying to do. I've run a few LLMs on my 64gig Mac but anything image or video related is ridiculously slow compared to even old NVidia on my PC.
Have you been using an MTP setup? I've been having pretty good luck with the Qwen models with the built in MTP heads via https://mtplx.com at around Q4. My main driver rig is an M5 Max MBP work got for me a few weeks ago. Hitting 1000-1200 TPS prefill on Qwen 3.8-Flash-Next there.
A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner.
However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
I have not spun up my own homelab, but, I have researched it and the electricity costs can be mitigated to a large extent.
One... the GPUs can be massively clocked down during idle state, to the point where the fans can be shut off as well. The machines themselves can be shut down and wait for a magic wake-on-LAN packet if needed.
Two... even under load, the GPU cores can be significantly underclocked with very little performance loss. The GPU and VRAM/HBM clocks are independent, and the GPU is largely bottlenecked on VRAM/HBM so you can just drop the GPU speed. A commonly reported figure I saw was, basically, 350W nominal cards being underclocked to consume "only" 200W under load with ~5-10% perf loss.
This all assumes you're using discrete GPUs and not AIO systems like a Mac Studio which is going to be pretty efficient just by design; they idle at 35W or so and in practice max out at a few hundred watts. I believe DGX Spark and Strix Halo are similar.
I must stress that this is second-hand anecdata here, admittedly, but my understanding is that it can be pretty manageable.
> There is no reason to believe that equivalent level model output will be more expensive in 12 months
It's almost never a drop-in replacement, and having to check and adjust integrations and workflows with new models gets old fast. My task was perfectly solved by the old model, I don't need a newer, "better" one - especially at higher prices ("more cost-effective" my foot). Local models lets one choose a model and freeze the downstream integrations forever, without being forced on the 6/8-month upgrade treadmill by aggressively short, scarcity-driven hosted model-deprecation schedules.
The big providers are losing on average tens of billions a year on these services, so yes prices must go up. Even Moore’s won’t help in the medium-term due to shortages and difficulty/reluctance to vastly increase capacity.
I thought the discussion was about "frontier labs" in the cloud versus open at home.
Others running open models in the cloud is a nice third alternative, but not a solution for those that need frontier models, which I believe continue to need training as well as inference. These are currently heavily subsidized and hardware constrained several years out.
They said "equivalent level model". It's reasonable to assume that open source models in 12 months will be just as good as frontier models today.
And this sub-discussion is more about whether there will continue to be some cloud model that's both more capable and cheaper than local, not so much about what kind of cloud model it is.
I would think it's the opposite: there's increasing chatter about how long the AI labs will subsidize cheap subscriptions. I suspect that at some point one of them will bump up prices, and the rest will follow. That'll signal the end of cheap subscriptions and prices will steadily creep up.
I'm not saying that local AI is necessarily economically justified right now; but it's certainly quite reasonable to think that in 4 years the subscriptions will be significantly more expensive.
Surely you cannot possibly believe there is no reason.
Don't get me wrong: I hope you are right, and I am generally optimistic about the future of AI.
But do you really think, in a world filled with examples of big software companies repeatedly taking away or hamstringing capabilities we've taken for granted, that you can just count on a big tech company hosting cheap inference on incredibly powerful models forever? Surely we have learned by now that these companies do not exist to provide a public service to us, and the government cannot always be counted on to have the best interests of the citizens in mind.
I mean, how many times have we seen this in just the past decade or two?
- Consistent attempts to pass legislation weakening or banning the use of encryption
- Exorbitant Reddit API pricing (still salty about the death of the amazing Apollo app)
- Google fighting against sideloading on android
- US gov't issuing export control directive to suspend access to Fable/Mythos
- US lawmakers considering ways to regulate adoption of open weight models
- Chinese officials considering restricting overseas access to their most advanced models
I can absolutely see a much more restricted, closed down, and expensive future due to a combination of government regulations (regardless of which nation is doing it) and big companies rug-pulling as the check comes due on all the billions of dollars spent to get here.
This is such a tired argument and it seems to be parroted every single time someone talks about local models on hacker news.
Yes, of course the most economical path is to hand over all your data and become fully dependent on a cloud provider who is already operating as scale, hoping that they won't change/remove models, hamstring capabilities, or raise prices.
If this were a thread about hosting your own email or blog or cloud photos, you'd have plenty of people out here telling you how easy it is to do it yourself instead of relying on Gmail for email or WordPress/Medium/Substack for blogging, or iCloud for cloud photos.
And yet, without fail, every single thread about self hosting local models seems to have some copy/paste form of this cost-savings argument.
Where is the appreciation for this cool thing GP built? Where is the appreciation for the desire to figure out how to host your own version of the incredible capabilities that were not available merely a few years ago? And why, on this site of all places, would someone advocate trading all of the knowledge and independence gained from learning how to host something like this ourselves in favor of throwing it all over the wall to Google?
It's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!".
> It's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!".
Because there are more privacy guarantees there, depending on the provider. "But what if they violate their contract!" is some pretty tin-foil hat stuff.
How is this any different than a business running their website out of the cloud, assuming you are using a provider with appropriate contractual terms?
You can care about tracking and ads but still be comfortable storing your backups in the cloud, and many have been for quite awhile, even sometimes without encryption - that is totally different than e.g. Meta actively trying to track you and understand your relationship graph and your purchases etc.
I'm not sure about the tin-foil-hattedness of worrying about them violating their contract. But that's by-the-by. It is definitely not tin-foil-hat to worry about the data being taken in a breach.
Depends on how you use it. You can use AWS in a way that protects your data even if Amazon is compromised. You can even run LLM inference that way, but it doesn't seem at all common.
I would say tech workers have very weak class consciousness and understanding of power structures. Those who care about not ceeding power is a tiny minority compared to say among MDs and lawyers and their guilds/unions.
I don’t understand the willingness to give up privacy so easily, particularly if you are developing something that you plan to monetize somewhere down the road.
I’m pretty sure that all of those disclaimers that all the AI model makers have for you to sign off on to say that they’re not responsible for anything that might go wrong if your work gets copied accidentally and used someplace else wink wink?
You know the lawsuits for that particular aspect are incoming in the future…
Yeah I use it for hobby only. I would use cloud models to develop new stuff, I don't really care if it gets lost because if I were to publish it it would be FOSS anyway. I hate entrepreneurism and monetisation so I'm happiest being a salaried employee with many hobbies :)
But no way whatsoever I'm uploading my personal files, photos, emails, chats into cloud AI. No way.
I think the price is beyond that though. AFAICT, there's not competitively fast image or video generation on Mac. To buy it will cost me $4000-$12000. So I rent.
Yes but I also get a full fledged computer in the deal. I can sell it later. I can use it for all sorts of things like games and browsing and video editing. Paying for Gemini for other tasks is also in the mix but at the end of 4 years I get...nothing.
> They are things that I would not be comfortable sending a cloud provider
It's also an old machine that the commenter already has; it's intellectually dishonest to compare it to the price of a brand new, 4-iteration-newer machine.
Yup and in my opinion it is a great use case for it. Faster to generate my own extension than to trust something already in the store. Not critical enough to matter if it failed and very low threat surface.
This could potentially have very interesting consequences for ML models. It will depend if their training data is considered part of the source code.
If there was a loss prevention specific video analytic that flagged a person’s behavior as abnormal, would the person have a right to audit the source code for the CNN and/or the training data that was used in the development of that analytic?
As someone working on such analytics it could become a real adventure to comply with that. My dataset came from customers that agreed to shared with positive/negative examples with me but not necessarily for me to share publicly. The privacy of the people in the shared examples would also need to be considered.
Wouldn't that just mean you'd have to make sure the data doesn't contain private identifying information to start with? Seems like a win-win for the people in datasets and the people who are auditing the code.
"Wouldn't that just mean you'd have to make sure the data doesn't contain private identifying information to start with"
That condition is not easy.
It is very hard, to have data about people related stuff, without private identifying information - especially because now there is face recognition and co.
If the person does not have the right to audit the model - i.e. determine how and why exactly it flagged them - with consequences wrt government interaction with that purpose, I would argue that it's a violation of due process. If it's impossible to meet that standard with ML in a satisfactory way, then perhaps ML should not be used in those contexts at all?
Backblaze is great for cloud backups. They do not backup network connected drives but will backup drives connected locally.
I'm not pretending a network drive is local, but actually mirror the important data from my NAS to a locally connected 14TB USB drive. It stays connected all the time.
I have some cron jobs that run rsync scripts, but the data that needs to be backed up rarely changes. This gives me 30TB on my NAS, of which, 14TB are backed up in backblaze for $60/year.
I can make this work in my situation because the items I want backed up are less than the working space I want on my NAS.
My experience with backblaze has been dark patterns and and customer hostile practices. They get recommended in every thread about backups, and I started using them because of these recommendations. I’m not a customer any more and I wouldn’t use them again.
Rather than accusing them of "dark patterns and customer hostile practices", list specifically what they did. I've had good experiences with their B2 service.
The best thing I did was get a hobby that got me out of the house. For me it is drone racing, but it could be jogging, swimming, or even go fly a kite.
Once I get back into the house, I've burned off some energy and can focus on my side project. For whatever reason if I sit in front of the computer all weekend I have a hard time getting started.
HUVRdata is looking for an experienced python/full stack developer to join our team. HUVR is expanding its data and analytics platform for drone based inspections. Our customer use drones to get a new view of the world. We provide them the tools and processing to get a new view of their data.
The ideal person has several successful projects that they've worked on before and has the right attitude to solve problems they've never seen before.
Our development team works with a very high level of autonomy. We trust each other to make the right decisions and help each other when needed.
We use Python and Google App Engine. If you haven't ever touched Google App Engine, don't worry, it's just a WSGI web server.
We use Node.js, NoSQL, mapreduce, and imagemagick, ffmpeg for image processing. Familiarity with these would be a plus.
If you are excited about what we are doing, send a resume and an introduction with an example of your work. Make sure it is something you're proud of (github, blog, Olympic medal, etc).
Insomnia seems to be stuck in an "Install update" loop. Each time it opens I get prompted to update and restart. I update and restart and then get updated again. The cycle continues.
I'm uninstalling and re-downloading it but just wanted to let you know.
As an Austin resident I think the real lesson to learn is how poor of a job Uber did of gaining supporters. I went from being a vocal supporter to an almost hostile opponent based on their behavior.
They spent a reported $8 million in very confusing ads. They sent text messages and push notifications to customers. All of this felt very disingenuous and deceptive. I will be sad to see them go but they are their own worst enemy right now.
I agree that many Austinites were turned off by the sheer amount of spam sent by the Uber/Lyft-backed campaign. Others, I believe, were simply angry about the cost of a single-issue special election. [1]
Yet others, I believe, were furious that those two generously funded companies were attempting to insert themselves into local politics and override a regulatory scheme enacted by the local elected government. And if Uber and Lyft could buy this election, what next? It's been said that Uber and Lyft needed to make an example out of Austin, but the reverse is true, too: Austin needed to make an example out of Uber and Lyft.
It was a very bad campaign, which makes me suspect that they did it on purpose. If they won this campaign, they'd have to win again in every other city that tries this. If they lose, pull out of Austin for a while, and make sure it's national news, other cities might think twice.
They basically turned it into "here's the legislation we wrote, if you don't pass it, we're leaving immediately", and then spammed everyone with text messages, INCREDIBLE amounts of junk mail, and door-to-door canvasing. They turned a lot of people off, big time. (Where's the story about that?)
I'm really hopeful some competitor can show up and eat their lunch, before they come slinking back. The way they behaved turned me from a fan into someone who actively distrusts them.
I'd be happy to go through some of your options even if you don't use us. We're a startup that does cloud based security video hosting. https://eagleeyenetworks.com
Let me know if you want to talk.
mcotton <at> eagleeyenetworks.com
I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS.
I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details).
It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automatically releases the models from RAM after 5 minutes of inactivity so it is pretty seamless to leave running in the background.
All the details are stored in a simple webapp with a SQLite db that I can search through.