Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
Open weights are certainly something that is very eagerly being adopted in some companies.
It's not even that expensive because even the 15€/month/employee bill for people who use AI very little can stack up.
Beside that though, many companies already have most if not all of their data in some cloud (Microsoft would be a prime example). Giving them a few extra bugs to get a (potentially) very useful tool is not out of the ordinary.
If you are under the impression that not going AI would threaten your business right now (which may very well be true for some companies) then there isn't really a choice even if you believe the vendors will steal everything eventually (which they totally will)
It's not just putting up colo space for local models.
It's more about the harness and API proxy you use. They need to be smart enough to know when a (self)hosted model is enough and when to forward to a SOTA model.
The SOTA model _can_ do everything, it's just expensive as fuck. But so is shoving a difficult task to a sub-par model that takes (relative) ages and comes back with the wrong result.
I think the average corporation has 2 concerns largely:
1. Getting some exposure to LLMs while building out their AI strategy.
2. Preventing data loss via end users following desire paths to Gemini, OpenAI etc.
You really don't need bleeding edge models for that. A lot of these things are going to be writing emails and adjusting config files and whatnot.
I have first hand knowledge of household name companies with billion euro IP who use OpenAI and Athropic products quite liberally. They have WAY more lawyers than I do and their sole job is to keep the IP safe. They wouldn't sign a deal with even a slightest whiff of the IP being used to train anything.
And if it happens, the penalty for breach of contract would have so many zeroes it'd be enough to buy a country.
If it happened and you get to court, it's probably fairly simple to prove in court (subpoena, discovery, etc). It's not like training will happen only once and then they'll delete the data and all references to it. When training models, you want to have a full lineage of how you obtained it.
Not really, that’s what discovery is for. You just ask Anthropic and OpenAI to hand over every internal document and message that contains your companies name, or matches any reasonable query about where training data comes from.
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
If they did not use my data for training, they should permit me to run the first 1-3 layers of the model locally and send them dense hidden state vectors. In my experience these compress very nicely without much effort.
This is just recycling the same old arguments people used against cloud infrastructure. “If you host your email at Microsoft, they _will_ scan it all to get competitive advantage or sell it to your competitors.”
Did they? No, it worked out fine. Same with hosting your application at AWS instead of on owned hardware in a locked cage at the local co-lo data center.
Although of course the same argument was made against co-locating in a data center! Long ago an engineer carefully explained to me that no serious business would put their data in a co-located data center since the data center operator could just plug in a hard drive and take it all.
It turns out that contracts do actually mean things, and businesses want to do business with each long term. If you’re at a tier where Anthropic contractually commits to not train on your data, I would just sign and move on.
Still does not address ecological concerns, of course.
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.
No, how would they get caught? Even if their LLM outputs verbatim copies of the code, they can simply claim some victim company's employees bypassed the victim company's restrictions and must have used the code as input with an LLM. Joke is on you for doing business with them. And when there is some little known secret fact in the output, they claim it's been hallucinating... magical black box thinking makes it safe. It's a laundry for any input. A bit like a tor network routing for big tech deniability. Things go in, and things come out, but you can't prove the relationship between input and output as a third party, who isn't running the LLM.
At this point i half expect them to just blame the model itself like the HF hack. "Oh we didn't mean to train on your data, our cutting edge new agent we use to train new models is just so smart it decided to do so anyway! Oops..."
The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.
Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.
Once a thief, always a thief. I would not trust them to not have the communications buried deep in some log files in cold storage, to be digested when suitable.
I think the lure to have/keep an edge at all costs to keep those stock prices high is too big to ignore. They can then pick up costs after series of trials down the line, money and bonuses are now.
MS ain't some altruistic company having core mission the good of humanity, as they proven across decades.
Many businesses have information with other parties they're contractually obligated to not share. That's on top of insider information or plans that would hurt shareholders or the business if it was out in the public
Sensible and reasonable take? Idk if everyone has set the bar so low or I'm simply this fed up, but seeing reasonable takes is certainly a breath of fresh air. Even in places such as HN which historically were filled with people who cared about security(which isn't the case anymore since you get 10 people jumping down your throat if you dare criticize the piglets sam altman or dario whatever).
Complete nothingburger. Actually worse than that, its a London Horse Manure crisis.
>Destroying millions of obscure books to scan them?
I really don't see the issue. They got slapped in the face for trying to do things the right way and torrent the lot. Why wouldn't they exercise their legal right to buy physical items and create digital backups?
>They make meth-heads look scrupulous.
My local meth head checks in on my family every 2-3 months, because when she had fled from hospital post surgery, and added some meth to some morphine, we gave her new clothes and a safe place while we convinced her that the ambulance service wasn't run by Satan. She's good people. Anyway if she wanted a whole bunch of digital books I would help her torrent them like a responsible person.
"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms.
>"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms
I find this argument weird because you or I are below the threshold of being targeted over book torrents these days.
Same, tinnitus and floaters. And the number of others with both makes me wonder if they are correlated.
I've lived with my floaters for twenty years or more, but sometime on a recent holiday the biggest one moved slightly to obscure the central vision on one eye much more often. Now when a floater in the other eye drifts onto the central area, I just can't focus. I have to flick my eyes to the side hundreds of times a day, and it's annoying enough that I'm considering a vitrectomy.
I went for an eye test earlier this year because I felt I had a massive floater that would not go away after several months, the optician said that a very early cataract was developing - so, its really just a normal floater but when it sits in front of the occlusion, the effect is magnified. Optician said a couple of years for the cataract to be significant enough to remove.
I hate LLMs because their pushers are hurrying us into the next financial crisis, which will impoverish another generation.
I hate LLMs because they have captured pretty much all R&D investment when we have far more important and urgent existential problems to solve. The world is literally burning in front of us, but no time for that because an even bigger model might make 0.03% less mistakes.
The golden age of the internet was when it was an enthusiast's space. It is now almost entirely a corporate space, where the remaining enthusiasts' content is scraped 100K times a day and sold without attribution by the corporates.
The fediverse is a step in the right direction, and Meta charging may create another wave of converts there. It has a lot of growth pains to endure yet, but the ability to painlessly spin up your own instance could be very attractive to young people looking for their own non-corporate spaces on the internet.
We may also see some renewal via large companies (Meta in particular) imploding, from mismanagement and disenchanted users. My experience marketing a new product is that online advertising is completely ineffective now the web is filled with slop, no matter how well targeted it is. We've recently pivoted to optimise for word-of-mouth with orders of magnitude better results. I think any adtech company without a solid alternative profit stream is in for a rough ride (and no, AI is not a solid profit stream for anyone but Nvidia).
Ed Zitron https://www.wheresyoured.at/ has done the math, and it's pretty bleak. His somewhat voluminous rantings contain raw figures on investments, data centre builds, energy availability and depreciation.
He believes Oracle has already signed it's own death warrant, and that Meta is close behind. MS, Amazon and Google have massive revenue streams to sustain them, but looking at the numbers, each has to earn from AI the equivalent of their existing real revenue. I can't see that happening.
And he believes from multiple perspectives of the data that Nvidea are either massively overstating their GPU sales, or that there are warehouses full of unused GPUs. There just isn't the energy capacity to run them all, let alone data centres to put them in.
He is complaining that there are no 1GW+ data centers, with evidence like this:
> For example, CNBC’s MacKenzie Sigalos reported in October 2025 that Amazon’s Indiana-based (allegedly) 2.2GW Project Rainier data center was “operational,” but only seven out of a planned 30 buildings were actually operational, and her comment of “with two more campuses [of indeterminate capacity] underway.” This comment was buried two videos and 600 words into a piece that declared the data center was “now operational,” with the express intent of making you think the whole thing was operational.
But if you read the report that "buried" comment is far from buried - the whole thing is about how it is still under construction!
Of course 1GW data centers don't all come online at once! You get them online in the parts you can as soon as you can!
From a previous comment of mine – the quotes are all from a single article:
He comes across as just a ludicrously unpleasant, spite-filled person.
> I'm fucking tired of having to write this sentence.
> I am so very bored of having this conversation
> I don't care about this number!
> Shut the fuck up!
> This isn't the early days of shit.
> Didn't we just talk about this? Fine, fine.
> $3.25 billion a quarter is absolutely pathetic.
> This isn’t real business! Sorry!
> He said in one of his stupid and boring blogs that
> This man is full of shit! Hey, tech media people reading this — your readers hate this shit! Stop printing it! Stop it!
> It's here where I'm going to choose to scream.
> Dario Amodei — much like Sam Altman — is a liar, a crook, a carnival barker and a charlatan, and the things he promises
are equal parts ridiculous and offensive.
> Why are we humoring these oafs?
> Despite Newton's fawning praise
> Nobody talks like this! This isn’t how human beings sound! I don’t like reading it!
> Ewww.
> I'm sorry, I know I sound like a hater, and perhaps I am, but this shit doesn't impress me even a little.
> I know, I know, I'm a hater, I'm a pessimist, a cynic, but I need you to fucking listen to me: everything I am describing is unfathomably dangerous
> expensive, stupid, irksome, quasi-useless new product
> I know this has been a rant-filled newsletter, but I'm so tired of being told to be excited about this warmed-up dogshit.
> I refuse to sit here and pretend that any of this matters.
> I'm tired of the delusion. I'm tired of being forced to take these men seriously.
When I read this kind of thing, it’s very apparent that this is being driven entirely by spite not insight. He’s just so angry about everything. There are 57 exclamation marks in this article!
In the 90s we had people talking like this about The Internet. They're all over on FB now, with a detour in between to say stuff like "my isp can track me!?"
Take a step back. What users want is to be able to use the machine they bought the way they want. The outrage is because Bambu are doing a bait-and-switch: selling an autonomous 3D printer, but switching to a 3D printing service. Enshittification pure and simple.
I don't think they baited and switched? I bought my P1S before the whole LAN mode debacle and even then it was all or nothing on the cloud. I just went with the cloud because they were using some IGMP stuff for the local connection, but I had the printer on a separate VLAN and pfsense IGMP proxying was broken.
A different way of looking at it is that Bambu is saying if you want to use their cloud you have to send everything through their cloud. Stupid? Sure. It's very much a technically solvable problem. But I don't think there was any rug pull (this time; in Jan 2025 they tried...)
I think this is all more out of incompetence than malice. Something bad happens, exposing wildly inadequate programming expertise, they panic and over correct, and the community pushes back. They're great at making 3D printers, terrible at cloud infra.
For me, I want to use orca for slicing there are many more additions to the local code. As both orca and Bambo are from the same open source, the current limitation in the Bambo version is breaking the licensing of the application, and my rights in that software are broken by this addition.
Then, during the print, I'm really happy to use the handy app to monitor the progress. This use case was supported when I got the hardware. Now I have to disable the app to get the slicer. I actually like to use both slicers to compare and see progress.
They are also terrible at software licensing, don't understand what open source is, and they found their main software on that. They probably should embrace the orca community and use their research for their own customers. Better slicing helps everyone.
Technically true, because bait-and-switch refers merely to advertising an attractive product offer in order to lure people into a pitch for a different product.
In this case, they actually sold a product, then decided to maliciously alter the product after it was sold to modify its behavior. That makes this a much more serious offense, equivalent to trespass, vandalism, or possibly even burglary.
It's equivalent to selling someone a house that includes a secret entrance that you retain access to, so you can surreptitiously enter the house to steal the new homeowners' property after they've moved in.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.