Hacker Newsnew | past | comments | ask | show | jobs | submit | stratos123's commentslogin

Similarly to this, OpenAI either took 3 months to notice that their agents breached an Australian Medicare website back in June, or sat on this information for three months without telling them.

> it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

As I understand it, there was supposed to be traffic; the sandbox allowed GET requests. So perhaps some sophisticated alarm could have noticed it (an anomaly detector? some clever heuristic that looks at domains?) but not a naive one.


> The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.

Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.


It's not even a secret operation. They have a SuperPAC named "Leading the Future" that exists to spread propaganda promoting deregulation of AI development. They've been caught, among other things, making a "news website" with LLMs pretending to be reporters (with human names and everything), which reached out to people asking for interviews and then wrote hit jobs on them.

https://www.modelrepublic.org/articles/reporters-ai-bots-ope...

https://twitter.com/FournesMaxime/status/2047697265280639459...


> Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

You'd have either learned to, or left long ago.


METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.


As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.

Not anyone, no. From an information theory standpoint, that it's possible at all to complete these 700 bits, only implies that anyone logically omniscient could. It's entirely possible that an LLM is capable enough to infer these 700 bits, and the human reader isn't.

Right but where is this information coming from, it can't come from the sender because an idea that can be conveyed in 300 bits cannot contain 1000 bits of information. You're just using the LLM to translate the idea into something more legible.

Alternatively those 700 bits are information the LLM added, but where is that information coming from? Is it noise? Random facts? Random lies? And who is the receiver even talking with if most of what they read is something the sender didn't know?


> you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer.

It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't. As an existing example, consider formulating a conjecture and having an LLM prove it. But I agree; if I wanted to read an LLM's output I'd simply ask it myself rather than read someone's supposedly-human writing.


Lets do a thought exercise:

You are stranded in a desert island. You start writing a message "Help, I am..." and pass at that point.

Somebody finds the message. They can no doubt come up with plausible continuations like "Help, I am Robinson Crusoe" or "Help, I am hungry" but they cannot create information. No matter how smart and how long you stare at the message, that is not going to tell you what the original person would have written.

Isn't it from Claude Shannon that information lowers uncertainty? Infinite regurgitation or massaging of data does not create new information. You will get the information form the LLM, not from that original person.


Once you've determined what the information-content of a message is, then you can apply information theory to it. But different receivers can derive a different amount of information from the same message.

Consider, for example, that if somebody doesn't know English at all, then before receiving the message, their best guess at what it is is some probability distribution over all English characters (or sounds, depending on what we assume them to know), and after knowing the first part is "Help, I am" that distribution might not change much at all. Therefore, they derived very little information from this message.

Going in the opposite direction: keeping fixed the knowledge someone starts with, there is an upper limit to how sure they could be (even if they are logically omniscient) in completing the message (that is, a lower limit on the entropy of their probability distribution) - this is what you're talking about in your example. But this limit only becomes important under these constraints - for example, knowing more about the person who wrote the message can let you predict it better, and if predictor A isn't logically omniscient, predictor B can do better than it with the same prior knowledge, just by being smarter than A.


What do you think about the prompt “The first 6 digits of pi are 3.1415”?

I think the prompt shows that you can't count to 6. :)

Also, "pi" is a unique constant name. You've conveyed far more than 10^[5|6] information in the subject phrase.


>It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't.

Then what's the point of the original writer?


I think you missed the import bit: "transfer of information from your brain to my brain." Including some third-party necessarily adds more that was not in your brain. So by using an LLM, you are transferring what is in your brain + some sludge that may or may not be useful for the other person's brain.

You seem to be implying that achievements of internal models are exaggerated, but that's rather implausible. The public does have access to, for example, Opus and Fable, and so we know what those models are capable of - finding real vulnerabilities in multiple codebases, for example. If you extrapolate from these capabilities one more generation, you'll get pretty much the same feats that the internal models are claimed to be capable of - so why should we doubt those claims? It's not like they're claiming that their internal models developed psychic powers and learned to teleport - the claim is pretty much just "we have models a few months ahead of what we're making available, and in those months they've been improving at the same rate as usual".

You have just word smithed their claims into something palatable. Good job. Why do you feel the need to do that for them?

This is a shallow dismissal.

You're claiming the frontier labs are lying about the capabilities of their next generation models. A reasonable person will expect that statement to be backed up by some verifiable evidence, given that we are several generations of models into this process and the capabilities are consistently increasing, often much more radically than people expected (remember "stochastic parrots"?)

Your post claims that they're lying and then throws out a bunch of fear, uncertainty, and doubt about what they're doing behind the scenes.

If I were to steelman your argument, you're probably saying that the AI had a support system around it of people and training data and feedback that allowed it to achieve the breakthroughs that the labs are claiming. That actually seems perfectly reasonable, but in my mind it does not invalidate the advances they're announcing.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: