Surprisingly, there's no category name yet for these things. "AI Agents" -- tons of things are agents that aren't in the category of OpenClaw, Hermes, Grok Bot, etc. "AI Bots" -- everything's a bot. Can we think of a good category name for them?
I was using Claude yesterday and the "advice" it was giving me quickly became confused and irrelevant to the prompt, even though there was not much text in the context window. Speculation: since Claude re-reads the entire chat at every turn, the "minimal" text revisions required by watermarking quickly compound such that even Claude can't follow the discussion.
It’s true in the limit, but I think we are nowhere near that limit. A friend recently started a bio startup and automated parts of mouse experiments enabling higher throughput on in vivo experimentation. This is not commonplace, and there is a lot of room for further automation here.
As anyone who works with agents daily can attest, 1) you can use agents to help with hypothesis refinement, bridging into areas adjacent to your expertise, etc. 2) once you have a rigorous /goal definition you can parallelize and let the agent crank.
It seems pretty obvious to me that with the right actuators and sensors you can apply this to real physical research loops too. (To be clear, this is not easy; a lot of bench work is Métis and needs experts in the loop at every stage.)
To your point, you can’t make plants grow faster but you can increase research throughput by enabling a researcher to have 10x or 100x as many experiments going at once.
This is a valid take but at this moment, AI is not trying to solve this fundamental dynamic. But it can still accelerate the process by aggressive exploration of the solution space which cannot be done even with an army of human researchers. Many ideas can be relatively quickly verified (and discarded if needed) by proper simulation even before real world experimentation, but we don't have enough capacity to process all potential ideas. If you can build a good model for simulation and establish a robust methodologies, we can use some ideas which never had a chance before.
> ...banning the use of these models by US businesses does nothing to address this risk, because bad actors are unlikely to be legitimate US businesses. It would protect US AI companies from competition, but that has never been my goal.
But on the other, he argues:
> All sufficiently capable models, open and closed, should go through mandatory safety testing.
Models that don't pass safety testing would be banned. Darius does not appear to be against banning models. He wants the government to have a regulatory body that has the ability to ban models. Then Anthropic can do regulatory capture of that agency and control what models are permitted to be released.
Also, during this mandatory safety testing, models would be blocked from use, and by the time the testing was done (probably years) the models would be obsolete.
"everyone knows the security threat is a pretext." On what planet? Anthropic itself made a big stink about Mythos being able to hack every app out there, and very dangerous as a result. Many reports have confirmed this.
For those too lazy to watch someone talk on video for ages to make a point:
The link is to a famous YouTuber called PewDiePie and he uses a local LLM to parse his email, to save time with that. They have an autoreply system and get notified about urgent matters.
The article says "A senior C.I.A. official was arrested last week after investigators found hundreds of gold bars worth over $40 million stashed in his Virginia residence, a small fortune that he apparently brought home from work, according to court papers."
The same article also says:
"When the C.I.A. conducted a review of where the gold and currency were stashed, the agency was “unable to locate the gold bars or significant amounts of the foreign currency,” according to court papers."
The NY Times fact checkers don't seem to have seen this article.
These LLMs are great for now, but they have to go by their training materials. And if people stop creating new ways to code, new languages, new coding patterns, etc, then the code LLMs produce will be stuck in 2026 forever.