I'm having a tough time using Opus to produce blog-grade writing. Claude produces jargon heavy, market-y language that no human would ever write, much less want to read. This issue exists even when I ask it to write internal reports or short memos. Opus is not working for writing in any task.
PS: I am not interest in the model's coding skills, only its ability to write or partner with me in writing.
Someone who until yesterday did not seem bothered by his technology being possibly used to bomb elementary girls school in another country seems to suddenly care about the repression of citizens in yet another country.
No, we don't buy your virtue signaling. And we certainly don't need your better-than-thou opinions on this year's "nightmare scenarios".
Palantir's Maven system integrates LLM models in decision making for identifying targets. It was using Claude when it identified targets in Iran, including the Minab elementary school
> Two sources confirmed to NBC News that Palantir’s AI systems, which draw in part on large language model technology, were used to identify targets. (Palantir’s CEO, Alex Karp, said he “can’t go into specifics” when asked about this on CNBC, but said that Claude was still integrated into Palantir’s systems used in the Iran war.)
This had nothing much to do with AI in particular though. The issue was that there had previously been a military installation at the same location, but the US military did not update its database. Apparently, one analyst had flagged the issue for review in 2019, but that was never done before the strike.
The point is that the same thing could have happened just as easily if you had selected targets based on the faulty data by any other means than an AI model.
No, the reports specifically mention that the AI target selection pipeline drastically increased the workload and expectations of the humans in the loop, leading to spending far less time on deciding whether to attack, iirc a few seconds per targeting decision.
Of course some overworked analyst is going to start neglecting the details and erring on the side of bomb it if the AI summary is dangerous sounding.
It’s gross negligence to be offloading this type of analysis to claude. Do you decide what’s factual based on the google AI summary? They’re deciding to end thousands of lives with about the same amount of rigor, including those schoolchildren. Have some humanity.
That's no excuse. The US has 161 public schools on military bases, more than any other country.
The Minab elementary school was painted in pastels and had prominent murals on the play area.
If anything, your comment just highlights how much this has to do with AI. When "outdated data" can lead to 168 dead schoolchildren, you can really see the consequences of relying on AI for decision making.
You act like we intentionally targeted a school. We didn't. We targeted a military facility that was in the exact location the school was in. Check the NYT analysis of it
There was a _school_ in thee exact location the school was in. Besides, when the expectation was to leave most of target selection to automated systems, one could arguably say that there was intentional targeting, due to the lack of human oversight.
"But the AI did it!" -- never did, never will, it cannot have agency, AGI is a lie, etc etc.
once upon a time you would have complained about cannons being used instead of good ol' arrows and swords. Get with the times. AI in war is inevitable and it never going to go away
Anthropic's ToS restricts the DoW from using their models for 1. Fully autonomous weapons systems, and 2. Mass domestic surveillance. There is no carve-out regarding who the autonomous weapons systems would be used against.
> Two sources confirmed to NBC News that Palantir’s AI systems, which draw in part on large language model technology, were used to identify targets. (Palantir’s CEO, Alex Karp, said he “can’t go into specifics” when asked about this on CNBC, but said that Claude was still integrated into Palantir’s systems used in the Iran war.) Brad Cooper, head of the US Central Command, has boasted that the military is using AI in Iran to “sift through vast amounts of data in seconds” in order to “make smarter decisions faster than the enemy can react”.
Iran had the courtyard painted in bright pastel pink/blue colors with murals and playground markings to clearly identify it as an elementary school. The Pentagon claimed they had "outdated intelligence data"
The Iranian strikes didn’t involve autonomous weapons systems. The “AI” component was used for target selection.
The actual aiming and firing of the weapon was performed by humans. In theory the target selection was also vetted by humans, but humans relying on exactly the same data that resulted in the “AI” systems misidentification of the target.
I put “AI” in scare quotes there, because these systems are really data processing pipelines, rather than LLM style AI systems.
I think the fully autonomous weapons systems restriction is not a very serious one and is very easy to circumvent.
What if you can hire a human to push a single approve button? And what if that human's job is to front load approvals by pressing the approve button a few thousand times, every morning?
Yeah, but then responsibility for the consequences of action of the system falls onto person pressing the button. I hope that no one would be willing to take such responsibility without second thought.
What part of placing a human in the loop requires telling that human what they’re approving, or even that they are approving something? Pressing a button could be as simple as following orders—or giving orders.
What does "Fully autonomous" mean? If someone presses a single button to authorize an autonomous drone to enter a theatre and conduct autonomous strikes for 4 hours without any further approval is it truely "fully autonomous?"
Yes. A munition that hangs around waiting for targets and decides whether it should attack without checking with a human. That’s a fully autonomous weapon.
Can you provide your definition of an autonomous weapon? It feels at the moment you are taking the absolute extreme position of only “autonomous” if a human was never involved. By that definition a human deciding to build an autonomous weapon automatically makes the weapon not autonomous…
Do you understand that real politics sometimes means you need to doublespeak? You don't get the outcomes you want by speaking your direct intentions. Don't be a fool to think that what someone says is what they believe while playing these high stakes games with actors of differing motives and value systems (the current white house)
So you are saying its fine if Claude is used to bomb school children, and murder people in Venezuela, and spy on the entire world (apart from US citizens), as long as they get to kiss up to the current administration? They could have cut all ties for real, and actually had a back bone, instead they bend over backwards for whatever the current wanabee god king wants. Im sure it has nothing to do with their valuation and trying to become another trillion dollar company.
> So you are saying its fine if Claude is used to bomb school children, and murder people in Venezuela, and spy on the entire world (apart from US citizens), as long as they get to kiss up to the current administration?
It's absolutely not "fine". I'm saying maybe there's a complex region-beta paradox where we can't get to the other side of it unless some actors enter undesirable terrain (by their own measure), don't exit the stage, and play moves they'd prefer not to.
He was directly asked in a Bloomberg interview whether Claude was used in the bombing of the girl’s school in Iran, and his answer was “We don’t know”.
His redline was autonomous weapons, not the death of 100 innocent girls.
No, he refused the DoW demand to use Anthropic models without limits. But Anthropic still agreed to military usage of their models, including for strike planning.
No, only for domestic surveillance and autonomous weapons. It can be used to evaluate targets as long as a human pulls the trigger. But the DoD created a really strange position for itself, they use Claude, but also designed it as a supply chain threat, which should mean they cannot use it
“We are also left with no further information about the potential use of artificial intelligence in the targeting decision. The CEO of Anthropic admitted in an interview with Bloomberg that he does not know if or how their artificial intelligence software may have been used to select targets for U. S. attacks in Iran, but he was quick to shift blame to the military, in particular to an unnamed ‘human’ making final targeting decisions.
Yeah Dario is just flat out disgusting in term of how shamelessly hypocritical he is.
Do people actually believe that he gives a shit about the well being of the Chinese people? If the U.S. starts a war with China start bombing Chinese cities Dario would absolutely jump onboard supporting it. He'd probably make Claude to add DeepSeek and Moonshot HQ to the targeting list lmao.
He is super pro-Israel as well, and never once has he brought up the risk of the Israeli government using AI to control and repress people in other countries.
He is also 100% onboard with working with Palantir, who has the explicit goal of using AI for population control and repression and building out a surveillance state.
Meanwhile the world's most repressive government is North Korea, and obviously they don't even need AI to achieve that.
If you talk to people in China they'd laugh their ass off at Dario's notion that somehow they are all getting oppressed by DeepSeek or Kimi.
The emptiness of AI companies' waxing poetic about the future of humankind is laid bare by simply looking at what they actually do, and who they do business with. Actions speak louder than words, and they've driven the worth of their words into the dirt many times over.
> Meanwhile the world's most repressive government is North Korea, and obviously they don't even need AI to achieve that.
The country that exists solely because of China? That North Korea?
I mean I get your point, all of these guys are elitist authoritarians who will do anything for a buck, but I wouldn't bring up the DPRK in this discussion.
No, it exists because 1) China shoved back UN forces from the Yalu and 2) their allies, now in charge of everything north of the 38th parallel, aimed a bunch of artillery at Seoul in case there was ever a thought of the ROK and its allies reunifying the country through force, or any other mechanism the Kim family didn't approve of.
The nuclear part is really just gilding the lily. Of course, China and Russia have helped North Korea circumvent the sanctions that were supposed to punish them for their nuclear program, but it ultimately all goes back to the Chinese support of Kim Il-Sung during the Korean War.
The bombing of the school was horrible, but it was a mistake. I could see essentially the opposite argument: Anthropic should be more involved as to prevent mistakes in the future with better technology. Imagine Claude asking the decision makers at CENTCOM "this looks like a school, are you sure you want this added as a target?"
> The bombing of the school was horrible, but it was a mistake
Please don't say this. It wasn't a mistake.
The very FIRST strike packages of the war are very clearly vetted. If you are involved in military planning, you know this.
They doubled tapped it after seeing people flee inside the save the children.
The goal here was to teach IRGC a lesson and demoralize Iranians, which clearly backfired.
Furthermore, US has been contentiously bombing civilian buildings, bridges, hospitals throughout the war. They have also bombed water desalination plants and other key infrastructure.
You have the US president threatening to use nuclear weapons and end the Persian civilization.
Ignoring the obvious joke about "random incidents" in schools regularly in America, is it not the case that the country with the most capacity to retaliate anywhere in the world would not be "demoralized"? I see the possibility that the effect of such terrorism might be different from place to place.
I find it totally implausible that the president or any subordinates intentionally targeted a school. As I recall it was located next to a military base. A mistake seems more likely.
You don't get to start a completely unprovoked and illegal war of aggression, bomb tons of targets with little care, kill 100 schoolchildren, and then say, "But it was a mistake."
The DoW head sees half of his own country’s population as enemies, Islam as an evil religion of hate, and the work he does as striking down the enemies of God. I heavily doubt that this man gives a single thought to children of muslims.
It's a mistake in the sense that if you go and rob a bank while shooting wildly, you only hit the customers "by mistake." That's not an excuse.
The US is bombing both military and civilian targets, by the way. Trump has been very open about this with his talk about "power-plant day." He considers attacking civilian targets to be a legitimate way of pressuring Iran to surrender.
Hegseth deliberately closed the unit that used to ensure it wont happen. The efforts to avoid civilian causaulities were deemed woke, weak and literally unmanly.
And administration repeatedly threatened to destroy civilian targets.
You give too much credit. Many think of it as an acceptable cost of waging war, not something to waste regret on. Some few think of it as a positive (they chant death to America over there).
I've never heard anyone say it was a positive. While you're absolutely right that many would say that it was a casualty of war and that we have to expect some amount of mistakes to be made that does not mean there is not regret.
Many of my conservative acquaintances find the death of those kids to be regrettable in the way you come back to find your car over the line regret your parking job. Certainly they do not feel the tragedy nor any responsibility.
The some few i mentioned who see it as a positive are the ones who say nuke Iran, the only good Iranian is a dead Iranian etc. They might not come out and say I’m glad those schoolchildren are dead because they’ll know they’ll get pilloried. But make no mistake in what they believe.
I find this really hard to believe. People are sick of the Islamic Republic regime, not the person on the street. If anyone implies the two are the same thing then they're being intellectually lazy but I bet if you pressed you'd still find them capable of nuance. People do have some humanity even if little patience for a terror exporting regime
Fox News et al repeat that their populace chants death to America in the streets, as a means to justify the war to their viewership. It’s not a stretch to assume that a small number of consumers of that media would take that message to heart. Of course I don’t have to make such an assumption, I’ve met them.
Yes they routinely call for death to us but I'm just pointing out that I have very rarely heard us call for death to their people.. only to the regime. In fact I really illuminates the difference. the drfenders of the regime like the gloss over this fact. A lot of people in the wwst like to make this about domestic politics instead of justice and I suppose for those people they wouldn't say they're defending the regime they're just trying to score points. So be it there's going to be people who do that I get it.
Is there a specific list of changes they made to the system prompt? They're claiming they removed 80% of it. That's quite substantial. It would be good to know what the model knows to do by training and what we need to avoid over-specifying in our system prompts.
Saying that "give Claude judgment" is too vague for agent implementors. Given the lack of specific details, my takeaway is that we need to go and review all context and rework prompts from prompts/descriptions from scratch until they pass the evals again.
I spent some time running mitmproxy and watching the system prompts and it's what drove me to codex. the main issue for me was their system prompt wrapped the CLAUDE.MD with a "IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task." https://github.com/anthropics/claude-code/issues/18560
Anyhow- if anyone is sufficiently curious and has access- just tell the agent to setup an mitmproxy to watch the traffic and see what the system prompt looks like.
It definitely makes me uneasy given the types of behaviours we've been hearing about from these models. Letting them use their judgement can go horribly wrong once a misaligned behaviour is triggered. It loosely translates into a relaxation of guardrails.
I worry that the ability of the model to reach similar benchmark scores to Fable is more to do with this "letting the agent off the hook", allowing it to explore a wider (but riskier) set of avenues to solve the problem than it is due to it getting genuinely better at the direct problem solving.
A while back, I saw a similar feature land in Codex (I'm using the VS Code plugin) but it got removed quickly. What is the chance that an LLM recommended this same idea to the Claude PM or lead? I see LLMs across different providers converging on similar ideas or biases frequently.
I was just looking at the same thing and how flawed it is that we tie parameter count with "intelligence". GLM-5.2 is my go to day to day model because of how darn good it is. I had no idea it had a substantially lower parameter count over deepseek v4.
The full [Kimi K3] model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
That's a quickstart page for using the model on the platform not a page about the model. I am skeptical you are correct that it said something about model license earlier.
Not the person you're responding to, just a person who still has the original version of the page open in their browser. Quoting from it:
"Kimi K3 is the first open-source model to reach the 2.8-trillion-parameter scale. It is the latest step in Kimi's continued push of model-scale boundaries: in 9 of the past 12 months, Kimi models have set new records for open-source model scale."
The page has definitely changed.
(I'm not sure why you would be skeptical of somebody recollecting something they probably read only half an hour earlier.)
The K3 marketing popup when I look at the Kimi Code page says "Kimi K3 Open Frontier Model". So, if it's not going to be open, they haven't told the whole team, yet.
Moonshot (true to their name?) has always lead in terms of releasing the largest among open weight LLMs.
> Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.
Think Moonshot, as a spin-out, can expect backing from its former parent, Alibaba? I don't think they would be particularly worried about finances, if the Kimi K series continues to outperform the Qwen Max series (which seems to be the case; while Kimi is also super popular in China).
Fable reportedly use 20T paramteres, 1T=1000B. Opus is probably 10T. That said, these are estimates based on model preformance and scope of general knowledge breadth.OpenAI, Anthropic, and Google have not openly report their model sizes.
Chinese models are way behind on the mode size race due to lack of abudent AI infrustructures. That said, it seems Chinese models are going pretty well on a seprate route. They manage to achieve 80-90% performance with 1/10 of the model size. This is some what related to the diminishing reward situation described in the scaling law. I think it can also be attributed to their persistent research in this direction. Thinking and DSA (deepseek attention) were both developed and opensourced by Chinese labs then adopted worldwide.
Kimi has almost no advantage over Zhipu (Z.AI), so the performance boost likely comes from the number of parameters. The 2.8T model may not be as large as Fable, so Fable’s performance may also stem from the number of parameters. Or perhaps they quickly distilled Fable or Mythos. Distilling Mythos has a significant barrier to entry, and since Fable was released not long ago, is this even feasible? It outperforms Fable in several tests—how did Distilling achieve these results? Is this some kind of cross-vendor RSI (Recursive Self-Improvement) or RDI(Recursive Distillation Improvement)?
I'm not sure what the economics of building a new runtime and ecosystem from scratch are but it seems we're already in a phase where individual developers are creating software which previously took a whole team. And its only getting started...
Is that impressive? If an LLM is just spewing out code which it has been trained on extensively then to me it's not that impressive at all. I want to see individual developers creating new software which previously took a whole team. Things not done previously. I am very pro AI btw, but the constant barrage of false hype is really tiresome.
Congratulations to the team for pulling off this feat while doing a responsible migration (looking at you, Bun).
Quick question: How does this affect downstream tools like tsdown and esbuild, which need to build the TypeScript codebase? Can I use TS 7 and current tsdown together?
> How does this affect downstream tools like tsdown and esbuild, which need to build the TypeScript codebase?
esbuild doesn't rely on TypeScript at all, so there's no issue there.
With tsdown on the other hand, it depends on if you use --isolatedDeclarations. If not, you can install TypeScript 6 side-by-side (instructions for this are on the blog)
> It’s worth calling out that workflows that use Vue, MDX, Astro, Svelte, and others will likely not yet be able to leverage TypeScript 7. Similarly, specialized type-checking within templates like Angular will also likely not use TypeScript 7. This is mainly because TypeScript 7 does not yet expose a stable programmatic API, and so tools (such as Volar) which embed TypeScript into their own compilers and language services can only currently rely on TypeScript 6.0. We expect this to be a point-in-time issue, as we are committed to providing a solution here. We will be actively working with the maintainers of these projects to ensure TypeScript 7 supports these workflows.
Not the op, but this TS migration started long before AI was able to help. It was done slowly and carefully, as a project supporting millions of users should. And the benefits are very clear.
Bun’s port was a vibe coding fever dream that happened from one day to the next, with much looser motive, and yet to be proven reliable.
Bun's migration to Rust was nothing more than a marketing stunt to sell more Claude subs under the impression it can perform this kind of work at scale, assuming that most who were convinced by it wouldn't look under the hood at what really took place.
It has its merits as a proof of concept that could eventually be cleaned up and released properly later, but I can't see it any other way.
Too many see it as this miraculous one-shot and are using it as a blueprint to justify more layoffs and buzzword salad in their boisterous LinkedIn announcements about how they're "completely overhauling their strategy" in engineering. Hogwash.
The irony is that the blog post actually points out as pain points the reasons many of us assert languages like Zig are out of place in the 21st century.
A nice collection of heap-use-after-free crash, use-after-free crash, crash and out-of-bounds read, memory leak, double-free crash, race condition crash.
Bun is infrastructure. Why would I want my infrastructure to be unstable? (By the way, 10,000 unsafe blocks last I checked, though the number is going down somewhat.)
I've never used bun on production for this very reason. But nevertheless, tens of thousands of people and businesses do.
And I'm not sure how you're responding to my comment. The parent said "this is a marketing stunt" derogatorily, as if it's slop that doesn't work. This is already the canary build, it's more stable than the current stable, and is actively in production products in wide use.
The parent is objectively wrong, whether or not I personally use Bun.
Marketing could of course be one of the main motivations, but it's not a "stunt". More like a marketing achievement, I guess? Stunt implies smoke and mirrors, and bun rewrite is quite real.
Probably cheaper than doing it by hand, however, that's the short term "port x to y" cost, the longer term cost (or benefit) is a lot harder to calculate.
I don't think irresponsible is the right word, but it has drastically reduced Bun's appeal. All the tools we use have a brand to them, and Bun basically changed their brand overnight to "reckless" in my eyes.
bun has never been fit for production, at least not for load bearing business apps. it’s been haunted by segfault bug reports since the early days, and i personally hit at least one a week when im doing lots of bun stuff. im excited for bun with less segfaults
It's unknowable, because the PR is unreviewable. The Bun migration PR is larger than any model ever made can fit into context. You just have to pray that test coverage is sufficient to catch all of the possible errors, which it almost certainly isn't.
The initial state has no bearing on whether the process was responsible or not. That's measuring along a different axis. If the bun rewrite lands and it breaks someone's app, that's bad no matter whether there's more or fewer bugs in the final state. The important metric in a rewrite of software that's used in production is stability.
It's really not even close to being the same. In the best case, a bug means your app crashes on the new version. In the worst case, something more insidious happens like opening a security vulnerability (say, TLS isn't handled correctly or HTTP headers are mishandled in a way that allows SSRF or request smuggling) or a previously linear time operation is accidentally quadratic (leading to DoS).
You can apply the same FUD to the old version. Your argument basically is “all change brings risk” which is true but doesn’t add any useful insight. It’s always easy to complain and warn about change causing problems while ignoring the problems of the status quo. The “everything is fine” meme in action.
Sure, all change brings risk. But this is change where:
1. The change isn't made by a human.
2. The change wasn't fully reviewed by humans or machines. It's not currently possible for a machine to review the whole thing as one.
3. It's a full rewrite. This isn't a ten thousand line change, it's multiple orders of magnitude more than that.
You're literally making the argument that all risk of any size from any change of any size is equivalent, so just don't worry about it. If you relied on this software before, good luck convincing yourself it's fine to rely on this software now: it's literally not the same software anymore.
No I’m not making that argument, you’re the one making that claim as if it’s the only alternative to your position.
My position is more nuanced. A) what does the test coverage look like B) how is the deployment managed.
I suspect B is going to be my biggest issue - normally you’d deploy this slowly over time to monitor problems and whatnot. But ultimately the real test is seeing how it actually performs in the wild and kinds of problems people report. But you can always keep using the zig version if you wanted. So ultimately it’s a lot of consternation over a nothing burger. You can laugh at them if they screw up the release, but it’s a bold attempt at trying something legit. It took Microsoft 2 years of many engineer hours migrating typescript to Go. If it takes significantly less calendar time and human time, you could reasonably even evaluate what a Rust based typescript looks like vs Go if you wanted to for an order of magnitude cheaper.
There were a lot of products and ideas that felt great during ZIRP but stopped making sense post-ZIRP. That's my take.
reply