Hacker Newsnew | past | comments | ask | show | jobs | submit | try-working's commentslogin

I actually define frontier as share of inference market spend. Anthropic's is falling.

Anthropic cannot IPO because AGI is the wrong strategy, and the market is rejecting the AGI approach.

Fable and Astra are what we currently call frontier models, but to be more specific they are generalist models, built in pursuit of AGI. The strategy is to have one single model that does everything, whether it's writing code or doing research, etc. Fable is a single, massively sized models that is intended to do specialist work across every domain.

The issue with this is first of all that it is the contradiction of a generalist doing specialist work, and that contradiction creates the present situation with model profiles that ensure that these models will rarely be chosen in a pool of models like V4.1 Flash that can now do GPT 5.4-level work.

We are seeing this reflected in the market where companies and individual developers are moving away from frontier models toward models with better cost profiles. In a sense, the market is killing Anthropic's dreams of AGI.

https://x.com/trydotworks/status/2098618997230985375?s=20


Fable and Astra are what we currently call frontier models, but to be more specific they are generalist models, built in pursuit of AGI. The strategy is to have one single model that does everything, whether it's writing code or doing research, etc. Fable is a single, massively sized models that is intended to do specialist work across every domain.

The issue with this is first of all that it is the contradiction of a generalist doing specialist work, and that contradiction creates the present situation with model profiles that ensure that these models will rarely be chosen in a pool of models like V4.1 Flash that can now do GPT 5.4-level work.

We are seeing this reflected in the market where companies and individual developers are moving away from frontier models toward models with better cost profiles. In a sense, the market is killing Anthropic's dreams of AGI.

And there we have it: the reason I say that Anthropic is no longer a frontier lab is because after this July, we have proof that their strategy and the models they put out do not match what us, the users and the market, needs and doesn't fit the work we need models to do. As a result, Anthropic's market share is dropping rapidly, and how can you be a frontier lab when you're losing every day, for months, without end in sight?

https://x.com/trydotworks/status/2098618997230985375


Anthropic is slowing down by themselves, no government intervention is needed. Their ARR has been crashing since July. https://x.com/trydotworks/status/2098618997230985375?s=20

Because they've not been great to their customers. I've stopped using them completely.

if you'd rather have your router running on your own machine, check out the router i'm building: https://role-model.dev/

that can be achieved by storing user traces and then running evals on own model vs competitor model. you don't need to route live customer requests to a competitor for this.

For agentic tasks where the model outputs tool calls that run on the customer's computer, you can't just store and eval later, because then the execution environment is no longer available.

It's a genuine Chinese news article written by a real person. I don't know why they choose to use these types of images, it truly makes it look like slop. I think it's because it's originally a WeChat article posted on their website.


For what it's worth the original article is in Chinese: https://www.qbitai.com/2026/09/483600.html

The article linked here seems to be a (bad) auto-translated version.

It's not uncommon for images to be incorporated in an article like this, as they typically get shared around WeChat and correspond better with the 'culture'.


How do we explain this little blob of sloptext right in the middle:

> The second task was more straightforward: both models were asked to search for flight tickets in a browser at the same time.

> I'm unable to open a browser or interact with live websites like Google Flights. I can only process text and don't have real-time browsing capabilities. To find this flight yourself, here's what you'd do: 1. Go to google.com/flights 2. Enter Singapore (SIN) → Beijing 3. Select one-way, set departure date to 5 days from today 4. Filter by "Nonstop" and "Economy" 5. Check the results — note that flights to Beijing may land at either Capital (PEK) or Daxing (PKX). Exclude any Daxing arrivals. 6. Compare prices and pick the cheapest nonstop option landing at PEK. If you'd like, I can help you think through typical price ranges, airline options on this route, or general tips for finding cheap flights.

> Among them, StartLux-V1.0-27B-Preview completed the search in about 95 seconds, finding an Air China ticket priced at $299.


they're quoting the models chain of thought.


Did you even read what it just said above? They're testing the model's capabilities so of course they'd have whatever text the model outputted in the article, that's the entire point.


There's no such thing as an innate most qualified person. Zuck was never the most qualified to be CEO of Meta. He is now, because of what he's done. At the same time there's an infinite number of jobs he is poorly qualified for.

Strict qualification can only be measured in a extreme niche cases. You founded an area of research, so you're the most qualified to work on that at Google. You have a massive Linux project, then maybe Torvalds or his core maintainers are the only ones who fit.

But outside of that, people become qualified because of what they do, and before they did it, they were unqualified.

As an example, I'm sure there are many people at OAI and Ant who are today world experts at research and training and infra. But they became that by just doing the job. Take their same resume from 5 years ago and they may not even reach the interview if they applied for the same role. The what does qualified mean, how do you measure and define it?

Countries need to develop their own talent and individuals need to develop themselves.


This is cool. I would like to see it compared to Pi, OpenCode, DSH in your benchmarks.


DeepSeek is consistently the fastest model, and Kimi the slowest. GPT in the middle. Anthropic is probably on the slow side as well.


Agree with the exception of Haiku. Haiku is insanely fast if it works for your task


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: