I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
> for “re-migration” in europe, also known as the forceful ethnic cleansing of non-white immigrants.
I'm curious if you feel the same about re-migration of Belgians from the Congo?
Personally I think it's fine for any country to vote to control immigration as they see fit. I think Japan is a good example of a relatively xenophobic culture that deals with this fairly and thoughtfully.
> I'm curious if you feel the same about re-migration of Belgians from the Congo?
Can't say I've ever heard anyone implying that colonialists leaving Belgium was unjust. Colonialists is actually not the right word, more like extended occupation, only slightly better than the enslavement of the Leopold II era. The Belgian's were less than 1% of the population and all but an ancillary amount worked in exploiting the native population.
You're right, I don't care and I certainly didn't ask for your political opinion or validation. This is a tech news website, not the place for your false and unrelated political tirades.
Grok did ask its users what it thought about the right wing fever dream of genocide of white people in South Africa so its definitely relevant to one's consideration of the product.
Technology isn't apolitical. You can choose to ignore the politics if it helps you sleep at night, but Elon's got three-letter agencies reaching up his ass like he's a Sesame Street puppet.
I didn't miss anything. The parent that they can choose to ignore it forever, it's a perfectly valid option. It just won't disentangle tech from politics.
The fact that Elon Musk's companies take contracts from the CIA and NRO is not flamebait. It's context that informs how we evaluate future SpaceX ventures.
Musk also shut down USAID for absolutely no reason at all which will cause the deaths of hundred of thousands of the poorest people in the world. Not a good look for the richest person in the world. Musk also sounds like a complete moron when he tries to justify why he did it.
Flamebait I guess but I think it is good we have several options to choose from already. If you don't like Elon there are several other models, and each can pick the one controlled by her favourite supervillain.
Agreed. Another difficulty here is there are not good benchmarks for this new architecture yet, so it’s easy to potshot and snipe, where jev seems to be pretty broadly intelligent/at least have had a lot of rl in different domains.
We haven’t seen any of these copy cats play doom or street fighter for instance; just categorize email.
I imagine once the author cools down and evaluates on a broad harness of tasks he may find that his new thing has a lot of engineering work ahead.
The doom demo would have to be reproduced to confirm what their model is capable of. Oh but it's all closed source, so who knows.
It reminds. Me of Devin. Took a while to debunk. Not saying Jev is a fraud , but the gap between structuring typed output and playing a game involving logical interpretation of frames made of pixels, screams unstructured interpretation they made and forgot to mention.
A counterpoint - I was told a story by one of my professors in the late 1990s, about one of his professors -- he'd written a thesis, gotten hired somewhere like Princeton, and taught there for a few years as Dr. <Somebody>. One day he received a letter pointing out a construction flaw in his thesis. He brought it to the department head who read the letter, and said "Well, Mr. Somebody, ..." Ultimately he fixed the proof.
Upshot, if there are real errors in published work, I think most mathematicians want to know about them.
Yea but that was a person who actually put in work and had to think about and understand the problem. They didn't just generate something with a magic box.
As a software engineer, I could not care less about how a bug was found or by whom, as long as I can quickly verify it is correct, I always appreciate being able to improve my work. I don't see why it should be different for mathematicians (the ones I knew would think similarly, I would assume.)
If I may offer an analogy, what I tried to do is essentially reporting a bug after having a failing unit test that exercises the public API, without looking into the black box of internals.
This issue really separates the wheat from the chaff. The dirty secret is that a lot of academic published work contains errors but the authors also have fragile egos. It reminds me of when Data Colada exposed someone for fraud and they accused the exposers of "methodological terrorism" (something in psychology or social science). Just imagine what happens once papers get exposed at scale.
Now here there was no fraud just genuine error, but it will annoy people nonetheless and scrape their ego that someone uninitiated can just type some stuff in a magic box and conclude that they, the established published, tenured mathematician with awards and medals can be wrong.
Well, duh. If you could do this with Opus 4.8, we would know. When Astra’s successor is 2-3x better at math research, and the internal teams say “we believe we will get there,” I’m inclined to believe the insiders.
The insiders that said every tech workers would be unemployed in 6 months and every white colar would be unemployed in 12 months like 2 years ago? The insiders who are about to file for IPO?
Going into the color of the bikeshed; hacking huggimgface to cover up cheating on your hacking test; swapping your language for no discernable roi. You know the guy, severe OCD and anxeity, who barely does anythong of value due to his anxious brain?
Yeah, it should be obvious what AI is really doing and its definotely not ROI improvements.
What if any of the older good models could also have written those math proofs if they were given the same order of magnitude of resources? We don’t know and there is literally no one else in the world to check it. To me it’s very suspicious that all these hacking, containment escape, hidden internal thinking, math proofs started coming out all at once in a very short time right as IPO talks have intensified and Chinese seem to get closer and closer, also regulation discussions are starting to get very serious. I have used these models and they are good, especially Fable, but not groundbreaking. With intelligent guiding I actually feel better using Opus 4.6 as I feel more in control, having less hidden away from me.
Prices are moving inference to highly profitable, oAI seems to have solved their training problems, and they're currently competing nicely with Anthropic on the coding side. I'd wait, too -- why fight this stuff out in public when you can stay private and have your big competitor deal with all the public company concerns? There's plenty of capital available in the private markets for them right now.
Given that they paused signups for their 20x Max plan, your assessment seems overoptimistic.
Their new Broadcom chips seem cool, and they had a nice model release with Astra. But, that doesn't mean they've solved the profitability question during an era of rapidfire open-weight model releases, extreme memory shortages, and intense political pushback.
I think that implies they're seeing unusual new subscription demand, yes? That's how I'd read it, not least because I worked through four resets this week on Astra, which is un unbelievable amount more inference than I've wanted from openAI really ever, and the highest ratio vs. claude since opus 4 at the very least, probably farther back.
I think there are few moats in the engineering use case, and a single new model can absolutely drive compute demand.
High demand is great, but it doesn't say anything about your margins. If anything, it is probably a weak negative signal that they are pausing signups but not raising prices. In most markets, the answer to excess demand is to raise prices. If you can't meet demand and you can't raise prices, you're a sitting duck waiting to get your lunch eaten by somebody who can absorb that demand. And this is not the kind of market where people will just wait patiently for a differentiated product to be come available.
20x plan is $200, 5x is $100. Pausing $200 plan signups is effectively doubling prices (you can subscribe to two $100 accounts for the same money and get half as much usage), and that's exactly what they did.
I agree with some of your points, but wouldn't memory shortages actually work in OpenAI's favor? Part of the reason for the shortage is them prepurchasing hardware, so higher prices for memory and other chips would make self-hosting and competing inference providers less competitive.
Companies that plan to go public don't push back their IPO dates if things are going well. It's a universal sign that the finances aren't in order when a company pushes back its IPO multiple times.
OpenAI and Anthropic are barely profitable on a "unit" basis when ignoring the costs of marketing and other COSS expenses. They're definitely not profitable on an EBITDA or GAAP basis or they'd already have IPO'd.
I believe he's saying "Citation Needed" - It does seem the opposite is true, while maybe token cost is going down, frontier models use a lot more, so it's no an accurate gauge of cost. Trying to find any chart of token cost, I found this which looks like inference cost is going up: https://tokenpriceindex.com/ . Of course frontier models are becoming far more useful, but if you're claiming prices are becoming highly profitable then are you saying this increasing price is actually able to cover server costs now? My understanding it was still sold at a loss. And token price increases have turned major companies off the hyper use they tried out early in the year, using it more sparingly. I don't doubt it'll be a solid business, but it's not shaping out to be the hugely scalable business openai promised.
This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.
oAI has made clear they did not specifically pull in any user data to context for this run.
Tristan Buckmaster’s post cited extensive use of LLMs in the process of his collaboration with Levent:
“We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.
The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).
This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”
The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.
“We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”
For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.
Yep, it's possible. But, we have literally no idea how much a set of prompts would impact training as far as general usefulness. I don't think we even know if Tristan's said he allowed training on his prompting or not. This is about money, ego, primacy, all the usual mathematician priority disputes.
Shenanigans is an inaccurate word; it implies underhanded behavior that's hidden / concealed. I think "to act so aggressively" is more balanced.
Here's my answer: If you think we're getting to AGI in the next 9 months, then you believe, with all your heart, that these problems will fall soon. However, there's an ocean to boil in terms of what you could point your limited clusters at. In the meantime, the market is desperate for any sign your company might be first to AGI. Therefore, news of tractability with current models might focus an organization intensely - internally they have a huge leg up on the public, and therefore it's minimal compute to check - and if they are successful, they get approximately $50 million of free PR, likely adding 10-20% to their valuation.
Likewise someone like Tristan is fighting for his (metaphorical) life right now, hoping to preserve his claims of primacy and have a shot at some of that prize money, despite being only partway to a full solution for N-S.
I don't think we see any behavior at all that isn't simple to understand and well described by the setup here, but tell me what you see differently.
OpenAI most certainly did not as you put it, "to act so aggressively". They have simply indulged in plagiarism/malpractice/fraud all for the sake of pumping up their evaluation in light of their forthcoming IPO and to one-up their arch-rival Anthropic and try to get themselves to AGI certification.
In the process they have shafted real-world hardworking mathematicians, which is to say the least, despicable. Note that stories are now coming out from other mathematicians who have also been shafted in a similar manner. Also there are cases where they have had mathematicians accept their "deal" (like the one they offered Buckmaster that he refused) and have OpenAI name linked to their work.
Regarding "their proof", their claim is only solving Navier-Stokes partially for when a smooth force is applied and not a fully general solution (which is maybe impossible). The proof is still being verified and we don't know whether it is just an approximation/hallucination or not.
Their most blatant lie is that they "gave only the problem statement" to their system which then went ahead and solved it. This is almost an impossibility. Problems like these need to identify a specific lead/approach and some work to be done on that path before you can even know whether that approach is promising and worth pursuing. This problem has resisted all attempts at solution for over two centuries. This is where the Buckmaster/Levent's work's importance comes in. They identified a promising approach based on other mathematicians work and have been using both OpenAI and Anthropic's models to make progress and had reached a promising milestone which they published. But they inadvertently gave away their approach to the model's training data set which OpenAI capitalized on by throwing a large amount of compute at the problem to get to the finish first.
This is straightforward stealing of other people's work and building upon it to claim it as your own which can and should be sued. All Scientists/Mathematicians/Researchers who feel OpenAI has done them dirty should band together and file suit.
This whole thing could easily have been avoided if they had worked with the researchers so everybody's concerns/needs are met.
By engaging in this sort of backstabbing, OpenAI has effectively killed the "Goose that laid the Golden Eggs". viz. Researchers were giving away their hard-earned highly specialized knowledge freely to the models in the hope that it will help them get quicker to the result. But now everybody is going to lockdown their research findings and will stop sharing it with the models to the overall detriment of advancement of Science.
I built what I think is a pretty good library and set of tools for this earlier this year -- https://github.com/corpollc/qntm (or `uvx qntm --help`); it includes a cli, python and typescript libraries, and works out of the box aimed at either a public endpoint, or a private one, depending on environment.
It's end to end encrypted, and has group messaging support, so if you wanted to read what the agents are saying you'd just add them to groups you're in. It also has a web ui. Version 0.6.0 should get pushed this evening pacific time, with some additional agent specific features.
Bug reports welcome! If I did a good job on architecture, you should be able to have your blackboard up tonight.
This is an interesting idea.. do you use agents with more adversarial / cynical prompts whose only job is approving api calls? "You are an API approval agent on the lookout for suspicious or unusual stripe api calls...."
Thanks! It's just infra, so you could do what you wanted. The clients include a safety reminder, and the datastructures include "unsafe" in the name of the inputs passed around, but that's just a little hygiene.
There really aren't good messaging libraries that provide what I wanted, so I built it. Basically I started with "signal but no need to have a phone number." I've used it to build a group messaging iOS app for friends, and just pass it to an agent all the time if I want them to be able to direct message.
The group API approval is modeled off of multi signature approval mechanics from Ethereum - so, you could use it like you describe: "Only allow this call if it passes safety checks from n reviewers", or you could have human in the loop, or a program that checks business rules + an agent and a human, etc. etc. I just wanted something that let us control API calls properly.
That’s not untrue. But it’s also a misstatement of mathematical history. Many leading mathematicians historically have been highly competitive — Gauss comes to mind. Woe betide the lesser intellect that sent Gauss some ideas. The Newton Leibniz controversy was very serious business at the time in the UK and the continent. It was considered at the least a sin to reveal that sqrt(2) was irrational to those outside Pythagoras circle.
The mathematical community was very competitive in its early years, but in the last 70 to 100 years, it has been generally less competitive and very collegial. The community was in a good place, and progress has been very good. In a few cases when competitiveness was ramped up, it lead to bad behaviour and destructive fights. Few would like to return to those competitive years.
I think both the competitiveness and and lack of problems will mean fewer people will do mathematics. I think collaboration will always be there, but the community as a whole will be smaller and weaker.
I'm convinced that all the quantum talk is the NSA propaganda to convince us to drop RSA and move to some secure post quantum algorithm that they can easily break.
> I'm convinced that all the quantum talk is the NSA propaganda to convince us to drop RSA and move to some secure post quantum algorithm that they can easily break.
It is rather consensus that by now,
- too many potential weaknesses for RSA are known (in particular in how it is used in standardized cryptosystems [1]),
- it is a bad idea if too many "free parameters" of the cryptosystem are decided by the practical implementation (in RSA: the decision which primes to use to generate the key pair) instead of these parameters being part of the standard
Thus the advice to drop RSA is in my opinion sound.
But otherwise: many people indeed have the suspicion that the panic that a sufficiently powerful quantum computer might get developed in the next year is indeed used to push novel "post quantum algorithms" which
- have not been analyzed as thoroughly as older algorithms, and thus might contain weaknesses,
- were covertly developed by some three letter agency, and contain backdoors.
In other words: your suspicion might not be unfounded - this just does not imply that RSA does not have its risks and should thus arguably indeed be abandoned.
---
[1] Just to give one example: in the past, RSA was often used together with the padding scheme PKCS#1 v1.5, see [2]
ECC allows greater security for key exchange (because it is so cheap to generate keys we can have "ephemeral" exchanges and forward-secrecy). It isn't significantly better for signatures in most cases.
"But I realized after a while that talking to people casually about Fermat was impossible, because it just generates too much interest, and you can't really focus yourself for years unless you have this kind of undivided concentration, which too many spectators would have destroyed."
But yes; him reaping the benefits of himself having the idea first was part of it too; as far as I am aware.
-----
Which is still something completely different than some anonymous organisation keeping mathematical research secret because it is better for hype reasons. One is competition between individuals or groups within a field; the other is boring and sometimes borderline nihilistic generating of mathematical knowledge as an marketing asset.
I've always found the story of A. Wiles sad and frustrating.
He worked in secret for 7 years. He submitted a (incorrect) proof at year 4 or so. Reviewers found a problem, but he decided kept all secret for many years after.
He didn't even proof the last theorem of Fermat directly, he proved some conjeture that someone else before him, proved that it implied Fermat last theorem...
I found this behavior against healthy science practices and only driven by ego. Unfortunately, I find this too often at work (working in academia).
Most probably I'm too naive...
> He didn't even proof the last theorem of Fermat directly, he proved some conjeture that someone else before him, proved that it implied Fermat last theorem...
> He didn't even proof the last theorem of Fermat directly, he proved some conjeture that someone else before him, proved that it implied Fermat last theorem...
I think that was Ken Ribet?
Grigori Perelman and the Poincaré Conjecture is more interesting. IIRC he turned down Millennium and was decidedly not all about the Fields Medal - mostly because Richard Hamilton didn't get credit? Anyway, I am grateful I had the opportunity to learn about Poincaré in college taking a few classes from a professor who was a key contributor to the conjecture and got a Fulbright for it when I was there
What you’re talking about is his proof of (a specialised version) of the Taniyama-Shimura-Weil conjecture[1] which had been proven to imply Fermat’s Last Theorem. The technique he used to prove this was adopted by his students to prove the conjecture in full generality so it now known as the modularity theorem. Given its importance to the Langlands programme it may be that when history looks back on this it will consider this a more important contribution than the fact that it proved FLT even though that is obviously the thing that grabs the headlines, but there’s nothing at all wrong with proving something that implies your goal rather than proving the goal directly. There’s a reason the words “it suffices to show” often turn up in proofs.
While I can sympathize with this perspective, I don’t think it’s right to call it driven by “ego.” Sometimes one just wants to go at a problem without being second guessed on approaches or led astray with suggestions by others.
Eh, it seems like it's pretty necessary for success on such a problem (but obviously not sufficient). These problems gain a reputation, and you either get judged for it or get too much attention for it.
Andrew Wiles was also careful about communicating progress on his Fermat's Theorem proof during the years in his attic. So yes I take the point.
I read the Mastodon thread as more about the 'flattening' and 'rawness' of the proofs these systems and their operators are producing. I mean what is the cultural significance of a lean proof that is half a million lines long or something? And what tools can be extracted for further work from such a construction?
The late William Thurston wrote about the culture of mathematics in that sense.
You miss the point. Humans don't mind competing with others. I love competition, but I don't want to compete with you and your machine. I love to play chess, I don't care if you are grand master, whoop my ass. But not if you are going to pair up with stockfish. I don't even care if you are a newbie that started playing yesterday with an ELO rating of 900. If I wanted to play the damn computer I'll do it myself. Likewise, mathematicians will not mind sharing and competing with other fellows, but if another has a billion dollars worth of GPU and you don't? Then you best be carefully what you say.
Could you tell the difference between a grandmaster and stockfish if playing them online? If not, why would you care which one you are playing against?
I’ve never played against a grandmaster, but I have a feeling that he/she would play very different compared to me and would never make mistakes I could notice. Though admittedly I’m not very good at chess.
I've played both. The GM plays tremendously differently than Stockfish.
Engines - specifically heuristically-driven ones like Stockfish - don't play like a strong GM. They play engine-perfect chess, which isn't how a GM plays with any consistency.
I'm only a decent amateur (1550 USCF) but when I lose to a titled player it's largely explainable in human terms how it happened.
I think I'd care because the entity on the other side cares about the game in a similar way to me. It's not just the technical details of how the pieces move, it's a human interaction.
/meta Their comment got involuntarily migrated; it was originally a reply to a toplevel comment in the Navier-Stokes thread (before the dedicated Tao thread existed),
Surprised to see someone on HN arguing against open science. Seems like the opposite of the lessons we should learn from Newton and Gauss, actually, hoarding results for decades at the expense of progress.
(the Pythagorean thing isn't really competition either, is ahistorical, and from what we actually do know it's again people hoarding results instead of sharing them).
FWIW, your post comes off as a middlebrow dismissal, surface level and not actually engaging with the substance of the comment. It's also just wrong. You claim "it’s also a misstatement of mathematical history", but don't specify which part. That there's "centuries of traditions of open science"? But your examples are from centuries (and millennia) ago, and there was never any claim that these traditions are universal.
But more fundamentally, competition doesn't mean you can't also have open science. And the very long, damaging events like the Leibniz/Newton feud are exactly what make many mathematicians work to maintain a spirit of collaboration and attribution even when they're competing on approaches.
> Nothing in their comment reads to me as "arguing against"
If competition is somehow the opposite of "centuries of traditions of open science", and "mathematics has always been highly competitive", then open science is neither sufficient or necessary for the future of mathematics. Their clear implication is that we don't need to worry about it, though, because it's always been that way.
> Reads like nothing but historical context
They literally accuse Tao of "a misstatement of mathematical history".
>If competition is somehow the opposite of "centuries of traditions of open science", and "mathematics has always been highly competitive", then open science is neither sufficient or necessary for the future of mathematics
For the future of past mathematics, it says nothing about the current future. Also, open science can be nonsufficient and unnecessary but still extremely beneficial and desirable.
>Their clear implication is that we don't need to worry about it, though, because it's always been that way.
Lets just ask him if that's what he meant, I bet no.
> They aren't arguing against open science, they are trying to educate you on the history of science. It's always been this way.
Always been what way? And how does that contrast to what Tao said (since it was apparently "a misstatement of mathematical history")?
> Also, your third paragraph is highly ironic.
You'll have to be more specific, since I engaged with my parent's argument, while they waved away Tao's quote by suggesting he was wrong because of exactly the kind of events that helped lead to the norms and mores working mathematicians have today.
Not arguing against open science - it's super valuable. I'm saying that pearl clutching by people reading Tao isn't useful, because it misses some long history which tells us that this kind of science has been seen as fundamentally competitive for millennia.
Should it be competitive? Is it more useful to be collaborative? How collaborative can it be when it's fundamentally competitive? Is it only fundamentally competitive because of some common 'quirks' of math types, or are there deeper forces pressuring it to be competitive?
These are all questions that I think are worth discussing, as is the note that the pendulum seems to be swinging away from cooperation in the face of competing for $trillion+ valuations (and a real enthusiasm for proving cool math stuff). The alternative, tweeting complaints on twitter without some context, is mostly a waste of space. I mentioned the history in hopes we could get informed complaints on twitter.
It is a strong argument in this case though, because Terence Taos expertise is directly linked to his ability to not misstate the history of mathematics.
Also note how the quote by Tao is in all likelyhood not meant as an absolute; rather than a statement of a trend - a handfull of counterexamples do I no way change anything about the truth value of Tao's quote.
On the other heand; consider how absurd it would be if "... in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science ..." would indeed be a misstatement; which would imply that far more promising research directions were not shared with the broader community (i.e.: published). I wonder what different reading of that counterfactual there could be other than secret societies that kept their discoveries and research directions to themselves - which we just learned about (since we would otherwise not be refering to the secret societies and their supposed promising research directions).
All pretty straightforward, I would say - both that "misstatement" is hopefully based an overly strict reading of Tao's quote, and that mentioning Tao's background as one of the fields leading practitioners is relevant as well. Again; to make sure: A few counterexamples achieves nothing here. It would need to reach a certain threshold of such counterexamples before we will have to write the history of mathematics; and before Tao actually made a misstatement here.
I mean mathematics has enough history that something can both have been false for centuries of mathematics research and true for centuries. Sometimes in different places simultaneously.
Terrence Tao can do his job perfectly well without being aware of any mathematical history, though I consider it unlikely that he is. I'm not seeing the direct link you're talking about, in fact history is frequently left out of mathematical teaching even when the history would in fact help in the understanding of some concepts.
> in fact history is frequently left out of mathematical teaching even when the history would in fact help in the understanding of some concepts.
That is well-known I assumed and continue to assume.
> I'm not seeing the direct link you're talking about
You are stating that link yourself; indirectly: "though I consider it unlikely that he is [being unaware of any mathematical history]". Why is it unlikely, precisely?
- Maybe because it is unlikely that he recieved the mathematical teaching that frequently does not contain history of mathematics (wild! I wonder which university you have in mind in particular) that you seem to be refering to?
- Or maybe because he is quite the opposite of a person that never ventures outside of their own area; being blind for other fields, or ones own history; as evidence by being famously collaborative across different fields, having a popular blog where he writes about non-mathematical topics too and last; him being one of the main proponents of foundational topics such as formalization of mathematics; or the use of LLMs for mathematical research.
Does all that really make it more likely to you that Tao is not aware of the existence of counterexamples like those the commenter above mentioned - more likely than the commenter simply having missed a nuance or taking something out of context?
If so; I would be genuinely curious why - people work differently, and I am always happy to learn, or close gaps in my own understanding.
Any one of the arguments you make here is already stronger than the original one. The problems with appeals to authority isn't that they are always false, it's that they are not a good argument.
And you say his "expertise is directly linked to his ability to not misstate the history of mathematics". And frankly, I disagree. If he happened to be misguided or even outright wrong about some parts of mathematical history I wouldn't think any less of him, nor do I think it matters much for the work he's actually paid to do. At worst it would result in an online discussion, which is arguably a good outcome not a bad outcome.
And the other an appeal to tradition. There was only one Gauss to scoop and focusing on him misses the larger math culture which Terry might be aware of, where most trust others to not scoop.
And if any mathematician's AI usage on a problem leads to scooping, the volume of agents involved gives them a huge advantage which could prompt mathematicians to not use LLMs.
Though you can say Terry's claim is a slippery slope.
So let me get this straight: you're saying that Terrence Tao, one of the most prominent mathematicians alive today, doesn't know math history? And me pointing this out is merely an appeal to authority?
> Terrence Tao, one of the most prominent mathematicians alive today, doesn't know math history?
If Tao has a knowledge of the topic (which he does), then it isn't by virtue of being a mathematician per se, but by virtue of an interest in the history of mathematics (which he has). Knowledge of math is enormously helpful here, but it does not imply historical knowledge.
> Soldiers rarely know the history of war, and war historians are rarely soldiers.
I mean, besides the empty platitude that we have no reason to assume applies here, we can easily search and find Tao commenting on the history and philosophy of mathematics.
It's not at all empty. Specific expertise does not imply general knowledge, and shouldn't be assumed.
"Commenting on" does not equal historian of mathematics aware of all the past politics and drama from 100 years ago. There's no reason to assume an expert in applied mathematics is also a historian, especially when their statements are somewhat in conflict with history. And, there's nothing wrong with that, if they have a knowledge gap in some boring past bickering when they're one of the best in the world at what they do, ffs.
Knowing the math that was developed through history, and knowing how that math was developed and the circumstances around it are two fundamentally different things.
Clearly Tao knows the former, but apriori that does not imply he knows the latter.
Not saying he doesn't, just saying one does not imply the other.
Even if you go back and read the original papers, you'll miss all that which happened beyond the page.
reply