I want to send you my money, but your expensive watch I want to buy is f**ing massive and my wrist is normally sized. I don't really understand why this is still an issue since it's been 4 years.
Thanks,
- Guy with slightly below average wrist circumference (40th %ile)
I just want the extra (vs non-Ultras) button THAT'S ALREADY THERE to allow mapping to "lap". Blew my mind that wasn't possible when I went to try one at the apple store.
Not that I could find in settings... the settings menu let me start a workout but that was all I saw. Maybe it just wasn't discoverable... Just looked at the online user manual and the only mention of lap control is in the stopwatch app (not workouts). It does mention secondary actions, but doesn't elaborate... if that actually does control the lap, Apple really sucks at doc.
:shrug:
Plus, I really want a mini version. The Ultra is MASSIVE to the point it's ridiculous.
Every company gets a limited number of innovation tokens. Where they choose to spend them is up to them. Some companies spend them on the model harness, some, like DS, spend them on the model architecture etc.
Nothing in there contradicts. A business that works well today may not work well tomorrow; it's management's job to change the business to better meet future conditions.
> or do you think that you'll need fewer humans because AI will be doing more of the work?
My bet is this one. To be fair, it isn't exactly an unpopular opinion. If I were starting a business today I'd plan to hire far fewer programmers than I would have just 2-3 years ago.
> Distillation is not illegal by every definition of the word
Note that Anthropic (and USG) alleges [0] not only that Kimi was distilled, but that they actively circumvented measures intended to stop distillation. There are multiple ways that's illegal, including:
- Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
- Economic espionage: 18 U.S.C. §1831 criminalizes obtaining a trade secret through theft, fraud, or deception while intending that it will benefit a foreign entity.
- Trade-secret misappropriation: if Anthropic could argue industrial-scale querying reconstructed proprietary aspects of Fable (like by showing it produces similar outputs, as others have done) then it's illegal under 18 U.S.C. §1832.
- California computer-access statute §502 bars knowingly accessing a computer system and, without permission, taking, copying, or using its data.
- Computer Fraud and Abuse Act protects against the case where restrictions against an activity are circumvented (like Kimi is alleged to have done).
> There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
A lack of prosecution does not make something legal. There is also the scale/commercialization thing, which isn't an issue with random tiny HF datasets/models. Remember: Kimi also sells K3 inference.
> kimi architecture is vastly different than that of fable
How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
> US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Cool. The difference is that one of those things is legal (because they chose to open-source) and one of those things is illegal theft of trade secrets (because it was stolen).
> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
1) this has nothing to do with other labs, just Moonshot (and Z.ai, MiniMax, DS)
2) slandering or not it happens to be completely true, so, there's that
All of the models stole the entirety of written knowledge on the internet to train. They are being sued for the few cases where we have some proof of what they did because of some whistleblowers, all the rest will just go unpunished. They breached Github TOS, robot.txt's, copyright, patents every form of IP protection under the sun from a billion sources. It's just ridiculous for the thieves to cry about someone else stealing from them.
For me, when it specifically comes to copying, I don't think it's bad to copy a copier. (And by that I mean Anthropic has no valid complaints against Moonshot. Any valid complaints from anyone in the original corpus are valid against both of them now.)
In this way, it is different from literal theft. Stealing money/objects from a thief and keeping them is not justified.
It's a little different in this case, since 1) not all the data Ant used was stolen and 2) they did contribute significantly to the value of the stolen good.
An analogy might be a baker stole 20% of the flour used to bake their special bread, which was then stolen. Both thefts are obviously wrong and bad.
The exact same argument could be made in Moonshot's favor.
Being a bit tongue in cheek, one could also argue that by releasing their models, Moonshot is contributing much more value than Anthrophic. Did Prometheus not create an immense amount of value, when he took fire from the hands of the Gods and gave it to humans?
I think any analogy with physical theft is too different from data to apply to this comparatively subtle case. Especially when we get into the details of just using the output of the model to train on.
The baker stole 100% of the flour to make the bread. He also stole the water and the salt and the yeast and the heat for his oven. What he didn't steal was the time he put into crafting a recipe for bread and the time he sat around waiting for the oven to bake it. Now, is that loaf stolen property? Hard to say. But the baker is undoubtedly a thief. He should be tried and forced to pay restitution out of his ill-gotten profits for sure. If we can't do that, the next step is pitchforks and guillotines.
Not OP, but copyright law is an absolute joke. No, I don’t care one whit that someone’s TOS was violated. In fact, I find it hilarious. And it’s not “theft.”
"I think we shouldn't have IP protection at all" is a totally valid position to hold, but that's not the law is. OP said it didn't violate the law, and it does.
Some laws are very obviously unworthy of consideration, with broad consensus from the public. See what happened when Napster came out. Literally no one cares about some red-faced RIAA suit flicking spittle over some shared Metallica albums.
Same thing here. This whole situation is just comical.
It's a rational position to care about the law, but insist on a queue when related parties are involved.
In this case: resolve the theft claims against the US frontier labs, and only then let them make claims against third parties. It would be totally unreasonable for (say) OpenAI to extract a settlement from Moonshot and use that to pay its own claims. Ordering matters.
If the law was applied uniformly, I would support its continued uniform application. In the last 10 years, I don't see it being applied fairly at all, I see an oligarchy, a criminal and corrupt government and rich and powerful entities getting away with anything. The most minimally competent legal system would ask the AI companies, show us the list of all the data you've used to train and lets hash out the copyright - instead we have to pray someone leaks one tiny piece of what they trained on and then sue for that. Open-weights models are the closest thing we have to justice in the world where the legal system no longer provides justice, because at least the model trained on all of our data is given back to all of us.
I find that I care more when copyright violations cause actual harm to the copyright owner. Let's say there's an American kid who can't speak Japanese but wants to keep up with a weekly manga. He downloads a bootleg translation and shares it among his friend group. That is a copyright violation, but meh. If he hadn't gone the illegal route, he'd more likely just not read it at all. There's very little chance he'd pay for a subscription and learn Japanese.
Now, if that kid were to print the bootleg translation and sell it to schoolmates, that's worth a slap on the wrist. The kids willing to pay would likely have paid for official copies.
When these LLM labs download our works, feed them into their models, and sell the output to people that used to pay for our work, that's worth a very hard slap. I honestly have less of a problem with the open models.
I do agree that two wrongs don't make a right, the terms of service generally gives cooperation the power to sever the contract, but it does not make things illegal in the literal sense. The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.
When I said "Does this matter?" I specially meant that distillation in itself, the data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
And what I explicitely pointed out that focusing so much on distillation is an attack on open research and claiming that the majority of advancements are thanks to US labs which is simply not true (at least not anymore this was somewhat true during deepseek R1 era), but that in itself was inspired by open research.
> How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
Because anthropic would be the first ones to make that information public and the architecture is unique to kimi... They made it, they wrote papers on it, it's their research.
P.S. none of the quoted laws apply here since no trade information is stolen, the one about circumventing distillation protection might hold up in court although unlikely.
> The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.
Agree, and this is exactly what Anthropic is alleging.
> data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
It's important to note this is NOT what happened. Anthropic was able to trace data directly back to employees at the company: "We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff."
> none of the quoted laws apply here since no trade information is stolen
There is a lot of work showing Kimi models produce similar outputs to Anthropic models, which constitutes trade information. This is not dissimilar to past and ongoing IP suits against Anthropic and OpenAI by showing the models would recreate images of Mickey Mouse/NYT articles etc.
For the record, I'm a researcher myself and I'm well aware how competent the researchers are at the open-source labs/how much they've contributed. But that's not at issue here, my disagreement with you is specific to your arguments about legality; you're conflating what you think should be legal with what actually is legal.
This is mostly just to reiterate myself as the original question was "Does this matter?"
Everything else is simply justifying why it shouldn't, the specifics don't really matter as there is no legal framework to stop china from continuing to distill models and anthropic has proven they cannot use software solutions to stop it either as distillation is still a problem. But I do still believe it wouldn't hold up in court either way as stopping companies from generating training data which was trained on the entire human knowledge corpus is just stealing from thieves and making it 'open' once again so the argument only gets weaker.
edit: to add, the mickey mouse / nyc was because anthropic trained on LICENSED works, not apple to oranges. The original work it was reciting was licensed and not licensed BY anthropic.
What precedents can you cite and specific examples of their applicability. That is, what would Anthropic's lawyers take to court? You can't say because there's nothing there that couldn't be ripped apart by the least legally capable community known to man, HN. That's why no lab has succeeded in a suit anything like what you're claiming could happen. The only reason Anthropic or any other lab would pursue this is political or commercial. They're either looking for help from officials or they're trying to establish a particular market position.
>Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
This is true, but Kimi also has a variety of defenses. Kimi can't raise unclean hands if Anthropic systematically violated others' terms of use, but it can raise copyright misuse (which is similar in some respects to unclean hands) as well as lack of standing to enforce restrictions in the contract due to the third party beneficiary principle (i.e., Kimi would argue that Anthropic cannot sue Kimi for derived IP that rightfully belongs to third parties whose terms of use were violated by Anthropic, and the proper party to sue Kimi, if any, would be those third parties). That latter argument usually fails in small-scale cases (ProCD) but has been successful in larger ones where the alternative would be anticompetitive.
>"A lack of prosecution does not make something legal"
Plainly who gives a flying fuck. The US can claim whatever rules they want and so can China or any other country. On international level all those rules are artificial constructs unless they can be enforced. China can just say for example that they do not recognize copyrights /patents / whatever so it is "legal" for them.
This is illegal in China too, there's just an enforcement asymmetry. I understand what you're saying is de facto true, I'm just taking issue with people saying either
1) its not illegal (it is)
2) it shouldn't be illegal because Anthropic stole training data (thats not how the law works)
I am a practical man. From what I see laws are mostly for common folks and often do not even serve real justice. The higher one goes and the amount of money / power involved the more the laws bend and on international level the only law that matters is the size of one's club and willingness to use it. And when the country with supposedly biggest one starts crying I find it laughable.
> Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
Ah yes, I remember when Anthropic crawlers abided by the TOS of the websites they slurped up.
All your other points are downstream from this, which makes them pretty tenuous. Labs don't think that ToS or other explicit wishes of content providers apply to them, but they expect everyone else to abide by theirs.
Probably won't have to wait that long. Prism released Bonsai 27B (https://huggingface.co/prism-ml/Ternary-Bonsai-27B-mlx-2bit) as a ternary model a few days ago, its just ~7GB and runs at 44+ t/sec on an m4 max laptop. That's already in the ballpark of active parameter count of most 200B+ models, so we will get a model like this whenever Prism feels like releasing one.
It is debatable if we will actually need that many parameters though, since recursive nets like HRM (https://huggingface.co/sapientinc/HRM-Text-1B) don't need to parametrize as heavily.
We're too easily conflating parameter count with capability. That Bonsai 27B you're running is at 2-bit quantization. Is it really better than the best 10-18B models?
1. yes. https://www.alphaxiv.org/abs/2607.bonsai-27b table 14 shows that bonsai retains roughly 95% of the fp16 27b model's average performance and outperforms post-training quantization at a similar bit width. it doesn't directly compare against every top 10-18b model, but it is clearly still performing like a large model.
2. quantization != native low precision training. a model trained in native ternary should generally outperform a full-precision model quantized after the fact.
even if a ternary model only retains 90-95% of the performance of its fp16 equivalent, who cares? if a 200b ternary model retains most of the capability of the 200b fp16 model while using a fraction of the memory and bandwidth, it can be substantially less efficient per parameter and still dominate a smaller fp16 model under the same hardware budget.
I know that's what the paper says the benchmarks say, but these models feel significantly worse than the base model when you start using them for real tasks.
Even the Q4 quant which they put in between their Bonsai models and the FP16 in the benchmarks has a tendency to go into doom loops and get lost compared to even Q5 or Q6.
I don't know how much of this is due to benchmaxxing (putting the benchmarks into the post-training loop) or cherry picking benchmarks to look good. If you spend a lot of time using local models you learn to take vendor provided benchmarks with a huge heap of doubt. Everything looks amazing in the benchmarks these days.
Models are not like a fps counter where losing a few percentage has no actual impact. These percentages may be the difference between a model that writes code and one that goes into loops, and think rm -rf . is a good idea.
There is a reason why most models try to stay in the FP4 or higher range, because the reduced accuracy can have major consequences.
You are better off with a 8b FP4+ model then a 27b Q2 model.
You're conflating two concepts: native bit width and post-training quantization.
Consider two models: one is 16B and trained natively in 2 bits; one is 8B and trained natively in FP4. These models have the same total number of weight bits, but one has twice as many parameters. There is no real reason the FP4 one should be better just because it's FP4. It might be, but that is an empirical question, not a general rule.
Post-training quantization is another thing entirely. Taking a model trained at higher precision and forcing it down to 2 bits is going to hurt performance, often very badly. But this was never my point.
But do you need to run every small problem through a 10B-30B model?
We're smashing ants with hammers most of the time. We're asking frontier Opus/Fable models to classify text and build frontend code.
Once we start dissecting these problems into smaller discreet tasks and having the big reasoning models do the tough stuff, we suddenly have an economical system. Not for the company hoping for a big IPO, but for the end user.
I do about 10 google search queries for every 1 opus/gpt prompt. For google, I don't actually open pages anymore 9 out 10 times; I rely on the AI summary. It's fast and accurate; the trick is that you learn where the boundary is of what you can ask it. Querying information the small model is great at.
Then there might be slow, batch tasks. I can see myself getting 1T of slow RAM one day (in a few years?) and having a slow onsite GLM5.2 doing batch jobs that would be wasteful of my subscription limits, plus sensitive but boring things, such as bookeeping and general admin.
I'd like to to read all my email and al quarterly reporting. But that would have to be a good local model, probably a model simmilar to whatever google search uses, which seems just correct unless you throw serious challenges at it.
The big hammers buy you more confidence and need less supervision. In pure task execution you _might_ smash the ant with a small surgical hammer, but if you absolutely need it smashed, that's when people still reach for the big hammer. It buys more confidence.
> do you need to run every small problem through a 10B-30B model? ... We're asking frontier Opus/Fable models to classify text
Actually probably yes: text analysis (magazine articles) by LLMs in the ~30b .. ~120b range failed miserably (and also randomly - the rare cases of proper interpretation occurred among the failure cases) with the main public models of around one year ago, tried extensively.
So, yes, you can employ an ~80IQ only if you will expect the related quality.
I don’t think you need a 10-30b model for most smartphone use cases.
But I meant to counter gp’s claim that “I can run a 27b model on an iPhone” is kind of pointless and disingenuous. Yes I’m sure someone will come up with a way to run a “27b model” at 0.1 bit quantization on an Apple Watch pretty soon misses the whole point of saying a model is “27b” in capability.
Achieving a parameter count is not the point. And is almost meaningless
Its somewhat good, the prism team’s webgpu demo gave it a couple dozen “kernels” written in Fable 5 and it calls them for almost everything procedural
I feel like these things are experiencing convergent evolution to be like biological brains. The large parameters are merely potentially large parameters and they keep having more and more and smaller active layers, which are themselves quantized down. This is seems analogous to the chemical spiking of neurons and inactive layers of a brain in power and efficiency.
There’s a good eval floating around somewhere and tl;dr they’re awesome but the benchmarks are cooked, you’re better off with Qwen 8B Q4 than 27B 1b or ternary.
Thanks for being skeptical, I maintain a llama.cpp-based client and it’s frustrating how high expectations are for local AI bc the median effort level means people mostly assemble their expectations and understanding via marketing soundbites
agreed!! in my heart i really wanted to say by the end of 2026 but wanted to add some wiggle room in case they start to ban open source AI development.
The one time in which I saw Juergen Schmidhuber in perfect nervous control, "coolness" they may say westward, was when he replied to one member of the audience, "The same observation was made when they invented fire: oh, it's dangerous. But in the end, now it's here (shrugh)".
There is a proposal in the USA to restrict LLM access. This will only have us depend more and more on open source models and their providers. And cause a drain of research in those areas in which it will be impeded.
The sooner the USG figures out a standard process for approving releases the better. There are many differing opinions on how much to regulate AI, but I think we can all agree ad-hoc policy sucks.
I'm quite curious what Tim Cook's legacy will end up being.
There is no question many of Apple's business experienced significant, impressive growth during his tenure. Amazing capital efficiency.
There is also no question Apple lost product velocity. Few new products were launched, and those that were had mixed success.
Tim was, at the end of the day, an elite financial operator. Apple shareholders were lucky to have him. Customers like myself probably have mixed opinions, and it remains to be seen how he set the company up for the future.
I'm just pointing out product velocity slowed. I'm far from the first person to say it, it's just a fact. In the five years before Cook we got first generation Apple TV, iPhone, iPad, and MacBook Air. Your list spans 14 years.
One could add the Vision Pro, MacBook Neo, Mac Studio, HomePods, and so on to the list as well.
The reality is everyone just wants another hit product like the iPhone, but its success was based on it being a personal convergence device. You can't really create a second carryable/wearable convergence device and expect it to be wildly successful at the level of the iPhone without it killing off the iPhone.
So far that revolutionary approach by third parties has not succeeded against the iPhone, and the evolutionary approach apple takes with the iPhone means there is no clear inflection point anywhere in the future where the phone form factor goes away.
Yes, a very successful CEO and he secured a great legacy. I was skeptical when Jobs stepped down, but under Cook innovation did continue, but primarily in hardware.
> Few new products were launched, and those that were had mixed success.
Tim oversaw the launch of the Apple Watch, Airpods, Airtags, Apple Pay, the Beats acquisition (which lead to Apple Music) and the launch of the M series chips.
He's had quite a few product launches under his belt, many of them company-defining products.
The M series transition was perfectly executed, but that trajectory was set up before Jobs left when they went all-in on in-house semiconductor design.
Apple released their first in-house ARM processor 16 years ago, and the M series is descendent from that lineage and acquisitions that got them started in that business such as PA Semi and Intrinsity.
Cook absolutely deserves credit for the successful desktop ARM transition, but building ARM processors in-house was in no way something he directed as CEO.
Jobs was likely very burned out on IBM failing to deliver a 3Ghz PowerPC G5 and one with a low enough TDP for a PowerBook.
So he switches to Intel because he needs chips, but the vulnerability still exists, and it's what happened again after the Skylake launch and the ensuing 4 years of terrible Macs designed for silicon that didn't exist.
Steve saw the danger, and probably acquired PA Semi because of it as well as the fact that PA Semi actually did deliver a power efficient PowerPC G5, even if it was a bit late.
Steve had the vision. Cook executed it very well. They both deserve credit.
To me, Tim Cook has turned Apple into a company that is both “doing amazingly well” and “in urgent need of a radical change in direction” at the same time.
FaceID, AirPods, Apple Silicon, Vision Pro (though it was flop was a good try). Overall, I would actually place Tim above Steve in terms of business, although maybe not from a Human Computer Interaction design novelty perspective
Tim Cook is a businessman who made the company bigger than Jobs could.
But he is not an aestheticist as much as Jobs was. See how Cook has been destroying the faces of iPhones and Macs, which had a huge dent or what is ironically called a "dynamic" island on the top of the screen. Back of iPhones is desparetely ugly.
Also he has not been presenting what makes us exicted. Apple's Siri is forgotten so that he has to rely on Google's Gemini instead of developing their own. While Samsung's Galaxy has been deploying its 7th foldable phone, Apple has done none. Leaks are usual so we can tell what he will show at its annual conference well before he acutually does and it gives us no surprise at all.
In a short term, "what Cook's Apple has innovated?" -- I guess zero. Rather, deteriorated.
As a long-standing user who started computer life with Performa 5220, keep using Macs as main machines and now run M3 MAX Macbook Pro to develop web apps, current Apple is never what I think it should be.
Making the company bigger is great. But what about their products and services? These are also where Cook has been leading to. He seems to forget Job's aphorism, "Stay hungry, stay foolish."
I don't think this is true. Apple Watch is basically in a market of its own. iPad might have existed before Cook but he turned it into something people actually use for stuff. Vision Pro may not be a financial success but the tech is impressive and it's clear that work will pay off in the near term in other wearables. Apple Silicon is a phenomenal success. Apple TV is no longer a hobby and he's been at the helm while they've developed their entire services business. AirPods rule the headphone market. Not mention the numerous Mac variants he presided over.
Many of their acquired pro tools, and pretty much all of their server hardware and software, though much of that started before Cook took over. Plus the Mac Pro missteps were on his watch, as well as the current cancellation. Apple seems more and more unwilling to invest in niche hardware like the Mac Pro, except where they see it pushing the platform forward, like the Vision Pro.
Methinks this post conflates “rent seeking” with “return on investment” just a tad.
Economic rent is the extra money you can charge for owning a scarce resource. ML models are not waterfront real estate, they are IP. Other people can make more models if they can/want to.
Now, whether IP should be legally protected is a totally separate question, and while we in the West tend to assume the answer is obvious geohot would certainly not be the first person to suggest broadly applying private property rights to information makes questionable sense.
Also, the defining feature of capitalism is that it encloses what was previously common.
Land used not to be owned (feudal lordship was functionally different than private ownership.) Then, society shifted, land became private, and that was the beginning of rent. This is enclosure.
The whole concept of IP is to explicitly extend this process to ideas -- they are not free, they are owned, and I have to pay you to use them. This is also enclosure, precisely.
The "rent" in "rent-seeking" does not refer to "rent" it refers to "economic rent."
Totally different concept. But don't take my word for it:
> "Rent-seeking" is an attempt to obtain economic rent (i.e., the portion of income paid to a factor of production in excess of what is needed to keep it employed in its current use) by manipulating the social or political environment in which economic activities occur, rather than by creating new wealth.[0]
> In economics, economic rent is any payment to the owner of a factor of production in excess of the costs needed to bring that factor into production. [1]
They started working on humanoid robots because Musk always has to have the next moonshot, trillion-dollar idea to promise "in 3 years" to keep the stock price high.
As soon as Waymo's massive robotaxi lead became undeniable, he pivoted to from robotaxis to humanoid robots.
Pretty much. They banked on "if we can solve FSD, we can partially solve humanoid robot autonomy, because both are robots operating in poorly structured real world environments".
Obviously both will exist and compete with each other on the margins. The thing to appreciate is that our physical world is already built like an API for adult humans. Swinging doors, stairs, cupboards, benchtops. If you want a robot to traverse the space and be useful for more than one task, the humanoid form makes sense.
The key question is whether general purpose robots can outcompete on sheer economies of scale alone.
I agree that each would be made slightly better with a more integrated system. But you could handle all of them in my hundred year old house with the form factor it was designed for: a humanoid. Probably pretty soon here for cheaper than each could be handled separately by more integrated systems.
For new builds, a laundry/utility room that includes the dishwashing and other "housekeeping" facilities is a no-brainer when there is a custom robot built to use those facilities as well as maneuver around the rest of the house.
For old/retrofit renovations it also makes sense, but otherwise, yes, a human-form robot makes sense.
The question is which is a better investment for any robot manufacturer in 2026?
The drop in demand for Tesla's clapped out model range would have meant embarrassing factory closures, so now they're being closed to start manufacturing a completely different product. Bait and switch for Tesla investors.
I wonder how long they'll be closed for "modifications" and whether the Optimus Prime robot factories will go into production before the "Trump Kennedy Center" is reopened after its "renovations".
Just reading your description, it sounds like there are two variables:
1. Prompt adherence: how well the models follow your stated strategy
2. Decision quality: how well models do on judgment calls that aren’t explicitly in the strategy
Candidly, since you haven’t shared the strategy, there’s no way for me to evaluate either (1) or (2). A model’s performance could be coming from the quality of your strategy, the model itself, or an interaction between the two, and I can’t disentangle that from what you’ve provided.
So as presented, the benchmark is basically useless to me for evaluating models (not because it’s pointless overall, but because I can’t tell what it’s actually measuring without seeing the strategy).
That's a fair point. You're right that without seeing the strategy, you can't fully disentangle what drives the differences.
But the strategy itself isn't really the point. Since every model gets the exact same prompt and the exact same market data, the only variable is the model. So relative performance differences are real regardless of what the strategy contains. If Model A consistently outperforms Model B under identical conditions, that tells you something meaningful about the model.
And honestly, that blend of prompt adherence and decision quality is how people actually use LLMs in practice. You give it instructions and context, and you care about the result.
You're right though that the strategy being private limits what outsiders can evaluate. It's something I'm thinking about.
To be more specific: the prompt defines a trading philosophy and tells models what to look for in the charts. But the actual read and the decision is entirely on the model. Using your framing — it's closer to "here's inspiration, now maximize money" than "implement this exact strategy."
Which means improvisation within that framework is exactly what's being measured.
I want to send you my money, but your expensive watch I want to buy is f**ing massive and my wrist is normally sized. I don't really understand why this is still an issue since it's been 4 years.
Thanks,
- Guy with slightly below average wrist circumference (40th %ile)