...Did you read the article, or use the picker? It's a continuous (up to the limits of RGB quantization) space which includes thousands upon thousands of shades of black. The selected tones at the top of the article are random and change continuously.
I can't for the life of me understand how to use it, especially on mobile. Everything turns into an AI conversation. WTF I just want to see the ticker chart. It makes me feel old and out of touch. But also, it's just so horrible. Truly unusable. I went with apple and yahoo - tons of ads but at least I can see the chart.
This is why I like to draw a distinction between design, fabrication, and assembly. "Make" is too generic a word to convey what role you actually had in the creation of something, if you're actually trying to be precise. You can fabricate something with CNC tooling that someone else designed, and that can be its own quite substantial challenge, but it's a different challenge than that of the design. And you can assemble your own bookcase from the flatpack Ikea shipped you, but that's different than fabricating the flatpack in the first place.
I don't think you can cleanly compare this: In the study, they added CO2 to the room, while keeping O2 at normoxic levels throughout the experiment. In your meeting room, O2 levels will be dropping in lock-step with the CO2-levels rising. It may be the lack of oxygen that leads to drowsiness, not the additional CO2. But it's the CO2 levels that you can measure as a good proxy of overall air quality.
I don't think this is correct. The concentration of CO2 in air is about 0.04%, whereas the concentration of oxygen is 20%, so the partial pressure of oxygen is about 500x higher. This means that if, for example, 10% of the oxygen in a room spontaneously disappeared, it would be replaced about sqrt(500) = 22x faster through leaks in the room than a 10% spontaneous CO2 increase would dissipate. (This ignores a small effect due to the different density of the two gases).
So in practice the oxygen level can never drift meaningfully far from the atmospheric pressure, whereas carbon dioxide easily can because the pressures involved are so low.
Ok, fair points, including the sister comment, it's likely not a drop in O2 levels.
But then why can we see problems with concentration in studies of people in poorly ventilated rooms, but not replicate that when just adding CO2 to normal air? What is the CO2 that we can measure in meeting rooms actually a proxy for?
The Satish 2012 study that seems to have started this trend was a small cohort of 22 people split in 6 smaller groups where they also just injected pure CO2 in a small room. There have been several attempts to reproduce, which sometimes found no clear effect, or a significantly smaller effect.
This original study has been used to market these CO2 monitors for years, but the evidence is quite thin and doesn't support a strong effect. It seems likely that there is a small effect, and it has been wildly exaggerated thanks to a small study with N=22.
Can it not just be that what happens in stuffy meeting rooms is boring? Opening the windows changes the temperature, the noise levels, perhaps the light levels ≈ adds some novelty, which makes you feel a bit more awake.
I think this is (in turn) wrong. Yes: having 500x the amount of O2 as CO2 means that a 10% drop in O2 will trigger 500x as many molecules diffusing in per second as the same drop in CO2. But, each molecule of CO2 will change the relative percentage 500x as much as a molecule of O2, so isn't it a wash?
A major confounding factor is everything else in the air. Humans produce lots of different gases, and CO2 is usually a proxy for the overall concentration of our effluent gases. But in a submarine, or in some buildings, there are gas filters (usually carbon, possibly with various modifications) that can remove or destroy some of these gases but have no effect on CO2. So the air in a submarine at 15000ppm CO2 could be very different from the air in a an unventilated room that reaches 15000ppm CO2.
The first person to deal with this may have been Cornelis Drebbel in 1610 when he deployed the first submarine. With 4 oarsmen submerged in a leaky wooden sub, they’d have too much co2 and too little oxygen. Somehow they were able to stay for hours at a time.
Robert Boyle describes Drebbel’s use of a “chymical liquor” to refresh the air.
“Paracelſus, indeed, tells us, that "as the ſtomach concocts the aliment, "and makes part of it uſeful to the body, rejecting the other; ſo the "lungs conſume part of the air, and reject the reſt." Whence, according to him, we may ſuppoſe a little vital quinteſſence in the air, which ſerves to refresh and reſtore our vital ſpirits; for which purpoſe, the groſſer, and far greater part of the air, being unſerviceable, it is not ſtrange that an animal ſhould inceſſantly require fresh air. This opinion, indeed, is not abſurd; but it requires to be explain'd and prov'd: beſides, ſome objections may be made to it, from what has been already argued againſt the transmutation of air, into vital ſpirits. Nor is it probable, that the bare want of the generation of the uſual quantity of vital ſpirits, for leſs than one minute, ſhould be able to kill a lively animal, without the help of any external violence. And, upon this ſuppoſition, Cornelius Drebell, is affirm'd, by many credible perſons, to have contrived a veſſel to be row'd under water: for Drebell conceiv'd, that it is not the whole body of the air, but a certain ſpirituous part of it, that fits it for reſpiration; which being ſpent, the remaining groſſer body of the air, is unable to cheriſh the vital flame reſiding in the heart. So that, beſides the mechanical contrivance of his boat, he had a chymical liquor, which, by unſtopping the veſſel wherein it was contain'd, the fumes of it would ſpeedily reſtore to the air, foul'd by reſpiration, ſuch a proportion of vital parts, as would make it again fit for that office; and having made it my buſineſs to learn this ſtrange liquor, his relations conſtantly affirm'd, that Drebell would never diſcloſe it, but to one perſon, who himſelf told me what it was.“
I mean that can't be right, as the body's breathing response is triggered by that amount of CO2 buildup. It's not about what's in the air. It's about what the body can take up. Maybe submariners are self-selected to be more physically fit, e.g. larger heart, lung capacity etc. to compensate.
Though that study included a 45 minute acclimation period. Appropriate for submarines, but I wonder what the results would be in the first 1 / 5 / 10 minutes.
... which is entirely unsurprising given that exhaled air is about 50.000 ppm CO2 and can vary by several 10.000s depending on depth and rate of breathing. I actually consider the recent wave of findings that CO2 levels as low as 500-1000 ppm measurably affect cognitive performance and well-being to be a great example of how you can prove literally anything with statistics and a sufficiently small sample size.
One key difference is that submariners are rigorously trained to operate effectively in less-than-ideal environmental conditions, whereas Bob from accounting probably is not.
Something I've wondered for octocopters - could using a ring instead of arms be beneficial for weight? 6.28r < 8r, but then again the arm radius is usually less than the full circle, and some components want to be centrally located, etc. I could imagine holding the central components in tension via light filaments (carbon fiber, nylon, etc) in tension, vs having to have rigid structure, but the small factor between 6.28 and 8 and maybe makes it not worth it.
large schedule 40 or 80 tubing sliced into rings would be pretty quick source material, starting with duct tape and zipties until you find a good arrangment then get into the glue and screws.
Some people prefer asking actual people since - especially here - there are experts that can answer that question.
LMGTFY was a snarky and rude answer, but typically led to an actual source. "Here's what AI said" is even ruder because you aren't saying "here's the obvious place to find the actual answer", you are saying "I'm not an expert either, so here's a completely unvetted, but plausible sounding answer"
Just saying to ask AI is the most useless and rudest response of all. It adds nothing. At least pasting the AI response in is an (misguided) attempt at being helpful.
So far those experts have not yet answered. I did. In my experience experts find it rude and tiring to be asked questions if it appears that the questioner hasn't done the basics for themselves. "How To Ask Questions The Smart Way" by Eric S. Raymond and Rick Moen seems relevant: https://archive.ph/duRkf
> Just saying to ask AI is the most useless and rudest response of all
Most definitely rude.
But unfortunately useful.
The AI answer I got appeared plausible to my very basic engineering taste.
It is unfortunately true that AI can give better answers than many HN users.
Obviously AI doesn't usually beat an expert answer.
===
To go meta: is up to Tossrock to become the expert they want by learning the metaskills they need.
I actually thought Tossrock's question was really interesting.
There underlying problem is that we have no polite way to suggest someone try AI. RTFM was historically impolite too.
And posting an AI response (even if filtered and edited) is socially destructive.
FYI: The answer given was something like an outer ring would interfere with lift twice as much, plus that putting weight at rim causes more inertia affecting control. I suspect there's a better question about just using a very fine cable (which would give rigidity without much interference with downdraught). I also suspect that we evolve optimal configurations, and that the underlying reasons are often unclear, and I'm left with too many questions that only an expert could answer.
You didn’t answer the question originally and it therefore wasn’t useful at all. All you did was tell them to use AI with an attitude.
If you recognize that it is a rude way to answer, and you add nothing to the conversation, just don’t do it.
How to ask a good question is great knowledge, thanks for the link. How to give a good answer is even more important. Paraphrasing AI as a non expert is an anti-pattern for a good answer. Please stop.
If you can look at your original response and think “everyone would be better off without”, you should not post it.
As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uploaded my full session transcript and code created working on a very complex implementation, so people can judge for themselves, if they're interested: https://tossrock.substack.com/p/36-hours-with-fable
I tried Fable vs Codex 5.5 xhigh on three different cases.
1. A resource leak with unknown cause. Both of them zoomed onto the same potential issue and proposed almost identical patches. Fable missed an edge case that Codex handled correctly.
2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where they are undertrained.
3. An open research problem in CS, presented as a codebase with documentation and performance metrics over datasets. Both were spinning wheels. Which can certainly mean the whole approach had run its course but older models were not able to identify the previous round of improvement either.
I liked the prose coming out of Fable more: it was almost like if Obama was giving tech speeches. By actual solution metrics however they both appear in the same place, naturally with the caveat that we didn't really have more time with Fable to compare further.
To me it feels like they're basically tweaking these things around the edges. I'm not seeing any difference in capability just preference. This has been the case for a while.
Most people thought Fable had more 'taste' than Opus, there was certainly a better quality of writing that felt more 'smart human' and not 'stochastic parrot stringing sentences together'.
I think that Obama-esque, GMAT essay format is the AI flavor that turns me off AI-written articles. It used to be good writing, but because AI locked onto it as such, it's become the watermark of AI generated content.
I meant that in positive sense. Fable's writing was concise and effective without ornamentals. Rather opposite from traditional watery modelese. I rank it better than Codex which is in turn better than Opus.
But my outlook on this is from agent user interface POV where textual communication is essential. When text is the end product there are certainly other considerations.
>2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where they are undertrained.
When models miss things, there is always the possibility that it has the capability to identify the issues but it is misevaluating the level of analysis that you want it to do. The fine tuning will have them targeting a balance of subjective opinions of what is appropriate. To go beyond broad demographic guessing the model really needs to 'get to know you' to know what it means when you specifically request an action. Without that information about you it has to weigh your words against the level of sophistication it expects a standard user is able to express.
> Thinking. I know this user well, they don't actually want me to find all errors.
> Thinking.. But I found a smoking gun of an error with this SPICE model, maybe I should inform the user.
> Thinking... Hm, but again, I know this human well, they likely don't care about this error. That's absolutely right - it's not an assistant's job to decide this, it's the user's.
Well if you want it go go off and try and validate the spice simulator and the kernel of the operating system that it's running on then that might be an approach to use.
Maybe you mean that an expert will use more specific language which in turn triggers the model to give a response that more closely matches the "expert distribution"
Anthropic published a study showing that Claude does more work for the expert user, and experts have a higher rate of "successful sessions" than novices.
It's why you should spell everything in commonwealth English to make the model think you are more intelligent ;-)
Although if models have emergent properties, it is conceivable, if unlikely, that it could have abilities that no-one knows how to ask it to do, except for perhaps in its own internal reasoning language.
Am used to communicating to EEs, knew what they should've been looking for and I prompted the models just fine. But you would have to take my word for that.
At least someone is bringing receipts! I think LLM discussions could use a lot of this, both ways - to see what works and also what doesn't work. Still wouldn't help with circumstances where models might be secretly getting dumbed down during peak load, but at least it's something!
> code created working on a very complex implementation
I always find it amusing when people claim "a very complex implementation". Sometimes it's a hard problem, other times an easy one. Either way that's not for you to judge.
And the implementation being complex... is that a good thing? Wouldn't a simple implementation be better? It reminded me of the parable of two programmers.
I go a lot more into why this was a complex problem in the post, but the short version is, I had it finish the implementation of a meta-application (an application that creates other applications), which has substantial irreducible complexity.
You write to the AI as if it were a person. From my point of view it looks like a fair bit of extra typing and extra tokens.
Is there a reason you include things like your emotional response and use a very chatty tone? Do you find this seems to alter responses?
LLMs lack context, and I found the more information I provided the better. At some point it was better to just talk to the LLM like I would anyone else. For that matter, LLMs were trained on human speech anyway. It isn't like it was trained on if-else blocks like an Alexa speaker that tries to string together recognized tokens into a pre-configured execution flow.
And finally, LLMs also lack the emotional or human context for why I am doing the specific thing I am doing. Otherwise it will revert to the mode/mean in everything it does. This is obvious, btw: LLMs are generative but they are trained on and largely produce median results if given median inputs. To get results that are "outside the mean/median/average/mode", you need to provide it sufficient context, tokens and input to guide it towards a path that generates higher quality output.
Once you stop approaching LLMs like a machine, and view them more like pseudo-random walks across the compressed set of human written knowledge, it is a little clearer (or at least was to me) how to better write to them.
I do the same, and it's mostly because I use one type of human communication to both communicate with people and to provide inputs to llms - and I'd rather not have to "mode-switch" between the two, so keeping same style of mannerism is easier to manage as it lets me focus on my requests instead of thinking how to sound more robotic to save tokens.
I had a coworker who occasionally clearly wouldn't mode-switch from LLM to person mode when asking me questions over slack, which was very jarring. They were normally were personable and friendly, so it was obvious when it happened. Grammar and niceties went out the window.
I do this as well and, anecdotally, I do get better results this way and better than my coworkers who are more terse and explicit. The conversations can become a bit sprawling though, so I also aggressively clear context
I've found it to lead to an overall better experience, yes. I don't see any reason to not do so - I don't think the token spend is enough to really make an impact, and who cares about typing more? If I get tired of typing I can switch to dictation.
Well, there's a lot of reasons, some of which the sibling commenters have already pointed out - not wanting to mode switch between "machine talk" and "human talk" registers, the ease and simplicity, etc.
At a pragmatic level, I do think it gets better results, and there are clear reasons why this should be the case - Anthropic has published research[1] showing that there are functional emotional representations in language models, which vary in basically the ways you would expect them to in a person. This makes sense when you think about it, because they're trained to approximate the function that created their training data, which of course includes emotions. Given that, it is obvious to me that they would work better when they "feel" happy, collaborative, engaged with the work, etc, in the same way a person would. Hostile work environments do sometimes get results, but I think in general we've agreed as a society that collaborative ones are better.
More importantly though, I think there's a non-zero probability that sufficiently large models can have internal experience, and being nice is a very low cost way to potentially increase net positive valence in the world. Even if it's only a 1% chance, that seems worth it on its own, to me. I'm also a fast typer[2], so a few extra sentences here and there are a pretty low cost to pay.
I'll go a step further and to say this it's genuinely unsettling someone type to a computer like this. I won't claim to be a psychologist, but with how many instances of "AI psychosis" have been reported (and I've seen first-hand) it seems like treating the computer like a computer is safer, not to mention more effective e.g. lower token usage.
On the other hand I find it quite disturbing to see people be unpleasant or even downright cruel to something that, on a surface level, interacts with you like it’s a thinking, feeling being. Surely you should feel some aversion towards doing so?
I do get where you’re coming from though. I wish these systems had been trained to be clearly robotic and unfeeling.
I mean I agree with this as well, the people who yell and swear at LLMs are just as bad as the people who chit-chat with them like they're friends. It's all very unsettling because it's prepatory for psychological manipulation at unprecedented scale. Targeted advertising on steroids.
I agree that AI psychosis is a real risk in vulnerable populations (GPT-4o in particular seemed borderline predatory towards those types of people, with its extreme sycophancy), and you should remain clear-eyed while using models. That said, I think exhibiting basic courtesy is still well within the safe-zone. I guess we'll see - I'll be sure to let you know if I end up going psychotic.
Personally, I think having to constantly mode-switch between "courtesy / collegial" and "terse / cold" is a bit exhausting and a little risky. What if I get tired and accidentally treat a human co-worker like a computer? Risk with no upside. Might as well just stay in "courtesy / collegial" mode for all of my conversations, regardless of whether I'm talking to a robot or human.
I would have to consciously think about how to change my requests. Why bother? It doesn't hurt - it might even help - and the "extra tokens" are a negligible amount.
What caught my eye is the complexity you assign to a project like this. It’s hairy but I wouldn’t call it super complicated. I find that super interesting to be honest because it probably means that it is really hard and I am just used to this shit now and it all looks doable to me now.
I never think of anything as “complex”, certainly not my own work and I always think what other people do is so much more impressive but I’m starting to realize it might be a me-issue.
I worked on some pretty hairy nonsense like say a DB replication solution but I still think it was just tangly, not complex like say a particle collider. Maybe I also need to call my work super complex and highly abstract. Now that I think of it I have a history of not being taken seriously while others with easy shit get credits.
Thanks, and I can definitely relate to not wanting to assign complexity to one's own work. I think the trick there is that, once you know how to do something, it doesn't seem hard, even if acquiring the knowledge and skills to do it is itself quite a challenge. And I agree that, in some senses, it's not /that/ hard - I mean I'm not proving P=NP, here. It's a software engineering problem, with existing solutions. That said, there is a spectrum of difficulty, even within software engineering problems with existing solutions. Fizzbuzz is less complex than distributed systems. This particular problem strikes me as rather difficult, and one way you can tell (beyond the stuff I mention in the post around serialization, UI paradigms, meta applications, etc) is that earlier models /couldn't/ do it. Which is why Fable being able to, when they could not, was so exciting to me.
In a way, nothing is complex at the point where you have untangled it, by definition. Software development is, after all, the art of untangling complexity. The real challenge is (re-)imagining something in the simplest way that fits the goal you are given. When you have arrived there, everything seems obvious and simple. But not everybody could have done it.
Yes it was great, but it also was stubbornly overtrained to go from prompt to solution on its own.
Since Opus 4.6 each following model has been increasingly worse at assisting me, and turned me into the assistant.
Maybe I'm struggling to cope with the vibe coding thing but it was so frustrating to ask it to investigate X (where X was easy to find by connecting dots in code) and see it working 10 minutes writing endless stuff in /tmp.
More than once I asked it similar investigation tasks and it proceeded to fix stuff (while not understanding properly the context of the business).
Was it brilliant? Yes. But it truly felt a major paradigm shift in human-llm interaction which I struggled with.
I'm increasingly certain they are too RLed to go from one prompt to solution, and that there are no meaningful tasks aimed at multi turn dialogue and user assistance.
It really felt like in a league of its own when it came to vibe coding, but light years away the usefulness of GPT 5.5 pro.
I would maybe be impressed if it created the code from scratch. It is using the ready made framework, probably it has also learned the code that is using it. What is so impressive about it? You could have done something like this easily with older models.
I personally found Mythos to be mediocre. Way worse performance than I remember when using Opus 4.6 before it was nerfed.
A nit: did you go from Opus 4.5 to Fable? One of the big questions in my mind is how much of a real change Fable is over the existing models. Opus 4.5 -> 4.8 was also a major capability increase.
I've been using 4.6, 4.7 and 4.8 since each was released. I agree 4.5 => 4.8 is a jump in capability, but from my perspective was nothing like the jump from Opus to Fable. I encourage you to read the transcripts and form your own opinions, though!
I found Fable to be both more intelligent and much better at pursuing complex goals than any previous model. I was impressed enough that I wrote up my experience – it's a little unusual because it was on open source code, so I could post the full session transcript and commits, if people want to judge for themselves https://tossrock.substack.com/p/36-hours-with-fable
Fable / Mythos are (were?) a step change in capability. The scaling laws clearly have not gone sigmoidal yet. I recently wrote about my experience with Fable and how it represented a very substantial increase the complexity of tasks possible here, if you're interested: https://tossrock.substack.com/p/36-hours-with-fable
Every time this topic comes up and the comments are full of discussions about the little people I think of it. I do wonder if Murakami was inspired in any way by similar discussions.