That's what I tend to do, but since foldr/foldl' is so ubiquitous in Haskell it would be nice if I could just remember the argument order of the callback. kccqzy's explanation (in particular "it replaces the comma") might just help me do that :)
Essentially all of economic theory is aimed at explaining, not predicting. The distinction between the two goals [0] is sometimes under-appreciated within the profession, and almost always under-appreciated outside of it.
Most predictive tools in economics and finance have “surprisingly” little economic content; but once you understand the distinction between the two goals, it should be unsurprising that predictive models tend to make few economic assumptions, relying rather on general statistical techiniques or on econometrics that incorporate a minimum of theory [1]. From that understanding comes the humbling realization that predicting the future is quite difficult in a context in which the relevant processes are continually seeking an equilibrium that often implies unpredictability. [2]
I’m not an economist, but I do a lot of applied financial-economic modeling. State-of-the-art LLMs are really, really terrible at economic intuition. They will hinder, not help, in formulating an economic model, which is a process of coming up with a set of modeling assumptions that lead to a useful (implicitly, tractable) model. LLMs are, however, quite good at math, and I’ve found them very useful in iterating through different sets of modeling assumptions to identify those that lead somewhere useful. Not having to work out all of the mathematical details myself, and thereby avoiding getting lost in the weeds and being better able to maintain a higher-level perspective on what I’m trying to accomplish, has accelerated my work immensely. But it’s a process of leading the LLM by the nose the whole time and asking it to fill in the details.
I should note, thought, that if you indotend “AI” to mean more than LLMs, them yes, there is starting to be a lot of good work done on predictive economic models that use specialized neural networks as black-box functions to compute model quantities that are otherwise difficult to come up with, just as is also happening in applied physics and other fields.
1. Many explanatory economic models refer to quantities that are fundamentally or practically unobservable or unidentifiable. Much of economics is built on models that were designed to provide a formal, logical basis for understanding the economic world, which is often quite unintuitive. (For example, many intelligent people uneducated in economics exhibit intuitions opposite of basic economic ideas like opportunity cost or comparative advantage.) Models of this sort have been very influential in determining the trajectory of economic thought, but they are often effectively impossible to calibrate to the real world.
2. The most influential and effective economic ideas fall into a third class: ideas that have created their own reality by shaping the way people think in a way that gives rise to the results the models explain or predict. This phenomenon is most evident in finance, where ideas like the various forms of the efficient market hypothesis, the CAPM, and the Black–Scholes model and its follow-one have arguably provided a framework that has reshaped the ways financial practitioners behave to such an extent that financial markets now conform much more closely to what the models describe than was formerly the case. Donald MacKenzie’s book An Engine, Not a Camera is an excellent study of this phenomenon: https://mitpress.mit.edu/9780262633673/an-engine-not-a-camer...
>Much of economics is built on models that were designed to provide a formal, logical basis for understanding the economic world, which is often quite unintuitive.
I disagree. It's purposefully unintuitive.
>(For example, many intelligent people uneducated in economics exhibit intuitions opposite of basic economic ideas like opportunity cost or comparative advantage.)
Most people don't believe in comparative advantage. They believe in something that economists can explain away as comparative advantage.
All unconsumed fixed size investments will result in something that is mathematically the same as comparative advantage. This is the intuitive view that people have. You go to university and get a 5 year degree. Now your cost basis for work that suits your expertise is much lower than for work that is out of expertise. A worker buys an expensive machine, now the cost basis for hiring the guy with the machine is lower than buying your own machine.
This also explains why specialization emerges: All specialization is basically a form of an investment that has some residual left over results that can be monetized in the future. If there was no residual it would be as if you forgot your education and at that point the investment is fully consumed and you turn back into a non-specialized worker.
All of this is incredibly intuitive, but economists instead insist on an invisible "factor" [0] to drive efficient production.
[0] The "factor" concept implies comparative advantage exists first rather than emerges as a result of past decisions.
Yes, Ricardo originally introduced the idea of comparative advantage in the context of international trade, in a model in which different countries had different endowments of resources.
Nothing that imtringued wrote above suggests he understands comparative advantage, which is the idea that it is relative productivity, not absolute productivity, that should determine what one specializes in. That’s precisely what I meant about people finding the concept unintuitive.
I think the joke is that Air Marshal Kenneth Porter was chief of the RAF’s maintenance command, and also not a pilot. [0]
Edit: fergal_reid’s reply below is correct. Porter did serve several years at the start of his career as a pilot, in the early 1930s, before spending the bulk of it as a signals officer. Holden also had his wings, but never served as a pilot. I read Porter’s comment as a joke between two RAF officers who were not “really” pilots, but now I’m not sure how to read it.
No biological process is 100% precise. DNA copying being imperfect is the driver of evolution, which is the prime example of this. What I assume they're referring to is that their ribosomes just "make fewer mistakes" when creating proteins from mRNA.
These "mistranslated proteins" are normally infrequent and benign, and their material is recycled eventually. But that whole process wastes energy, so making a "better ribosome" makes all cellular processes more energy efficient and allows the energy budget to be reallocated.
The obvious solution is to run things in reverse, inputting the AI-generated output to recover the prompt that generated it.
Most generative models can be run in reverse by algorithms that already exist [0], but you have to have the model weights. For closed-weight models, or for a process that can handle unknown models, you’d have to do some engineering.
But do we have the technology to build models that back out the prompt from suspected AI output? Yes.
0. I don’t mean that most neural networks are invertible functions. They’re not. But you can do backprop in reverse, from output to input, to train a model to generate an input to the original model that best predicts its output.
Right, that’s why I wrote, “I don’t mean that most neural networks are invertible functions.”
For a neural network that is not bijective, you can obtain an input that maps to a desired output by the following algorithm.
1. Start with a trained neural network. (The weights will not change throughout this procedure.)
2. Pick a random input.
3. Given an output for which you want to compute an associated input, feed the input into the network to compute the output.
4. Compute the loss of the computed output relative to the target output (e.g., mean-square error). If the loss is sufficiently small, you’ve found an input that maps to an output close to your target output and you’re done.
5. Otherwise, compute the gradient of the loss with respect to the input (e.g., by backprop).
6. Update the input according to a gradient-update rule. And go back to Step 3.
In theory, you can recover a “representative” prompt for the output of an LLM in this manner. For outputs that could have been generated by a large set of disparate prompts, obviously this won’t work well.
The inflection point was 2012, when AlexNet [0], a deep convolutional neural net, achieved a step-change improvement in the ImageNet classification competition.
After seeing AlexNet’s results, all of the major ML imaging labs switched to deep CNNs, and other approaches almost completely disappeared from SOTA imaging competitions. Over the next few years, deep neural networks took over in other ML domains as well.
The conventional wisdom is that it was the combination of (1) exponentially more compute than in earlier eras with (2) exponentially larger, high-quality datasets (e.g., the curated and hand-labeled ImageNet set) that finally allowed deep neural networks to shine.
The development of “attention” was particularly valuable in learning complex relationships among somewhat freely ordered sequential data like text, but I think most ML people now think of neural-network architectures as being, essentially, choices of tradeoffs that facilitate learning in one context or another when data and compute are in short supply, but not as being fundamental to learning. The “bitter lesson” [1] is that more compute and more data eventually beats better models that don’t scale.
Consider this: humans have on the order of 10^11 neurons in their body, dogs have 10^9, and mice have 10^7. What jumps out at me about those numbers is that they’re all big. Even a mouse needs hundreds of millions of neurons to do what a mouse does.
Intelligence, even of a limited sort, seems to emerge only after crossing a high threshold of compute capacity. Probably this has to do with the need for a lot of parameters to deal with the intrinsic complexity of a complex learning environment. (Mice and men both exist in the same physical reality.)
On the other hand, we know many simple techniques with low parameter counts that work well (or are even proved to be optimal) on simple or stylized problems. “Learning” and “intelligence”, in the way we use the words, tends to imply a complex environment, and complexity by its nature requires a large number of parameters to model.
Thanks for posting a through and accurate summary of the historical picture. I think it is important to know the past trajectory to extrapolate to the future correctly.
For a bit more context: Before 2012 most approaches were based on hand crafted features + SVMs that achieved state of the art performance on academic competitions such as Pascal VOC and neural nets were not competitive on the surface. Around 2010 Fei Fei Li of Stanford University collected a comparatively large dataset and launched the ImageNet competition. AlexNet cut the error rate by half in 2012 leading to major labs to switch to deeper neural nets. The success seems to be a combination of large enough dataset + GPUs to make training time reasonable. The architecture is a scaled version of ConvNets of Yan Lecun tying to the bitter lesson that scaling is more important than complexity.
Comparing Deep Learning with neuroscience may turn out to be erroneous. They may be orthogonal.
The brain likely has more in common with Reservoir Computing (sans the actual learning algorithm) than Deep Learning.
Deep Learning relies on end to end loss optimization, something which is much more powerful than anything the brain can be doing. But the end-to-end limitation is restricting, credit assignment is a big problem.
Consider how crazy the generative diffusion models are, we generate the output in its entirety with a fixed number of steps - the complexity of the output is irrelevant. If only we could train a model to just use Photoshop directly, but we can't.
Interestingly, there are some attempts at a middle ground where a variable number of continuous variables describe an image: <https://visual-gen.github.io/semanticist/>
If you think a 2 year old is doing deep learning, you're probably wrong.
But if you think natural selection was providing end to end loss optimization, you might be closer to right. An _awful lot_ of our brain structure and connectivity is born, vs learned, and that goes for Mice and Men.
Why not both? A pre-trained LLM has an awful lot of structure, and during SFT, we're still doing deep learning to teach it further. Innate structure doesn't preclude deep learning at all.
There's an entire line of work that goes "brain is trying to approximate backprop with local rules, poorly", with some interesting findings to back it.
Now, it seems unlikely that the brain has a single neat "loss function" that could account for all of learning behaviors across it. But that doesn't preclude deep learning either. If the brain's "loss" is an interplay of many local and global objectives of varying complexity, it can be still a deep learning system at its core. Still doing a form of gradient descent, with non-backpropagation credit assignment and all. Just not the kind of deep learning system any sane engineer would design.
I don't know what you mean by end to end loss optimization in particular, but if you mean something that involves global propagation of errors e.g. backpropagation you are dead wrong.
Predictive coding is more biologically plausible because it uses local information from neighbouring neurons only.
Modern systems like Nano Banana 2 and ChatGPT Images 2.0 are very close to "just use Photoshop directly" in concept, if not in execution.
They seem to use an agentic LLM with image inputs and outputs to produce, verify, refine and compose visual artifacts. Those operations appear to be learned functions, however, not an external tool like Photoshop.
This allows for "variable depth" in practice. Composition uses previous images, which may have been generated from scratch, or from previous images.
> If only we could train a model to just use Photoshop directly, but we can't.
It is probably coming, I get the impression - just from following the trend of the progress - that internal world models are the hardest part. I was playing with Gemma 4 and it seemed to have a remarkable amount of trouble with the idea of going from its house to another house, collecting something and returning; starting part-way through where it was already at house #2. It figured it out but it seemed to be working very hard with the concept to a degree that was really a bit comical.
It looks like that issue is solving itself as text & image models start to unify and they get more video-based data that makes the object-oriented nature of physical reality obvious. Understanding spatial layouts seems like it might be a prerequisite to being able to consistently set up a scene in Photoshop. It is a bit weird that it seems pulling an image fully formed from the aether is statistically easier than putting it together piece by piece.
> If only we could train a model to just use Photoshop directly, but we can't.
They're obviously more general purpose but LLMs can also be used to drive external graphics programs. A relatively popular one is Blender MCP [1], which lets an LLM control Blender to build and scaffold out 3D models.
Indeed. I would add a third factor to compute and datasets: the lego-like aspect of NN that enabled scalable OSS DL frameworks.
I did some ML in mid 2000s, and it was a PITA to reuse other people code (when available at all). You had some well known libraries for SVM, for HMM you had to use HTK that had a weird license, and otherwise looking at experiments required you to reimplement stuff yourself.
Late 2000s had a lot of practical innovation that democratized ML: theano and then tf/keras/pytorch for DL, scikit learn for ML, etc. That ended up being important because you need a lot of tricks to make this work on top of "textbook" implementation. E.g. if you implement EM algo for GMM, you need to do it in the log space to avoid underflow, DL as well (gorot and co initialization, etc.).
I think your post may have more acronyms than any other post I have ever read on hn. Do you have a guide to which specific things you are talking about with each acronym? Deep Learning and Machine Learning are obvious but some of the others I can’t follow at all - they could be so many different things.
> but I think most ML people now think of neural-network architectures as being, essentially, choices of tradeoffs that facilitate learning in one context or another when data and compute are in short supply, but not as being fundamental to learning.
I feel like you are downplaying the importance of architecture. I never read the bitter lesson, but I have always heard more as a comment on embedding knowledge into models instead of making them to just scale with data. We know algorithmic improvement is very important to scale NNs (see https://www.semanticscholar.org/paper/Measuring-the-Algorith...). You can't scale an architecture that has catastrophic forgetting embedded in it. It is not really a matter of tradeoffs, some are really worse in all aspects. What I agree is just that architectures that scale better with data and compute do better. And sure, you can say that smaller architectures are better for smaller problems, but then the framing with the bitter lesson makes less sense.
> Intelligence, even of a limited sort, seems to emerge only after crossing a high threshold of compute capacity. Probably this has to do with the need for a lot of parameters to deal with the intrinsic complexity of a complex learning environment.
Real intelligence deals with information over a ludicrous number of size scales. Simple models effectively blur over these scales and fail to pull them apart. However, extra compute is not enough to do this effectively, as nonparametric models have demonstrated.
The key is injecting a sensible inductive bias into the model. Nonparametric models require this to be done explicitly, but this is almost impossible unless you're God. A better way is to express the bias as a "post-hoc query" in terms of the trained model and its interaction with the data. The only way to train such a model is iteratively, as it needs to update its bias retroactively. This can only be accomplished by a nonlinear (in parameters) parametric model that is dense in function space and possesses parameter counts proportional to the data size. Every model we know of that does this is called "a neural network".
That’s not a meaningful technical obstacle. If you wanted to, you could just take the output of the model and use it at each iteration of the training phase to perform (badly) whatever task the model is intended to do.
The reason noone does this is you don’t have to and you’ll get much better results if you first fully train and then apply the best model you have to whatever problem. Biological systems don’t have that luxury.
> I think most ML people now think of neural-network architectures as being, essentially, choices of tradeoffs that facilitate learning in one context or another when data and compute are in short supply, but not as being fundamental to learning.
Is this a practical viewpoint? Can you remove any of the specific architectural tricks used in Transformers and expect them to work about equally well?
I think this question is one of the more concrete and practical ways to attack the problem of understanding transformers. Empirically the current architecture is the best to converge training by gradient descent dynamics. Potentially, a different form might be possible and even beneficial once the core learning task is completed. Also the requirements of iterated and continuous learning might lead to a completely different approach.
> Even a mouse needs hundreds of millions of neurons to do what a mouse does.
Under the very light assumption that a mouse doesn’t have neurons it doesn’t need, a mouse needs whatever number of neurons it has to do what a mouse does, so that’s not saying much.
That page also says 71 million for the house mouse. So what is it that a mouse does that reptiles do not do that requires them to have that much larger a brain? Caring for their children?
Mice seem to have quite a good representation of the 3d environment around them and motor skills. I had one in my flat run off an jump through an approx 1 x 2 inch hole 6 inches off the ground and about 10 inches from where it jumped from. Humans would probably have a job with that and I've not seen a lizard say seem to have similar ability to know its way around.
I daresay I don't think animals actually need some number or neurons. There's probably just a trade off between more giving better results versus being heavier and more energy consuming.
Mice do a hell of a lot more socialization than lizards, and mammalian socialization is more complex per individual (more competition, feinting, theory-of-mind-like strategies) than the eusocial insect strategies of "my body is the swarm, I just happen to be the limb I have direct control over".
> The conventional wisdom is that it was the combination of (1) exponentially more compute than in earlier eras with (2) exponentially larger, high-quality datasets (e.g., the curated and hand-labeled ImageNet set) that finally allowed deep neural networks to shine.
I'd thought it was some issue with training where older math didn't play nice with having too many layers.
Sigmoid-type activation functions were popular, probably for the bounded activity and some measure of analogy to biological neuron responses. They work, but get problematic scaling of gradient feedback outside their most dynamic span.
My understanding of the development is that persistent layer-wise pretraining with RBM or autoencoder created an initiation state where the optimization could cope even for more layers, and then when it was proven that it could work, analysis of why led to some changes such as new initiation heuristics, rectified linear activation, eventually normalizations ... so that the pretraining was usually not needed any more.
One finding was that the supervised training with the old arrangement often does work on its own, if you let it run much longer than people reasonably could afford to wait around for just on speculation contrary to observations in CPU computations in the 80s--00s. It has to work its way to a reasonably optimizable state using a chain of poorly scaled gradients first though.
> Even for billion-parameter theories, a small amount of vectors might dominate the behaviour.
We kinda-sorta already know this is true. The lottery-ticket hypothesis [0] says that every large network contains a randomly initialized small network that performs as well as the overall network, and over the past eight years or so researchers have indeed managed to find small networks inside large networks of many different architectures that demonstrate this phenomenon.
Nobody talks much about the lottery-ticket hypothesis these days because it isn’t practically useful at the moment. (With the pruning algorithms and hardware we have, pruning is more costly than just training a big network.) But the basic idea does suggest that there may be hope for interpretability, at least in the odd application here or there.
That is, the (strong) lottery-ticket hypothesis suggests that the training process is a search through a large parameter space for a small network that already (by random initialization) exhibit the overall desired network behavior; updating parameters during the training process is mostly about turning off the irrelevant parts of the network.
For some applications, one would think that the small sub-network hiding in there somewhere might be small enough to be interpretable. I won’t be surprised if some day not too far into the future scientists investigating neural networks start to identify good interpretable models of phenomena of intermediate complexity (those phenomena that are too complex to be amenable to classic scientific techniques, but simple enough that neural networks trained to exhibit the phenomena yield unusually small active sub-networks).
Sandvault [0] (whose author is around here somewhere), is another approach that combines sandbox-exe with the grand daddy of system sandboxes, the Unix user system.
Basically, give an agent its own unprivileged user account (interacting with it via sudo, SSH, and shared directories), then add sandbox-exe on top for finer-grained control of access to system resources.
Means a lot coming from you - thanks for taking the time to post, and for taking the time to make the Homebrew formula. (I am also a fan of the author's (webcoyote's) other work.)
OK, let’s survey how everybody is sandboxing their AI coding agents in early 2026.
What I’ve seen suggests the most common answers are (a) “containers” and (b) “YOLO!” (maybe adding, “Please play nice, agent.”).
One approach that I’m about to try is Sandvault [0] (macOS only), which uses the good old Unix user system together with some added precautions. Basically, give an agent its own unprivileged user account and interact with it via sudo, SSH, and shared directories.
I use KVM/QEMU on Linux. I have a set of scripts that I use to create a new directory with a VM project and that also installs a debian image for the VM. I have an ./pull_from_vm and ./push_to_vm that I use to pull and push the git code to and from the vm. As well as a ./claude to start claude on the vm and a ./emacs to initialize and start emacs on the vm after syncing my local .spacemacs directory to the vm (I like this because of customized emacs muscle memory and because I worry that emacs can execute arbitrary code if I use it to ssh to the VM client from my host).
I try not to run LLM's directly on my own host. The only exception I have is that I do use https://github.com/karthink/gptel on my own machine, because it is just too damn useful. I hope I don't self own myself with that someday.
I'm mainly addressing sandboxing by running stuff in Claude Code for web, at which point it's Anthropic's problem if they have a sandbox leak, not mine.
It helps that most of my projects are open source so I don't need to worry about prompt injection code stealing vulnerabilities. That way the worst that can happen would be an attack adding a vulnerability to my code that I don't spot when I review the PR.
And turning off outbound networking should protect against code stealing too... but I allow access to everything because I don't need to worry about code stealing and that way Claude can install things and run benchmarks and generally do all sorts of other useful bits and pieces.
Containers here, though I don't run Claude Code within containers, nor do I pass `--dangerously-skip-permissions`. Instead, I provide a way for agents to run commands within containers.
These containers only have the worker agent's workspace and some caching dirs (e.g. GOMODCACHE) mounted, and by default have `--network none` set. (Some commands, like `go mod download`, can be explicitly exempted to have network access.)
I also use per-skill hooks to enforce more filesystem isolation and check if an agent attempts to run e.g. `go build`, and tell it to run `aww exec go build` instead. (AWW is the name of the agent workflow system I've been developing over the past month—"Agent Workflow Wrangler.")
This feels like a pragmatic setup. I'm sure it's not riskless, but hopefully it does enough to mitigate the worst risks. I may yet go back to running Claude Code in a dedicated VM, along with the containerized commands, to add yet another layer of isolation.
The interesting thing in that thread is how many people have landed on isolation as a workaround while still lacking a real control plane on top of it. Containers reduce blast radius, but they don’t answer approvals, policy, or auditability. That’s the gap I keep seeing in these setups. I've found a software, called Daedalab, that instead of sandboxing AI puts deterministic control on agents actions.
The sandboxing options are set when you connect the MCP to the agent, not by the agent passing params about its own sandbox.
There’s a misconception about the right security boundary for agents. The agent code needs secrets (API keys, prompts, code) and the network (docs, other use cases). Wrapping the whole agent in a container puts secrets, network access, and arbitrary agent cli execution into the same host OS.
If you sandbox just the agent’s CLI access, then it’s can’t access its own API keys/code/host-OS/etc.
But I use that specifically to run 'user-emulation' stories where an agent starts in their own `~/` environment with my tarball at ~/Downloads/app.tar.gz, and has to find its way through the docs / code / cli's and report on the experience.
There's an intermediate step, which is to use a combination of claude code sandboxing (bubblewrap), plus some pre tool hooks to look for sketchy commands, but it's still interactive and probably not the right longterm approach.
reply