I urge you to reconsider your beliefs. You are missing something important because you are thinking in terms of low-dimensional statistics. Deep learning doesn't just fit data, it finds features (abstractions) of the data.
I asked the original commenter to confirm my understanding of what they were saying.
I find it a fascinating alternative view to what is largely well understood (your counter-point).
The thing that stood out is the comment that it’s an approximation of the original data generator (humanity). Early approximations were poor (GPT 2-3, to an extent GPT-4).
I’m not so sure I can reject the hypothesis that such an approximation can be found.