Possibly. I personally think it's the type of data and scale that're the primary differentiators. The use of characters is a fundamental flaw because characters are synthetic entities. Instead the models should be based on raw sensory data types, such as pixels and waveforms, and iterate from there on something close to the existing architecture.
The implementation details are not clear, not the goals.
I never said that the feature has to be coded explicitly. I said it has to be there.