This was the argument made by the proponents of LLMs as a route to AGI and there seems to be some force to it - but the models don't seem to have worked out how words are encoded as letters, even if that might have been the best way to comprehend elements of the corpus.