Here is a nice nugget from the paper I am reading: "The double descent phenomenon seen in neural networks [39], in which, as model size increases, test loss goes down, then up due to overfitting and then down again, has been proposed as an example of emergence through the breaking of scaling..."