Now at ~27 tok/s

Nick Fisher

So far I've managed to double Gemma 4 E2B throughput to a respectable 23 tok/s on my Mac Mini. That lags the MPS/GPU implementation, but there ANE is very power efficient, so it's a much lighter weight pathway for inference.