A20 pro devices take too long on CoreML model specialization

Hi, I'm a ML developer specializing in local audio related models.

I upgraded from a 16 pro max to an 18 pro max.

I'm not sure what the reason is, but model specialization (happens only on the first load, cpu and neural engine compute units) is sometimes 2x - 5x longer on the 18 pro max vs the 16 pro max. Both devices are on iOS 27.

For example, I have a stem separation model running, and the 18 pro max took 171 seconds (including inference) and the 16 pro max took 37 seconds (including inference). 18 pro max inference is ~55x RTFx and the 16 pro max is ~40 RTFx for this particular model, but I've noticed it as well in TTS models etc. Please look into this, I know the 18 pro max has double the neural engine cores, so I'm not sure if the complexity makes specialization take longer.

Subsequent loads are great, same high rtfx for the 18 PM (only 0.5s load), and this is where it beats the 16 PM handily. I attached 2 images showing the 18 PM on a first load (model specialization likely) vs a subsequent run.

Please look into this. (Note this is not just my implementation, I tested a range of different models, all with the same outcome).

A20 pro devices take too long on CoreML model specialization
 
 
Q