On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc.
I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones:
iPhone 18 Pro iPhone 17 Pro M3 Pro Mac
cold predict cold predict cold predict
Split-einsum attention, 4 layers 53.4 s 21.6 ms 0.8 s 111.4 ms 1.0 s 105.8 ms
Q.K einsum only 5.1 s 24.0 ms 1.2 s 36.8 ms 1.4 s 41.8 ms
Q.K einsum + softmax 36.1 s 61.5 ms 1.7 s 141.0 ms 2.0 s 128.3 ms
Scores keys-last (softmax last axis) 10.5 s 109.0 ms 0.6 s 198.4 ms 0.7 s 180.2 ms
scaled_dot_product_attention 1.4 s 140.4 ms 0.4 s 184.9 ms 0.5 s 148.1 ms
Cached loads take 0.01–0.03 s everywhere.
What I observed:
Compile time grows linearly with the number of attention layers.
Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed.
During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help.
Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit).
The specialization cache doesn't survive reinstalling the app.
In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes).
Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices.
Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0.
Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.
3
4
49