Post

Replies

Boosts

Views

Activity

iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro
On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc. I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones: iPhone 18 Pro iPhone 17 Pro M3 Pro Mac cold predict cold predict cold predict Split-einsum attention, 4 layers 53.4 s 21.6 ms 0.8 s 111.4 ms 1.0 s 105.8 ms Q.K einsum only 5.1 s 24.0 ms 1.2 s 36.8 ms 1.4 s 41.8 ms Q.K einsum + softmax 36.1 s 61.5 ms 1.7 s 141.0 ms 2.0 s 128.3 ms Scores keys-last (softmax last axis) 10.5 s 109.0 ms 0.6 s 198.4 ms 0.7 s 180.2 ms scaled_dot_product_attention 1.4 s 140.4 ms 0.4 s 184.9 ms 0.5 s 148.1 ms Cached loads take 0.01–0.03 s everywhere. What I observed: Compile time grows linearly with the number of attention layers. Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed. During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help. Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit). The specialization cache doesn't survive reinstalling the app. In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes). Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices. Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0. Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.
3
4
49
5h
Core AI on iOS/macOS 27: seven issues found while porting an SDXL pipeline
We ported a complete Stable Diffusion XL (Lightning) pipeline to Core AI which we previously had on Core ML. The pipeline comprises text encoders, a ViT-H image encoder, ControlNet, the UNet and the VAE decoder. It works very well: on iPhone 17 Pro and M1 iPad Pro it generates 1.4–1.6x faster than our Core ML pipeline at the same image quality. However, along the way, we hit seven issues. Each one is filed with a small self-contained repro. We're posting them here together for visibility, with our workarounds, in case others run into the same problems. Ahead-of-time compilation (coreai-build compile) Each of these gives wrong results on the Neural Engine with no error. The same .aimodel specialized on the device is correct. FB25067476: A 3×3 fp16 conv with a large output returns a wrong bottom half (16.7 dB vs 60.7 dB for the top half). This depends on the chip: 256×512×512 fails on M3 Pro and A19 Pro, while 256×384×384 fails only on M1. Our workaround: split such convs by output channels so that no single output exceeds about 128×384×384. FB25067497: A nearest 2× upsample followed by a 3×3 fp16 conv returns output unrelated to the correct result (correlation 0.02) on all three chips. The upsample alone and the conv alone are correct. Our workaround:* write the upsample as repeat_interleave(4) + pixel_shuffle(2). FB25067530: A 3×3 stride-2 conv with 8-bit palettized weights is wrong (16 dB) on all three chips. The same conv with fp16 weights, or palettized with stride 1, is correct. Our workaround:* a 2×2 pixel unshuffle followed by a 2×2 conv. Caching and loading FB25066222: The cache for ahead-of-time compiled .aimodelc files is keyed by the outer program only (main.hash), not the weights, so a weights-only model update silently runs the old weights. On iOS, the Neural Engine program cache is also keyed by function name per app, and survives AIModelCache.deleteAll() and deleting the app. Our workaround: append a hash of the model's content to every function name. FB25066664: On macOS, ahead-of-time builds of transformer models load 3–4× slower than on-device specialization, and parallel loads finish one at a time. Precompiling therefore made our Mac first launch slower (1152 s vs about 650 s), not faster. Neural Engine runtime (Mac) FB25067019: The built-in GroupNorm returns wrong results on the Mac's Neural Engine for large inputs (GroupNorm(32, 256) on 1×256×256×256: 21.7 dB), while the GPU and CPU are correct (84 dB). Our workaround: compute the means and variance explicitly. FB25064649: A ViT-H vision tower with an 8-bit palettized 14×14/stride-14 patch conv aborts the process on the Mac's Neural Engine (MPSGraph ANE error -19). Our workaround: keep the patch conv in fp16, or use a pixel unshuffle followed by a 1×1 conv. Environment macOS 27.0.1 (26A434), iOS/iPadOS 27.0.1 (24A446) Xcode 27.0, Metal Toolchain 27.1.266.1 (coreai-build 3600.83.1), coreai-torch 0.4.3 Tested on M3 Pro, iPhone 17 Pro and M1 iPad Pro We're happy to provide more data. We'd also like to hear whether others see the ahead-of-time issues on other chips.
0
3
64
2d
iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro
On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc. I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones: iPhone 18 Pro iPhone 17 Pro M3 Pro Mac cold predict cold predict cold predict Split-einsum attention, 4 layers 53.4 s 21.6 ms 0.8 s 111.4 ms 1.0 s 105.8 ms Q.K einsum only 5.1 s 24.0 ms 1.2 s 36.8 ms 1.4 s 41.8 ms Q.K einsum + softmax 36.1 s 61.5 ms 1.7 s 141.0 ms 2.0 s 128.3 ms Scores keys-last (softmax last axis) 10.5 s 109.0 ms 0.6 s 198.4 ms 0.7 s 180.2 ms scaled_dot_product_attention 1.4 s 140.4 ms 0.4 s 184.9 ms 0.5 s 148.1 ms Cached loads take 0.01–0.03 s everywhere. What I observed: Compile time grows linearly with the number of attention layers. Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed. During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help. Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit). The specialization cache doesn't survive reinstalling the app. In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes). Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices. Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0. Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.
Replies
3
Boosts
4
Views
49
Activity
5h
Core AI on iOS/macOS 27: seven issues found while porting an SDXL pipeline
We ported a complete Stable Diffusion XL (Lightning) pipeline to Core AI which we previously had on Core ML. The pipeline comprises text encoders, a ViT-H image encoder, ControlNet, the UNet and the VAE decoder. It works very well: on iPhone 17 Pro and M1 iPad Pro it generates 1.4–1.6x faster than our Core ML pipeline at the same image quality. However, along the way, we hit seven issues. Each one is filed with a small self-contained repro. We're posting them here together for visibility, with our workarounds, in case others run into the same problems. Ahead-of-time compilation (coreai-build compile) Each of these gives wrong results on the Neural Engine with no error. The same .aimodel specialized on the device is correct. FB25067476: A 3×3 fp16 conv with a large output returns a wrong bottom half (16.7 dB vs 60.7 dB for the top half). This depends on the chip: 256×512×512 fails on M3 Pro and A19 Pro, while 256×384×384 fails only on M1. Our workaround: split such convs by output channels so that no single output exceeds about 128×384×384. FB25067497: A nearest 2× upsample followed by a 3×3 fp16 conv returns output unrelated to the correct result (correlation 0.02) on all three chips. The upsample alone and the conv alone are correct. Our workaround:* write the upsample as repeat_interleave(4) + pixel_shuffle(2). FB25067530: A 3×3 stride-2 conv with 8-bit palettized weights is wrong (16 dB) on all three chips. The same conv with fp16 weights, or palettized with stride 1, is correct. Our workaround:* a 2×2 pixel unshuffle followed by a 2×2 conv. Caching and loading FB25066222: The cache for ahead-of-time compiled .aimodelc files is keyed by the outer program only (main.hash), not the weights, so a weights-only model update silently runs the old weights. On iOS, the Neural Engine program cache is also keyed by function name per app, and survives AIModelCache.deleteAll() and deleting the app. Our workaround: append a hash of the model's content to every function name. FB25066664: On macOS, ahead-of-time builds of transformer models load 3–4× slower than on-device specialization, and parallel loads finish one at a time. Precompiling therefore made our Mac first launch slower (1152 s vs about 650 s), not faster. Neural Engine runtime (Mac) FB25067019: The built-in GroupNorm returns wrong results on the Mac's Neural Engine for large inputs (GroupNorm(32, 256) on 1×256×256×256: 21.7 dB), while the GPU and CPU are correct (84 dB). Our workaround: compute the means and variance explicitly. FB25064649: A ViT-H vision tower with an 8-bit palettized 14×14/stride-14 patch conv aborts the process on the Mac's Neural Engine (MPSGraph ANE error -19). Our workaround: keep the patch conv in fp16, or use a pixel unshuffle followed by a 1×1 conv. Environment macOS 27.0.1 (26A434), iOS/iPadOS 27.0.1 (24A446) Xcode 27.0, Metal Toolchain 27.1.266.1 (coreai-build 3600.83.1), coreai-torch 0.4.3 Tested on M3 Pro, iPhone 17 Pro and M1 iPad Pro We're happy to provide more data. We'd also like to hear whether others see the ahead-of-time issues on other chips.
Replies
0
Boosts
3
Views
64
Activity
2d