Hi,
I’m training an Object Tracking reference object in Create ML using Extended training mode + All Angles for use with visionOS high-frame-rate object tracking.
The training ran successfully for approximately 48 hours and reached 97.4%. At that point, Create ML automatically paused.
The issue is now completely reproducible: whenever I click Resume, training runs for approximately 3 seconds and then automatically returns to Paused at exactly 97.4%.
I captured the system logs while reproducing the issue. MLRecipeExecutionService crashes with:
LayerVariable.swift:116: Fatal error:
The new value must have the same shape as the current value
([40, 1, 3, 3]), but it has [96, 1, 3, 3].
Immediately afterwards, Create ML reports an interrupted XPC connection to MLRecipeExecutionService.
Shortly before the crash, the service also logs:
Detector training already finished but its loss file was removed from the cache; reporting loss as 0.
and:
Error cleaning up detector images:
NSCocoaErrorDomain Code=4
NSPOSIXErrorDomain Code=2 "No such file or directory"
I would really like to avoid discarding ~48 hours of training if the existing tracker/detector checkpoints can still be recovered.
Thanks!
0
0
38