Ten FSKit issues found building a network file system module (all filed with minimal repros)

While building an SMB 2/3 client as an FSKit file system module (FSUnaryFileSystem + FSVolume, the macOS 27 Handler protocols), I ran into a number of framework-level issues. I have filed each one with a title starting "FSKit:" so they are easy to find, and every report has a minimal reproduction attached: a small in-memory FSKit module (no network, no disk, no cache of its own), so none of them depend on SMB. All were measured on macOS 27.0 (26A5406e and 26A5416b) with Xcode 27.0 beta 5.

Summaries below in case anyone else is hitting these.

FB24419773: renameatx_np with RENAME_SWAP returns success but destroys the destination file. On any FSKit volume a RENAME_SWAP is performed as an ordinary clobbering rename: rc=0, but the destination's contents are silently lost instead of exchanged. The module cannot refuse it because renameItem receives no flags; a swap and a plain overwriting rename look identical. (RENAME_EXCL works correctly.)

FB24419825: a negative lookup is cached permanently. Once anything gets ENOENT for a name on an FSKit volume, the kernel serves that ENOENT for the life of the vnode. If the file is created later (for example by another machine on a network volume), it stays unopenable by that name indefinitely, while ls of the same directory lists it. There is no API through which a module can report that a name now exists.

FB24419858: a data-cache grant from openItem can be applied after the module has already invalidated. The grant in FSOpenItemResult is applied asynchronously after the module's reply, and an invalidation issued in that window succeeds (setCacheState returns no error) and is then overwritten by the stale grant. The result is a kernel cache no future event will invalidate; readers see stale data.

FB24419870: synchronize(flags:) is never called on a URL-backed volume. fsync(2), fcntl(F_FULLFSYNC), fcntl(F_BARRIERFSYNC) and sync(8) all return success with zero calls reaching the module, so durability is reported and never established. A packet capture of the same SMB share shows five SMB2 FLUSH requests through Apple's smbfs and zero through an FSKit module.

FB24419894: FSItemSetAttributesRequest.consumedAttributes is never observed, and wasAttributeConsumed(.changeTime) answers about the wrong attribute. Consuming everything and consuming nothing are indistinguishable to the caller (chmod returns 0 either way), even though the setAttributes documentation says the upper layers will detect unsupported attributes. Separately, wasAttributeConsumed answers YES for changeTime when only accessTime was consumed, and never answers correctly about changeTime itself; this part reproduces by constructing the request directly, no file system needed.

FB24419911: restrictsOwnershipChanges = true does not reject non-superuser chown. The property is documented as "the volume rejects a chown(2) from anyone other than the superuser", but on an -o owners mount a non-root chgrp is delivered to the module's setAttributes anyway, so every module has to enforce the policy itself.

FB24419932: a failed activate wedges the resource URL. After a module's activate throws once, every later mount of the same URL string fails with "Resource busy" (fskitd logs "Can't start new task, resource state is 5"), while the same volume under a different URL spelling mounts fine. For a network module the ordinary trigger is one wrong password. Recovery requires killing both fskitd and the extension process.

FB24419964: enumeration cannot report extended-attribute presence. FSItem.Attributes has no per-item "has xattrs" field, so one cold ls -l of a 500-entry directory costs about 2,000 FSKit boundary crossings: an xattr call per entry plus a "._name" AppleDouble sidecar lookup per entry, and each of those ENOENTs is then pinned by FB24419825. Suggestion: a per-entry hasExtendedAttributes flag so getattrlistbulk can be satisfied from the enumeration.

FB24419974: no byte-range lock operations. flock(2) and fcntl(2) locks on an FSKit mount stay kernel-local and never reach the module, so advisory locks cannot coordinate between clients of a network file system. Suggestion: an optional lock-operations handler.

FB24419979: no ACL or security descriptor operations. ls -le, chmod +a, acl_get_file(3) and cp -p with ACLs cannot work on any FSKit volume; a network server's real ACLs are invisible behind synthesized mode bits. The nearest surface, FSVolumeAccessCheckHandler, can only be asked yes/no questions about a descriptor the module has no way to provide. Suggestion: an optional ACL-operations protocol.

If any of these are biting you too, duplicate feedbacks referencing the FB numbers above genuinely help with prioritization.

Answered by DTS Engineer in 902954022

While building an SMB 2/3 client as an FSKit file system module (FSUnaryFileSystem + FSVolume, the macOS 27 Handler protocols)

To start with, as a general comment, the reality is that FSKit isn't currently at a place where it can properly support a network file system. It's certainly closer this year than it was last year, but it's not really "there" yet. As usual, I can't talk about our future plans, but the reality is that FSKit's basic function is to replace the in-kernel VFS layer, a task which it clearly cannot yet do. For most of the bugs I don't specifically mention, the answer boils down to "yes, that's a VFS feature FSKit does not yet support".

Looking at a few issues that I wanted to directly address:

FB24419773: renameatx_np with RENAME_SWAP returns success but destroys the destination file.

This is entirely a bug. In the long run, FSKit needs an API to support this properly, but until then (and the default behavior), VOL_CAP_INT_RENAME_SWAP should be 0, which would mean RENAME_SWAP never reached your extension.

FB24419894: FSItemSetAttributesRequest.consumedAttributes is never observed, and wasAttributeConsumed(.changeTime) answers about the wrong attribute. Consuming everything and consuming nothing are indistinguishable to the caller (chmod returns 0 either way), even though the setAttributes documentation says the upper layers will detect unsupported attributes.

First off, to clarify, "unsupported" and "failed" are totally different cases, with the behavior of unsupported attributes being varying depending on the attribute. For many attributes, the unsupported behavior is that the operation succeeds without actually changing anything.

Separately, wasAttributeConsumed answers YES for changeTime when only accessTime was consumed, and never answers correctly about changeTime itself; this part reproduces by constructing the request directly, no file system needed.

...and, yes, that's a bug.

FB24419858: a data-cache grant from openItem FB24419825: a negative lookup is cached permanently. FB24419932: a failed activate wedges the resource URL.

Those are bugs as well.

FB24419911: restrictsOwnershipChanges = true does not reject non-superuser chown. The property is documented as "the volume rejects a chown(2) from anyone other than the superuser", but on an -o owners mount a non-root chgrp is delivered to the module's setAttributes anyway, so every module has to enforce the policy itself.

...and so is this, though it (r.177650457) appears to actually be a complicated side effect of doing the pathconf man page ("return 1") says instead of what POSIX requires ("return _PC_CHOWN_RESTRICTED").

FB24419870: synchronize(flags:) is never called on a URL-backed volume. fsync(2), fcntl(F_FULLFSYNC), fcntl(F_BARRIERFSYNC) and sync(8) all return success with zero calls

FYI, the complication here is that both fcntl calls (and I think the other two as well) are actually built on the generic VFS IOCTL system. Implementing that generic mechanism for FSKit is slightly broken/strange, since the main problem it solves (the VFS layer has a difficult time communicating with user space) is a problem FSKit simply does not have. Putting that another way, if you want to send commands to your FSKit extension, routing those commands through the kernel is a bit silly. FSKit will need to address this issue, but it’s going to require more work than it would otherwise appear.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Accepted Answer

While building an SMB 2/3 client as an FSKit file system module (FSUnaryFileSystem + FSVolume, the macOS 27 Handler protocols)

To start with, as a general comment, the reality is that FSKit isn't currently at a place where it can properly support a network file system. It's certainly closer this year than it was last year, but it's not really "there" yet. As usual, I can't talk about our future plans, but the reality is that FSKit's basic function is to replace the in-kernel VFS layer, a task which it clearly cannot yet do. For most of the bugs I don't specifically mention, the answer boils down to "yes, that's a VFS feature FSKit does not yet support".

Looking at a few issues that I wanted to directly address:

FB24419773: renameatx_np with RENAME_SWAP returns success but destroys the destination file.

This is entirely a bug. In the long run, FSKit needs an API to support this properly, but until then (and the default behavior), VOL_CAP_INT_RENAME_SWAP should be 0, which would mean RENAME_SWAP never reached your extension.

FB24419894: FSItemSetAttributesRequest.consumedAttributes is never observed, and wasAttributeConsumed(.changeTime) answers about the wrong attribute. Consuming everything and consuming nothing are indistinguishable to the caller (chmod returns 0 either way), even though the setAttributes documentation says the upper layers will detect unsupported attributes.

First off, to clarify, "unsupported" and "failed" are totally different cases, with the behavior of unsupported attributes being varying depending on the attribute. For many attributes, the unsupported behavior is that the operation succeeds without actually changing anything.

Separately, wasAttributeConsumed answers YES for changeTime when only accessTime was consumed, and never answers correctly about changeTime itself; this part reproduces by constructing the request directly, no file system needed.

...and, yes, that's a bug.

FB24419858: a data-cache grant from openItem FB24419825: a negative lookup is cached permanently. FB24419932: a failed activate wedges the resource URL.

Those are bugs as well.

FB24419911: restrictsOwnershipChanges = true does not reject non-superuser chown. The property is documented as "the volume rejects a chown(2) from anyone other than the superuser", but on an -o owners mount a non-root chgrp is delivered to the module's setAttributes anyway, so every module has to enforce the policy itself.

...and so is this, though it (r.177650457) appears to actually be a complicated side effect of doing the pathconf man page ("return 1") says instead of what POSIX requires ("return _PC_CHOWN_RESTRICTED").

FB24419870: synchronize(flags:) is never called on a URL-backed volume. fsync(2), fcntl(F_FULLFSYNC), fcntl(F_BARRIERFSYNC) and sync(8) all return success with zero calls

FYI, the complication here is that both fcntl calls (and I think the other two as well) are actually built on the generic VFS IOCTL system. Implementing that generic mechanism for FSKit is slightly broken/strange, since the main problem it solves (the VFS layer has a difficult time communicating with user space) is a problem FSKit simply does not have. Putting that another way, if you want to send commands to your FSKit extension, routing those commands through the kernel is a bit silly. FSKit will need to address this issue, but it’s going to require more work than it would otherwise appear.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Thanks for the response, Kevin! I appreciate you taking the time to address these issues, even though I might be disappointed that FSKit isn't yet up to the task.

The current FSKit documentation points to things like "a network resource you identify with a URL" or "some sort of network address for a remote file system" or "to network connections, and beyond." Maybe an update to the documentation would steer would-be adopters appropriately. For me, it sounds like I'll just have to be patient.

The current FSKit documentation points to things like "a network resource you identify with a URL" or "some sort of network address for a remote file system" or "to network connections, and beyond."

Yeah, it's a tricky one. The challenge here is that FSKit is at a point where it is genuinely useful, even in a network context. For example, I know of one company with a very large[1] source code base that uses it as the front end to their source code archive system. FSKit lets the client present a coherent view of resources that are actually stored in a different hierarchy, and most of the issues you've raised don't matter because the access is read-only and relatively static.

Of course, that's very different than smb where the file system is a “complete” solution and the main reason to pay for a third-party solution is that you're supporting functionality that our driver doesn't.

[1] Large enough that managing everything through a single version control system wouldn't really be viable, particularly since much of the code is old enough that multiple control systems have come and gone.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

The challenge here is that FSKit is at a point where it is genuinely useful, even in a network context.

Indeed, and I can see that based on your example. Thanks again for engaging here and for providing very useful responses.

For anyone following a similar path, my attempt to serve SMB over NFS: NFSv4.1: racing open/unlink/recreate of the same filename can leave processes unkillable.

Thanks for sharing your findings in this thread. I've been building an FSKit module for a source control filesystem and I hit many of the same things. I have some findings to share related to FB24419825 (the negative lookup that stays pinned for the lifetime of the directory vnode) in case they might be helpful, and also for visbility because if we end up depending on these workarounds, they would need to stay usable until FB24419825 is fixed.

A workaround that I found is to revoke the parent directory instead of the file. setCacheState(for:cacheMode:coherencyType:action: .revoke) on the directory that contains the pinned name drops the directory's child dentries, including the pinned negative. Revoking the file itself doesn't work, the call fails with kIOReturnBadArgument (-536870206). As far as I can tell, revoke only succeeds on items the kernel has already looked up, and a file behind a pinned negative can never be looked up, since the pin answers every lookup for that name. The pin belongs to the parent directory, which the kernel has looked up, so revoking the parent works.

That workaround doesn't work at the volume root, revoking the root always fails with kIOReturnBadArgument, meaning a file created at the root after something looked for it stays unopenable by name for the life of the mount, even though it shows up in listings.

However, I came across a second workaround here. Opening the pinned name with O_CREAT|O_EXCL will clear the pin. This holds on 27.2, including at the root (unlike the first workaround). Without O_CREAT, the open fails with ENOENT from the kernel's cache. With O_CREAT, the kernel tries to create the file instead, so it skips lookupItem and sends createItem to the module. The module returns EEXIST because the name now exists in its tree, and after that the name opens normally.

This second workaround has two problems, though. The module has to know exactly which names appeared, and if a name is gone again by the time the open runs, O_CREAT|O_EXCL creates an empty file in its place.

Ten FSKit issues found building a network file system module (all filed with minimal repros)
 
 
Q