5.2.8 Names and Case

Every string in the LCS interface — key names, value names, hive names, layer names, paths — is UTF-8. Invalid UTF-8 is rejected with EINVAL before parsing, routing, folding, layer resolution or source dispatch. Null bytes are rejected in all of them.

The one thing that is not a string is value data, which is opaque bytes.

5.2.8.1 Lengths are byte counts #

Every configured length limit is measured in UTF-8 bytes, not Unicode scalar values and not display characters. MaxPathComponentLength (default 255) bounds one component or one value or layer name; MaxTotalPathLength (default 16383) bounds a whole path; MaxKeyDepth (default 512) bounds nesting.

A syscall path arrives as a null-terminated C string, is copied under a hard bound, and has its terminator stripped before anything is measured — so the terminator is not part of the length. Ioctl and RSI strings are length-delimited and need no terminator; a terminator byte included in the length is a null byte and is therefore invalid.

5.2.8.2 Separators #

Backslash is canonical. Forward slash is accepted on input and treated as a separator wherever a separator is recognised, so a component can never contain either. There is no string-rewriting step: normalisation is a property of how paths are split rather than a transformation applied to them.

5.2.8.3 Case folding #

Comparison is case-insensitive and storage is case-preserving. The algorithm is Unicode Simple Case Folding — the C and S status entries of CaseFolding.txt, with the full (F) and Turkic (T) entries excluded. It is a fixed one-to-one codepoint mapping with no locale input, applied after decoding from UTF-8, never to raw bytes.

The Unicode version is pinned at 16.0. The table is generated by pkm/tools/lcs/generate_casefold_table.py and checked in with the digest of the source data, so adopting a newer Unicode version means regenerating the table deliberately. It is not something that happens by updating a dependency.

This gives practical compatibility with Windows' RtlCompareUnicodeString without claiming byte-identical behaviour in every edge case.

Unicode normalisation is not performed. NFC and NFD forms of the same visual character are different names, matching Windows.

Case folding is what "identity" means throughout: a layer's identity is its folded name, a hive route's identity includes its folded name, and a duplicate is a folded-equal duplicate. RoleA and rolea are one layer, not two.

Two comparisons in the kernel are ASCII-only rather than folded: the check for whether a layer name is base on two of its call sites, and KACS's duplicate check when parsing private layer names into a token. For the literal string base the two agree; for arbitrary names they do not.

Edit this page