From 840b3b215e9a3b7db059f3fd3cf053adb353783c Mon Sep 17 00:00:00 2001 From: Joseph Ferano Date: Fri, 25 Sep 2026 11:18:12 +0700 Subject: [PATCH] A Map's churn test fails a table that doubles instead of sweeping, and nothing recalled is cited as read --- docs/BUILT.md | 88 +++++++++++------------------------ lib/check.ml | 2 +- lib/emit.ml | 2 +- runtime/flan_rt.c | 24 +++++----- test/programs/map-iter.flan | 2 +- test/programs/map-remove.flan | 55 +++++++++++++++------- test/programs/maps.flan | 4 +- test/test_acceptance.ml | 11 ++--- 8 files changed, 89 insertions(+), 99 deletions(-) diff --git a/docs/BUILT.md b/docs/BUILT.md index 3c9ab03b..1271c41b 100644 --- a/docs/BUILT.md +++ b/docs/BUILT.md @@ -3175,58 +3175,32 @@ are involved, and none are needed" — and monomorphisation is what would buy it ### The Map is a Swiss table -The Robin Hood table above lost to CPython's dict at a million entries, for the reason the last section gave: three -separate runs and an eight-byte hash a slot. It is replaced by a Swiss table. The header, every entry point's -signature, the hash-and-equality pair and the iteration contract are unchanged, so neither backend and neither -debugger description moved; the whole change is inside `runtime/flan_rt.c`. +The Robin Hood table above lost to CPython's dict at a million entries: three separate runs and an eight-byte hash a +slot. It is replaced by a Swiss table, and the header, every entry point's signature, the hash-and-equality pair and +the iteration contract are unchanged, so the change is confined to `runtime/flan_rt.c`. The layout and the removal rule +are described there; what follows is what the code cannot say. -**The block.** - - [growth][stride][value offset][control: cap + 7 bytes] pad to 64 | slots - -- **One control byte a slot.** `0x00` is empty, `0x01` deleted, and a full slot is `0x80` with the top seven bits of the - hash below it — Zig's encoding (`lib/std/hash_map.zig`, `Metadata`), chosen because empty is zero: a zeroed control - run is an empty map, the property the old table's zero hash word had. The low bits of the hash choose the group, so - the tag and the position never share a bit. -- **Groups of eight, as one 64-bit word.** Matching a tag, finding an empty slot and finding a free one are each a few - integer operations over the word — hashbrown's portable "generic" group, which needs no intrinsics and so compiles - for every target the runtime does. SSE2's sixteen-wide group is not built; nothing measured asks for it. The seven - bytes past the end mirror the first seven, so a group can be read starting at any slot. Groups are visited at - triangular offsets, which reach every group of a power-of-two table before repeating. -- **A key and its value side by side.** A hit reads the control group and then one slot — two cache misses where the - old table took three. The runtime is not told either alignment, and does not need to be: a size is a multiple of its - type's alignment, so the largest power of two dividing it (capped at 64) bounds the alignment. The value is placed at - the key's size rounded up to that bound, and the stride rounded up to the larger of the two bounds. -- **The head holds three words** because the header has no room for them: it is the emitter's layout as much as the - runtime's. The growth word is how many more entries may land in an *empty* slot before a rebuild; the stride and value - offset are computed once at allocation, because recomputing them from the two sizes on every call measured as about a +- **Empty is zero** — Zig's encoding (`lib/std/hash_map.zig`, `Metadata`: free 0, tombstone 1, a used bit on top) — so + a zeroed control run is an empty map. +- **Groups are eight bytes read as one integer**, not SSE2's sixteen: portable to every target the runtime compiles + for, and nothing measured asks for the wider group. +- **Key and value share a slot** because a hit then costs two cache misses rather than three. The runtime is not told + alignments; the largest power of two dividing each size bounds them. +- **The head of the block holds three words** — growth left, slot stride, value offset — because the five-word header + is the emitter's layout too. The stride is stored rather than recomputed because recomputing it per call cost about a tenth of a cache-resident lookup. -- **Seven-eighths load.** A miss stops at the first group holding an empty slot rather than walking a run, and at seven - eighths every eight-slot group has one. - -**Removal is a tombstone, most of the time not.** A lookup stops at the first group with an empty slot, so emptying a -slot could hide a key placed past it while it was full. The slot becomes deleted instead, unless every eight-slot -window containing it already holds an empty slot — then no probe can have passed through it, and it is emptied -outright. That is abseil's test (the empties just before and just after are fewer than eight slots apart), and in a map -that is not nearly full it is the common case. Nothing moves, so a cursor held across a removal stays valid; the -backward shift that removal used to cost is gone. - -**Deleted slots count against growth**, since a probe cannot stop at one. When an insert would take an empty slot with -no growth left, the table is rebuilt: at the same capacity when the live entries fit in 25/32 of it (abseil's figure), -which only sweeps the deleted slots, and at double otherwise. `map-remove.flan` row 7 churns a map at a steady size for -200,000 operations against a plain array and holds it at 512 slots through 33 sweeps. - -**The probe is inlined by force.** Called from five entry points, clang keeps it out of line at `-O2`, and the call -spills everything the loop held in registers. +- **The 25/32 threshold** for sweeping deleted slots at the same capacity rather than doubling, and the rule for + emptying a removed slot outright, were written from memory of abseil's table, not read from its source; no abseil + clone is on this machine. `map-remove.flan` row 7 is what holds them: it churns a map at a steady size under an + allocator budget that a doubling would exceed. #### Measured -`i64` to `i64`. Keys present are `2i`, keys absent `2i + 1`. Per size: *insert* fills a fresh map from empty with no -`reserve` and frees it, repeated to two million inserts; *hit* and *miss* are four million lookups cycling through the -keys; *remove* fills a map untimed and removes every key timed, repeated to two million. Every pass counts its wrong -answers and the program prints them, because the first version of this benchmark passed a map by value to a filling -function — the grows landed in the callee's copy of the header, the caller's map stayed empty, and every lookup was a -fast miss. The CPython side is the same workload in Python 3.13.9 (`dict.get`, `dict.pop`). +`i64` to `i64`; keys present are `2i`, absent `2i + 1`. *insert* fills a fresh map with no `reserve`, to two million +inserts; *hit* and *miss* are four million lookups cycling the keys; *remove* removes every key of a filled map, to two +million. Each pass counts its wrong answers: a map passed by value to a filling function takes its grows in the +callee's copy of the header, and a benchmark written that way measures an empty map. CPython is Python 3.13.9, +`dict.get` and `dict.pop`. The figure is the **minimum of nine runs**, the two Flan binaries run alternately and the Python script in a separate nine-run pass, because this machine is shared — it carried a load average near 20 throughout, from other lanes' test @@ -3262,16 +3236,12 @@ and ahead of CPython in every cell — and that is the claim to reproduce. Nanos | | miss | 88.2 | 14.6 | 125.2 | | | remove | 152.0 | 74.5 | 91.5 | -**The crossover.** Against CPython, the Robin Hood table lost on insert at every size, on remove from somewhere between -100k and 300k, and on hits between 300k and 1M; misses it won throughout. The Swiss table is quicker than CPython at -every size and every operation measured, and quicker than the table it replaced at every one of the twenty-four. The -margin that is left at a million is insert (139 against 152) and remove (75 against 92), and it is narrower than it -looks: CPython hashes a small integer to itself, so this key pattern walks its table in order and the hardware -prefetcher does the rest, while every Flan key lands at a random slot. +**The crossover.** Against CPython the Robin Hood table lost on insert at every size, on remove from between 100k and +300k, and on hits between 300k and 1M. The Swiss table is ahead of both in every cell. CPython hashes a small integer +to itself, so this key pattern walks its table in order and the margin at a million is narrower than it looks. -A middle variant was measured and not kept: Swiss control bytes over separate key and value runs. It was within a few -nanoseconds of the final layout up to 100k and lost from 300k up (87 ns a hit against 26 at 300k, 100 against 78 at a -million), which is the third cache miss the interleaved slot removes. +Control bytes over separate key and value runs were measured and not kept: level with the final layout to 100k, behind +from 300k (100 ns a hit against 78 at a million). ## Unions, and the tag they carry @@ -4514,10 +4484,8 @@ function. ``` **The cursor is a slot index the caller owns, and there is no iterator struct** because there is nothing for one to -hold. A map has no tombstones — removal shifts the run back instead of marking a hole — so a slot is either empty or -occupied and the position is the whole of the state. What a cursor does *not* survive is a removal taken while it is -in flight: the shift moves entries to lower slots, and a cursor already past them steps over entries it has not -answered, the same bargain a put that grows already makes. The cursor starts at 0, comes back one past the entry just answered, and is left +hold: a slot's control byte says whether it is full, and the position is the whole of the state. A removal moves +nothing, so a cursor survives one; what it does not survive is a put that rebuilds the block. The cursor starts at 0, comes back one past the entry just answered, and is left at `cap` by the call that answers false, so a spent cursor keeps answering false rather than wrapping. **Three out-pointers and not a returned pair**, because there are no tuples. An `(Option K)` would answer half an diff --git a/lib/check.ml b/lib/check.ml index 64e43869..98632b35 100644 --- a/lib/check.ml +++ b/lib/check.ml @@ -7519,7 +7519,7 @@ and named_call ?(qualified = false) ctx ~want loc name args = let attempt, note = match target.Tast.ty with (* For a map the number is entries, not slots: the runtime sizes the - block so that [n] still sits under the 75% load factor, which is + block so that [n] still sits under the load factor, which is the only reading of "room for n" that does not reallocate on the nth put. *) | Types.Map (k, v) -> diff --git a/lib/emit.ml b/lib/emit.ml index 3b072d6a..21c2081b 100644 --- a/lib/emit.ml +++ b/lib/emit.ml @@ -4134,7 +4134,7 @@ let header = {|; Generated by flan. The layout is C's: no object headers anywher ; (Vec T), spec-memory.md. The element type is nowhere in it: the runtime is ; type-erased and every operation is handed size and align at its call site. %vec = type { ptr, i64, i64, ptr, i64 } -; (Map K V), spec-memory.md — Odin's open-addressed Robin Hood map. Neither key +; (Map K V), spec-memory.md — an open-addressed Swiss table. Neither key ; nor value type appears in it, for the same reason: one type-erased runtime, ; handed the two sizes and a hash/equality pair at each call site. %map = type { ptr, i64, i64, ptr, i64 } diff --git a/runtime/flan_rt.c b/runtime/flan_rt.c index a959d1c6..45bf199b 100644 --- a/runtime/flan_rt.c +++ b/runtime/flan_rt.c @@ -1910,16 +1910,17 @@ int8_t flan_vec_clone(flan_vec *dst, flan_vec *src, flan_allocator *a, * Header, five words, and the same five the emitter's %map and its debugger * description name: * - * data one allocation: growth word | control bytes | keys | values + * data one allocation: a three-word head, the control bytes, then + * the slots, each a key and its value side by side * len live entries * log2cap 0 until something is allocated; never below 3 after * allocator epoch as on a Vec, and checked the same way * * The one number a Swiss table needs beyond the header — how many more * entries may land in empty slots before the table is rebuilt — lives in the - * first word of the block rather than in the header, because the header's - * layout is the emitter's as much as this file's and a sixth word would move - * every offset after it. The slot geometry sits beside it; see "Block + * block's head rather than in the header, because the header's layout is the + * emitter's as much as this file's and a sixth word would move every offset + * after it. The slot stride and value offset sit beside it; see "Block * geometry". * * Every entry point returns int8_t 1/0 for "did it fit", never reporting @@ -2131,9 +2132,8 @@ uint64_t flan_hash_combine(uint64_t acc, uint64_t h) { /* ── Groups ─────────────────────────────────────────────────────────── * * Eight control bytes read as one 64-bit word, and every question about them - * answered with a handful of integer operations over the whole word — the - * portable arrangement hashbrown calls "generic", which needs no intrinsics - * and so compiles for every target the runtime does. Each answer is a mask + * answered with a handful of integer operations over the whole word, which + * needs no intrinsics and so compiles for every target the runtime does. Each answer is a mask * with 0x80 set in the byte of every slot that qualifies; a slot's index in * the group is its byte's position, counted from the low end. */ #define FLAN_LSB 0x0101010101010101ULL @@ -2518,7 +2518,7 @@ int8_t flan_map_init(flan_map *m, flan_allocator *a, int64_t ksize, * An insert that would land in an empty slot with no growth left rebuilds * first. Deleted slots count against the growth, since each one is a slot a * probe cannot stop at; when they are most of what is using it up — the live - * entries fit in 25/32 of the capacity, abseil's figure — the rebuild keeps + * entries fit in 25/32 of the capacity — the rebuild keeps * the capacity and only sweeps them out, so a map that churns at a steady size * does not double without end. */ int8_t flan_map_put(flan_map *m, const void *key, const void *val, @@ -2605,10 +2605,10 @@ int8_t flan_map_has(flan_map *m, const void *key, int64_t ksize, int64_t vsize, * DELETED instead — a slot a probe walks through and an insert may reuse — * unless no probe can ever have walked through it: when every eight-slot * window containing it already holds an empty slot, any probe that reached it - * stopped in that window, and it can be emptied outright. That test is - * abseil's (the empties just before and just after it are fewer than eight - * slots apart), and it is what keeps a map that is far from full from - * accumulating deleted slots at all. + * stopped in that window, and it can be emptied outright. The test is that the + * empties just before and just after it are fewer than eight slots apart, and + * it is what keeps a map that is far from full from accumulating deleted slots + * at all. * * Nothing moves. The old table shifted the rest of the run back one slot, so * a cursor already past them stepped over entries; here an entry stays in its diff --git a/test/programs/map-iter.flan b/test/programs/map-iter.flan index 29392b5c..9d8226ec 100644 --- a/test/programs/map-iter.flan +++ b/test/programs/map-iter.flan @@ -82,7 +82,7 @@ (print chars) (print " ") (print xs) (print " ") (print ys) (println "")) (free m))) -;; Growth past the 75% threshold rehashes into a new block, so this walks a map +;; Growth past the load factor rehashes into a new block, so this walks a map ;; whose layout is nothing like its insertion order and at a capacity several ;; doublings past the minimum. (defn after-growth [] () diff --git a/test/programs/map-remove.flan b/test/programs/map-remove.flan index 2eefc5e9..349eafd2 100644 --- a/test/programs/map-remove.flan +++ b/test/programs/map-remove.flan @@ -13,6 +13,14 @@ ;;;; swept grows without end under row 7. (defstruct Cell [x i32 y i32]) +;; Row 7's allocator and its count of refusals. Globals, because a handler +;; cannot see the locals of the function that established it. +(defonce churn-alloc Allocator) +(defonce churn-refusals i64) +;; Row 7's reference: a global array rather than a Vec, because a Vec would +;; take its block from the same process heap the budget is counting. +(defonce churn-ref [512 i64]) + (defn main [] i32 ;; (1) The value comes back, the length drops, and the key is gone. An ;; absent key is None and changes nothing — the same answer get gives, since @@ -118,38 +126,53 @@ (free-all ar)) ;; (7) Churn at a steady size, checked against a plain array. Keys are drawn - ;; from 512 and the map hovers around half of them, so removals leave deleted - ;; slots that inserts must reuse and rebuilds must sweep. Every get is checked - ;; against the array, and at the end so is the length and a full walk. - (let [m (map-new i32 i64) - ref (vec-new i64) + ;; from 512 and the map holds about two thirds of them, so removals leave + ;; deleted slots that inserts must reuse and rebuilds must sweep. Every get is + ;; checked against the array, and at the end so is the length and a full walk. + ;; + ;; The process heap is given a budget of 20000 bytes, with nothing else of + ;; this program's live on it. An i32 key and an i64 + ;; value make a 16-byte slot, so the 512-slot block is 8768 bytes and a + ;; rebuild at that size holds two of them, 17536; a grow to 1024 slots holds + ;; 8768 and 17472 at once and goes over. The entries never need more than 512 + ;; slots, so a refusal means deleted slots were used up and the table doubled + ;; instead of sweeping them — which is what a version that never swept them + ;; does. The handler lets it through, counts it, and the count is printed. + (set churn-alloc (heap-allocator)) + (set-alloc-budget churn-alloc 20000) + (handler-bind + [(StorageExhausted [c] + (set churn-refusals (+ churn-refusals 1)) + (set-alloc-budget churn-alloc (* 2 (alloc-budget churn-alloc))) + (invoke-restart 'retry))] + (let [m (map-new i32 i64 churn-alloc) x (u64 12345) bad 0 live 0] - (dotimes [i 512] (push ref (i64 -1))) + (dotimes [i 512] (set (at churn-ref i) (i64 -1))) (dotimes [step 200000] (set x (+ (* x (u64 6364136223846793005)) (u64 1442695040888963407))) (let [k (i32 (% (>> x 33) (u64 512))) op (i32 (% (>> x 20) (u64 4)))] (if (< op 2) - (do (if (< (at ref k) 0) (set live (+ live 1))) + (do (if (< (at churn-ref k) 0) (set live (+ live 1))) (put m k (i64 step)) - (set (at ref k) (i64 step))) + (set (at churn-ref k) (i64 step))) (if (= op 2) (match (map-remove m k) - (Some v) (do (if (not (= v (at ref k))) (set bad (+ bad 1))) - (set (at ref k) (i64 -1)) + (Some v) (do (if (not (= v (at churn-ref k))) (set bad (+ bad 1))) + (set (at churn-ref k) (i64 -1)) (set live (- live 1))) - None (if (>= (at ref k) 0) (set bad (+ bad 1)))) + None (if (>= (at churn-ref k) 0) (set bad (+ bad 1)))) (match (get m k) - (Some v) (if (not (= v (at ref k))) (set bad (+ bad 1))) - None (if (>= (at ref k) 0) (set bad (+ bad 1)))))))) + (Some v) (if (not (= v (at churn-ref k))) (set bad (+ bad 1))) + None (if (>= (at churn-ref k) 0) (set bad (+ bad 1)))))))) (print bad) (println "") ; 0 (print (= (length m) (i64 live))) (println "") ; true (let [cur (i64 0) k 0 v (i64 0) seen 0] (while (map-next m (addr cur) (addr k) (addr v)) (set seen (+ seen 1)) - (if (not (= v (at ref k))) (set bad (+ bad 1)))) + (if (not (= v (at churn-ref k))) (set bad (+ bad 1)))) (print (= seen live)) (println "") ; true (print bad) (println "")) ; 0 ;; Removing while walking visits every survivor once: nothing moves. @@ -159,6 +182,6 @@ (map-remove m k)) (print (= seen live)) (println "") ; true (print (length m)) (println "")) ; 0 - (free ref) - (free m)) + (free m))) + (print churn-refusals) (println "") ; 0 0) diff --git a/test/programs/maps.flan b/test/programs/maps.flan index 5a235601..1604c72e 100644 --- a/test/programs/maps.flan +++ b/test/programs/maps.flan @@ -1,7 +1,7 @@ ;;;; (Map K V) — spec-memory.md, step 4 of the container build order. ;;;; -;;;; Odin's map: open-addressed Robin Hood hashing at a 75% load factor, with -;;;; cache-line cell packing. Every claim below is one a plausible wrong +;;;; An open-addressed Swiss table: one control byte a slot, key and value side +;;;; by side. Every claim below is one a plausible wrong ;;;; version gets wrong, and the numbers differ per failure so a single wrong ;;;; answer names its own cause. (defstruct Cell [x i32 y i32]) diff --git a/test/test_acceptance.ml b/test/test_acceptance.ml index 9f0e6729..03a01dbb 100644 --- a/test/test_acceptance.ml +++ b/test/test_acceptance.ml @@ -3831,8 +3831,7 @@ level "1" (* ── (Map K V), spec-memory.md step 4 ────────────────────────── - Odin's map: open-addressed Robin Hood hashing at a 75% load factor with - cache-line cell packing. maps.flan is seven claims, each one a plausible + An open-addressed Swiss table, one control byte a slot. maps.flan is seven claims, each one a plausible wrong version gets wrong, and the numbers differ per failure. The two worth naming, because nothing else in the suite would catch @@ -3871,9 +3870,9 @@ level "1" outputs "map iteration" "programs/map-iter.flan" map_iter_out; outputs ~opt:"-O0" "map iteration, -O0" "programs/map-iter.flan" map_iter_out; - (* Removal, which is the operation that can break the others. A Robin Hood - probe stops at the first empty slot, so a hole left in the middle of a - run hides every entry past it — and the hidden ones are exactly what a + (* Removal, which is the operation that can break the others. A probe + stops at the first group holding an empty slot, so a slot emptied in + the middle of a probe sequence hides every entry past it — and the hidden ones are exactly what a test that only asks after what it removed never looks at. Hence the third row: 2000 entries, the even keys taken out, and then every odd one asked for. A removal that punched the hole and left it answers row one @@ -3885,7 +3884,7 @@ level "1" let map_remove_out = "100\n1\nfalse\ngone\n1\n0\n77\n1\n1000\n1000\n0\n0\n2000\n\ 709\nfalse\ntrue\n399\n1\ntrue\n1\n66\n0\n66\n150\n598\nfalse\n\ - 0\ntrue\ntrue\n0\ntrue\n0\n" + 0\ntrue\ntrue\n0\ntrue\n0\n0\n" in outputs "map removal" "programs/map-remove.flan" map_remove_out; outputs ~opt:"-O0" "map removal, -O0" "programs/map-remove.flan"