diff --git a/TODO.org b/TODO.org index 97deca9a..4c0d3a22 100644 --- a/TODO.org +++ b/TODO.org @@ -1271,10 +1271,13 @@ value an unsigned compare waves through. lock the agent's listener thread may hold inside =dlopen=, so a program that should die could hang. Unconditionally rather than only under =--dev=. -** TODO A Map's bounds check and the stale-container failure still die -Deliberate for the stale case — the region the container lived in was released and -there is no frame to go back to that would not read freed memory. The Map path was -left alone rather than converted half-way. +** DONE A Map's bounds check and the stale-container failure still die +CLOSED: [2026-09-25] +A Map has no bounds check: =get= and =map-remove= answer =None= for an absent key +and nothing on the map path indexes by a number, so there was nothing to convert. +The stale-container failure keeps dying — the region was released and there is no +frame to go back to that would not read freed memory. Rules out a signalled +condition on the stale path. ** DONE Six trap paths park instead of killing the session CLOSED: [2026-09-18] @@ -1334,15 +1337,17 @@ docs/BUILT.md, "Three amendments to a frozen spec". ** DONE Map removal costs a backward-shift loop Removal landed, with the loop the spec predicted as its cost. Deferring it was what had kept the implementation free of tombstones and of Odin's -backward-shift loop; taking it is taking the loop. +backward-shift loop; taking it is taking the loop. Superseded by the Swiss table, +which removes by tombstone and moves nothing. -** NEXT The Map is slower than CPython's dict at a million entries -Decided 2026-09-25: build the Swiss-table layout, one metadata byte a slot. Measured against the current Map at small, game and large sizes before and after. -Keys, values and hashes are three separate runs, so a lookup that misses -everything costs three cache misses where a compact dict costs two. One byte of -metadata a slot — the Swiss-table arrangement — is the known answer and is not -built. The crossover is somewhere between ten thousand and a million and nobody -has found it. +** DONE The Map is slower than CPython's dict at a million entries +CLOSED: [2026-09-25] +The Map is a Swiss table: one control byte a slot, key and value side by side, +groups of eight probed as one 64-bit word, seven-eighths load, removal by +tombstone with a same-capacity rebuild to sweep them. Header, entry points and +iteration contract unchanged. Rules out Robin Hood, the backward shift and +separate key and value runs; SSE2 groups are not built. See docs/BUILT.md, "The +Map is a Swiss table". ** DONE dyn maps and interned keywords CLOSED: [2026-09-20] diff --git a/docs/BUILT.md b/docs/BUILT.md index 692a80e1..117ad3ab 100644 --- a/docs/BUILT.md +++ b/docs/BUILT.md @@ -3088,6 +3088,10 @@ restriction under it. ### The three properties that were the point +**Superseded.** The table below is no longer a Robin Hood map: it is a Swiss table, and "The Map is a Swiss table" +further down says what replaced what and why. This section and the next are kept for the reasoning they record; the +hash-and-equality pair, the key restrictions and the header are unchanged. + Odin's header states them and they are why it was the thing to follow (`base/runtime/dynamic_map_internal.odin`). - **Open-addressed Robin Hood hashing at a 75% load factor.** No buckets and no per-entry allocation: one block holds @@ -3104,14 +3108,11 @@ Odin's header states them and they are why it was the thing to follow (`base/run ### Two departures from Odin, both deliberate -**There are still no tombstones, now that removal exists.** A slot is empty or occupied and nothing else, and -`flan_map_remove` keeps it that way by shifting the run back over the hole rather than marking it. Odin does the -opposite, and this paragraph used to say otherwise: read out of -`base/runtime/dynamic_map_internal.odin`, `map_erase_dynamic` sets a tombstone bit and leaves the repair to the next -insert, which is why Odin's *insert* carries a backward-shift loop and its every lookup tests for a tombstone. The -trade is the usual one — erase is O(1) there and the shift is here, and the lookups, which outnumber the removals, -pay nothing. What `spec-memory.md` still defers is the rest of its sentence: move-aware lookup and owned entries. A -removed value is copied out, and nothing is dropped. +**Removal was a backward shift, and is now a tombstone.** While the table was Robin Hood, a slot was empty or occupied +and `flan_map_remove` shifted the run back over the hole. The Swiss table has a third control state, deleted, and +empties a slot outright only where no probe can have passed through it; see "The Map is a Swiss table". What +`spec-memory.md` still defers is unchanged: move-aware lookup and owned entries. A removed value is copied out, and +nothing is dropped. **The header does not tag the capacity into the data pointer.** Odin stuffs `log2cap` into the low six bits because its `Raw_Map` must be three words. This header already carries an allocator, a generation and an epoch, so the tagging @@ -3119,10 +3120,10 @@ would buy nothing, cost a mask on every access, and — the part that actually m block being 64-byte aligned. Alignment is *requested*; an arena whose base is not cache-aligned now gives a slower map rather than a wrong one. -The header is six words, 48 bytes, the same as a `Vec`'s and for the same reason: a layout that changes with a build +The header is five words, 40 bytes, the same as a `Vec`'s and for the same reason: a layout that changes with a build flag can disagree across the reload boundary. - data len log2cap allocator gen epoch + data len log2cap allocator epoch ### The hash and equality pair, and why most key types do not get one @@ -3232,7 +3233,7 @@ Measured on this machine, `i64` to `i64`, against CPython 3.13's dict on the sam left out. Both are waiting on memory there, and this layout waits longer: keys, values and hashes are three separate runs, so a lookup that misses everything takes three cache misses where a compact dict takes two, and the hash run is a full eight bytes a slot. Cell packing buys probe locality, which is a win while the hash run is resident and a loss -once nothing is. One byte of metadata a slot — the Swiss-table arrangement — is the known answer and is not built. +once nothing is. One byte of metadata a slot — the Swiss-table arrangement — is the known answer, and the next section is it being built. The path from 35 ns to 18 ns (the cache-resident floor, at 500 entries) was **profiled, and the first two guesses were both wrong**: the per-slot cell division and the block-size divisions were each replaced first and neither moved the @@ -3243,10 +3244,80 @@ twice per lookup; and the seed stopped being a five-multiply avalanche on the cr again immediately after. `64/size` is a table, which is Odin's `Map_Cell_Info` by another route — Odin precomputes it per type because the probe loop must not divide, and here the sizes arrive as ordinary arguments. -What remains at 18 ns is the type erasure itself: a non-inlinable call into the runtime and two non-inlinable indirect +What remained at 18 ns was the type erasure itself: a non-inlinable call into the runtime and two non-inlinable indirect calls to the pair. That is the trade `spec-memory.md` chose deliberately — "It is type-erased on purpose… No generics are involved, and none are needed" — and monomorphisation is what would buy it back, at the cost the spec declined. +### The Map is a Swiss table + +The Robin Hood table above lost to CPython's dict at a million entries: three separate runs and an eight-byte hash a +slot. It is replaced by a Swiss table, and the header, every entry point's signature, the hash-and-equality pair and +the iteration contract are unchanged, so the change is confined to `runtime/flan_rt.c`. The layout and the removal rule +are described there; what follows is what the code cannot say. + +- **Empty is zero** — Zig's encoding (`lib/std/hash_map.zig`, `Metadata`: free 0, tombstone 1, a used bit on top) — so + a zeroed control run is an empty map. +- **Groups are eight bytes read as one integer**, not SSE2's sixteen: portable to every target the runtime compiles + for, and nothing measured asks for the wider group. +- **Key and value share a slot** because a hit then costs two cache misses rather than three. The runtime is not told + alignments; the largest power of two dividing each size bounds them. +- **The head of the block holds three words** — growth left, slot stride, value offset — because the five-word header + is the emitter's layout too. The stride is stored rather than recomputed because recomputing it per call cost about a + tenth of a cache-resident lookup. +- **The 25/32 threshold** for sweeping deleted slots at the same capacity rather than doubling, and the rule for + emptying a removed slot outright, were written from memory of abseil's table, not read from its source; no abseil + clone is on this machine. `map-remove.flan` row 7 is what holds them: it churns a map at a steady size under an + allocator budget that a doubling would exceed. + +#### Measured + +`i64` to `i64`; keys present are `2i`, absent `2i + 1`. *insert* fills a fresh map with no `reserve`, to two million +inserts; *hit* and *miss* are four million lookups cycling the keys; *remove* removes every key of a filled map, to two +million. Each pass counts its wrong answers: a map passed by value to a filling function takes its grows in the +callee's copy of the header, and a benchmark written that way measures an empty map. CPython is Python 3.13.9, +`dict.get` and `dict.pop`. + +The figure is the **minimum of nine runs**, the two Flan binaries run alternately and the Python script in a separate +nine-run pass, because this machine is shared — it carried a load average near 20 throughout, from other lanes' test +suites — and a minimum measures the program rather than the neighbours. Both Flan binaries are `-O2`. The magnitudes +carry that load: across three nine-run passes the Robin Hood column moved by as much as half (a 100k hit read 51, 83 +and 35 ns). What held is the ordering — the Swiss table ahead of the Robin Hood table in every cell of every pass, +and ahead of CPython in every cell — and that is the claim to reproduce. Nanoseconds per operation: + +| Entries | Op | Robin Hood | Swiss | CPython dict | +|---|---|---|---|---| +| 16 | insert | 74.5 | 30.5 | 46.3 | +| | hit | 20.0 | 9.8 | 83.1 | +| | miss | 16.1 | 8.0 | 100.5 | +| | remove | 23.3 | 15.4 | 69.0 | +| 1k | insert | 80.9 | 30.0 | 57.4 | +| | hit | 22.2 | 9.7 | 124.6 | +| | miss | 15.9 | 7.8 | 128.3 | +| | remove | 24.5 | 14.8 | 91.9 | +| 10k | insert | 112.4 | 45.5 | 85.5 | +| | hit | 29.6 | 10.5 | 134.9 | +| | miss | 26.5 | 8.5 | 125.3 | +| | remove | 33.2 | 17.6 | 99.7 | +| 100k | insert | 204.4 | 47.6 | 123.6 | +| | hit | 35.3 | 16.1 | 133.6 | +| | miss | 30.4 | 14.4 | 128.4 | +| | remove | 40.6 | 22.9 | 98.4 | +| 300k | insert | 199.4 | 73.0 | 128.7 | +| | hit | 69.9 | 25.9 | 118.0 | +| | miss | 39.8 | 10.9 | 138.2 | +| | remove | 116.8 | 36.9 | 94.3 | +| 1M | insert | 454.2 | 138.5 | 151.5 | +| | hit | 147.1 | 78.4 | 120.3 | +| | miss | 88.2 | 14.6 | 125.2 | +| | remove | 152.0 | 74.5 | 91.5 | + +**The crossover.** Against CPython the Robin Hood table lost on insert at every size, on remove from between 100k and +300k, and on hits between 300k and 1M. The Swiss table is ahead of both in every cell. CPython hashes a small integer +to itself, so this key pattern walks its table in order and the margin at a million is narrower than it looks. + +Control bytes over separate key and value runs were measured and not kept: level with the final layout to 100k, behind +from 300k (100 ns a hit against 78 at a million). + ## Unions, and the tag they carry `defdata` parsed and its shape was checked long before this; naming the type (`check.ml:312`) and constructing a @@ -4488,10 +4559,8 @@ function. ``` **The cursor is a slot index the caller owns, and there is no iterator struct** because there is nothing for one to -hold. A map has no tombstones — removal shifts the run back instead of marking a hole — so a slot is either empty or -occupied and the position is the whole of the state. What a cursor does *not* survive is a removal taken while it is -in flight: the shift moves entries to lower slots, and a cursor already past them steps over entries it has not -answered, the same bargain a put that grows already makes. The cursor starts at 0, comes back one past the entry just answered, and is left +hold: a slot's control byte says whether it is full, and the position is the whole of the state. A removal moves +nothing, so a cursor survives one; what it does not survive is a put that rebuilds the block. The cursor starts at 0, comes back one past the entry just answered, and is left at `cap` by the call that answers false, so a spent cursor keeps answering false rather than wrapping. **Three out-pointers and not a returned pair**, because there are no tuples. An `(Option K)` would answer half an @@ -4501,11 +4570,10 @@ that is written through on the way out. **It is the one map entry point that carries neither a hash nor an equality function.** Walking asks nothing about a key. The two sizes are still there, because the runtime is type-erased and the block geometry is computed from them. -**The layout, restated, because it is the thing to get wrong here.** `data` is *one* allocation laid out -keys | values | hashes | scratch, each run cell-packed to a cache line — the arrangement the Valgrind lane described -while explaining why a probe overrun is not observable. A key is reached through `flan_cell_at` and never as -`ks + i * ksize`. The hashes are the exception `flan_map_clone` already relies on: an 8-byte element packs 8 to a -64-byte cell with nothing left over, so `g.hs[i]` is the right index and a flat one. +**The layout, restated, because it is the thing to get wrong here.** `data` is *one* allocation: a three-word head, +the control bytes, then the slots, each slot a key and its value side by side at the stride the head records. A slot is +full when its control byte has the top bit set, and a key is reached as `slots + i * stride`, never as +`ks + i * ksize`. **Order is block order**, which is the hash's order and not the insertion's, and it changes when the map grows. `programs/map-iter.flan` is therefore written entirely in sums, counts and lengths — every claim in it is order-free, @@ -5732,10 +5800,10 @@ IR (which is also why `--no-bounds-checks` never reached it), so both grew a tra `Emit`'s `Rt` arm guards those two symbols and no others — they are the only ones in that family that can transfer; everything else there is arithmetic over a container header. -**A `Map`'s bounds and a `Vec`'s stale-allocator check still die.** `flan_vec_stale_fail` is a different kind of +**A `Vec`'s stale-allocator check still dies, and so does a `Map`'s.** `flan_vec_stale_fail` is a different kind of failure — the region a container lived in was released, and there is no frame to go back to that would not read freed -memory — and the map path was left alone rather than converted half-way. Written down here so it is a known edge -rather than a discovery. +memory. A `Map` has no bounds check to convert: `get` and `map-remove` answer `None` for an absent key and nothing on +the map path indexes by a number, so the stale check is the only map failure that dies. ## An address answers with a type, and nothing grew a tag word diff --git a/lib/check.ml b/lib/check.ml index 4f7c5afd..0a094997 100644 --- a/lib/check.ml +++ b/lib/check.ml @@ -7629,7 +7629,7 @@ and named_call ?(qualified = false) ctx ~want loc name args = let attempt, note = match target.Tast.ty with (* For a map the number is entries, not slots: the runtime sizes the - block so that [n] still sits under the 75% load factor, which is + block so that [n] still sits under the load factor, which is the only reading of "room for n" that does not reallocate on the nth put. *) | Types.Map (k, v) -> diff --git a/lib/emit.ml b/lib/emit.ml index 606fc40b..4e513fad 100644 --- a/lib/emit.ml +++ b/lib/emit.ml @@ -334,7 +334,7 @@ let rec ll (t : Types.t) = the Vec's address — so the shape is here only so that a slot, a struct field and a copy in the IR are the right number of bytes. *) | Types.Vec _ -> "%vec" - (* data + len + log2cap + allocator + gen + epoch. Six words, exactly as the + (* data + len + log2cap + allocator + epoch. Five words, as many as the Vec's, and read here for exactly the same reason: nothing in this file touches a field of one — every operation is a runtime call taking the map's address — so the shape exists only so that a slot, a struct field @@ -4379,7 +4379,7 @@ let header = {|; Generated by flan. The layout is C's: no object headers anywher ; (Vec T), spec-memory.md. The element type is nowhere in it: the runtime is ; type-erased and every operation is handed size and align at its call site. %vec = type { ptr, i64, i64, ptr, i64 } -; (Map K V), spec-memory.md — Odin's open-addressed Robin Hood map. Neither key +; (Map K V), spec-memory.md — an open-addressed Swiss table. Neither key ; nor value type appears in it, for the same reason: one type-erased runtime, ; handed the two sizes and a hash/equality pair at each call site. %map = type { ptr, i64, i64, ptr, i64 } diff --git a/runtime/flan_rt.c b/runtime/flan_rt.c index 36ad3679..c64df81c 100644 --- a/runtime/flan_rt.c +++ b/runtime/flan_rt.c @@ -2037,12 +2037,13 @@ int8_t flan_vec_clone(flan_vec *dst, flan_vec *src, flan_allocator *a, /* ── (Map K V), spec-memory.md ────────────────────────────────────────── * - * Odin's map, followed deliberately: open-addressed Robin Hood hashing at a - * 75% load factor, cache-line cell packing, and pointer-width integers through - * the probe loop (base/runtime/dynamic_map_internal.odin, whose header states - * the same three). One type-erased runtime over (key size, value size) and a + * A Swiss table: open addressing with one control byte a slot, probed eight + * slots at a time. One type-erased runtime over (key size, value size) and a * compiler-emitted hash and equality pair, exactly as the Vec runtime is one - * over (size, align). + * over (size, align) — Odin's Map_Info arrangement, which this keeps. What it + * does not keep is Odin's table (base/runtime/dynamic_map_internal.odin), a + * Robin Hood map with an eight-byte hash a slot; docs/BUILT.md, "The Map is a + * Swiss table", has the measurements that replaced it. * * Why the shape matters, since the obvious question is whether this is another * Python dict. Python's algorithm is fine. What makes it slow is that every @@ -2051,62 +2052,46 @@ int8_t flan_vec_clone(flan_vec *dst, flan_vec *src, flan_allocator *a, * a key is raw bytes inside the block and the hash and comparison are compiled * concretely per key type. That is most of the gap before any cleverness. * - * Robin Hood, in one paragraph. Every occupied slot has a probe distance: how - * far it sits from the slot its hash wanted. On insert, if the element already - * in a slot is closer to its desired slot than the element being placed, the - * two swap and the poorer one carries on down the run. Distances even out, the - * worst case collapses towards the average, and a lookup may stop the moment - * it is further from home than the occupant it is looking at — which is the - * early exit in flan_map_find and is why a miss costs about what a hit does. + * The control byte, in one paragraph. Seven bits of the key's hash and one bit + * saying the slot is full. A lookup compares the seven bits of eight slots at + * once and calls the equality function only where they agree, which is one + * slot in 128 by chance, so a probe reads the control run and nothing else + * until it has a candidate. The control run is one byte a slot — a million + * entries is a megabyte of it — so a miss is usually one cache miss and a hit + * two: the control group, then the key. * - * Cache-line cells, in one more. A flat [capacity]K array lets one key straddle - * two cache lines, so a probe that walks four slots can touch five lines. A - * Map_Cell packs as many Ks as fit in 64 bytes and pads the remainder, so no - * key ever straddles a line and a linear probe walks memory in the order the - * prefetcher expects. Keys, values and hashes are three separate blocks, so a - * probe — which reads hashes and only then one key — touches hash lines and - * nothing else until it has a candidate. + * Header, five words, and the same five the emitter's %map and its debugger + * description name: * - * Header, six words, the same as flan_vec's and for the same reason (a layout - * that changes with a build flag can disagree across the reload boundary): - * - * data one allocation: keys | values | hashes | scratch + * data one allocation: a three-word head, the control bytes, then + * the slots, each a key and its value side by side * len live entries - * log2cap 0 until something is allocated; never 1 or 2 after + * log2cap 0 until something is allocated; never below 3 after * allocator epoch as on a Vec, and checked the same way * - * Odin stuffs log2cap into the low six bits of the data pointer because its - * Raw_Map must be three words. This header already carries an allocator and - * an epoch, so the bit-stuffing would buy nothing and cost a - * mask on every access — and, more usefully, not tagging means correctness - * never depends on the block being 64-byte aligned. It is requested as 64, and - * cell packing pays off when the request is honoured, but an arena whose base - * is not cache-aligned gives a slower map rather than a wrong one. + * The one number a Swiss table needs beyond the header — how many more + * entries may land in empty slots before the table is rebuilt — lives in the + * block's head rather than in the header, because the header's layout is the + * emitter's as much as this file's and a sixth word would move every offset + * after it. The slot stride and value offset sit beside it; see "Block + * geometry". * * Every entry point returns int8_t 1/0 for "did it fit", never reporting * failure any other way — the condition, the restart and the message are the * compiler's job (Check's alloc_guard). */ -#define FLAN_MAP_CACHE_LINE 64 -#define FLAN_MAP_LOAD_FACTOR 75 -#define FLAN_MAP_MIN_LOG2 3 /* 8 slots */ +#define FLAN_MAP_ALIGN 64 +#define FLAN_MAP_MIN_LOG2 3 /* 8 slots: one group */ +#define FLAN_MAP_GROUP 8 +#define FLAN_MAP_HEAD 24 /* growth, slot stride, value offset */ -/* The hash word. Zero means the slot is empty, which is what makes a - * zeroed hash block an empty map. There is still no tombstone now that - * removal exists: the only two states a slot has are empty and occupied, and - * flan_map_remove restores that by shifting the run back rather than by - * marking the hole. Odin marks it and repairs on the next insert, which is - * why its insert has a second loop and its every lookup tests for a - * tombstone; neither is here. What spec-memory.md still defers is the rest of - * that sentence — move-aware lookup and owned entries — so a removed value is - * copied out and nothing is dropped. - * - * The top bit is set on every stored hash so that a hasher answering 0 does - * not read as an empty slot. It is the highest bit, so the desired slot and - * the probe distance — which use only the low log2cap bits — are unchanged by - * it, and no hash needs rewriting when the capacity changes. */ -typedef uint64_t flan_map_hash; -#define FLAN_MAP_OCCUPIED ((uint64_t)1 << 63) +/* The control byte. Empty is zero, so a zeroed control run is an empty map — + * the same property the old table's zero hash word had, and Zig's encoding + * (lib/std/hash_map.zig, Metadata: free 0, tombstone 1, used bit on top). A + * full slot is the high bit and the top seven bits of the hash. */ +#define FLAN_CTRL_EMPTY ((uint8_t)0x00) +#define FLAN_CTRL_DELETED ((uint8_t)0x01) +#define FLAN_CTRL_FULL ((uint8_t)0x80) /* The pair the compiler emits per key type. [size] is the key's size, passed * so that the flat hasher and comparator below can serve every key whose @@ -2297,183 +2282,146 @@ uint64_t flan_hash_combine(uint64_t acc, uint64_t h) { return flan_mix64(acc ^ (h + 0x9e3779b97f4a7c15ULL + (acc << 6) + (acc >> 2))); } -/* ── Cell geometry ──────────────────────────────────────────────────── +/* ── Groups ─────────────────────────────────────────────────────────── * - * Odin precomputes these into a Map_Cell_Info so the probe loop never divides. - * They are derived from the element size alone — alignment cannot matter, - * because a cell starts on a 64-byte boundary and no Flan type is aligned - * above that — so they are derived once on entry to each operation and kept in - * locals, which is the same trade with one less thing for the checker to pass - * and get wrong. */ -/* 64/size for every size a cell can pack, as a table rather than a division. - * - * This is Odin's Map_Cell_Info by another route. Odin precomputes - * elements_per_cell and size_of_cell into a static per-type record because the - * probe loop must not divide; the same number is wanted here and the call site - * cannot hand it over, because the sizes reach this runtime as ordinary i64 - * arguments rather than as a compile-time record. A 64-entry table is one load - * and needs nothing added to the calling convention. - * - * It is worth the lines: the geometry is recomputed on every lookup, three - * times over (keys, values, hashes), and three divisions there measured as a - * fifth of the whole operation. */ -static const uint8_t flan_epc_table[64] = { - 1, 64, 32, 21, 16, 12, 10, 9, - 8, 7, 6, 5, 5, 4, 4, 4, - 4, 3, 3, 3, 3, 3, 2, 2, - 2, 2, 2, 2, 2, 2, 2, 2, - 2, 1, 1, 1, 1, 1, 1, 1, - 1, 1, 1, 1, 1, 1, 1, 1, - 1, 1, 1, 1, 1, 1, 1, 1, - 1, 1, 1, 1, 1, 1, 1, 1, -}; + * Eight control bytes read as one 64-bit word, and every question about them + * answered with a handful of integer operations over the whole word, which + * needs no intrinsics and so compiles for every target the runtime does. Each answer is a mask + * with 0x80 set in the byte of every slot that qualifies; a slot's index in + * the group is its byte's position, counted from the low end. */ +#define FLAN_LSB 0x0101010101010101ULL +#define FLAN_MSB 0x8080808080808080ULL -static int64_t flan_cell_epc(int64_t size) { - if (size <= 0 || size >= FLAN_MAP_CACHE_LINE) return 1; - return (int64_t)flan_epc_table[size]; +static uint64_t flan_group_load(const uint8_t *p) { + uint64_t x; + memcpy(&x, p, 8); +#if defined(__BYTE_ORDER__) && __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__ + x = __builtin_bswap64(x); +#endif + return x; } -/* log2 of [epc] when it is a power of two, and -1 when it is not. - * - * epc is 64/size clamped to at least 1, so the only powers of two it can ever - * be are these seven. A switch over them is a handful of compares the branch - * predictor gets right every time; a loop looking for the bit measured *worse* - * than the division it was replacing, which is why this is written out. */ -static int flan_log2_epc(int64_t epc) { - switch (epc) { - case 1: return 0; - case 2: return 1; - case 4: return 2; - case 8: return 3; - case 16: return 4; - case 32: return 5; - case 64: return 6; - default: return -1; - } +/* Exactly the zero bytes. The shorter (x - LSB) & ~x form is approximate — + * a borrow out of a zero byte can mark the byte above it — and the removal + * rule below counts positions, so it wants the exact one. */ +static uint64_t flan_group_zero(uint64_t x) { + return ~(((x & ~FLAN_MSB) + ~FLAN_MSB) | x) & FLAN_MSB; } -/* Rounding to a cache line is a mask, not a division: the generic - * flan_align_up divides, and this sits on the lookup path. */ +static uint64_t flan_group_match(uint64_t g, uint8_t ctrl) { + return flan_group_zero(g ^ (FLAN_LSB * ctrl)); +} + +static uint64_t flan_group_empty(uint64_t g) { return flan_group_zero(g); } + +/* Empty or deleted: the two states without the full bit. */ +static uint64_t flan_group_free(uint64_t g) { return ~g & FLAN_MSB; } + +static int64_t flan_mask_first(uint64_t m) { + return (int64_t)(__builtin_ctzll(m) >> 3); +} + +static int64_t flan_mask_last_gap(uint64_t m) { + return (int64_t)(__builtin_clzll(m) >> 3); +} + +/* ── Block geometry ─────────────────────────────────────────────────── + * + * [growth][stride][value offset][control: cap + 7 bytes] pad to 64 | slots + * + * The control run carries seven bytes past the end that mirror its first + * seven, so a group can be read at any slot without wrapping — a group + * starting at slot cap - 3 reads three real bytes and five mirrored ones. + * + * A slot is a key and its value side by side, so a hit reads the control + * group and then one slot: two cache misses in a map too large for the cache, + * where separate key and value runs would be three. The run starts on a + * 64-byte boundary. + * + * The runtime is not told either type's alignment, and does not need to be: a + * type's size is a multiple of its alignment, so the largest power of two + * dividing the size bounds it. The value is placed at the key's size rounded + * up to the value's bound, and the stride is rounded up to the larger of the + * two — an i64 key and an i64 value make a 16-byte slot, and an i32 key with + * an i64 value makes the same one with four bytes of padding. The cost of + * over-estimating is padding and never a misaligned load. + * + * The stride and the value offset are computed once, when the block is + * allocated, and kept in its head. Recomputing them from the two sizes on + * every call measured as about a tenth of a cache-resident lookup. */ static int64_t flan_map_round(int64_t x) { - return (x + (FLAN_MAP_CACHE_LINE - 1)) & ~(int64_t)(FLAN_MAP_CACHE_LINE - 1); + return (x + (FLAN_MAP_ALIGN - 1)) & ~(int64_t)(FLAN_MAP_ALIGN - 1); } -static int64_t flan_cell_size(int64_t size) { - return flan_map_round(flan_cell_epc(size) * size); +static int64_t flan_map_ctrl_bytes(int64_t cap) { + return flan_map_round(FLAN_MAP_HEAD + cap + (FLAN_MAP_GROUP - 1)); } -/* The bytes a [count]-element run of cells occupies, rounded to a cache line - * so the next block starts on one too. - * - * Division-free in the common case, and that matters more here than anywhere - * else in this file: flan_map_blocks calls this five times and is itself - * called once per lookup, so a division here is five divisions on the hot - * path — which measured as the difference between a 39ns lookup and a 12ns - * one, far outweighing the per-slot indexing the cell shift covers. */ -static int64_t flan_run_bytes(int64_t epc, int64_t cell, int64_t shift, - int64_t count) { - int64_t cells, bytes; - if (count < 0 || count > FLAN_BYTES_UNREPRESENTABLE - epc) - return FLAN_BYTES_UNREPRESENTABLE; - cells = (shift >= 0) ? ((count + epc - 1) >> shift) : ((count + epc - 1) / epc); - /* One predictable branch on a path that is otherwise two shifts. The count - * is a capacity the program asked for, so the product is only bounded by - * what the checker knows the element to be — see flan_mul_bytes. */ - if (!flan_mul_bytes(cells, cell, &bytes)) return FLAN_BYTES_UNREPRESENTABLE; - return bytes; +typedef struct flan_map_geom { + uint8_t *ctrl; + uint8_t *slots; + int64_t stride; + int64_t voff; +} flan_map_geom; + +/* The largest power of two dividing [size], at most 64: an upper bound on the + * alignment of any type that size, since a size is a multiple of its + * alignment. A zero size answers 1. */ +static int64_t flan_map_align_bound(int64_t size) { + int64_t b = size & -size; + if (b <= 0) return 1; + return b > FLAN_MAP_ALIGN ? FLAN_MAP_ALIGN : b; } -static int64_t flan_cells_bytes(int64_t size, int64_t count) { - int64_t epc = flan_cell_epc(size); - return flan_run_bytes(epc, flan_cell_size(size), flan_log2_epc(epc), count); +static int64_t flan_map_up(int64_t x, int64_t a) { return (x + a - 1) & -a; } + +/* Where the value sits in a slot, and how far apart slots are. */ +static void flan_map_slot(int64_t ksize, int64_t vsize, int64_t *stride, + int64_t *voff) { + int64_t ka = flan_map_align_bound(ksize), va = flan_map_align_bound(vsize); + *voff = flan_map_up(ksize, va); + *stride = flan_map_up(*voff + vsize, ka > va ? ka : va); } -/* log2 of [epc] when it is a power of two, and -1 when it is not. - * - * This is the difference between a probe that costs a shift and one that costs - * two 64-bit integer divisions, and it is measurable: with the divisions in - * place a cache-resident i64 lookup took 39ns, and without them 12ns. Odin - * does not need it because its static path resolves elements_per_cell at - * compile time and its dynamic path special-cases 1 and 2; here the number is - * always a run-time value, so the compiler cannot turn the division into a - * shift and something has to. - * - * It is a power of two whenever the element size is, which is every primitive, - * every pointer, and a struct whose size rounds to one — so the fallback is - * the rare path rather than the common one. */ - -/* Slot [i] of a cell-packed run. [epc], [cell] and [shift] are hoisted by - * every caller that walks, which is why they are parameters rather than - * recomputed here — recomputing [shift] per slot would cost more than the - * division it removes. */ -static uint8_t *flan_cell_at(uint8_t *base, int64_t size, int64_t epc, - int64_t cell, int64_t shift, int64_t i) { - if (epc == 1) return base + i * cell; - if (shift >= 0) - return base + (i >> shift) * cell + (i & (epc - 1)) * size; - return base + (i / epc) * cell + (i % epc) * size; -} - -/* The four blocks. Keys, values and hashes get one run each; the scratch is - * two more keys and two more values, which is where the Robin Hood swap keeps - * the element in flight. Odin allocates the same two, for the same reason: the - * swap is a memcpy between type-erased buffers and there is no local of the - * right type to hold one. */ /* Saturating rather than refusing, because this is called for its number by - * the geometry as well as by the allocation, and the geometry has nowhere to - * put a failure. FLAN_BYTES_UNREPRESENTABLE out of here is the only value the - * allocation is allowed to see and not attempt: flan_map_alloc turns it into - * the StorageExhausted an impossible request deserves. */ + * the dev registry as well as by the allocation. FLAN_BYTES_UNREPRESENTABLE + * out of here is the only value the allocation is allowed to see and not + * attempt: flan_map_alloc turns it into the StorageExhausted an impossible + * request deserves. */ static int64_t flan_map_block_size(int64_t ksize, int64_t vsize, int64_t cap) { - int64_t total = 0; - int64_t parts[5]; - int i; - parts[0] = flan_cells_bytes(ksize, cap); - parts[1] = flan_cells_bytes(vsize, cap); - parts[2] = flan_cells_bytes((int64_t)sizeof(flan_map_hash), cap); - parts[3] = flan_cells_bytes(ksize, 2); - parts[4] = flan_cells_bytes(vsize, 2); - for (i = 0; i < 5; i++) - if (!flan_add_bytes(total, parts[i], &total)) - return FLAN_BYTES_UNREPRESENTABLE; + int64_t stride, voff, sb, total; + if (cap < 0 || cap > ((int64_t)1 << 40)) return FLAN_BYTES_UNREPRESENTABLE; + if (ksize < 0 || vsize < 0 || ksize > ((int64_t)1 << 40) + || vsize > ((int64_t)1 << 40)) + return FLAN_BYTES_UNREPRESENTABLE; + flan_map_slot(ksize, vsize, &stride, &voff); + if (!flan_mul_bytes(stride, cap, &sb)) return FLAN_BYTES_UNREPRESENTABLE; + if (!flan_add_bytes(flan_map_ctrl_bytes(cap), sb, &total)) + return FLAN_BYTES_UNREPRESENTABLE; return total; } -/* Everything an operation needs to walk the block, computed once on entry. - * - * It used to be five calls to flan_cells_bytes, each recomputing the element - * geometry it had just been asked for, on a function called once per lookup. - * Gathering it into one struct is the single largest win in this file after - * the hash: the arithmetic is the same, it simply happens once. */ -typedef struct flan_map_geom { - int64_t kepc, kcell, vepc, vcell; - int64_t kshift, vshift; - uint8_t *ks; - uint8_t *vs; - flan_map_hash *hs; - uint8_t *sk; - uint8_t *sv; -} flan_map_geom; - static void flan_map_geometry(const flan_map *m, int64_t ksize, int64_t vsize, int64_t cap, flan_map_geom *g) { uint8_t *p = (uint8_t *)m->data; - int64_t hsize = (int64_t)sizeof(flan_map_hash); - int64_t hepc = flan_cell_epc(hsize), hcell = flan_cell_size(hsize); - int hshift = flan_log2_epc(hepc); - g->kepc = flan_cell_epc(ksize); g->kcell = flan_cell_size(ksize); - g->vepc = flan_cell_epc(vsize); g->vcell = flan_cell_size(vsize); - g->kshift = flan_log2_epc(g->kepc); - g->vshift = flan_log2_epc(g->vepc); - g->ks = p; - p += flan_run_bytes(g->kepc, g->kcell, g->kshift, cap); - g->vs = p; - p += flan_run_bytes(g->vepc, g->vcell, g->vshift, cap); - g->hs = (flan_map_hash *)p; - p += flan_run_bytes(hepc, hcell, hshift, cap); - g->sk = p; - p += flan_run_bytes(g->kepc, g->kcell, g->kshift, 2); - g->sv = p; + const int64_t *head = (const int64_t *)p; + (void)ksize; (void)vsize; /* read once, at allocation: see above */ + g->ctrl = p + FLAN_MAP_HEAD; + g->slots = p + flan_map_ctrl_bytes(cap); + g->stride = head[1]; + g->voff = head[2]; +} + +static uint8_t *flan_map_k(const flan_map_geom *g, int64_t i) { + return g->slots + i * g->stride; +} + +static uint8_t *flan_map_v(const flan_map_geom *g, int64_t i) { + return g->slots + i * g->stride + g->voff; +} + +static int64_t *flan_map_growth(const flan_map *m) { + return (int64_t *)m->data; } /* The same epoch check a Vec does, and it runs in every build for the same @@ -2505,17 +2453,11 @@ static flan_allocator *flan_map_adopt(flan_map *m) { /* The seed, derived from the block address exactly as Odin derives it: two * maps with the same keys then disagree about which slot is which, so an * adversarial insertion order against one is not an insertion order against - * the other. It changes on every grow, which is why hashes are recomputed - * there rather than carried over. */ + * the other. It changes on every rebuild, which is why every key is rehashed + * there. One multiply, not a full avalanche: whatever it returns is fed to + * the hasher, which mixes properly. The block is 64-byte aligned when the + * allocator honours the request, so the low six bits carry nothing. */ static uint64_t flan_map_seed(const flan_map *m) { - /* One multiply, not a full avalanche. This is recomputed on every lookup and - * all it has to do is decorrelate two maps from each other: whatever it - * returns is fed to the hasher, which mixes properly. A splitmix here was - * five dependent multiplies on the critical path of every probe, for - * mixing that happens again immediately afterwards. - * - * The block is 64-byte aligned when the allocator honours the request, so - * the low six bits carry nothing and are shifted out before multiplying. */ return (((uint64_t)(uintptr_t)m->data >> 6) * 0x9e3779b97f4a7c15ULL); } @@ -2523,186 +2465,190 @@ static int64_t flan_map_cap(const flan_map *m) { return m->data ? ((int64_t)1 << m->log2cap) : 0; } -/* 75% of capacity, as fixed-point integer arithmetic. Robin Hood wants a - * maximum load factor under 100% and 75% is where Odin sets it. */ +/* Seven eighths. A group probe tolerates a higher load than Robin Hood's + * 75%, because a miss stops at the first group holding an empty slot rather + * than walking a run, and at seven eighths every group of eight has one. */ +static int64_t flan_map_threshold_of(int64_t cap) { return cap - cap / 8; } + static int64_t flan_map_threshold(const flan_map *m) { - return (flan_map_cap(m) * FLAN_MAP_LOAD_FACTOR) / 100; + return flan_map_threshold_of(flan_map_cap(m)); } -/* How far this element is from the slot its hash wanted. Odin's identity: - * (slot - hash) & mask is the same number as (slot + cap - desired) & mask, - * with fewer operations, because desired is hash & mask. */ -static int64_t flan_map_distance(uint64_t hash, int64_t slot, int64_t mask) { - return (int64_t)(((uint64_t)slot - hash) & (uint64_t)mask); +/* The low bits choose the group and the top seven are the tag, so the two + * never share a bit at any capacity below 2^57. */ +static uint8_t flan_map_tag(uint64_t h) { + return (uint8_t)(FLAN_CTRL_FULL | (h >> 57)); } -/* Place one element, already hashed, into a map known to have room. This is - * Odin's swap_loop and nothing else: with no tombstones there is no second - * loop, and the load factor guarantees an empty slot is reached. */ -static void flan_map_place(flan_map *m, uint64_t h, const void *ikey, - const void *ival, int64_t ksize, int64_t vsize) { - flan_map_geom g; - int64_t cap = flan_map_cap(m), mask = cap - 1; - int64_t pos = (int64_t)(h & (uint64_t)mask), dist = 0; - uint8_t *k, *v, *tk, *tv; +/* Writing a control byte writes its mirror too, when it has one. For a slot + * at or past 7 the second index is the slot itself, so the store is simply + * repeated; below 7 it is the mirrored copy past the end. No branch. */ +static void flan_map_set_ctrl(uint8_t *ctrl, int64_t mask, int64_t i, + uint8_t c) { + ctrl[i] = c; + ctrl[((i - (FLAN_MAP_GROUP - 1)) & mask) + (FLAN_MAP_GROUP - 1)] = c; +} - flan_map_geometry(m, ksize, vsize, cap, &g); - /* The element in flight lives in scratch slot 0; slot 1 is the swap - * temporary. Both are inside the block, so nothing here touches the stack - * with a size only known at run time. */ - k = flan_cell_at(g.sk, ksize, g.kepc, g.kcell, g.kshift, 0); - v = flan_cell_at(g.sv, vsize, g.vepc, g.vcell, g.vshift, 0); - tk = flan_cell_at(g.sk, ksize, g.kepc, g.kcell, g.kshift, 1); - tv = flan_cell_at(g.sv, vsize, g.vepc, g.vcell, g.vshift, 1); - flan_copy_small(k, ikey, ksize); - if (vsize > 0) flan_copy_small(v, ival, vsize); +/* The probe sequence. Groups are visited at triangular offsets — 0, 8, 24, + * 48 and on — which for a power-of-two number of groups reaches every one of + * them before repeating, so a probe that has not found an empty group has + * not yet looked everywhere. */ +/* The first slot, along [h]'s probe sequence, that is empty or deleted. The + * threshold guarantees one exists. */ +static int64_t flan_map_find_free(const uint8_t *ctrl, int64_t mask, + uint64_t h) { + int64_t pos = (int64_t)(h & (uint64_t)mask), stride = 0; for (;;) { - uint64_t eh = g.hs[pos]; - if (eh == 0) { - flan_copy_small(flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, pos), - k, ksize); - if (vsize > 0) - flan_copy_small(flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, pos), - v, vsize); - g.hs[pos] = h; - return; - } - /* The Robin Hood swap: the occupant is richer — closer to home — than the - * element in flight, so the poorer one takes the slot and the richer one - * carries on. This is what keeps the variance down. */ - if (dist > flan_map_distance(eh, pos, mask)) { - uint8_t *kp = flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, pos); - uint8_t *vp = flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, pos); - uint64_t th; - flan_copy_small(tk, k, ksize); - flan_copy_small(k, kp, ksize); - flan_copy_small(kp, tk, ksize); - if (vsize > 0) { - flan_copy_small(tv, v, vsize); - flan_copy_small(v, vp, vsize); - flan_copy_small(vp, tv, vsize); - } - th = h; h = g.hs[pos]; g.hs[pos] = th; - dist = flan_map_distance(h, pos, mask); - } - pos = (pos + 1) & mask; - dist++; + uint64_t f = flan_group_free(flan_group_load(ctrl + pos)); + if (f) return (pos + flan_mask_first(f)) & mask; + stride += FLAN_MAP_GROUP; + pos = (pos + stride) & mask; } } -/* The slot holding [key], or -1. The middle test is the Robin Hood early exit: - * this probe is further from home than the occupant is, and Robin Hood - * maintains that no element is ever further from home than one it passed, so - * the key cannot be further along. A miss therefore costs about what a hit - * does, which is the property the ordering buys. */ -static int64_t flan_map_find_g(flan_map *m, const void *key, int64_t ksize, - int64_t vsize, flan_hash_fn hash, flan_eq_fn eq, - flan_map_geom *gp) { +/* Place a key known not to be in the map, already hashed, into a map known + * to have room. Used by the rebuild and by clone, where nothing needs looking + * up first. */ +static void flan_map_place(flan_map *m, uint64_t h, const void *key, + const void *val, int64_t ksize, int64_t vsize) { flan_map_geom g; - int64_t cap, mask, pos, dist = 0; + int64_t cap = flan_map_cap(m), mask = cap - 1, at; + flan_map_geometry(m, ksize, vsize, cap, &g); + at = flan_map_find_free(g.ctrl, mask, h); + if (g.ctrl[at] == FLAN_CTRL_EMPTY) (*flan_map_growth(m))--; + flan_map_set_ctrl(g.ctrl, mask, at, flan_map_tag(h)); + flan_copy_small(flan_map_k(&g, at), key, ksize); + if (vsize > 0) flan_copy_small(flan_map_v(&g, at), val, vsize); +} + +/* The slot holding [key], or -1. The probe stops at the first group with an + * empty slot: an insert takes the first free slot along the same sequence, so + * a key that is present was placed before any empty its probe could reach. + * + * [free_at], when not NULL, receives the first empty-or-deleted slot the + * probe passed — which is where an insert of this key goes, so a put that + * misses does not walk the sequence a second time. + * + * Inlined by force: it is called from five entry points, and at -O2 that is + * enough for clang to keep it out of line, which costs a call and a spill of + * everything the probe had in registers on every lookup. */ +static inline __attribute__((always_inline)) int64_t +flan_map_find_g(flan_map *m, const void *key, int64_t ksize, int64_t vsize, + flan_hash_fn hash, flan_eq_fn eq, flan_map_geom *g, + uint64_t *hout, int64_t *free_at) { + int64_t cap, mask, pos, stride = 0; uint64_t h; + uint8_t tag; + const uint8_t *ctrl; void *xfer = NULL; - if (!m->data || m->len == 0) return -1; + /* The hash first: it is an indirect call, and everything computed before it + * would have to survive it. */ + h = hash(key, flan_map_seed(m), ksize, &xfer); cap = flan_map_cap(m); mask = cap - 1; - flan_map_geometry(m, ksize, vsize, cap, &g); - if (gp) *gp = g; - h = hash(key, flan_map_seed(m), ksize, &xfer) | FLAN_MAP_OCCUPIED; + flan_map_geometry(m, ksize, vsize, cap, g); + ctrl = g->ctrl; + if (hout) *hout = h; + if (free_at) *free_at = -1; + tag = flan_map_tag(h); pos = (int64_t)(h & (uint64_t)mask); for (;;) { - uint64_t eh = g.hs[pos]; - if (eh == 0) return -1; - if (dist > flan_map_distance(eh, pos, mask)) return -1; - if (eh == h - && eq(key, flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, pos), - ksize, &xfer)) - return pos; - pos = (pos + 1) & mask; - dist++; + uint64_t grp = flan_group_load(ctrl + pos); + uint64_t hits = flan_group_match(grp, tag); + while (hits) { + int64_t at = (pos + flan_mask_first(hits)) & mask; + if (eq(key, flan_map_k(g, at), ksize, &xfer)) return at; + hits &= hits - 1; + } + if (free_at && *free_at < 0) { + uint64_t f = flan_group_free(grp); + if (f) *free_at = (pos + flan_mask_first(f)) & mask; + } + if (flan_group_empty(grp)) return -1; + stride += FLAN_MAP_GROUP; + pos = (pos + stride) & mask; } } -/* Allocate a block for 2^log2cap slots and zero the hashes. Only the hash run - * needs zeroing — a key or value slot is never read without its hash saying it - * is live — so the keys and values are left as the allocator returned them. */ +/* Allocate a block for 2^log2cap slots and zero the control run. Only the + * control run needs zeroing — a key or value slot is never read without its + * control byte saying it is full — so the keys and values are left as the + * allocator returned them. */ static int8_t flan_map_alloc(flan_map *m, flan_allocator *a, int64_t log2cap, int64_t ksize, int64_t vsize) { int64_t cap = (int64_t)1 << log2cap; int64_t bytes = flan_map_block_size(ksize, vsize, cap); void *p; flan_fail_bytes = bytes; - flan_fail_align = FLAN_MAP_CACHE_LINE; + flan_fail_align = FLAN_MAP_ALIGN; flan_fail_id = (int64_t)(intptr_t)a; /* A block whose size does not fit in the count is a request no allocator can * answer, and asking anyway would hand a wrapped number to a proc that might * take it. The failure is the allocator's own. */ if (bytes == FLAN_BYTES_UNREPRESENTABLE) return 0; - p = a->proc(a, FLAN_ALLOC_ALLOC, NULL, 0, bytes, FLAN_MAP_CACHE_LINE); + p = a->proc(a, FLAN_ALLOC_ALLOC, NULL, 0, bytes, FLAN_MAP_ALIGN); if (!p) return 0; m->data = p; m->log2cap = log2cap; m->len = 0; + memset(p, 0, (size_t)flan_map_ctrl_bytes(cap)); { - flan_map_geom g; - flan_map_geometry(m, ksize, vsize, cap, &g); - memset(g.hs, 0, - (size_t)flan_cells_bytes((int64_t)sizeof(flan_map_hash), cap)); + int64_t *head = (int64_t *)p; + head[0] = flan_map_threshold_of(cap); + flan_map_slot(ksize, vsize, &head[1], &head[2]); } return 1; } -/* Double the capacity and reinsert. The seed moves with the block, so every - * hash is recomputed rather than carried over — which is what a fresh seed per - * block is for. The old block is released only after the last read of it. */ -static int8_t flan_map_grow(flan_map *m, int64_t want, int64_t ksize, - int64_t vsize, flan_hash_fn hash) { +/* Rebuild into a fresh block of 2^log2cap slots: a grow when the capacity is + * larger, and a sweep of the deleted slots when it is the same. The seed moves + * with the block, so every key is rehashed. The old block is released only + * after the last read of it, and a failed allocation leaves the map exactly as + * it was — map-exhausted.flan retries against it. */ +static int8_t flan_map_rebuild(flan_map *m, int64_t log2cap, int64_t ksize, + int64_t vsize, flan_hash_fn hash) { flan_allocator *a = flan_map_adopt(m); flan_map fresh; - int64_t log2cap = FLAN_MAP_MIN_LOG2, old_cap = flan_map_cap(m); + int64_t old_cap = flan_map_cap(m); flan_map_geom g; - int64_t i, moved; + int64_t i; void *xfer = NULL; - /* Smallest power of two whose 75% threshold still holds [want]. */ - while ((((int64_t)1 << log2cap) * FLAN_MAP_LOAD_FACTOR) / 100 < want) { - if (log2cap >= 40) return 0; - log2cap++; - } - if (log2cap <= m->log2cap && m->data) return 1; - fresh.data = NULL; fresh.len = 0; fresh.log2cap = 0; fresh.alloc = a; fresh.epoch = (int64_t)a->epoch; if (!flan_map_alloc(&fresh, a, log2cap, ksize, vsize)) return 0; if (m->data) { flan_map_geometry(m, ksize, vsize, old_cap, &g); - moved = m->len; - for (i = 0; i < old_cap && moved > 0; i++) { + for (i = 0; i < old_cap; i++) { uint64_t h; - if (g.hs[i] == 0) continue; - h = hash(flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, i), - flan_map_seed(&fresh), ksize, &xfer) | FLAN_MAP_OCCUPIED; - flan_map_place(&fresh, h, - flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, i), - flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, i), - ksize, vsize); - fresh.len++; - moved--; + if (!(g.ctrl[i] & FLAN_CTRL_FULL)) continue; + h = hash(flan_map_k(&g, i), flan_map_seed(&fresh), ksize, &xfer); + flan_map_place(&fresh, h, flan_map_k(&g, i), flan_map_v(&g, i), ksize, + vsize); } if (a->caps & FLAN_CAN_FREE) a->proc(a, FLAN_ALLOC_FREE, m->data, - flan_map_block_size(ksize, vsize, old_cap), 0, - FLAN_MAP_CACHE_LINE); + flan_map_block_size(ksize, vsize, old_cap), 0, FLAN_MAP_ALIGN); } m->data = fresh.data; m->log2cap = fresh.log2cap; - m->len = fresh.len; - /* Every key and value moved, so any pointer into the old block is stale — - * the same word, bumped for the same reason, as a Vec's reallocation. */ return 1; } +/* The smallest capacity whose threshold holds [want] entries, never below + * the map's own. 0 when no capacity the block size can describe does. */ +static int64_t flan_map_log2_for(const flan_map *m, int64_t want) { + int64_t log2cap = FLAN_MAP_MIN_LOG2; + while (flan_map_threshold_of((int64_t)1 << log2cap) < want) { + if (log2cap >= 40) return 0; + log2cap++; + } + if (m->data && log2cap < m->log2cap) log2cap = m->log2cap; + return log2cap; +} + int8_t flan_map_init(flan_map *m, flan_allocator *a, int64_t ksize, int64_t vsize, const uint8_t *loc, int64_t loclen) { if (!a) flan_null_alloc_fail(loc, loclen); @@ -2720,34 +2666,62 @@ int8_t flan_map_init(flan_map *m, flan_allocator *a, int64_t ksize, /* The upsert. spec-memory.md: it either inserts or replaces, and returns Unit * — there is no Result and no ignorable error code, because a put that put * nothing and said nothing is the outcome the StorageExhausted rule exists to - * make impossible. */ + * make impossible. + * + * An insert that would land in an empty slot with no growth left rebuilds + * first. Deleted slots count against the growth, since each one is a slot a + * probe cannot stop at; when they are most of what is using it up — the live + * entries fit in 25/32 of the capacity — the rebuild keeps + * the capacity and only sweeps them out, so a map that churns at a steady size + * does not double without end. */ int8_t flan_map_put(flan_map *m, const void *key, const void *val, int64_t ksize, int64_t vsize, flan_hash_fn hash, flan_eq_fn eq, const uint8_t *loc, int64_t loclen) { - int64_t at; + int64_t at = -1, free_at = -1, mask; flan_map_geom g; - uint64_t h; + uint64_t h = 0; flan_map_check(m, loc, loclen); - at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g); - if (at >= 0) { - /* Replace. The key already in the block compares equal to the one handed - * in, so it is left alone: overwriting it would be a no-op for every - * bytewise key and a question nobody has asked for the others. */ - if (vsize > 0) - flan_copy_small( - flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, at), val, vsize); + if (m->data) { + at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g, &h, &free_at); + if (at >= 0) { + /* Replace. The key already in the block compares equal to the one + * handed in, so it is left alone: overwriting it would be a no-op for + * every bytewise key and a question nobody has asked for the others. */ + if (vsize > 0) flan_copy_small(flan_map_v(&g, at), val, vsize); + return 1; + } + } + + if (!m->data + || (*flan_map_growth(m) == 0 && g.ctrl[free_at] == FLAN_CTRL_EMPTY)) { + int64_t log2cap; + if (m->data && m->len * 32 <= flan_map_cap(m) * 25) + log2cap = m->log2cap; + else + log2cap = flan_map_log2_for(m, m->data ? flan_map_threshold(m) + 1 + : m->len + 1); + if (log2cap == 0) { + flan_fail_bytes = FLAN_BYTES_UNREPRESENTABLE; + flan_fail_align = FLAN_MAP_ALIGN; + flan_fail_id = (int64_t)(intptr_t)flan_map_adopt(m); + return 0; + } + if (!flan_map_rebuild(m, log2cap, ksize, vsize, hash)) return 0; + { + void *xfer = NULL; + h = hash(key, flan_map_seed(m), ksize, &xfer); + } + flan_map_place(m, h, key, val, ksize, vsize); + m->len++; return 1; } - if (!m->data || m->len + 1 > flan_map_threshold(m)) - if (!flan_map_grow(m, m->len + 1, ksize, vsize, hash)) return 0; - - { - void *xfer = NULL; - h = hash(key, flan_map_seed(m), ksize, &xfer) | FLAN_MAP_OCCUPIED; - } - flan_map_place(m, h, key, val, ksize, vsize); + mask = flan_map_cap(m) - 1; + if (g.ctrl[free_at] == FLAN_CTRL_EMPTY) (*flan_map_growth(m))--; + flan_map_set_ctrl(g.ctrl, mask, free_at, flan_map_tag(h)); + flan_copy_small(flan_map_k(&g, free_at), key, ksize); + if (vsize > 0) flan_copy_small(flan_map_v(&g, free_at), val, vsize); m->len++; return 1; } @@ -2761,81 +2735,68 @@ int8_t flan_map_get(flan_map *m, const void *key, void *out, int64_t ksize, int64_t at; flan_map_geom g; flan_map_check(m, loc, loclen); - /* The geometry the probe already built, rather than a second helping of the - * same arithmetic: it was a fifth of the operation, computed twice. */ - at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g); + if (!m->data || m->len == 0) return 0; + at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g, NULL, NULL); if (at < 0) return 0; - if (vsize > 0) - flan_copy_small( - out, flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, at), vsize); + if (vsize > 0) flan_copy_small(out, flan_map_v(&g, at), vsize); return 1; } int8_t flan_map_has(flan_map *m, const void *key, int64_t ksize, int64_t vsize, flan_hash_fn hash, flan_eq_fn eq, const uint8_t *loc, int64_t loclen) { + flan_map_geom g; flan_map_check(m, loc, loclen); - return (int8_t)(flan_map_find_g(m, key, ksize, vsize, hash, eq, NULL) >= 0); + if (!m->data || m->len == 0) return 0; + return (int8_t)(flan_map_find_g(m, key, ksize, vsize, hash, eq, &g, NULL, NULL) >= 0); } -/* Removal, by backward shift, which is what keeps this file tombstone-free. +/* Removal, and the one place a Swiss table needs a third state. * - * Odin's own erase (base/runtime/dynamic_map_internal.odin, map_erase_dynamic) - * marks a tombstone and leaves the repair to the next insert, which is why its - * insert carries a second loop this file has never had. Read rather than - * recalled: the note in this file that said Odin deletes by backward shift was - * describing its *insert*. The trade is the usual one — erase is O(1) there - * and the shift is here, and every lookup there pays a tombstone test this one - * does not. + * A lookup stops at the first group holding an empty slot, so emptying a slot + * could hide a key that was placed past it while it was full. The slot becomes + * DELETED instead — a slot a probe walks through and an insert may reuse — + * unless no probe can ever have walked through it: when every eight-slot + * window containing it already holds an empty slot, any probe that reached it + * stopped in that window, and it can be emptied outright. The test is that the + * empties just before and just after it are fewer than eight slots apart, and + * it is what keeps a map that is far from full from accumulating deleted slots + * at all. * - * The invariant Robin Hood lookups depend on is that no live element is ever - * separated from its home slot by an empty one: the probe stops at the first - * empty slot, so a hole left in the middle of a run would hide everything - * after it. So the hole walks forward: each following element that is not - * already home moves back one slot, and the walk stops at the first slot that - * is empty or whose occupant is already home — neither can be moved back, and - * neither can be hiding anything. + * Nothing moves. The old table shifted the rest of the run back one slot, so + * a cursor already past them stepped over entries; here an entry stays in its + * slot until a rebuild, and removal during iteration visits every surviving + * entry once. * * It releases nothing. A key and a value live inside the one block the map * allocated, so there is no per-entry allocation to hand back and nothing here * asks the allocator for anything — which is what makes removal from a map * backed by an arena, or by any allocator that refuses can-free, mean exactly - * what it means for a heap-backed one. The block is released only by free and - * by the grow that replaces it. + * what it means for a heap-backed one. * * [out] takes a copy of the value that was there, or is NULL when the caller - * does not want one. A cursor held across this is invalidated the way a put - * that grows invalidates one: the shift moves entries to lower slots, and an - * iteration resuming at a higher index would step over them. */ + * does not want one. */ int8_t flan_map_remove(flan_map *m, const void *key, void *out, int64_t ksize, int64_t vsize, flan_hash_fn hash, flan_eq_fn eq, const uint8_t *loc, int64_t loclen) { flan_map_geom g; - int64_t at, mask, pos; + int64_t at, mask; + uint64_t before, after; flan_map_check(m, loc, loclen); - at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g); + if (!m->data || m->len == 0) return 0; + at = flan_map_find_g(m, key, ksize, vsize, hash, eq, &g, NULL, NULL); if (at < 0) return 0; - if (vsize > 0 && out) - flan_copy_small( - out, flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, at), vsize); + if (vsize > 0 && out) flan_copy_small(out, flan_map_v(&g, at), vsize); mask = flan_map_cap(m) - 1; - pos = at; - for (;;) { - int64_t next = (pos + 1) & mask; - uint64_t eh = g.hs[next]; - if (eh == 0 || flan_map_distance(eh, next, mask) == 0) { - g.hs[pos] = 0; - break; - } - flan_copy_small(flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, pos), - flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, next), - ksize); - if (vsize > 0) - flan_copy_small( - flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, pos), - flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, next), vsize); - g.hs[pos] = eh; - pos = next; + after = flan_group_empty(flan_group_load(g.ctrl + at)); + before = flan_group_empty( + flan_group_load(g.ctrl + ((at - FLAN_MAP_GROUP) & mask))); + if (after && before + && flan_mask_first(after) + flan_mask_last_gap(before) < FLAN_MAP_GROUP) { + flan_map_set_ctrl(g.ctrl, mask, at, FLAN_CTRL_EMPTY); + (*flan_map_growth(m))++; + } else { + flan_map_set_ctrl(g.ctrl, mask, at, FLAN_CTRL_DELETED); } m->len--; return 1; @@ -2847,35 +2808,22 @@ int64_t flan_map_len(flan_map *m, const uint8_t *loc, int64_t loclen) { } /* The cursor step, and the whole of iteration. - * - * Everything else in this file addresses *one* entry: get, put and has each - * hash a key and probe. Nothing walked the block, so a map's keys and its - * values could not be read out at all, and this is the one function that - * changes it. * * The cursor is a slot index the caller owns, and the contract is the one a * slot index gives for free: it starts at 0, it is written back one past the * entry just answered, and a 0 answer leaves it at [cap] so calling again is * still 0. There is no iterator struct because there is nothing for one to - * hold — a map has no tombstones, so no state beyond the position is needed to - * know where to resume. + * hold: the control byte says whether a slot is full, and nothing else is + * needed to know where to resume. * * Invalidated by anything that moves the block, exactly as a Vec's slice is: - * a put that grows rehashes into a new block and every index before it means a - * different entry. A remove is the same hazard without the reallocation — its - * backward shift moves entries to lower slots, and a cursor already past them - * steps over entries it has not answered. The epoch check below catches a - * released arena and nothing catches either of these, which is the same - * bargain [slice] already makes. + * a put that rebuilds rehashes into a new block and every index before it + * means a different entry. A remove moves nothing and does not invalidate it. + * The epoch check below catches a released arena and nothing catches a + * rebuild, which is the same bargain [slice] already makes. * - * The layout is the one the geometry describes and is worth restating because - * it is the thing most likely to be got wrong here: [data] is *one* allocation - * laid out keys | values | hashes | scratch, each run cell-packed, so a key is - * reached through [flan_cell_at] and never by [ks + i * ksize]. The hashes are - * the exception the clone loop already relies on — an 8-byte element packs 8 - * to a 64-byte cell with nothing left over, so a flat index is the right - * index. Order is block order, which is the hash's order and not the - * insertion's; two maps holding the same entries may walk them differently. */ + * Order is block order, which is the hash's order and not the insertion's; + * two maps holding the same entries may walk them differently. */ int8_t flan_map_next(flan_map *m, int64_t *cursor, void *kout, void *vout, int64_t ksize, int64_t vsize, const uint8_t *loc, int64_t loclen) { @@ -2889,11 +2837,9 @@ int8_t flan_map_next(flan_map *m, int64_t *cursor, void *kout, void *vout, if (i >= cap) { *cursor = cap; return 0; } flan_map_geometry(m, ksize, vsize, cap, &g); for (; i < cap; i++) { - if (g.hs[i] == 0) continue; - memcpy(kout, flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, i), - (size_t)ksize); - memcpy(vout, flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, i), - (size_t)vsize); + if (!(g.ctrl[i] & FLAN_CTRL_FULL)) continue; + memcpy(kout, flan_map_k(&g, i), (size_t)ksize); + memcpy(vout, flan_map_v(&g, i), (size_t)vsize); *cursor = i + 1; return 1; } @@ -2901,14 +2847,26 @@ int8_t flan_map_next(flan_map *m, int64_t *cursor, void *kout, void *vout, return 0; } -/* Room for [n] entries without reallocating, which means a block whose 75% - * threshold is at least n. */ +/* Room for [n] entries without reallocating: a block whose threshold is at + * least n, and with enough growth left that the entries not yet in it will + * not use it up. A map carrying deleted slots may have the capacity and not + * the growth, and is rebuilt at the same capacity to sweep them. */ int8_t flan_map_reserve(flan_map *m, int64_t n, int64_t ksize, int64_t vsize, flan_hash_fn hash, const uint8_t *loc, int64_t loclen) { + int64_t log2cap; flan_map_check(m, loc, loclen); if (n <= 0) return 1; - if (m->data && n <= flan_map_threshold(m)) return 1; - return flan_map_grow(m, n, ksize, vsize, hash); + if (m->data && n <= flan_map_threshold(m) + && n - m->len <= *flan_map_growth(m)) + return 1; + log2cap = flan_map_log2_for(m, n); + if (log2cap == 0) { + flan_fail_bytes = FLAN_BYTES_UNREPRESENTABLE; + flan_fail_align = FLAN_MAP_ALIGN; + flan_fail_id = (int64_t)(intptr_t)flan_map_adopt(m); + return 0; + } + return flan_map_rebuild(m, log2cap, ksize, vsize, hash); } /* spec-memory.md's first release point, and the same rules the Vec's free @@ -2920,7 +2878,7 @@ void flan_map_free(flan_map *m, int64_t ksize, int64_t vsize, if (m->data && m->alloc && (m->alloc->caps & FLAN_CAN_FREE)) m->alloc->proc(m->alloc, FLAN_ALLOC_FREE, m->data, flan_map_block_size(ksize, vsize, flan_map_cap(m)), 0, - FLAN_MAP_CACHE_LINE); + FLAN_MAP_ALIGN); m->data = NULL; m->len = 0; m->log2cap = 0; @@ -2930,33 +2888,28 @@ void flan_map_free(flan_map *m, int64_t ksize, int64_t vsize, /* A deep, independent copy. It reinserts rather than copying the block: the * seed is derived from the block address, so a bytewise copy would be a map - * whose stored hashes disagree with its own seed and whose every lookup - * missed. Reinserting is also what makes the copy's layout independent of the - * original's insertion history. */ + * whose control bytes disagree with its own seed and whose every lookup + * missed. Reinserting also leaves the copy with no deleted slots. */ int8_t flan_map_clone(flan_map *dst, flan_map *src, flan_allocator *a, int64_t ksize, int64_t vsize, flan_hash_fn hash, const uint8_t *loc, int64_t loclen) { flan_map_geom g; - int64_t cap, i, moved; + int64_t cap, i, log2cap; void *xfer = NULL; flan_map_check(src, loc, loclen); if (!flan_map_init(dst, a, ksize, vsize, loc, loclen)) return 0; if (!src->data || src->len == 0) return 1; - if (!flan_map_grow(dst, src->len, ksize, vsize, hash)) return 0; + log2cap = flan_map_log2_for(dst, src->len); + if (!flan_map_rebuild(dst, log2cap, ksize, vsize, hash)) return 0; cap = flan_map_cap(src); flan_map_geometry(src, ksize, vsize, cap, &g); - moved = src->len; - for (i = 0; i < cap && moved > 0; i++) { + for (i = 0; i < cap; i++) { uint64_t h; - if (g.hs[i] == 0) continue; - h = hash(flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, i), - flan_map_seed(dst), ksize, &xfer) | FLAN_MAP_OCCUPIED; - flan_map_place(dst, h, flan_cell_at(g.ks, ksize, g.kepc, g.kcell, g.kshift, i), - flan_cell_at(g.vs, vsize, g.vepc, g.vcell, g.vshift, i), - ksize, vsize); + if (!(g.ctrl[i] & FLAN_CTRL_FULL)) continue; + h = hash(flan_map_k(&g, i), flan_map_seed(dst), ksize, &xfer); + flan_map_place(dst, h, flan_map_k(&g, i), flan_map_v(&g, i), ksize, vsize); dst->len++; - moved--; } return 1; } @@ -2984,7 +2937,7 @@ void flan_dev_reg_note_vec(flan_vec *v, int64_t size, const char *type, void flan_dev_reg_note_map(flan_map *m, int64_t ksize, int64_t vsize, const char *type, int64_t typelen) { if (!m || !m->data) return; - /* One block holding the hashes, the keys and the values, so the element + /* One block holding the control bytes, the keys and the values, so the element size is meaningless here and is passed as 0: an address inside it is "+n into" rather than "[i] of". */ flan_dev_reg_note(m->data, diff --git a/spec-memory.md b/spec-memory.md index ac5e57d2..c7a100dd 100644 --- a/spec-memory.md +++ b/spec-memory.md @@ -20,7 +20,7 @@ facility (see plan.org, "Managed classes"). | `[n T]` | n contiguous `T` | copies | no (inline) | — | | `[T]` | ptr + len | copies the *view* | no | — | | `(Vec T)` | ptr + len + cap | **moves** | yes | stored | -| `(Map K V)` | open-addressed, flat key/value arrays | **moves** | yes | stored | +| `(Map K V)` | open-addressed, one control byte a slot, key and value side by side | **moves** | yes | stored | - `[n T]` is a value. It lives wherever it is declared, copies on assignment and on pass-by-value, and is what `defconst colors [4 u32] ...` and diff --git a/test/programs/map-iter.flan b/test/programs/map-iter.flan index 29392b5c..9d8226ec 100644 --- a/test/programs/map-iter.flan +++ b/test/programs/map-iter.flan @@ -82,7 +82,7 @@ (print chars) (print " ") (print xs) (print " ") (print ys) (println "")) (free m))) -;; Growth past the 75% threshold rehashes into a new block, so this walks a map +;; Growth past the load factor rehashes into a new block, so this walks a map ;; whose layout is nothing like its insertion order and at a capacity several ;; doublings past the minimum. (defn after-growth [] () diff --git a/test/programs/map-remove.flan b/test/programs/map-remove.flan index eaf1def2..349eafd2 100644 --- a/test/programs/map-remove.flan +++ b/test/programs/map-remove.flan @@ -1,17 +1,26 @@ ;;;; (map-remove m k) — the operation a Map has been missing. ;;;; -;;;; Removal is the one map operation that can break the *other* ones: Robin -;;;; Hood lookups stop at the first empty slot, so a hole punched in the middle -;;;; of a probe run hides every entry after it, and the entries it hides are -;;;; found by no test that only removes and asks about what it removed. So the -;;;; rows below are mostly about the survivors. +;;;; Removal is the one map operation that can break the *other* ones: a +;;;; lookup stops at the first group of slots holding an empty one, so a slot +;;;; emptied in the middle of a probe sequence hides every entry placed past +;;;; it, and the entries it hides are found by no test that only removes and +;;;; asks about what it removed. So the rows below are mostly about the +;;;; survivors. ;;;; -;;;; Odin marks a tombstone and repairs on the next insert; this shifts the run -;;;; back and stays tombstone-free, which is the arrangement the rest of the -;;;; file already assumed. A version that punched the hole and left it passes -;;;; row 1 and fails row 3. +;;;; A removed slot is marked deleted, and emptied outright only where no probe +;;;; can have passed through it. A version that emptied every removed slot +;;;; passes row 1 and fails row 3; a version whose deleted slots were never +;;;; swept grows without end under row 7. (defstruct Cell [x i32 y i32]) +;; Row 7's allocator and its count of refusals. Globals, because a handler +;; cannot see the locals of the function that established it. +(defonce churn-alloc Allocator) +(defonce churn-refusals i64) +;; Row 7's reference: a global array rather than a Vec, because a Vec would +;; take its block from the same process heap the budget is counting. +(defonce churn-ref [512 i64]) + (defn main [] i32 ;; (1) The value comes back, the length drops, and the key is gone. An ;; absent key is None and changes nothing — the same answer get gives, since @@ -41,9 +50,8 @@ ;; (3) The survivors, which is the row that matters. 2000 entries is eight ;; grows, so the runs are long and interleaved; removing the even keys and - ;; then asking after every odd one is asking whether any probe run was cut. - ;; A backward shift that stopped one slot early loses entries here and - ;; nowhere else. + ;; then asking after every odd one is asking whether any probe sequence was + ;; cut. (let [m (map-new i32 i64)] (dotimes [i 2000] (put m i (* (i64 i) 3))) (let [taken 0] @@ -90,9 +98,7 @@ (print (length s)) (println "") ; 1 (free s)) - ;; (5) Iteration after removals answers exactly the survivors. The cursor is - ;; started fresh — a cursor held *across* a removal is invalidated by the - ;; shift, the same way a put that grows invalidates one. + ;; (5) Iteration after removals answers exactly the survivors. (let [m (map-new i32 i32)] (dotimes [i 100] (put m i i)) (dotimes [i 100] (if (= 0 (% i 3)) (map-remove m i))) @@ -118,4 +124,64 @@ (match (get t 299) (Some v) (do (print v) (println "")) None (println "?")) ; 598 (print (has-key? t 298)) (println ""))) ; false (free-all ar)) + + ;; (7) Churn at a steady size, checked against a plain array. Keys are drawn + ;; from 512 and the map holds about two thirds of them, so removals leave + ;; deleted slots that inserts must reuse and rebuilds must sweep. Every get is + ;; checked against the array, and at the end so is the length and a full walk. + ;; + ;; The process heap is given a budget of 20000 bytes, with nothing else of + ;; this program's live on it. An i32 key and an i64 + ;; value make a 16-byte slot, so the 512-slot block is 8768 bytes and a + ;; rebuild at that size holds two of them, 17536; a grow to 1024 slots holds + ;; 8768 and 17472 at once and goes over. The entries never need more than 512 + ;; slots, so a refusal means deleted slots were used up and the table doubled + ;; instead of sweeping them — which is what a version that never swept them + ;; does. The handler lets it through, counts it, and the count is printed. + (set churn-alloc (heap-allocator)) + (set-alloc-budget churn-alloc 20000) + (handler-bind + [(StorageExhausted [c] + (set churn-refusals (+ churn-refusals 1)) + (set-alloc-budget churn-alloc (* 2 (alloc-budget churn-alloc))) + (invoke-restart 'retry))] + (let [m (map-new i32 i64 churn-alloc) + x (u64 12345) + bad 0 + live 0] + (dotimes [i 512] (set (at churn-ref i) (i64 -1))) + (dotimes [step 200000] + (set x (+ (* x (u64 6364136223846793005)) (u64 1442695040888963407))) + (let [k (i32 (% (>> x 33) (u64 512))) + op (i32 (% (>> x 20) (u64 4)))] + (if (< op 2) + (do (if (< (at churn-ref k) 0) (set live (+ live 1))) + (put m k (i64 step)) + (set (at churn-ref k) (i64 step))) + (if (= op 2) + (match (map-remove m k) + (Some v) (do (if (not (= v (at churn-ref k))) (set bad (+ bad 1))) + (set (at churn-ref k) (i64 -1)) + (set live (- live 1))) + None (if (>= (at churn-ref k) 0) (set bad (+ bad 1)))) + (match (get m k) + (Some v) (if (not (= v (at churn-ref k))) (set bad (+ bad 1))) + None (if (>= (at churn-ref k) 0) (set bad (+ bad 1)))))))) + (print bad) (println "") ; 0 + (print (= (length m) (i64 live))) (println "") ; true + (let [cur (i64 0) k 0 v (i64 0) seen 0] + (while (map-next m (addr cur) (addr k) (addr v)) + (set seen (+ seen 1)) + (if (not (= v (at churn-ref k))) (set bad (+ bad 1)))) + (print (= seen live)) (println "") ; true + (print bad) (println "")) ; 0 + ;; Removing while walking visits every survivor once: nothing moves. + (let [cur (i64 0) k 0 v (i64 0) seen 0] + (while (map-next m (addr cur) (addr k) (addr v)) + (set seen (+ seen 1)) + (map-remove m k)) + (print (= seen live)) (println "") ; true + (print (length m)) (println "")) ; 0 + (free m))) + (print churn-refusals) (println "") ; 0 0) diff --git a/test/programs/maps.flan b/test/programs/maps.flan index 5a235601..1604c72e 100644 --- a/test/programs/maps.flan +++ b/test/programs/maps.flan @@ -1,7 +1,7 @@ ;;;; (Map K V) — spec-memory.md, step 4 of the container build order. ;;;; -;;;; Odin's map: open-addressed Robin Hood hashing at a 75% load factor, with -;;;; cache-line cell packing. Every claim below is one a plausible wrong +;;;; An open-addressed Swiss table: one control byte a slot, key and value side +;;;; by side. Every claim below is one a plausible wrong ;;;; version gets wrong, and the numbers differ per failure so a single wrong ;;;; answer names its own cause. (defstruct Cell [x i32 y i32]) diff --git a/test/test_acceptance.ml b/test/test_acceptance.ml index 7cda52a1..a8c6d82d 100644 --- a/test/test_acceptance.ml +++ b/test/test_acceptance.ml @@ -3968,8 +3968,7 @@ level "1" (* ── (Map K V), spec-memory.md step 4 ────────────────────────── - Odin's map: open-addressed Robin Hood hashing at a 75% load factor with - cache-line cell packing. maps.flan is seven claims, each one a plausible + An open-addressed Swiss table, one control byte a slot. maps.flan is seven claims, each one a plausible wrong version gets wrong, and the numbers differ per failure. The two worth naming, because nothing else in the suite would catch @@ -4008,9 +4007,9 @@ level "1" outputs "map iteration" "programs/map-iter.flan" map_iter_out; outputs ~opt:"-O0" "map iteration, -O0" "programs/map-iter.flan" map_iter_out; - (* Removal, which is the operation that can break the others. A Robin Hood - probe stops at the first empty slot, so a hole left in the middle of a - run hides every entry past it — and the hidden ones are exactly what a + (* Removal, which is the operation that can break the others. A probe + stops at the first group holding an empty slot, so a slot emptied in + the middle of a probe sequence hides every entry past it — and the hidden ones are exactly what a test that only asks after what it removed never looks at. Hence the third row: 2000 entries, the even keys taken out, and then every odd one asked for. A removal that punched the hole and left it answers row one @@ -4021,7 +4020,8 @@ level "1" into the wrong half of it. *) let map_remove_out = "100\n1\nfalse\ngone\n1\n0\n77\n1\n1000\n1000\n0\n0\n2000\n\ - 709\nfalse\ntrue\n399\n1\ntrue\n1\n66\n0\n66\n150\n598\nfalse\n" + 709\nfalse\ntrue\n399\n1\ntrue\n1\n66\n0\n66\n150\n598\nfalse\n\ + 0\ntrue\ntrue\n0\ntrue\n0\n0\n" in outputs "map removal" "programs/map-remove.flan" map_remove_out; outputs ~opt:"-O0" "map removal, -O0" "programs/map-remove.flan" diff --git a/test/test_valgrind.ml b/test/test_valgrind.ml index 9b8c4419..2e656c3c 100644 --- a/test/test_valgrind.ml +++ b/test/test_valgrind.ml @@ -275,10 +275,11 @@ let corpus = down: [check_at] is behind the flag and so is *half* of [check_slice] — the [hi <= len] compare — while its [lo <= hi] stays, because that one is the claim that a slice's length word is a count and not a bounds check at all. - A Vec's and a Map's bounds checks are not behind it either: they live inside - flan_vec_at and the map probe in flan_rt.c, are plain C, and run in every - build. So the flag lowers the guard on indexing a fixed array or a slice, - and on nothing else. These are the programs where that distinction reaches + A Vec's bounds check is not behind it either: it lives inside flan_vec_at, + is plain C, and runs in every build. A Map has no bounds check to lower — + its probe is masked to the capacity and an absent key is None. So the flag + lowers the guard on indexing a fixed array or a slice, and on nothing + else. These are the programs where that distinction reaches heap storage. *) let unchecked_subset = [ "programs/allocators.flan", [];