A Map is a Swiss table, one control byte a slot with key and value side by side

Removal marks a deleted slot unless no probe can have passed it, and a rebuild at the same capacity sweeps them. The header, the entry points and the iteration contract are unchanged. A Map has no bounds check, so the entry asking to convert one is closed with nothing to convert.
This commit is contained in:
Joseph Ferano 2026-09-25 10:56:56 +07:00
parent bf827dc55b
commit cbfac9f474
8 changed files with 603 additions and 500 deletions

View File

@ -1214,10 +1214,13 @@ value an unsigned compare waves through.
lock the agent's listener thread may hold inside =dlopen=, so a program that
should die could hang. Unconditionally rather than only under =--dev=.
** TODO A Map's bounds check and the stale-container failure still die
Deliberate for the stale case — the region the container lived in was released and
there is no frame to go back to that would not read freed memory. The Map path was
left alone rather than converted half-way.
** DONE A Map's bounds check and the stale-container failure still die
CLOSED: [2026-09-25]
A Map has no bounds check: =get= and =map-remove= answer =None= for an absent key
and nothing on the map path indexes by a number, so there was nothing to convert.
The stale-container failure keeps dying — the region was released and there is no
frame to go back to that would not read freed memory. Rules out a signalled
condition on the stale path.
** DONE Six trap paths park instead of killing the session
CLOSED: [2026-09-18]
@ -1269,15 +1272,17 @@ practice because the freed block still holds the bumped value. A question about
** DONE Map removal costs a backward-shift loop
Removal landed, with the loop the spec predicted as its cost. Deferring it was
what had kept the implementation free of tombstones and of Odin's
backward-shift loop; taking it is taking the loop.
backward-shift loop; taking it is taking the loop. Superseded by the Swiss table,
which removes by tombstone and moves nothing.
** NEXT The Map is slower than CPython's dict at a million entries
Decided 2026-09-25: build the Swiss-table layout, one metadata byte a slot. Measured against the current Map at small, game and large sizes before and after.
Keys, values and hashes are three separate runs, so a lookup that misses
everything costs three cache misses where a compact dict costs two. One byte of
metadata a slot — the Swiss-table arrangement — is the known answer and is not
built. The crossover is somewhere between ten thousand and a million and nobody
has found it.
** DONE The Map is slower than CPython's dict at a million entries
CLOSED: [2026-09-25]
The Map is a Swiss table: one control byte a slot, key and value side by side,
groups of eight probed as one 64-bit word, seven-eighths load, removal by
tombstone with a same-capacity rebuild to sweep them. Header, entry points and
iteration contract unchanged. Rules out Robin Hood, the backward shift and
separate key and value runs; SSE2 groups are not built. See docs/BUILT.md, "The
Map is a Swiss table".
** DONE dyn maps and interned keywords
CLOSED: [2026-09-20]

View File

@ -3013,6 +3013,10 @@ restriction under it.
### The three properties that were the point
**Superseded.** The table below is no longer a Robin Hood map: it is a Swiss table, and "The Map is a Swiss table"
further down says what replaced what and why. This section and the next are kept for the reasoning they record; the
hash-and-equality pair, the key restrictions and the header are unchanged.
Odin's header states them and they are why it was the thing to follow (`base/runtime/dynamic_map_internal.odin`).
- **Open-addressed Robin Hood hashing at a 75% load factor.** No buckets and no per-entry allocation: one block holds
@ -3029,14 +3033,11 @@ Odin's header states them and they are why it was the thing to follow (`base/run
### Two departures from Odin, both deliberate
**There are still no tombstones, now that removal exists.** A slot is empty or occupied and nothing else, and
`flan_map_remove` keeps it that way by shifting the run back over the hole rather than marking it. Odin does the
opposite, and this paragraph used to say otherwise: read out of
`base/runtime/dynamic_map_internal.odin`, `map_erase_dynamic` sets a tombstone bit and leaves the repair to the next
insert, which is why Odin's *insert* carries a backward-shift loop and its every lookup tests for a tombstone. The
trade is the usual one — erase is O(1) there and the shift is here, and the lookups, which outnumber the removals,
pay nothing. What `spec-memory.md` still defers is the rest of its sentence: move-aware lookup and owned entries. A
removed value is copied out, and nothing is dropped.
**Removal was a backward shift, and is now a tombstone.** While the table was Robin Hood, a slot was empty or occupied
and `flan_map_remove` shifted the run back over the hole. The Swiss table has a third control state, deleted, and
empties a slot outright only where no probe can have passed through it; see "The Map is a Swiss table". What
`spec-memory.md` still defers is unchanged: move-aware lookup and owned entries. A removed value is copied out, and
nothing is dropped.
**The header does not tag the capacity into the data pointer.** Odin stuffs `log2cap` into the low six bits because
its `Raw_Map` must be three words. This header already carries an allocator, a generation and an epoch, so the tagging
@ -3044,10 +3045,10 @@ would buy nothing, cost a mask on every access, and — the part that actually m
block being 64-byte aligned. Alignment is *requested*; an arena whose base is not cache-aligned now gives a slower map
rather than a wrong one.
The header is six words, 48 bytes, the same as a `Vec`'s and for the same reason: a layout that changes with a build
The header is five words, 40 bytes, the same as a `Vec`'s and for the same reason: a layout that changes with a build
flag can disagree across the reload boundary.
data len log2cap allocator gen epoch
data len log2cap allocator epoch
### The hash and equality pair, and why most key types do not get one
@ -3157,7 +3158,7 @@ Measured on this machine, `i64` to `i64`, against CPython 3.13's dict on the sam
left out. Both are waiting on memory there, and this layout waits longer: keys, values and hashes are three separate
runs, so a lookup that misses everything takes three cache misses where a compact dict takes two, and the hash run is a
full eight bytes a slot. Cell packing buys probe locality, which is a win while the hash run is resident and a loss
once nothing is. One byte of metadata a slot — the Swiss-table arrangement — is the known answer and is not built.
once nothing is. One byte of metadata a slot — the Swiss-table arrangement — is the known answer, and the next section is it being built.
The path from 35 ns to 18 ns (the cache-resident floor, at 500 entries) was **profiled, and the first two guesses were
both wrong**: the per-slot cell division and the block-size divisions were each replaced first and neither moved the
@ -3168,10 +3169,110 @@ twice per lookup; and the seed stopped being a five-multiply avalanche on the cr
again immediately after. `64/size` is a table, which is Odin's `Map_Cell_Info` by another route — Odin precomputes it
per type because the probe loop must not divide, and here the sizes arrive as ordinary arguments.
What remains at 18 ns is the type erasure itself: a non-inlinable call into the runtime and two non-inlinable indirect
What remained at 18 ns was the type erasure itself: a non-inlinable call into the runtime and two non-inlinable indirect
calls to the pair. That is the trade `spec-memory.md` chose deliberately — "It is type-erased on purpose… No generics
are involved, and none are needed" — and monomorphisation is what would buy it back, at the cost the spec declined.
### The Map is a Swiss table
The Robin Hood table above lost to CPython's dict at a million entries, for the reason the last section gave: three
separate runs and an eight-byte hash a slot. It is replaced by a Swiss table. The header, every entry point's
signature, the hash-and-equality pair and the iteration contract are unchanged, so neither backend and neither
debugger description moved; the whole change is inside `runtime/flan_rt.c`.
**The block.**
[growth][stride][value offset][control: cap + 7 bytes] pad to 64 | slots
- **One control byte a slot.** `0x00` is empty, `0x01` deleted, and a full slot is `0x80` with the top seven bits of the
hash below it — Zig's encoding (`lib/std/hash_map.zig`, `Metadata`), chosen because empty is zero: a zeroed control
run is an empty map, the property the old table's zero hash word had. The low bits of the hash choose the group, so
the tag and the position never share a bit.
- **Groups of eight, as one 64-bit word.** Matching a tag, finding an empty slot and finding a free one are each a few
integer operations over the word — hashbrown's portable "generic" group, which needs no intrinsics and so compiles
for every target the runtime does. SSE2's sixteen-wide group is not built; nothing measured asks for it. The seven
bytes past the end mirror the first seven, so a group can be read starting at any slot. Groups are visited at
triangular offsets, which reach every group of a power-of-two table before repeating.
- **A key and its value side by side.** A hit reads the control group and then one slot — two cache misses where the
old table took three. The runtime is not told either alignment, and does not need to be: a size is a multiple of its
type's alignment, so the largest power of two dividing it (capped at 64) bounds the alignment. The value is placed at
the key's size rounded up to that bound, and the stride rounded up to the larger of the two bounds.
- **The head holds three words** because the header has no room for them: it is the emitter's layout as much as the
runtime's. The growth word is how many more entries may land in an *empty* slot before a rebuild; the stride and value
offset are computed once at allocation, because recomputing them from the two sizes on every call measured as about a
tenth of a cache-resident lookup.
- **Seven-eighths load.** A miss stops at the first group holding an empty slot rather than walking a run, and at seven
eighths every eight-slot group has one.
**Removal is a tombstone, most of the time not.** A lookup stops at the first group with an empty slot, so emptying a
slot could hide a key placed past it while it was full. The slot becomes deleted instead, unless every eight-slot
window containing it already holds an empty slot — then no probe can have passed through it, and it is emptied
outright. That is abseil's test (the empties just before and just after are fewer than eight slots apart), and in a map
that is not nearly full it is the common case. Nothing moves, so a cursor held across a removal stays valid; the
backward shift that removal used to cost is gone.
**Deleted slots count against growth**, since a probe cannot stop at one. When an insert would take an empty slot with
no growth left, the table is rebuilt: at the same capacity when the live entries fit in 25/32 of it (abseil's figure),
which only sweeps the deleted slots, and at double otherwise. `map-remove.flan` row 7 churns a map at a steady size for
200,000 operations against a plain array and holds it at 512 slots through 33 sweeps.
**The probe is inlined by force.** Called from five entry points, clang keeps it out of line at `-O2`, and the call
spills everything the loop held in registers.
#### Measured
`i64` to `i64`. Keys present are `2i`, keys absent `2i + 1`. Per size: *insert* fills a fresh map from empty with no
`reserve` and frees it, repeated to two million inserts; *hit* and *miss* are four million lookups cycling through the
keys; *remove* fills a map untimed and removes every key timed, repeated to two million. Every pass counts its wrong
answers and the program prints them, because the first version of this benchmark passed a map by value to a filling
function — the grows landed in the callee's copy of the header, the caller's map stayed empty, and every lookup was a
fast miss. The CPython side is the same workload in Python 3.13.9 (`dict.get`, `dict.pop`).
The figure is the **minimum of nine runs**, the two Flan binaries run alternately and the Python script in a separate
nine-run pass, because this machine is shared — it carried a load average near 20 throughout, from other lanes' test
suites — and a minimum measures the program rather than the neighbours. Both Flan binaries are `-O2`. The magnitudes
carry that load: across three nine-run passes the Robin Hood column moved by as much as half (a 100k hit read 51, 83
and 35 ns). What held is the ordering — the Swiss table ahead of the Robin Hood table in every cell of every pass,
and ahead of CPython in every cell — and that is the claim to reproduce. Nanoseconds per operation:
| Entries | Op | Robin Hood | Swiss | CPython dict |
|---|---|---|---|---|
| 16 | insert | 74.5 | 30.5 | 46.3 |
| | hit | 20.0 | 9.8 | 83.1 |
| | miss | 16.1 | 8.0 | 100.5 |
| | remove | 23.3 | 15.4 | 69.0 |
| 1k | insert | 80.9 | 30.0 | 57.4 |
| | hit | 22.2 | 9.7 | 124.6 |
| | miss | 15.9 | 7.8 | 128.3 |
| | remove | 24.5 | 14.8 | 91.9 |
| 10k | insert | 112.4 | 45.5 | 85.5 |
| | hit | 29.6 | 10.5 | 134.9 |
| | miss | 26.5 | 8.5 | 125.3 |
| | remove | 33.2 | 17.6 | 99.7 |
| 100k | insert | 204.4 | 47.6 | 123.6 |
| | hit | 35.3 | 16.1 | 133.6 |
| | miss | 30.4 | 14.4 | 128.4 |
| | remove | 40.6 | 22.9 | 98.4 |
| 300k | insert | 199.4 | 73.0 | 128.7 |
| | hit | 69.9 | 25.9 | 118.0 |
| | miss | 39.8 | 10.9 | 138.2 |
| | remove | 116.8 | 36.9 | 94.3 |
| 1M | insert | 454.2 | 138.5 | 151.5 |
| | hit | 147.1 | 78.4 | 120.3 |
| | miss | 88.2 | 14.6 | 125.2 |
| | remove | 152.0 | 74.5 | 91.5 |
**The crossover.** Against CPython, the Robin Hood table lost on insert at every size, on remove from somewhere between
100k and 300k, and on hits between 300k and 1M; misses it won throughout. The Swiss table is quicker than CPython at
every size and every operation measured, and quicker than the table it replaced at every one of the twenty-four. The
margin that is left at a million is insert (139 against 152) and remove (75 against 92), and it is narrower than it
looks: CPython hashes a small integer to itself, so this key pattern walks its table in order and the hardware
prefetcher does the rest, while every Flan key lands at a random slot.
A middle variant was measured and not kept: Swiss control bytes over separate key and value runs. It was within a few
nanoseconds of the final layout up to 100k and lost from 300k up (87 ns a hit against 26 at 300k, 100 against 78 at a
million), which is the third cache miss the interleaved slot removes.
## Unions, and the tag they carry
`defdata` parsed and its shape was checked long before this; naming the type (`check.ml:312`) and constructing a
@ -4426,11 +4527,10 @@ that is written through on the way out.
**It is the one map entry point that carries neither a hash nor an equality function.** Walking asks nothing about a
key. The two sizes are still there, because the runtime is type-erased and the block geometry is computed from them.
**The layout, restated, because it is the thing to get wrong here.** `data` is *one* allocation laid out
keys | values | hashes | scratch, each run cell-packed to a cache line — the arrangement the Valgrind lane described
while explaining why a probe overrun is not observable. A key is reached through `flan_cell_at` and never as
`ks + i * ksize`. The hashes are the exception `flan_map_clone` already relies on: an 8-byte element packs 8 to a
64-byte cell with nothing left over, so `g.hs[i]` is the right index and a flat one.
**The layout, restated, because it is the thing to get wrong here.** `data` is *one* allocation: a three-word head,
the control bytes, then the slots, each slot a key and its value side by side at the stride the head records. A slot is
full when its control byte has the top bit set, and a key is reached as `slots + i * stride`, never as
`ks + i * ksize`.
**Order is block order**, which is the hash's order and not the insertion's, and it changes when the map grows.
`programs/map-iter.flan` is therefore written entirely in sums, counts and lengths — every claim in it is order-free,
@ -5657,10 +5757,10 @@ IR (which is also why `--no-bounds-checks` never reached it), so both grew a tra
`Emit`'s `Rt` arm guards those two symbols and no others — they are the only ones in that family that can transfer;
everything else there is arithmetic over a container header.
**A `Map`'s bounds and a `Vec`'s stale-allocator check still die.** `flan_vec_stale_fail` is a different kind of
**A `Vec`'s stale-allocator check still dies, and so does a `Map`'s.** `flan_vec_stale_fail` is a different kind of
failure — the region a container lived in was released, and there is no frame to go back to that would not read freed
memory — and the map path was left alone rather than converted half-way. Written down here so it is a known edge
rather than a discovery.
memory. A `Map` has no bounds check to convert: `get` and `map-remove` answer `None` for an absent key and nothing on
the map path indexes by a number, so the stale check is the only map failure that dies.
## An address answers with a type, and nothing grew a tag word

View File

@ -294,7 +294,7 @@ let rec ll (t : Types.t) =
the Vec's address — so the shape is here only so that a slot, a struct
field and a copy in the IR are the right number of bytes. *)
| Types.Vec _ -> "%vec"
(* data + len + log2cap + allocator + gen + epoch. Six words, exactly as the
(* data + len + log2cap + allocator + epoch. Five words, as many as the
Vec's, and read here for exactly the same reason: nothing in this file
touches a field of one — every operation is a runtime call taking the
map's address — so the shape exists only so that a slot, a struct field

File diff suppressed because it is too large Load Diff

View File

@ -20,7 +20,7 @@ facility (see plan.org, "Managed classes").
| `[n T]` | n contiguous `T` | copies | no (inline) | — |
| `[T]` | ptr + len | copies the *view* | no | — |
| `(Vec T)` | ptr + len + cap | **moves** | yes | stored |
| `(Map K V)` | open-addressed, flat key/value arrays | **moves** | yes | stored |
| `(Map K V)` | open-addressed, one control byte a slot, key and value side by side | **moves** | yes | stored |
- `[n T]` is a value. It lives wherever it is declared, copies on assignment and
on pass-by-value, and is what `defconst colors [4 u32] ...` and

View File

@ -1,15 +1,16 @@
;;;; (map-remove m k) — the operation a Map has been missing.
;;;;
;;;; Removal is the one map operation that can break the *other* ones: Robin
;;;; Hood lookups stop at the first empty slot, so a hole punched in the middle
;;;; of a probe run hides every entry after it, and the entries it hides are
;;;; found by no test that only removes and asks about what it removed. So the
;;;; rows below are mostly about the survivors.
;;;; Removal is the one map operation that can break the *other* ones: a
;;;; lookup stops at the first group of slots holding an empty one, so a slot
;;;; emptied in the middle of a probe sequence hides every entry placed past
;;;; it, and the entries it hides are found by no test that only removes and
;;;; asks about what it removed. So the rows below are mostly about the
;;;; survivors.
;;;;
;;;; Odin marks a tombstone and repairs on the next insert; this shifts the run
;;;; back and stays tombstone-free, which is the arrangement the rest of the
;;;; file already assumed. A version that punched the hole and left it passes
;;;; row 1 and fails row 3.
;;;; A removed slot is marked deleted, and emptied outright only where no probe
;;;; can have passed through it. A version that emptied every removed slot
;;;; passes row 1 and fails row 3; a version whose deleted slots were never
;;;; swept grows without end under row 7.
(defstruct Cell [x i32 y i32])
(defn main [] i32
@ -41,9 +42,8 @@
;; (3) The survivors, which is the row that matters. 2000 entries is eight
;; grows, so the runs are long and interleaved; removing the even keys and
;; then asking after every odd one is asking whether any probe run was cut.
;; A backward shift that stopped one slot early loses entries here and
;; nowhere else.
;; then asking after every odd one is asking whether any probe sequence was
;; cut.
(let [m (map-new i32 i64)]
(dotimes [i 2000] (put m i (* (i64 i) 3)))
(let [taken 0]
@ -90,9 +90,7 @@
(print (length s)) (println "") ; 1
(free s))
;; (5) Iteration after removals answers exactly the survivors. The cursor is
;; started fresh — a cursor held *across* a removal is invalidated by the
;; shift, the same way a put that grows invalidates one.
;; (5) Iteration after removals answers exactly the survivors.
(let [m (map-new i32 i32)]
(dotimes [i 100] (put m i i))
(dotimes [i 100] (if (= 0 (% i 3)) (map-remove m i)))
@ -118,4 +116,49 @@
(match (get t 299) (Some v) (do (print v) (println "")) None (println "?")) ; 598
(print (has-key? t 298)) (println ""))) ; false
(free-all ar))
;; (7) Churn at a steady size, checked against a plain array. Keys are drawn
;; from 512 and the map hovers around half of them, so removals leave deleted
;; slots that inserts must reuse and rebuilds must sweep. Every get is checked
;; against the array, and at the end so is the length and a full walk.
(let [m (map-new i32 i64)
ref (vec-new i64)
x (u64 12345)
bad 0
live 0]
(dotimes [i 512] (push ref (i64 -1)))
(dotimes [step 200000]
(set x (+ (* x (u64 6364136223846793005)) (u64 1442695040888963407)))
(let [k (i32 (% (>> x 33) (u64 512)))
op (i32 (% (>> x 20) (u64 4)))]
(if (< op 2)
(do (if (< (at ref k) 0) (set live (+ live 1)))
(put m k (i64 step))
(set (at ref k) (i64 step)))
(if (= op 2)
(match (map-remove m k)
(Some v) (do (if (not (= v (at ref k))) (set bad (+ bad 1)))
(set (at ref k) (i64 -1))
(set live (- live 1)))
None (if (>= (at ref k) 0) (set bad (+ bad 1))))
(match (get m k)
(Some v) (if (not (= v (at ref k))) (set bad (+ bad 1)))
None (if (>= (at ref k) 0) (set bad (+ bad 1))))))))
(print bad) (println "") ; 0
(print (= (length m) (i64 live))) (println "") ; true
(let [cur (i64 0) k 0 v (i64 0) seen 0]
(while (map-next m (addr cur) (addr k) (addr v))
(set seen (+ seen 1))
(if (not (= v (at ref k))) (set bad (+ bad 1))))
(print (= seen live)) (println "") ; true
(print bad) (println "")) ; 0
;; Removing while walking visits every survivor once: nothing moves.
(let [cur (i64 0) k 0 v (i64 0) seen 0]
(while (map-next m (addr cur) (addr k) (addr v))
(set seen (+ seen 1))
(map-remove m k))
(print (= seen live)) (println "") ; true
(print (length m)) (println "")) ; 0
(free ref)
(free m))
0)

View File

@ -3884,7 +3884,8 @@ level "1"
into the wrong half of it. *)
let map_remove_out =
"100\n1\nfalse\ngone\n1\n0\n77\n1\n1000\n1000\n0\n0\n2000\n\
709\nfalse\ntrue\n399\n1\ntrue\n1\n66\n0\n66\n150\n598\nfalse\n"
709\nfalse\ntrue\n399\n1\ntrue\n1\n66\n0\n66\n150\n598\nfalse\n\
0\ntrue\ntrue\n0\ntrue\n0\n"
in
outputs "map removal" "programs/map-remove.flan" map_remove_out;
outputs ~opt:"-O0" "map removal, -O0" "programs/map-remove.flan"

View File

@ -269,10 +269,11 @@ let corpus =
down: [check_at] is behind the flag and so is *half* of [check_slice] — the
[hi <= len] compare — while its [lo <= hi] stays, because that one is the
claim that a slice's length word is a count and not a bounds check at all.
A Vec's and a Map's bounds checks are not behind it either: they live inside
flan_vec_at and the map probe in flan_rt.c, are plain C, and run in every
build. So the flag lowers the guard on indexing a fixed array or a slice,
and on nothing else. These are the programs where that distinction reaches
A Vec's bounds check is not behind it either: it lives inside flan_vec_at,
is plain C, and runs in every build. A Map has no bounds check to lower —
its probe is masked to the capacity and an absent key is None. So the flag
lowers the guard on indexing a fixed array or a slice, and on nothing
else. These are the programs where that distinction reaches
heap storage. *)
let unchecked_subset =
[ "programs/allocators.flan", [];