COST.md, with the benchmarks measured and the corpus table still to come

This commit is contained in:
Joseph Ferano 2026-09-13 23:14:57 +07:00
parent 82c8b38dda
commit f9f476dc41

150
spike/x86/COST.md Normal file
View File

@ -0,0 +1,150 @@
# What the hand-written x86-64 backend costs
`survey.sh` has said for three handoffs that this backend agrees with LLVM on all 97 corpus programs it can build.
Item 7 of `HANDOFF-x86-rt.md` is the other half of that sentence — nobody had a number for what the agreement
costs — and it came with a list of suspects: a guard after every call, three frame temporaries per bounds check,
every intermediate in memory, `rep movsb` block copies, and an extra load per call site in a dev build. This is
the measurement. It does not change anything; two of the five suspects turn out not to matter, and the one that
matters most is not on the list.
Produced by `spike/x86/cost.sh` (the corpus, for size) and `spike/x86/bench.sh` (four purpose-written programs,
for speed). Both are documented in their own headers. Machine: 16-core x86-64, Fedora, clang as the assembler and
linker on both sides, and other work running on it throughout — which is why every time below is the *minimum* of
five or seven runs and why nothing here rests on a difference of a few percent.
## What is being compared, and against what
The interesting column is not the size of the executable. A Flan binary is mostly `flan_rt.o` and libc glue, the
same object on both sides, and it drowns the signal: over the corpus the whole file is only **RATIO_FILE×** bigger
through this backend and `.text` only **RATIO_TEXT×**, which would be a reassuring number and a meaningless one.
So the measurement is the sum of the sizes of the defined symbols the compiler *named itself* — everything called
`flan.<something>`. The runtime's C is `flan_<something>` with an underscore, so the two never collide, and a
runtime symbol has the same size on both sides (`flan_map_clone`, `0x4b3` either way), which is the check that
says the difference really is codegen and not a differently-linked runtime.
There is a third column, and it is what makes the second readable: **LLVM with `--debug`, which forces `-O0`**.
This backend has no optimiser at all, so measuring it against LLVM at `-O2` charges it for the whole of mem2reg,
inlining and constant folding. `-O0` is LLVM's instruction selection with none of that, which is the comparison
that says something about *this* backend rather than about the absence of a middle end. (`--x86 --debug` is
refused — the backend emits no DWARF — so the column exists on one side only. DWARF lands in `.debug_*` sections
and not in `.text`, checked, so it does not contaminate the symbol sums.)
## The corpus: size
CORPUS_TABLE
## Where the size goes
Every one of these is from the disassembly of the benchmark programs, which are small enough to read whole.
**Every intermediate goes through the frame, and so does every constant.** This is the big one and it is not one
feature, it is the shape of the whole backend. `(step acc 1)` in `b1-calls` compiles to:
movabs $0x1,%rax ; a 10-byte immediate ...
mov %rax,-0x30(%rbp) ; ... stored to a frame slot ...
mov -0x30(%rbp),%rsi ; ... and loaded back into the argument register
Three instructions and 24 bytes where LLVM writes `mov $1,%esi`, five. The loop bound gets the same treatment
*every iteration* — `movabs $0x1312d00` into a slot, sign-extended out of it, compared — because nothing is
hoisted. So does the loop condition: `cmp`/`setl`/`movzbq`/store a byte to the frame/reload it/`test`/`jne`,
seven instructions for what is `cmp`/`jge` anywhere else. This is most of the 2× against `-O0` and essentially
all of the difference on the programs at the bottom of the table, which have no calls, no bounds checks and no
aggregates in them at all.
**The guard after every call is four instructions and one dependent load.**
call 4009f8 <flan.step>
mov -0x18(%rbp),%r11 ; the condition frame, from its own slot
mov 0x0(%r11),%r11 ; ... dereferenced
test %r11,%r11
jne <unwind>
Roughly 25 bytes per call site. Real, cheap, and third in size behind the two above it — on a program that is
nothing *but* calls (`b1`) the whole backend is 2.8× LLVM `-O0`, and the guard is a minority of that.
**The bounds check is three frame temporaries, as suspected, and it costs code and not time.** The index is
widened, stored, reloaded, stored again, the limit goes to a third slot, and then `cmp`/`jb` — with the failure
path, its `.rodata` location string and its length, inline at the branch target. On `b2-bounds`, `flan.main` is
`0x4b1` with checks and `0x3c0` without: **241 bytes, a quarter of the function.** In time it is 112.6ms against
104.9ms over 20.5 million checked loads — **about 0.4ns, one cycle a check** — because the branch predicts
perfectly and the loads were going to memory anyway. That contradicts the way the handoff's list reads. The three
temporaries are a code-size item. They are not a speed item.
**`rep movsb` is real and it is the most expensive single instruction here.** A 64-byte struct copy lowers to
`lea`/`lea`/`movabs $0x40,%rcx`/`rep movsb`, and `b4-copy` runs 2 million of them in 20.8ms against LLVM `-O0`'s
7.2ms: **about 6.8ns of the difference per copy, some twenty cycles**, which is `rep movsb`'s startup cost and
almost none of it the 64 bytes. It is also the one place where the backend loses to `-O0` by a factor (2.9×) it
does not lose by on straight-line code, and the one item on the suspect list where a targeted fix — inline
16-byte moves under some size threshold — would pay for itself.
## The corpus: speed, and why there is barely any
Almost nothing. **Every program in `test/programs` runs in about 2.5 milliseconds, nearly all of it `execve` and
the dynamic loader**, and both backends produce the same 2.5 milliseconds. Best-of-five does not rescue a signal
that is not there. There is exactly one corpus program whose own code is a measurable part of its runtime, and it
is the right one:
| program | LLVM -O2 | LLVM -O0 | x86 | what it is |
|---|---|---|---|---|
| `recur` | ~1ms | 10ms | 60ms | a ten-million-iteration counting loop, written to prove `recur` is a jump |
Read that carefully, because the 25× against `-O2` is not a fact about this backend: LLVM folds the loop to its
answer and runs nothing. Against `-O0`, which also runs ten million iterations, it is **6×** — about 6ns an
iteration against 1ns, or roughly eighteen cycles for `i+1` and a compare. That is the frame-slot round trip
above, four or five times over, and it is the honest number.
## The benchmarks
Four programs in `spike/x86/bench/`, each written so that one suspected cost is most of what the program does.
They are in a subdirectory on purpose: `survey.sh` globs `spike/x86/*.flan` and a benchmark is not a case.
Times are best-of-seven, in milliseconds.
BENCH_TABLE
`b1-calls` and `b3-spill` are the pair to read together. `b1` is 20 million calls of a one-instruction function
and lands at 2.8× `-O0`; `b3` is no calls at all, six dependent arithmetic temporaries per iteration, and lands
at 4.7×. **The backend is worse at arithmetic than it is at calling**, which is the opposite of what the suspect
list implies, and it is because a call already costs enough that four extra instructions beside it disappear,
while an add that should be one instruction costs five.
The `-O2` column in `b1` and `b4` is 2ms — the loop is gone. That is a true fact about the toolchain Flan ships
and a useless one about code generation, which is the whole reason the `-O0` column exists.
## `--dev`, which is the one axis both backends pay
`SURVEY_FLAGS=--dev` still reports 97 MATCH, so the comparison is available. The suspected cost was the extra
load per call site — every cross-function call going through its indirection cell. **It is not measurable.** On
`b1-calls`, 20 million calls, x86 release is 104.3ms and x86 `--dev` is 100.9ms: the same number, and the dev
build is nominally the *faster* of the two, which is what a difference below the noise floor looks like. The load
is from a `.data` cell that is in L1 after the first call and the machine was already waiting on the frame.
What a dev build actually costs is something else entirely, and both backends pay it. `b1-calls` is a program
with two functions in it:
| build | own code |
|---|---|
| LLVM, release | 82 bytes |
| x86, release | 807 bytes |
| LLVM, `--dev` | 32,714 bytes |
| x86, `--dev` | 83,018 bytes |
**A dev build emits the entire prelude**, because anything might be redefined and so nothing may be dropped. That
is four hundred times the code for this program, and it dwarfs every item on the suspect list put together. It is
also not a backend cost — LLVM pays a 400× of its own — so it is `Reach`'s business and not `x86.ml`'s. The
backend's share of it is the same ~2.5× it charges everywhere else.
## What this says to the lane rewriting `lib/x86.ml`
Ranked by what the numbers actually support, and not by the order of the list in the handoff:
1. **Keep values in registers across a single expression.** Not a register allocator — just not routing every
constant and every subexpression through a frame slot, and not re-materialising a loop bound every iteration.
This is most of the 2× against `-O0`, most of `recur`'s 6×, and all of `b3`'s 4.7×.
2. **Inline small aggregate copies** instead of `rep movsb`. One instruction, twenty cycles, on a copy that is
four `movdqu` pairs.
3. **`flan_dev_reg_note` in a release build** (item 2 of the old handoff's list) is worth doing and is small.
4. **The call guard is fine.** Four instructions and 25 bytes, invisible in time. Leave it.
5. **The bounds check is fine on time and fat on code.** If it is ever worth touching, it is worth touching for
the 241 bytes — hoisting the failure path out of line would get most of that back without changing a cycle.
6. **The dev call cell is free.** Whatever the redefinition emitter costs, it does not cost this.