103 lines
7.3 KiB
Markdown
103 lines
7.3 KiB
Markdown
# Handoff — what the x86 backend costs, and keeping it from rotting
|
||
|
||
Branch: `dev-loop`. This lane is item 7 of `HANDOFF-x86-rt.md`'s "What remains" — code size and speed, measured
|
||
rather than guessed — plus the closing note under it, which asks for the survey to run on its own so a refusal
|
||
cannot sit unnoticed for a month again.
|
||
|
||
Nothing in `lib/` is touched. Another lane is rewriting `lib/x86.ml` at the same time, so everything here lives in
|
||
`spike/x86` and in `test/dune`, and anything this lane finds that would want a compiler change is written down at
|
||
the bottom rather than made.
|
||
|
||
## 1. The `@x86` alias — done, and it fails when it should
|
||
|
||
`test/dune` has a fourth alias beside `@sanitize` and `@valgrind`, shaped the same way and opt-in for the same
|
||
reason: it builds every corpus program twice and runs both, which is minutes where `dune test` is forty seconds.
|
||
|
||
dune build --root . @x86
|
||
|
||
It is a `(rule ...)` with no `(executable ...)` beside it, which is where it differs from its two neighbours. The
|
||
check already exists — `spike/x86/survey.sh` is what every handoff quotes its counts from — and a second
|
||
implementation of it in OCaml would be a second thing to drift, which is exactly the failure this alias is meant
|
||
to prevent. So the rule runs the script, and the script grew three environment variables to make that possible:
|
||
|
||
- `SURVEY_STRICT=1` turns the report into an exit status. A DIFFER is a wrong answer and a refusal by name is a
|
||
node the backend has stopped lowering; either fails the build. NOX86 and SKIP do not — the first is usually a
|
||
toolchain that is not installed on the machine, the second is the frontend refusing the program on both sides.
|
||
- `FLAN=<path>` hands it a compiler somebody else built. Without it the script runs `dune build` on `bin/main.exe`,
|
||
and a dune inside a dune action waits on a lock it will never get. The rule has `bin/main.exe` in its deps
|
||
instead, exactly as the `(tests ...)` stanza already does.
|
||
- `SURVEY_CORPUS` for where the programs are read from, which under dune is the build tree.
|
||
|
||
Standalone, with nothing set, the script is byte-for-byte the measurement it always was.
|
||
|
||
**Verified both directions.** `dune build --root . @x86` passes on a clean tree and prints the same
|
||
**97 MATCH / 0 DIFFER / 0 REFUSED / 0 NOX86 / 36 SKIP** the script prints standalone — the build tree and the
|
||
source tree agree about the corpus, which was not obvious in advance. Broken deliberately (a fake DIFFER injected
|
||
into the script, the rule narrowed to one program so the round trip was thirty seconds) it fails the dune build
|
||
with `x86 survey FAILED: 1 differ, 0 refused` and a nonzero status. Both edits reverted.
|
||
|
||
`dune test --root .` is still green and still about forty seconds. The new stanza is a rule on its own alias and
|
||
touches no `(modules ...)` list, so it cannot be pulled into `@runtest` by accident. The two
|
||
`clang: error: linker command failed` lines it prints are the `dev-robust` fixture doing its job and are called
|
||
out in `HANDOFF-x86-rt.md` §4 already.
|
||
|
||
## 2. The measurement — `spike/x86/COST.md`
|
||
|
||
`COST.md` is the deliverable and it carries the numbers and the reading of them; the raw rows are committed beside
|
||
it as `cost-corpus.tsv` and `cost-bench.tsv`. The headline, in one paragraph:
|
||
|
||
**This backend emits 3.84× the code LLVM does at `-O2`, and 1.92× what LLVM emits at `-O0`** — half the factor is
|
||
the optimiser Flan ships with and not the backend. Per program the second ratio is tight: median 2.21, quartiles
|
||
1.94 and 2.98. Speed is unmeasurable over the corpus (every program is 2.5ms of `execve`) except `recur`, which
|
||
loops ten million times and is **6× LLVM `-O0`**; on purpose-written loops the backend is 2.8× to 4.7×.
|
||
|
||
Of the five suspected costs the old handoff named, the measurement says: the frame-slot round trip on every
|
||
intermediate is most of everything and is the item worth fixing; `rep movsb` costs twenty cycles a copy and is
|
||
worth fixing cheaply; the bounds check's three temporaries cost 241 bytes and *one cycle*, so they are a code-size
|
||
item and not a speed item; the guard after every call is four instructions and invisible; and the dev build's
|
||
extra load per call site is not measurable at all. What a dev build does cost — 400× the code, because nothing may
|
||
be dropped when anything might be redefined — is `Reach`'s doing and both backends pay it.
|
||
|
||
The two scripts are how that was produced:
|
||
|
||
- `cost.sh` sweeps the corpus and prints a TSV: file size, `.text`, and **the sum of the `flan.*` defined symbols**,
|
||
which is the program's own code with the runtime excluded. `COST_FLAGS=--dev` is the `SURVEY_FLAGS` precedent.
|
||
- `bench.sh` runs four programs in `spike/x86/bench/`, each written so that one suspected cost is most of what the
|
||
program does, because every program in `test/programs` runs in about two and a half milliseconds of which nearly
|
||
all is `execve`.
|
||
|
||
They live in `spike/x86/bench/`, a subdirectory, *deliberately*: `survey.sh` globs `spike/x86/*.flan`, and a
|
||
benchmark landing in the survey would move the 97 that three handoffs quote.
|
||
|
||
## 3. One trap, paid for once
|
||
|
||
Two copies of `cost.sh` sharing a scratch directory silently corrupt each other's numbers. The first release sweep
|
||
was run, killed, and restarted; `pkill` missed a child, and the survivor kept writing the same `l` and `x` files
|
||
the new run was building and timing. The result was a row for `math` carrying `dev-locals`' binaries — which is
|
||
visible only because those two happened to collide *exactly*, and a collision between two programs of merely
|
||
similar size would have produced a plausible row nobody would question.
|
||
|
||
`cost.sh` now makes a `run.$$` subdirectory with a plain `mkdir`, so a second copy cannot land in the first one's.
|
||
The sweep behind `COST.md` was re-run clean afterwards. If you extend these scripts, keep that property.
|
||
|
||
## 4. Nothing was changed in `lib/`, and nothing needed to be
|
||
|
||
No finding here wants a compiler change that this lane was not allowed to make. The costs are all in
|
||
`lib/x86.ml`'s lowering and the next lane rewriting that file will meet them; `COST.md`'s last section says which
|
||
ones are worth its attention and which are not, which was the point of measuring rather than guessing.
|
||
|
||
## 5. What was not done
|
||
|
||
- **The `--dev` sweep over the whole corpus was not finished.** It was started, ran into the `dev-*` programs
|
||
(each of which waits out a twenty-second timeout twice), and was killed to free the machine for the clean
|
||
release re-run. The `--dev` numbers in `COST.md` come from `bench.sh`, which measures the same thing on four
|
||
programs and measures it better, because a corpus program's dev cost is invisible under `execve`. Re-running it
|
||
is `COST_FLAGS=--dev spike/x86/cost.sh` and about half an hour.
|
||
- **No corpus program was disassembled.** The attribution in `COST.md` is all from the four benchmarks, which are
|
||
small enough to read whole and were written so that each one is mostly a single suspected cost. The corpus
|
||
outliers are explained from the `-O0` column instead — which is what that column is for, and in `bounds`' case
|
||
it turns an alarming 11× into an ordinary 2.07×.
|
||
- **`LLVM -O0` is only reachable through `--debug`**, which also asks for DWARF. Checked and harmless: DWARF lands
|
||
in `.debug_*` and the metric sums `.text` symbols. A plain `-O0` flag on the CLI would be a `bin/main.ml` change
|
||
and this lane was not touching compiler sources.
|