The idiv in b3 is most of the denominator, and the two b2 numbers now agree

This commit is contained in:
Joseph Ferano 2026-09-13 23:28:27 +07:00
parent 219a42ed5b
commit 35c33c7b8b

View File

@ -111,7 +111,7 @@ nothing *but* calls (`b1`) the whole backend is 2.8× LLVM `-O0`, and the guard
widened, stored, reloaded, stored again, the limit goes to a third slot, and then `cmp`/`jb` — with the failure widened, stored, reloaded, stored again, the limit goes to a third slot, and then `cmp`/`jb` — with the failure
path, its `.rodata` location string and its length, inline at the branch target. On `b2-bounds`, `flan.main` is path, its `.rodata` location string and its length, inline at the branch target. On `b2-bounds`, `flan.main` is
`0x4b1` with checks and `0x3c0` without: **241 bytes, a quarter of the function.** In time it is 112.6ms against `0x4b1` with checks and `0x3c0` without: **241 bytes, a quarter of the function.** In time it is 112.6ms against
104.9ms over 20.5 million checked loads — **about 0.4ns, one cycle a check** — because the branch predicts 105.0ms over 20.5 million checked loads — **about 0.4ns, a cycle or two a check** — because the branch predicts
perfectly and the loads were going to memory anyway. That contradicts the way the handoff's list reads. The three perfectly and the loads were going to memory anyway. That contradicts the way the handoff's list reads. The three
temporaries are a code-size item. They are not a speed item. temporaries are a code-size item. They are not a speed item.
@ -149,17 +149,23 @@ Times are best-of-seven, in milliseconds.
| `b1-calls` | 20M calls of a one-instruction function | 2.1 | 36.8 | 104.3 | 2.8× | 200 → 807 | | `b1-calls` | 20M calls of a one-instruction function | 2.1 | 36.8 | 104.3 | 2.8× | 200 → 807 |
| `b2-bounds` | 20.5M bounds-checked array loads | 4.7 | 27.0 | 112.6 | 4.2× | 407 → 1390 | | `b2-bounds` | 20.5M bounds-checked array loads | 4.7 | 27.0 | 112.6 | 4.2× | 407 → 1390 |
| `b3-spill` | 5M iterations of a six-deep arithmetic tree | 22.2 | 29.0 | 137.3 | 4.7× | 248 → 1117 | | `b3-spill` | 5M iterations of a six-deep arithmetic tree | 22.2 | 29.0 | 137.3 | 4.7× | 248 → 1117 |
`b3`'s ratio is the one to distrust slightly: its expression ends in a `%`, which is an `idiv`, and an `idiv` is
twenty-odd cycles on every side. That is most of LLVM's own 22.2ms and a good part of its 29.0ms, so the
denominator is largely a hardware latency this backend cannot do anything about. The absolute gap — 108ms over
5 million iterations, about 21ns of extra work each — is the honest reading of that row.
| `b4-copy` | 2M copies of a 64-byte struct | 2.0 | 7.2 | 20.8 | 2.9× | 348 → 877 | | `b4-copy` | 2M copies of a 64-byte struct | 2.0 | 7.2 | 20.8 | 2.9× | 348 → 877 |
`b2` with `--no-bounds-checks` on both sides: LLVM 4.5ms, x86 105.0ms — the 7.6ms the check costs over 20.5 `b2` with `--no-bounds-checks` on both sides: LLVM 4.5ms, x86 105.0ms — the 7.6ms the check costs over 20.5
million of them, and the 241 bytes it costs in `flan.main`, are the whole of it. million of them, and the 241 bytes it costs in `flan.main`, are the whole of it.
`b1-calls` and `b3-spill` are the pair to read together. `b1` is 20 million calls of a one-instruction function `b1-calls` and `b3-spill` are the pair to read together, with the `idiv` caveat above in mind. `b1` is 20 million
and lands at 2.8× `-O0`; `b3` is no calls at all, six dependent arithmetic temporaries per iteration, and lands calls of a one-instruction function and lands at 2.8× `-O0`; `b3` has no calls at all and pays 21ns an iteration
at 4.7×. **The backend is worse at arithmetic than it is at calling**, which is the opposite of what the suspect for six dependent arithmetic temporaries. **The backend is worse at arithmetic than it is at calling**, which is
list implies, and it is because a call already costs enough that four extra instructions beside it disappear, the opposite of what the suspect list implies, and it is because a call already costs enough that four extra
while an add that should be one instruction costs five. instructions beside it disappear, while an add that should be one instruction costs five. `recur`, which is a
counting loop and nothing else, says the same thing on a corpus program: 6× LLVM `-O0`.
The `-O2` column in `b1` and `b4` is 2ms — the loop is gone. That is a true fact about the toolchain Flan ships The `-O2` column in `b1` and `b4` is 2ms — the loop is gone. That is a true fact about the toolchain Flan ships
and a useless one about code generation, which is the whole reason the `-O0` column exists. and a useless one about code generation, which is the whole reason the `-O0` column exists.
@ -193,7 +199,7 @@ Ranked by what the numbers actually support, and not by the order of the list in
1. **Keep values in registers across a single expression.** Not a register allocator — just not routing every 1. **Keep values in registers across a single expression.** Not a register allocator — just not routing every
constant and every subexpression through a frame slot, and not re-materialising a loop bound every iteration. constant and every subexpression through a frame slot, and not re-materialising a loop bound every iteration.
This is most of the 2× against `-O0`, most of `recur`'s 6×, and all of `b3`'s 4.7×. This is most of the 2× against `-O0` and most of `recur`'s 6×, and it is what `b3`'s 21ns an iteration buys.
2. **Inline small aggregate copies** instead of `rep movsb`. One instruction, twenty cycles, on a copy that is 2. **Inline small aggregate copies** instead of `rep movsb`. One instruction, twenty cycles, on a copy that is
four `movdqu` pairs. four `movdqu` pairs.
3. **`flan_dev_reg_note` in a release build** (item 2 of the old handoff's list) is worth doing and is small. 3. **`flan_dev_reg_note` in a release build** (item 2 of the old handoff's list) is worth doing and is small.