The sweep over test/programs named its own next two nodes. (at grid r c)
is one node with two indices and not two nodes — an array of arrays is
contiguous, so the second index walks into the element the first landed
on — and machine.flan is the program that says so.
Then unions: MakeCase, CaseField and Match. The payload offset comes from
lay_fields over the same two fields Emit.lay measures a union as, and a
case's field offsets from lay_fields over that case's own fields, so
there is still one layout calculator and this file is still a caller of
it. match reads the tag and compares, an Option reads an i8 at offset 0
and a declared union an i32, and everything past the tag and the binds is
shared — the arrangement emit.ml settled on, for the same reason.
An exhausted match falls through to ud2 rather than to whatever follows.
The checker proved it cannot happen; a defined SIGILL at the instruction
that fell through costs two bytes and is the cheap half of item 15's
question 4.
machine.flan, bytes2.flan, array-ctor.flan and destructure.flan all agree
with the LLVM build now.
The corpus sweep found it, and it found it the way item 15 said this work
fails: array-ctor.flan crashed, and the assembly around the crash read
correctly. (set (.x (at pts 0)) 1.5) went through lvalue, lvalue had no
case for At, and the fallback evaluates — so the store landed in a copy of
the element and the array kept its zeros.
emit.ml has this as addr's own At case. One line here, and the program
matches the LLVM build.
Two more programs beside the fizz: one for the internal calling
convention the fizz does not touch at all — a struct argument, a struct
return through the hidden pointer, f32 in the SSE half, eight integer
arguments so two go on the stack, and a slice by pointer — and one for
the rest of the core: a global with an initialiser, recursion, break,
continue, the bitwise family, unsigned shifts and the conversions both
ways. Both agree with LLVM.
al is now zero at every call this backend makes, including the three in
main that were reaching flan_rt_init, flan_argv and flan_exit without it.
Inert on a fixed callee; the point is that there is no exception to the
rule to remember.
x86.ml was an encoder and a frame model with nothing calling it. It now
lowers a whole Tast.program to an assembly file, and `flan build --x86`
hands that file to the same clang invocation the LLVM path uses, against
the same runtime objects. The flag is off by default; LLVM stays the
release backend and the default one.
Three programs, built both ways and compared by what they print and what
they exit with rather than by reading bytes: exit 0; a dotimes that
prints; and a fizz over a call, an if, a remainder and two string
literals. All three agree with the LLVM build.
The measurement decided the target. hist.ml over the fizz program shows
no Signal, no Handled, no RestartCase — a loop that prints does not drag
conditions in. What does is the bounds check and the allocator, and
neither is in the reachable set of a program that prints a number.
That is why there is no transfer guard here, and check_no_transfer is
what makes the omission sound rather than hopeful: if nothing reachable
can write the channel, no call can return with it set. It is a
whole-program property, so it is checked once per build and the build
stops with the node's name when it fails.
INCOMPLETE AND NOT WIRED IN. lib/x86.ml is not in lib/dune, so nothing
compiles it and nothing calls it; `dune test --root . -j 1` was green at
the tip this branched from and is unaffected, because no file the build
reads was changed. The module itself has never been type-checked.
What is here: the instruction encoder (integer and SSE, loads and stores
at every width, division, shifts, setcc, rip-relative addressing, rep
movsb), the layout bridge to Emit.lay, the frame allocator, and the
.rodata constant emitters. What is not here: the expression lowering,
the call sequence, the function prologue and epilogue, the assembly file
assembly, the build.ml flag and the differential harness. The header
comment is the design; the second half of the file is missing.
THE INTERNAL CONVENTION, which is the decision hardest to recover from
the code, and which is chosen rather than inherited:
- Scalars -- integers, bool, ptr, enum, handle, allocator, Fn -- in
SysV's integer registers rdi rsi rdx rcx r8 r9, then right to left
on the stack. bool is one byte, zero-extended on load.
- Floats in xmm0-xmm7, then on the stack.
- EVERY aggregate by pointer. An argument is a pointer to a copy the
caller made; a return is a hidden sret pointer in the FIRST integer
register with every other argument shifted along, and that same
pointer comes back in rax. Nothing is classified, nothing is split
across register classes, there is no eightbyte rule.
- The transfer channel is the last argument of all, a pointer, in the
integer sequence -- emit.ml's `signature` rule, unchanged. It is a
pointer to a pointer: main allocates one cell, stores null, and
threads its address down; a callee that transfers stores non-null
into it and every caller loads, tests and branches to its pad.
- Frame: every intermediate value is a frame temporary, bump-allocated
below rbp with a high-water mark, and the outgoing-argument area is
reserved once in the prologue. rsp is written exactly twice, by the
prologue's sub and by leave. So rsp % 16 == 0 at every call site is
a property of one rounded sub, and the spike's depth counter is not
needed -- its bug class is removed rather than guarded against.
WHY THE CONVENTION IS OURS TO PICK, confirmed rather than assumed: a dev
build compiled by this backend never emits a .ll at all, and a release
build never runs this backend, so no process holds code from both. The
only boundary that must match SysV exactly is C, and check.ml rejects an
aggregate in a `declare` while the generated shim flattens every struct,
so no Flan-emitted call ever hands C an aggregate. I found no path that
mixes the two backends in one process. I did NOT get far enough to test
that claim by running anything, so it stands on reading build.ml's
`executable` and emit.ml's `signature`, not on an experiment.
WHAT THE MEASUREMENT SAYS, and it is the one new fact this branch has.
spike/backend/hist.ml histograms Tast nodes over a program after Reach
prunes it. Item 15's four buckets undercount what a whole-program build
must do on day one:
- enum-compare.flan needs Str, Make, Field and Call before it prints
anything, because the prelude builds a slice to print one. Aggregates
are not a later row; they are in the first program.
- loops.flan carries Handled, RestartCase and Signal one each. The
"no plan" row is in the reachable set of a program that only loops,
so conditions cannot be deferred behind a whole-program flag.
- The text primitives (Bytes, I64ToBytes, WriteStdout) are C calls,
not instruction work, so they are cheap.
WHAT THE NEXT PERSON SHOULD DO FIRST, in order:
1. Finish the lowering as destination-driven: `eval f e ~dst` writes
e's value into [rbp+dst] and nothing is ever live in a register
across a statement. That is what makes aggregates and scalars one
code path and what keeps the frame model's promise.
2. Emit an assembly file -- .byte blobs with `call sym` and
`.long lbl - . - 4` for the few relocated fields -- and add the
flag to build.ml as FLAN_X86 plus an `opts` field, off by default.
Do not write an ELF writer; it produces no Flan progress and a bug
in it looks exactly like an encoding bug.
3. Copy test/test_sanitize.ml's shape for the differential harness.
There is no differential run yet, so nothing about correctness has
been demonstrated on this branch.
4. Bounds checks are implementable and should not be skipped:
flan_bounds_error(ptr, i64, i64, i64, ptr) and flan_slice_error
take the transfer channel, so they are an ordinary guarded call.
THE TWO LANGUAGE PREREQUISITES, unchanged and still not decided here.
Uninit is the one that bites: this backend gives whatever the stack slot
held, LLVM may reason from poison, and that is the one construct where
the two backends are supposed to differ. Division by zero, INT64_MIN/-1
and the float-to-int cast are the other three that x86 answers
differently from LLVM's "undefined" -- idiv raises SIGFPE where LLVM
says nothing, and cvttsd2si answers the integer indefinite value. The
Fn-value question -- body pointer or cell pointer -- is untouched: the
lowering here would have emitted direct calls, which means no
redefinition, and that is a gap to close before this backend is the dev
backend rather than an experiment.