20 Commits

Author SHA1 Message Date
9d5689ffa2 Every citation of a moved document now resolves from where it is written 2026-09-14 07:12:27 +07:00
c9c23c5079 Call frame information, which is the thing the assembler gets right
The debug-information section added in the last commit is about what GAS
gets wrong against a file with no instructions in it: .loc is flushed when
an instruction is assembled, and this backend assembles none. CFI is the
opposite case and worth recording beside it. Its advances come from frag
positions, so .cfi_def_cfa_offset interleaved with .byte comes out exact --
measured, readelf --debug-dump=frames on a .byte-only function gives the
right advances.

The content is a constant, and that is the header's claim about the frame
model paying for itself. rsp is written exactly twice, so: on entry the CFA
is rsp+8; push rbp makes it rsp+16 with the saved rbp at cfa-16; mov rsp,rbp
moves the rule onto rbp and it stays there for the whole body; after leave
rsp is rbp+8 and the CFA is rsp+8 again. Five directives, three sites --
emit_fn, emit_main and emit_globals_init, which have the same prologue.

emit_main has no closing rule because it has no epilogue: it leaves through
flan_exit and the ud2 after that is unreachable, so the rbp rule holds to
the last byte, which is what a backtrace out of anything main called wants.

gdb did not need this -- its prologue analyser already unwound out of
flan_bounds_error into flan.main with a line number, because push rbp;
mov rbp,rsp; sub rsp,N is the pattern it recognises. It is here because the
description is now stated rather than guessed, and because a break at the
very first byte of a function -- before the push -- now unwinds from a rule
rather than from a heuristic.

Gated on --debug so a release build's assembly stays byte-for-byte what it
was. That is conservative rather than principled: the description is correct
in every build and a release build is where a crash would most want it. What
stops it being unconditional is only that nothing measures the .eh_frame it
would add, and another lane is measuring backend cost right now.
2026-09-13 23:32:35 +07:00
201cd87bc9 Debug information for --x86, written out as bytes
--x86 --debug was refused in build.ml for want of DWARF. It emits it now:
a compile unit, a subprogram per function, and a line table, so gdb breaks
by Flan file and line and a backtrace names Flan source.

The reason it is written as data rather than as .loc directives is the
finding worth carrying forward, and HANDOFF-x86-debug.md leads with it.
GAS builds its line table out of dwarf2_emit_insn, which runs only when an
instruction is assembled, and this backend assembles none -- everything is
a .byte blob. A pending .loc therefore sits until the next .loc and is
flushed at whatever the location counter has reached by then: every row
comes out one statement late and the last statement of every function gets
no row at all. Measured on GAS 2.44, and interposing labels does not help.

So .debug_line is emitted here the way the instructions are. Every row is
a full DW_LNE_set_address on a label rather than an advance_pc with a
computed delta, because a delta would be a difference of two labels inside
a .uleb128 -- a value whose size changes the offsets after it, and whose
failure would look exactly like a DWARF bug.

The .file 1 directive at the top of the assembly is not about our table at
all: without it clang's integrated assembler generates a compile unit of
its own over the .s, whose rows land at the call mnemonics, and those
addresses are inside the functions we already describe. -gdwarf-4 goes
with it so the empty stub it leaves behind parses.

Rows are deduplicated on the byte counter as well as on the position.
lower recurses, so an outer form and the inner one that emits its first
byte both ask for a row at the same address; keeping the first is smaller
and is the better answer, since a debugger takes the last of a run. buf.n
already existed and was written but never read. This is its first reader.

No locals and no types, deliberately. A slot here is a bump-allocated
frame temporary whose offset is known but whose lifetime is not modelled
-- scoped reclaims temporaries and a later expression reuses the bytes --
so a DW_TAG_variable would be right at some addresses and confidently
wrong at others. That is the call build.ml already makes about wasm32's
member offsets.

Verified under gdb against test/programs/debug.flan, and against an LLVM
--debug build of the same program. Identical bar the parameter values,
which is the locals work. Transcript in the handoff.
2026-09-13 23:32:35 +07:00
853bc35be7 A release build stops calling into a registry that is switched off
emit.ml:1918 drops the allocation registry's notes in a release build --
the checker builds a Tast.Rt it cannot know is unwanted, because it does
not know whether this is a dev build. The x86 backend had no counterpart
and emitted the calls for real: correct, since the registry answers
nothing when it is disabled, but one call per container operation into a
function that returns immediately.

The guard is emit.ml's byte for byte, strict > 17 included: bare
flan_dev_reg_note is the runtime's own C entry point and is never a
Tast.Rt, and what check.ml builds is the _vec, _map and _pool wrappers.
emit.ml drops the note before the arguments are walked so that taking the
address of the container does not leave an escaped alloca behind; here the
arguments are not touched until call_rt, so answering () is already early
enough.

Measured, call sites of flan_dev_reg_note in the disassembly:

  vec.flan       46 -> 1     (--dev: 46)
  maps.flan      55 -> 1     (--dev: 55)
  registry.flan  36 -> 1     (--dev: 36)

The one left in each is not emitted code -- it is inside the runtime's own
flan_dev_reg_note_vec. An LLVM release build of vec.flan has the same one,
so the two backends now agree.

HANDOFF-x86-debug.md is the stub for item 6, which is next, and it leads
with the finding that changes that item's plan: .loc does not work against
a backend that emits .byte blobs, so the line table has to be written out
by hand.
2026-09-13 23:32:35 +07:00
b0f4fb73e1 The x86 backend stops leaving a divide by zero to the hardware
Item 3 of HANDOFF-x86-rt.md, which was blocked on the language decision rather
than on code. check_div and check_cast sit beside check_at and check_slice and
reuse bounds_call unchanged; flan_arith_error takes three extras, so the
channel lands in r9 and the argument registers are exactly full.

Three things differ from the LLVM side because the instruction set does: two
branches rather than one branch and a select, since there is no select and a
second compare on the cold path is free; the cast bounds compared in the
source's own precision rather than widened to a double, which is exact because
every bound is a power of two; and NaN excluded by choosing the direction of
each compare, because ucomis sets CF, ZF and PF together when either operand is
unordered.

Both programs now print byte-identical stdout, stderr and exit status through
either backend.
2026-09-13 23:07:23 +07:00
60d4c762e1 Merge branch 'worktree-agent-ab6daf83988bee428' into dev-loop 2026-09-13 22:50:40 +07:00
624a94af7b A smaller hunk, and the restart-case shape that is not a counterexample 2026-09-13 22:46:16 +07:00
939417446d Item 5's guard dropped, item 4's reached and correct 2026-09-13 22:42:42 +07:00
bd69f684ed An --x86 host reloads an --x86 module, checked by running it 2026-09-13 22:40:52 +07:00
69ea66deca An x86 redefinition emitter, and GOT addressing for the host's symbols 2026-09-13 22:38:48 +07:00
81f7e464ec The indirection cell on the x86 backend, and --x86 --dev with it
FnAddr (Fnval n) emitted the symbol, which is right for a whole-program build
and wrong the instant anything is redefined into it. It now reads the cell,
and so does every direct call, which is what emit.ml's body_of does and is the
half that matters: a redefinition is one store, and it has to reach call sites
that already exist.

What is emitted, all of it behind dev:

  - one cell per function in .data, .globl, initialised to the body this build
    compiled. Spelled exactly as Emit.cellname spells it, because the point of
    having one here is that an LLVM-built module binds
    @"flan.cell.<n>" = external global ptr against it. nm -D over the two
    builds of the same program gives identical sets of 68 cell symbols.
  - the cell load placed after the arguments, which emit.ml has as a
    load-bearing comment: a redefinition landing between two calls must not
    land in the middle of one. CallPtr stays the other way round.
  - the flan_dev_reg_enable constructor, which arms the allocation registry.

Not emitted: Emit.cellptr, the deeper spelling for a name the host was never
built with. It cannot arise in a whole-program build and belongs with the
redefinition module that would introduce one.

The --x86 --dev refusal is relaxed, and the argument is that flan dev never
reaches this fork: --x86 is read only by flan build, and the daemon builds
host and modules through Build.executable / Build.shared without it. So the
flag means a host whose call sites are redefinable, and nothing claims the
module that would redefine through them exists.

Two things were needed to believe any of that. First, the corpus with --dev on
both sides: 97 MATCH, 0 DIFFER, same as without it. Before the constructor was
added that read 96/1 — registry.flan asks (live? ...) and got four zeroes,
which is the whole of what a dev host does differently besides the cells.

Second, and the corpus cannot do this one: a dev build starts with every cell
pointing at the body this build compiled, so it prints what a release build
prints whether anything reads the cell or not. spike/x86/cells.sh preloads a
shared object whose constructor dlsyms flan.cell.twice and stores a different
body there -- the one store a redefinition ends in, done from outside, no
compiler involved. Both dev builds then print the new answer for a direct call
and for a function value, and both release builds are unchanged, which is what
says the change came from the indirection and not from symbol interposition.

One thing the later lane inherits, now written in both headers rather than
left to be discovered. x86.ml licenses its own calling convention on the
grounds that a dev build is compiled entirely here and a release build
entirely by LLVM, so the two never meet in one process. A cell an LLVM-built
module can store into is the first thing that could make that false: the
conventions agree on scalars and disagree on every aggregate, so an
Emit.redefinition module dlopened into an --x86 host would be right until the
first redefined function took or returned a struct. The answer is a
redefinition emitter here, not a classifier.
2026-09-13 21:17:39 +07:00
8ff8a71de8 The aggregate-return refusal cannot fire, and the reason is in check.ml
Item 17 left this as a loose end: flan_vec_as_slice returns a slice by value
and should hit the "aggregate return" refusal, bounds-condition.flan exercises
it in and out of bounds and matches, and nobody traced why.

It is the first of the two possibilities that report named — the refusal is
narrower than it reads, and nothing is going right by accident.
flan_vec_as_slice's Flan-level return type is Unit. check.ml builds it as
[rt loc Types.Unit "flan_vec_as_slice"] and flan_rt.c writes the two words
through a [void *out] parameter, so [is_void rty] answers first and the
[is_agg rty] test below it is never reached.

That is not one symbol's accident, it is the convention. Every aggregate-valued
runtime result crosses through an out-pointer the checker allocates; every
other [rt] builder in check.ml answers Unit, an Int, a Ptr, an Alloc or a
Handle. And the other user of this path, a [declare]d C function, is covered by
[crossable], which admits String and Slice only as a parameter and refuses an
aggregate return outright.

So there is no sret convention to build for Rt, and building one would be worse
than the refusal: the C boundary wants SysV classification — a 16-byte slice
comes back in rax:rdx — and not the hidden-pointer convention this backend uses
internally. There is no classifier in the file and nothing to test one against.
The line stays as a guard against those two rules changing, and now says which
rules and what the work would actually be.

With it goes the rest of item 16's claim that the container runtime is
unexercised. It is: Vec and Map through vec.flan, vec-of-vec.flan, maps.flan
and map-iter.flan, and Pool through registry.flan, handles.flan,
generics.flan and pool-stale-region.flan. All match.
2026-09-13 20:36:56 +07:00
bdecd2b8f2 slice-from-ptr on the x86 backend, and the check that has to be signed
Another lane landed (slice-from-ptr p n) while this backend was not looking,
and it arrived as two refusals rather than one: slice-from-ptr.flan and
bounds.flan both stopped building through --x86. Neither is a new obstacle —
a Slice _ is {ptr, i64} here exactly as it is in emit.ml, so the form is one
store of the pointer and one of the length and no new representation at all.

The half worth writing down is the check. There is nothing to compare the
length against — only the caller knows how many elements live behind that
pointer — so what is checked is that the promise is not absurd, and that test
is *signed*. check_slice's own compares are unsigned, and a negative i32
sign-extended to 64 bits is a huge unsigned value that an unsigned "hi <= len"
waves through; the result would be a slice about 2^64 long that reads as a
pass and faults somewhere else entirely.

Nothing in the corpus walks that path: every length in slice-from-ptr.flan is
a literal, and a negative literal is refused by check.ml before any code is
emitted. So spike/x86/p7-slice-from-ptr.flan takes the length as a parameter
and runs it through a restart-case, which puts the condition's low/high/length
on stdout and compares them against the LLVM build.
2026-09-13 20:27:44 +07:00
e0e5c1e645 The whole corpus goes through the hand-written backend
The transfer exit returned whatever the return temporary held where
emit.ml returns zero. Meaningless to a caller -- its guard sees the
channel set and never looks -- but main is a caller with no guard, and
what it finds in rax is the process exit status.

The survey compares stderr as well now, which is where every message
the new machinery produces goes: the bounds and slice errors, the
three restart refusals, the transfer failure. Each carries a location
this backend emits by hand as a .rodata label and a length in a
register, and an exit status of 134 with the wrong text beside it is
exactly the failure that reads as a match. It also walks spike/x86's
own probes.

p6-transfer.flan is the two re-propagation branches the corpus does
not reach. Every transfer in restarts.flan stops at a restart-case
inside the handler-bind's extent, so the handler frames never come off
on the transfer path; and in nested and shadowed the inner frame
offers the name, so a restart-case the transfer is not aimed at never
has to put the target back. allocators.flan already covers the third.

  89 MATCH  0 DIFFER  0 refused, over test/programs and spike/x86,
  comparing stdout, stderr and the exit status.

DISCUSS.md item 17 is the report.
2026-09-13 18:41:13 +07:00
63d9b87b7a Conditions on the x86 backend, and bounds checks with them
The transfer channel was the only thing between 41 programs and the
corpus. It is there now: a guard after every Flan call, a landing pad
per restart-case, handler-bind and with-allocator, a transfer exit per
function that runs its fdefers, and check_at and check_slice, which
could not exist until the guard did.

Measured by what the programs print and what they exit with, never by
reading bytes. spike/x86/survey.sh builds every program in
test/programs both ways and diffs stdout and the exit status; it did
not exist, so it is here too, and it is the progress meter.

  before  41 MATCH   1 DIFFER  41 refused by name
  after   83 MATCH   0 DIFFER   0 refused by name

The one DIFFER was bounds.flan, and it was the honest answer to
"--x86 is silently a --no-bounds-checks build". It is not one any
more: check_at and check_slice signal through the channel exactly as
emit.ml's do, so a bounds violation signals, a restart-case catches
it, and an unhandled one exits 134 on both backends. The transitional
refusal that would have said so retired before it was written.

check_no_transfer is not removed, it is narrowed to the one place the
argument still holds: a global's initialiser runs from
flan..init-globals, before main and before anything can handle
anything, so a transfer out of it has nowhere to go.

Four bugs, and three of them are the shape item 16 predicted -- code
that reads correctly and answers wrong, found by output and not by
objdump:

- The body fell through into the transfer exit, so every fdefer ran
  twice on a normal return. emit.ml cannot have this bug: its ret
  terminates the block.
- A Vec crossed to the runtime as the address of a *copy*, so pushes
  grew the copy and an in-bounds (at v 1) signalled against a length
  of zero.
- ucomis sets CF, ZF and PF together for a NaN, so sete answered true
  for (= x x) and the prelude's NaN test never fired: (/ 0.0 0.0)
  formatted as -9223372036854775808. Flan's comparisons are LLVM's
  ordered ones, so < and <= swap and =, != take a setnp beside them.
- A union read field 0 through the struct table and was refused by
  name rather than laid out as a tag and a payload.

And one that could not have been found later: emit_globals_init stored
a null *into* the channel slot rather than a cell address into it,
which is a null pointer for every callee to write through. Harmless
while nothing could transfer; a fault the first time a guard loaded
through it.
2026-09-13 18:05:08 +07:00
58b1f49cf2 A discarded value was being stored over the return address
edn.flan crashed by jumping into .rodata, several statements after the
mistake, and the assembly at the jump read correctly. Item 15 said this
is how hand-encoding fails, and it is: the crash and the cause were in
different functions.

The cause is one line of design. A form whose value is thrown away was
handed the sink, and the sink was spelled as an address — rbp+0. That is
the saved rbp, and rbp+8 is the return address, so a non-void form in
statement position stored its value straight over both. A 16-byte slice
did it in one rep movsb.

The sink is now compared by identity and never used as an address:
anything with a value that is handed it gets a frame temporary instead,
reclaimed immediately. The point is not the temporary, it is that the
store has somewhere legal to go.

edn.flan matches the LLVM build now — 60 lines of a hand-written EDN
reader, unions, options, nested collections and all.
2026-09-13 15:08:38 +07:00
fcdaa105af Unions and a two-index (at), which the corpus asked for by name
The sweep over test/programs named its own next two nodes. (at grid r c)
is one node with two indices and not two nodes — an array of arrays is
contiguous, so the second index walks into the element the first landed
on — and machine.flan is the program that says so.

Then unions: MakeCase, CaseField and Match. The payload offset comes from
lay_fields over the same two fields Emit.lay measures a union as, and a
case's field offsets from lay_fields over that case's own fields, so
there is still one layout calculator and this file is still a caller of
it. match reads the tag and compares, an Option reads an i8 at offset 0
and a declared union an i32, and everything past the tag and the binds is
shared — the arrangement emit.ml settled on, for the same reason.

An exhausted match falls through to ud2 rather than to whatever follows.
The checker proved it cannot happen; a defined SIGILL at the instruction
that fell through costs two bytes and is the cheap half of item 15's
question 4.

machine.flan, bytes2.flan, array-ctor.flan and destructure.flan all agree
with the LLVM build now.
2026-09-13 15:06:31 +07:00
786656dfee (at a i) on the left of a set has to reach the array
The corpus sweep found it, and it found it the way item 15 said this work
fails: array-ctor.flan crashed, and the assembly around the crash read
correctly. (set (.x (at pts 0)) 1.5) went through lvalue, lvalue had no
case for At, and the fallback evaluates — so the store landed in a copy of
the element and the array kept its zeros.

emit.ml has this as addr's own At case. One line here, and the program
matches the LLVM build.

Two more programs beside the fizz: one for the internal calling
convention the fizz does not touch at all — a struct argument, a struct
return through the hidden pointer, f32 in the SSE half, eight integer
arguments so two go on the stack, and a slice by pointer — and one for
the rest of the core: a global with an initialiser, recursion, break,
continue, the bitwise family, unsigned shifts and the conversions both
ways. Both agree with LLVM.

al is now zero at every call this backend makes, including the three in
main that were reaching flan_rt_init, flan_argv and flan_exit without it.
Inert on a fixed callee; the point is that there is no exception to the
rule to remember.
2026-09-13 15:03:41 +07:00
2155c41465 A whole program goes through the hand-written backend and runs
x86.ml was an encoder and a frame model with nothing calling it. It now
lowers a whole Tast.program to an assembly file, and `flan build --x86`
hands that file to the same clang invocation the LLVM path uses, against
the same runtime objects. The flag is off by default; LLVM stays the
release backend and the default one.

Three programs, built both ways and compared by what they print and what
they exit with rather than by reading bytes: exit 0; a dotimes that
prints; and a fizz over a call, an if, a remainder and two string
literals. All three agree with the LLVM build.

The measurement decided the target. hist.ml over the fizz program shows
no Signal, no Handled, no RestartCase — a loop that prints does not drag
conditions in. What does is the bounds check and the allocator, and
neither is in the reachable set of a program that prints a number.

That is why there is no transfer guard here, and check_no_transfer is
what makes the omission sound rather than hopeful: if nothing reachable
can write the channel, no call can return with it set. It is a
whole-program property, so it is checked once per build and the build
stops with the node's name when it fails.
2026-09-13 14:56:14 +07:00
66d315813f The encoder and the frame model for a hand-written x86-64 backend
INCOMPLETE AND NOT WIRED IN. lib/x86.ml is not in lib/dune, so nothing
compiles it and nothing calls it; `dune test --root . -j 1` was green at
the tip this branched from and is unaffected, because no file the build
reads was changed. The module itself has never been type-checked.

What is here: the instruction encoder (integer and SSE, loads and stores
at every width, division, shifts, setcc, rip-relative addressing, rep
movsb), the layout bridge to Emit.lay, the frame allocator, and the
.rodata constant emitters. What is not here: the expression lowering,
the call sequence, the function prologue and epilogue, the assembly file
assembly, the build.ml flag and the differential harness. The header
comment is the design; the second half of the file is missing.

THE INTERNAL CONVENTION, which is the decision hardest to recover from
the code, and which is chosen rather than inherited:

  - Scalars -- integers, bool, ptr, enum, handle, allocator, Fn -- in
    SysV's integer registers rdi rsi rdx rcx r8 r9, then right to left
    on the stack. bool is one byte, zero-extended on load.
  - Floats in xmm0-xmm7, then on the stack.
  - EVERY aggregate by pointer. An argument is a pointer to a copy the
    caller made; a return is a hidden sret pointer in the FIRST integer
    register with every other argument shifted along, and that same
    pointer comes back in rax. Nothing is classified, nothing is split
    across register classes, there is no eightbyte rule.
  - The transfer channel is the last argument of all, a pointer, in the
    integer sequence -- emit.ml's `signature` rule, unchanged. It is a
    pointer to a pointer: main allocates one cell, stores null, and
    threads its address down; a callee that transfers stores non-null
    into it and every caller loads, tests and branches to its pad.
  - Frame: every intermediate value is a frame temporary, bump-allocated
    below rbp with a high-water mark, and the outgoing-argument area is
    reserved once in the prologue. rsp is written exactly twice, by the
    prologue's sub and by leave. So rsp % 16 == 0 at every call site is
    a property of one rounded sub, and the spike's depth counter is not
    needed -- its bug class is removed rather than guarded against.

WHY THE CONVENTION IS OURS TO PICK, confirmed rather than assumed: a dev
build compiled by this backend never emits a .ll at all, and a release
build never runs this backend, so no process holds code from both. The
only boundary that must match SysV exactly is C, and check.ml rejects an
aggregate in a `declare` while the generated shim flattens every struct,
so no Flan-emitted call ever hands C an aggregate. I found no path that
mixes the two backends in one process. I did NOT get far enough to test
that claim by running anything, so it stands on reading build.ml's
`executable` and emit.ml's `signature`, not on an experiment.

WHAT THE MEASUREMENT SAYS, and it is the one new fact this branch has.
spike/backend/hist.ml histograms Tast nodes over a program after Reach
prunes it. Item 15's four buckets undercount what a whole-program build
must do on day one:

  - enum-compare.flan needs Str, Make, Field and Call before it prints
    anything, because the prelude builds a slice to print one. Aggregates
    are not a later row; they are in the first program.
  - loops.flan carries Handled, RestartCase and Signal one each. The
    "no plan" row is in the reachable set of a program that only loops,
    so conditions cannot be deferred behind a whole-program flag.
  - The text primitives (Bytes, I64ToBytes, WriteStdout) are C calls,
    not instruction work, so they are cheap.

WHAT THE NEXT PERSON SHOULD DO FIRST, in order:

  1. Finish the lowering as destination-driven: `eval f e ~dst` writes
     e's value into [rbp+dst] and nothing is ever live in a register
     across a statement. That is what makes aggregates and scalars one
     code path and what keeps the frame model's promise.
  2. Emit an assembly file -- .byte blobs with `call sym` and
     `.long lbl - . - 4` for the few relocated fields -- and add the
     flag to build.ml as FLAN_X86 plus an `opts` field, off by default.
     Do not write an ELF writer; it produces no Flan progress and a bug
     in it looks exactly like an encoding bug.
  3. Copy test/test_sanitize.ml's shape for the differential harness.
     There is no differential run yet, so nothing about correctness has
     been demonstrated on this branch.
  4. Bounds checks are implementable and should not be skipped:
     flan_bounds_error(ptr, i64, i64, i64, ptr) and flan_slice_error
     take the transfer channel, so they are an ordinary guarded call.

THE TWO LANGUAGE PREREQUISITES, unchanged and still not decided here.
Uninit is the one that bites: this backend gives whatever the stack slot
held, LLVM may reason from poison, and that is the one construct where
the two backends are supposed to differ. Division by zero, INT64_MIN/-1
and the float-to-int cast are the other three that x86 answers
differently from LLVM's "undefined" -- idiv raises SIGFPE where LLVM
says nothing, and cvttsd2si answers the integer indefinite value. The
Fn-value question -- body pointer or cell pointer -- is untouched: the
lowering here would have emitted direct calls, which means no
redefinition, and that is a gap to close before this backend is the dev
backend rather than an experiment.
2026-09-13 09:42:50 +07:00