flan/lib/types.ml
Joseph Ferano a4c6b996ff def re-runs its initialiser, and defvar is renamed defonce
The trio the author decided on 2026-09-20 is now all built: def is CL's
defparameter — its initialiser runs on every daemon re-run, unguarded, so
an edited initialiser repaints the same storage on C-c C-c plus re-run —
defonce (Clojure's name for CL's defvar, per the author) initialises once
behind the .init~once. flag, and defconst stays the image.

One parse arm reads both forms; the difference is Ast.reinit, carried to
Tast.global's grerun. Emit.startup_plan gives a def no guard flag, and
Check.check_global lifts every def initialiser — zero and literal
included — into global/<n>, so the host's startup reaches it through the
function cell and a re-evaluated def swaps it (Session's def_inits;
Emit.redefinition declares the cell for a non-sibling target). The old
defvar spelling is refused with the rename and both compiling spellings,
and every program, test, doc and editor list is swept — except sand.flan,
the author's live WIP, whose seven defvar lines are flagged in FIX.org
and keep its three dependent tests red on this branch.
2026-09-21 07:12:04 +07:00

269 lines
13 KiB
OCaml

(** Resolved types: what [Ast.texpr] means once names are looked up.
The AST's type expressions are surface syntax — [Tname "Ptr"] and
[Tapp ("Option", ...)] are just names there. Here they are the real thing,
and two types are the same type exactly when they are structurally equal.
Milestone 2 has no generics, so there is no unification and no substitution:
a type variable is parsed, carried, and rejected the moment a value would
have to have it. That rejection lives in [Check]; this module only names
the shape. *)
(* Machine integer types. Signedness and width are both part of the type, and
two of them are the same type only when both halves match. A value may move
to a type that cannot lose it — see [widens_to] at the bottom of this file,
FIX.org 2026-09-20 — and never the other way: narrowing is written or it
does not happen. *)
type ikind = I8 | I16 | I32 | I64 | U8 | U16 | U32 | U64
type fkind = F32 | F64
type t =
| Int of ikind
| Float of fkind
| Bool
| String
| Unit (* the zero-sized type, not C's void *)
| Never (* return, exit, error: no value at all *)
| Named of string (* a struct or data type, declared here *)
(* A C enum: an i32 at run time, but its own type, so a keyword at a call
site has something to resolve against and a plain integer does not fit. *)
| Enum of string
| Slice of t (* [T] ptr+len, non-owning *)
| Array of int64 * t (* [n T] inline, a value, copies *)
| Map of t * t (* (Map K V) *)
| Ptr of t (* (Ptr T) *)
(* [Allocator]: a builtin opaque type, the way [string] is a builtin
ptr+len. It is a [Types.t] case with no user-writable constructor, which
is what lets spec-memory.md's "procedure plus an opaque data pointer" be
expressed with none of milestone 5's function values — the procedure is a
C symbol the emitter names and no Flan type ever mentions it. At run time
it is a pointer to the runtime's [flan_allocator], never a copy of one:
the capability set and the epoch have to be shared by every container
made from it, and a copy would give each its own. *)
| Alloc
(* [(Vec T)]: ptr + len + cap + allocator, owning and move-only. One
type-erased runtime over (size, align) stands behind every instantiation,
so this is a container without generics — the concrete type is known only
at the call site, which is exactly where the two numbers are produced. *)
| Vec of t
| Option of t (* (Option T) *)
| Fn of t list * t (* (Fn [T ...] R) *)
| Var of string (* a type variable — milestone 5 *)
(* [dyn]: one machine word whose contents the runtime knows and this module
does not. It is a written type — [(defonce x dyn 5)] boxes the 5 — and it
is also what an unannotated [defn] parameter means, which is why it is a
case here and not a Named type the prelude declares: the checker has to
recognise it to choose the boxing and the dyn op lowering, and a name in a
table cannot be matched on.
Nothing about the representation is stated here on purpose. The word is
opaque to the compiler — runtime/flan_dyn.h owns which bits are a tag —
so that milestone 2 can change the encoding without touching Emit. *)
| Dyn
let signed = function
| I8 | I16 | I32 | I64 -> true
| U8 | U16 | U32 | U64 -> false
let bits = function
| I8 | U8 -> 8 | I16 | U16 -> 16 | I32 | U32 -> 32 | I64 | U64 -> 64
let bits_f = function F32 -> 32 | F64 -> 64
(* [int] and [float] are the two builtin aliases, and they are spelled here
rather than as prelude [defalias]es so that they are the machine type and
not a second name for it. The difference is visible at a cast: [Check]'s
[is_cast] asks these two functions whether a head names a primitive, and an
entry in the alias table is not consulted there — so a prelude alias would
give [(int x)] no reading while [(i32 x)] had one. Named here, every path
that already accepts [i32] accepts [int] without learning the word.
Only these two. The rest of the foreign spellings — [long], [double],
[uint], [str] — stay refusals that teach the Flan name; see
[Check.foreign_spelling] for why the line is drawn where it is.
The mapping is one-way on purpose: [ikind_name] and [fkind_name] below
still answer [i32] and [f32], so every message, every DWARF name and every
inspector line the user sees says the machine type, whichever spelling the
source used. *)
let ikind_of_name = function
| "i8" -> Some I8 | "i16" -> Some I16 | "i32" | "int" -> Some I32
| "i64" -> Some I64
| "u8" -> Some U8 | "u16" -> Some U16 | "u32" -> Some U32 | "u64" -> Some U64
| _ -> None
let fkind_of_name = function
| "f32" | "float" -> Some F32 | "f64" -> Some F64 | _ -> None
(* Every name the resolver accepts as a primitive type. The list exists so a
near-miss can be reported as the typo it is. [Unit] is on it because the
resolver still answers to that name -- [Cimport] builds [Tname "Unit"] for
C's void, and never goes through the parser -- but nobody writes it: unit
is spelled [()] in source, and [Parse.texpr] refuses the word.
[int] and [float] are on it for the same reason they are in the two
functions above: the places that ask this list — whether a [defonce]'s third
element is a type, whether [(vec-new int)] names an element type, whether a
[let] binding vector has an annotation wedged into it — must answer the
same for [int] as for [i32], or the alias is a type-position-only spelling
and the identity is a half one. *)
let primitive_names =
[ "i8"; "i16"; "i32"; "i64"; "u8"; "u16"; "u32"; "u64";
"f32"; "f64"; "bool"; "string"; "dyn"; "Unit"; "Never"; "Allocator";
"int"; "float" ]
let ikind_name k =
(if signed k then "i" else "u") ^ string_of_int (bits k)
let fkind_name = function F32 -> "f32" | F64 -> "f64"
(* Structural equality is the whole story for *identity*: no subtyping, no
variance, and nothing here bends to admit a conversion. Implicit widening
(below) is deliberately not expressed as a loosening of this function or of
[fits] — it is a separate predicate that every caller must pair with a
[Cast] on the value, so a node's type never lies about the bits it holds.
Written out rather than using [=] so that adding a case with a function or
a mutable field cannot silently break it. *)
let rec equal a b =
match a, b with
| Int x, Int y -> x = y
| Float x, Float y -> x = y
(* [Dyn] is equal to itself and to nothing else. Two dyn values may hold
different things at run time, which is the point of the type and is not
this function's question: this is identity of *static* types, and there is
one dyn type the way there is one string type. *)
| Bool, Bool | String, String | Unit, Unit | Never, Never | Dyn, Dyn -> true
| Named x, Named y | Enum x, Enum y -> String.equal x y
| Slice x, Slice y -> equal x y
| Array (n, x), Array (m, y) -> Int64.equal n m && equal x y
| Map (k, v), Map (k', v') -> equal k k' && equal v v'
| Ptr x, Ptr y -> equal x y
| Alloc, Alloc -> true
| Vec x, Vec y -> equal x y
| Option x, Option y -> equal x y
| Fn (ps, r), Fn (ps', r') ->
List.length ps = List.length ps'
&& List.for_all2 equal ps ps'
&& equal r r'
| Var x, Var y -> String.equal x y
| _ -> false
let rec to_string = function
| Int k -> ikind_name k
| Float k -> fkind_name k
| Bool -> "bool"
| String -> "string"
| Unit -> "()"
| Never -> "Never"
| Named n | Enum n -> n
| Slice t -> "[" ^ to_string t ^ "]"
| Array (n, t) -> Printf.sprintf "[%Ld %s]" n (to_string t)
| Map (k, v) -> Printf.sprintf "(Map %s %s)" (to_string k) (to_string v)
| Ptr t -> "(Ptr " ^ to_string t ^ ")"
| Alloc -> "Allocator"
| Vec t -> "(Vec " ^ to_string t ^ ")"
| Option t -> "(Option " ^ to_string t ^ ")"
| Fn (ps, r) ->
Printf.sprintf "(Fn [%s] %s)"
(String.concat " " (List.map to_string ps)) (to_string r)
| Var n -> n
| Dyn -> "dyn"
let is_numeric = function Int _ | Float _ -> true | _ -> false
(* Every integer kind, signed and unsigned, at every width — and nothing
else. This is [integer?]'s question: the bound that admits a body written
with %, the bitwise operators or the shifts, and that keeps the same body
from ever being instantiated at a float, where those operations either do
not exist or mean something different. *)
let is_integer = function Int _ -> true | _ -> false
(* The key types the first Map implementation admits (spec-memory.md, "Maps —
first implementation"): integers, enums, strings, fixed arrays, and value
structs composed recursively from those. Equality and hashing for them are
compiler-provided structural operations, so this is the whole of what the
emitted hash and equality pair has to cover — there is no dispatch to design
and no type class anywhere.
A struct is [Named], and whether its fields qualify cannot be decided here:
this module has no field table. [Check] finishes the job by walking them,
which is also where it emits the pair. Everything this does say no to says
no for a reason that will not change with a milestone: a [Ptr] or a [Slice]
key would hash an address, and hashing an address is a different operation
from hashing what it points at. *)
let rec keyable = function
| Int _ | Enum _ | Bool | String -> true
| Float _ -> false (* NaN /= NaN, and 0.0 and -0.0 differ bytewise *)
| Array (_, t) -> keyable t
| Named _ -> true (* [Check] decides, by walking the fields *)
| _ -> false
(* Ordering is defined on machine types and on nothing else — structs and
slices have no built-in [<], because an unconstrained type supports only
what every type supports (plan.org, Types). A string has no ordering
either: there is no true answer to whether one string is less than another
until the language picks a collation, and byte order is not it. *)
let is_comparable = function Enum _ -> true | t -> is_numeric t
(* Equality admits one type ordering does not: a string, grown in by the M2
queue's item 5 — bytewise, by content and not by address, so two
separately built strings with the same bytes are equal. *)
let is_equatable = function String -> true | t -> is_comparable t
(* [Never] is the type of an expression that does not produce a value: return,
an early-returning `some`, exit. It fits anywhere, and that is the only
place anything resembling subtyping exists. *)
let fits ~expected ~actual =
match actual with Never -> true | _ -> equal expected actual
(* ── Implicit widening, FIX.org 2026-09-20 ────────────────────────────
Which numeric types a value may move to without the program saying so.
One rule decides every entry: the conversion is admitted exactly when no
value of the source type can come out the other side as a different number.
Narrowing is not on this list and never will be — [(u32 x)] is how an i64
becomes a u32, because that one can lose.
Read out of that rule:
- Same signedness, strictly wider: i8→i16→i32→i64, u8→u16→u32→u64.
- Unsigned into a strictly wider signed: u8→i16, u8/u16→i32, u8/u16/u32→i64.
Every u32 fits in an i64, so nothing is lost. The mirror never holds:
signed into unsigned drops the negatives, at any width.
- Equal width across signedness (i32→u32, u32→i32) is refused for the same
reason — one of the two halves of the range has nowhere to go.
- f32→f64.
- Integer into float only where the float's significand covers the integer
exactly: f64 has 53 bits, so i8/i16/i32/u8/u16/u32 reach it and i64/u64 do
not (2^53+1 is not an f64); f32 has 24, so only i8/i16/u8/u16 reach it.
Odin is looser here and lets any integer into any float. This is the
tighter rule on purpose: a program that wants the lossy one writes (f64 x)
and says so, and loosening later adds programs where tightening later
would break them.
Nothing else participates. [Bool] is not a number, an [Enum] is its own type
whose whole point is that a bare integer does not fit it, [Dyn] crosses by
boxing and unboxing rather than by this, and a container is invariant: a
[Vec i32] is not a [Vec i64] and a [[i32]] is not a [[i64]], because the
elements would each have to be rewritten and a slice does not own its
bytes. *)
let widens_to ~(from : t) ~(into : t) =
match from, into with
| Int a, Int b ->
if signed a = signed b then bits b > bits a
else (not (signed a)) && signed b && bits b > bits a
| Float F32, Float F64 -> true
| Int a, Float b -> bits a <= (match b with F64 -> 32 | F32 -> 16)
| _ -> false
(* The type a binary operator's two operands meet at: whichever of the pair the
other one widens into, and nothing otherwise. That is total and it is not a
real lattice — (i32, u32) has no answer here, and inventing i64 for it would
be picking a type neither operand was written at. Equal types answer
themselves, so a caller can use this without checking for that first. *)
let join a b =
if equal a b then Some a
else if widens_to ~from:a ~into:b then Some b
else if widens_to ~from:b ~into:a then Some a
else None