The x86 backend stops leaving a divide by zero to the hardware
Item 3 of HANDOFF-x86-rt.md, which was blocked on the language decision rather than on code. check_div and check_cast sit beside check_at and check_slice and reuse bounds_call unchanged; flan_arith_error takes three extras, so the channel lands in r9 and the argument registers are exactly full. Three things differ from the LLVM side because the instruction set does: two branches rather than one branch and a select, since there is no select and a second compare on the cold path is free; the cast bounds compared in the source's own precision rather than widened to a double, which is exact because every bound is a power of two; and NaN excluded by choosing the direction of each compare, because ucomis sets CF, ZF and PF together when either operand is unordered. Both programs now print byte-identical stdout, stderr and exit status through either backend.
This commit is contained in:
parent
a431cddd3b
commit
b0f4fb73e1
@ -53,11 +53,31 @@ handler after it, because `SIGFPE` cannot be caught and resumed. Once the zero t
|
|||||||
Unsigned division gets the zero test only; there is no overflow case. Float division is **not** guarded at all:
|
Unsigned division gets the zero test only; there is no overflow case. Float division is **not** guarded at all:
|
||||||
IEEE `x / 0.0` is `inf`, which is defined and wanted — `rand-f32` in the prelude divides by a float constant.
|
IEEE `x / 0.0` is `inf`, which is defined and wanted — `rand-f32` in the prelude divides by a float constant.
|
||||||
|
|
||||||
## Order of work
|
## The x86 backend, item 3 of `HANDOFF-x86-rt.md`
|
||||||
|
|
||||||
1. `lib/prelude.ml`, `runtime/flan_rt.c`, `lib/emit.ml`, tests. Committed first and separately.
|
Landed too, in its own commit and after a rebase onto the dev-loop tip. `check_div` and `check_cast` sit next to
|
||||||
2. `lib/x86.ml`'s `prim` `Div`/`Rem`/`Cast` arms last, rebased onto the dev-loop tip, kept mechanical — this is item 3
|
`check_at` and `check_slice` and reuse `bounds_call` unchanged — it already spells the whole shape, the location string
|
||||||
of `HANDOFF-x86-rt.md`'s "What remains".
|
into rdi/rsi, the extras out of frame temporaries, the channel, the guard and the `ud2`. `flan_arith_error` takes three
|
||||||
|
extras, so the channel lands in r9 and the argument registers are exactly full.
|
||||||
|
|
||||||
|
Three things differ from the LLVM side, all of them because the instruction set does:
|
||||||
|
|
||||||
|
- **Two branches rather than one branch and a `select`.** There is no `select` here, and a second compare on the cold
|
||||||
|
path costs nothing. The ordinary path still pays one compare and one not-taken branch, which is what `emit.ml` pays.
|
||||||
|
- **The cast bounds are compared in the source's own precision** instead of widening the value to a double first. Every
|
||||||
|
bound is a power of two and therefore exact in an `f32` as well as an `f64`, so the two routes answer identically —
|
||||||
|
and the survey is there to say so, which is why `arith.flan` has an `f32` case.
|
||||||
|
- **NaN is excluded by choosing the direction of each compare.** `ucomis` sets CF, ZF and PF together when either
|
||||||
|
operand is unordered, so the low test jumps to the failure on "below" (which a NaN takes) and the high test swaps its
|
||||||
|
operands to ask "hi > v" (which a NaN answers false).
|
||||||
|
|
||||||
|
The `INT_MIN / -1` test compares against the *narrow* type's most negative value in a 64-bit register, which is sound
|
||||||
|
because `load_loc` has already widened both operands according to their own signedness. That case is where the two
|
||||||
|
backends used to diverge silently rather than both dying: x86 divided in 64 bits and truncated on the store, producing
|
||||||
|
`-2147483648` for an `i32`, where LLVM emitted poison. `arith.flan` has an `i32` case for exactly that reason.
|
||||||
|
|
||||||
|
`spike/x86/survey.sh` is 100 MATCH / 0 DIFFER / 0 REFUSED — the 98 that was the baseline plus the two programs this
|
||||||
|
change adds.
|
||||||
|
|
||||||
## What landed
|
## What landed
|
||||||
|
|
||||||
|
|||||||
153
lib/x86.ml
153
lib/x86.ml
@ -1752,6 +1752,140 @@ and check_slice f (base : loc) (ty : Types.t) (loc : Loc.t) (lo : Tast.expr)
|
|||||||
bounds_call f "flan_slice_error" loc [ a; b; c ];
|
bounds_call f "flan_slice_error" loc [ a; b; c ];
|
||||||
lbl f.b ok)
|
lbl f.b ok)
|
||||||
|
|
||||||
|
(* [emit.ml]'s [check_div] and [check_cast], item 3 of HANDOFF-x86-rt.md's
|
||||||
|
list, and the one item on it that was blocked on a language decision rather
|
||||||
|
than on code. That decision is in HANDOFF-arith.md: a divide or remainder by
|
||||||
|
zero, the one division that overflows, and a float to integer cast whose
|
||||||
|
value does not fit all signal ArithError, exactly as a bad index signals
|
||||||
|
BoundsError.
|
||||||
|
|
||||||
|
Why it could not be left to the hardware, which is the temptation here and
|
||||||
|
is what this backend did until now. `idiv` raises SIGFPE on both the zero
|
||||||
|
and the overflow case, and a SIGFPE cannot be caught and resumed — so there
|
||||||
|
is no version of this that tests afterwards, and nothing that dies with a
|
||||||
|
location. The other backend calls the same three situations undefined and
|
||||||
|
folds them to whatever it likes. Neither is a behaviour a program can be
|
||||||
|
written against, and the two disagreed, which is what a survey diff would
|
||||||
|
eventually have found the hard way.
|
||||||
|
|
||||||
|
These reuse [bounds_call] unchanged: it spells the whole shape — the
|
||||||
|
location string into rdi/rsi, the extra arguments out of frame temporaries
|
||||||
|
into rdx/rcx/r8/r9, the channel after them, the guard, and the [ud2] that
|
||||||
|
stands where [emit.ml] writes [unreachable]. flan_arith_error takes three
|
||||||
|
extras, so the channel lands in r9 and the register file is exactly full. *)
|
||||||
|
|
||||||
|
(* The codes are [emit.ml]'s, read from there rather than copied: they are an
|
||||||
|
agreement with flan_arith_fail in the runtime, and an agreement kept in two
|
||||||
|
places is an agreement that drifts. *)
|
||||||
|
|
||||||
|
(* rax holds the dividend and rcx the divisor, both already widened to 64 bits
|
||||||
|
by [load_loc] according to their own signedness — which is what lets the
|
||||||
|
overflow test compare against the *narrow* type's most negative value in a
|
||||||
|
64-bit register and mean it.
|
||||||
|
|
||||||
|
Two tests and two ways out rather than [emit.ml]'s single branch with a
|
||||||
|
[select], because there is no [select] here and a second compare on the cold
|
||||||
|
path is free. The ordinary path still pays one compare and one
|
||||||
|
not-taken branch, which is the same as over there.
|
||||||
|
|
||||||
|
Both tests are dropped when a literal divisor cannot trigger them. That is
|
||||||
|
not a micro-optimisation: (/ x 2) is the common case and would otherwise
|
||||||
|
carry a compare and a branch forever. *)
|
||||||
|
and check_div f (loc : Loc.t) ~is_rem (k : Types.ikind) ~lit =
|
||||||
|
if f.md.Emit.checks then begin
|
||||||
|
let need_zero = match lit with Some n -> Int64.equal n 0L | None -> true in
|
||||||
|
let need_ovf =
|
||||||
|
Types.signed k
|
||||||
|
&& (match lit with Some n -> Int64.equal n (-1L) | None -> true)
|
||||||
|
in
|
||||||
|
if need_zero || need_ovf then
|
||||||
|
scoped f (fun () ->
|
||||||
|
let so = ptmp f and sa = ptmp f and sb = ptmp f in
|
||||||
|
store_int f.b ~src:rax ~mm:(Frame sa) ~size:8;
|
||||||
|
store_int f.b ~src:rcx ~mm:(Frame sb) ~size:8;
|
||||||
|
let ok = new_label f "arith" and bad = new_label f "arithbad" in
|
||||||
|
let zcode = if is_rem then Emit.arith_rem_zero else Emit.arith_div_zero in
|
||||||
|
let ocode = if is_rem then Emit.arith_rem_overflow else Emit.arith_div_overflow in
|
||||||
|
if need_zero then begin
|
||||||
|
let nz = new_label f "arithnz" in
|
||||||
|
cmp_imm f.b ~dst:rcx 0;
|
||||||
|
jcc_lbl f.b ~cc:cc_ne nz;
|
||||||
|
imm_into f ~reg:rdx (Int64.of_int zcode);
|
||||||
|
store_int f.b ~src:rdx ~mm:(Frame so) ~size:8;
|
||||||
|
jmp_lbl f.b bad;
|
||||||
|
lbl f.b nz
|
||||||
|
end;
|
||||||
|
if need_ovf then begin
|
||||||
|
cmp_imm f.b ~dst:rcx (-1);
|
||||||
|
jcc_lbl f.b ~cc:cc_ne ok;
|
||||||
|
(* Through r11 rather than as an immediate: the most negative i64
|
||||||
|
does not fit the imm32 [cmp_imm] encodes, and one spelling for
|
||||||
|
every width beats a special case for the one that does not. *)
|
||||||
|
imm_into f ~reg:r11 (Int64.neg (Int64.shift_left 1L (Types.bits k - 1)));
|
||||||
|
cmp_rr f.b ~a:rax ~c:r11;
|
||||||
|
jcc_lbl f.b ~cc:cc_ne ok;
|
||||||
|
imm_into f ~reg:rdx (Int64.of_int ocode);
|
||||||
|
store_int f.b ~src:rdx ~mm:(Frame so) ~size:8
|
||||||
|
end
|
||||||
|
else jmp_lbl f.b ok;
|
||||||
|
lbl f.b bad;
|
||||||
|
bounds_call f "flan_arith_error" loc [ so; sa; sb ];
|
||||||
|
lbl f.b ok;
|
||||||
|
(* rdx is the high half of the dividend and [cqo] is what fills it, so
|
||||||
|
whatever the overflow test left there does not survive; rax and rcx
|
||||||
|
are untouched on this path and do not need reloading. *)
|
||||||
|
())
|
||||||
|
end
|
||||||
|
|
||||||
|
(* A float to integer cast whose value does not fit. xmm0 holds the value, in
|
||||||
|
the *source's* precision, and the bounds are compared in that same precision
|
||||||
|
rather than widened to a double first the way [emit.ml] does it: both bounds
|
||||||
|
are powers of two, so both are exact in an f32 as well as in an f64, and the
|
||||||
|
two tests therefore answer identically. Doing it here saves a conversion and
|
||||||
|
a second live xmm register.
|
||||||
|
|
||||||
|
The direction of each compare is chosen so that a NaN fails both. [ucomis]
|
||||||
|
sets CF, ZF and PF together when either operand is unordered, so the test
|
||||||
|
for the low end is written as "jump to the failure when below", which a NaN
|
||||||
|
takes, and the test for the high end swaps its operands and asks the same
|
||||||
|
question the other way round. A NaN cast to an integer is exactly as
|
||||||
|
undefined as 1e300 is and has no business walking through the guard. *)
|
||||||
|
and check_cast f (loc : Loc.t) (src : Types.fkind) (k : Types.ikind) =
|
||||||
|
if f.md.Emit.checks then begin
|
||||||
|
let f64 = (src = Types.F64) in
|
||||||
|
let n = Types.bits k in
|
||||||
|
let signed = Types.signed k in
|
||||||
|
let lo_f = if signed then ldexp (-1.0) (n - 1) else 0.0 in
|
||||||
|
let hi_f = if signed then ldexp 1.0 (n - 1) else ldexp 1.0 n in
|
||||||
|
let lo_i = if signed then Int64.neg (Int64.shift_left 1L (n - 1)) else 0L in
|
||||||
|
let hi_i =
|
||||||
|
if signed then Int64.sub (Int64.shift_left 1L (n - 1)) 1L
|
||||||
|
else if n = 64 then -1L
|
||||||
|
else Int64.sub (Int64.shift_left 1L n) 1L
|
||||||
|
in
|
||||||
|
let klo = float_const f lo_f ~f64 and khi = float_const f hi_f ~f64 in
|
||||||
|
scoped f (fun () ->
|
||||||
|
let so = ptmp f and sa = ptmp f and sb = ptmp f in
|
||||||
|
let ok = new_label f "fits" and bad = new_label f "nofit" in
|
||||||
|
fload f.b ~dst:1 ~mm:(Sym (klo, 0)) ~f64;
|
||||||
|
ucomis f.b ~f64 ~a:xmm0 ~c:1;
|
||||||
|
jcc_lbl f.b ~cc:cc_b bad;
|
||||||
|
fload f.b ~dst:1 ~mm:(Sym (khi, 0)) ~f64;
|
||||||
|
(* The operands the other way round, so that the code asked for is one a
|
||||||
|
NaN answers false to: this is "hi > v" and not "v < hi". *)
|
||||||
|
ucomis f.b ~f64 ~a:1 ~c:xmm0;
|
||||||
|
jcc_lbl f.b ~cc:cc_a ok;
|
||||||
|
lbl f.b bad;
|
||||||
|
imm_into f ~reg:rax (Int64.of_int Emit.arith_cast_range);
|
||||||
|
store_int f.b ~src:rax ~mm:(Frame so) ~size:8;
|
||||||
|
imm_into f ~reg:rax lo_i;
|
||||||
|
store_int f.b ~src:rax ~mm:(Frame sa) ~size:8;
|
||||||
|
imm_into f ~reg:rax hi_i;
|
||||||
|
store_int f.b ~src:rax ~mm:(Frame sb) ~size:8;
|
||||||
|
bounds_call f "flan_arith_error" loc [ so; sa; sb ];
|
||||||
|
lbl f.b ok)
|
||||||
|
end
|
||||||
|
|
||||||
and element f (base : loc) (ty : Types.t) (i : Tast.expr) : loc =
|
and element f (base : loc) (ty : Types.t) (i : Tast.expr) : loc =
|
||||||
let elem =
|
let elem =
|
||||||
match ty with
|
match ty with
|
||||||
@ -1949,6 +2083,18 @@ and prim f (e : Tast.expr) (p : Tast.prim) (args : Tast.expr list) dst =
|
|||||||
| Tast.Shl -> shl_cl f.b ~dst:rax
|
| Tast.Shl -> shl_cl f.b ~dst:rax
|
||||||
| Tast.Shr -> if signed then sar_cl f.b ~dst:rax else shr_cl f.b ~dst:rax
|
| Tast.Shr -> if signed then sar_cl f.b ~dst:rax else shr_cl f.b ~dst:rax
|
||||||
| Tast.Div | Tast.Rem ->
|
| Tast.Div | Tast.Rem ->
|
||||||
|
(* The guard goes *before* the instruction, which is the whole of why
|
||||||
|
it has to be emitted at all: `idiv` raises SIGFPE on both the zero
|
||||||
|
and the overflow case and a SIGFPE cannot be caught and resumed.
|
||||||
|
Integers only — this arm is already inside the non-float half —
|
||||||
|
because IEEE x / 0.0 is an infinity and is a defined answer. *)
|
||||||
|
(match t with
|
||||||
|
| Types.Int k ->
|
||||||
|
let lit =
|
||||||
|
match b.Tast.e with Tast.Int (n, _) -> Some n | _ -> None
|
||||||
|
in
|
||||||
|
check_div f e.Tast.loc ~is_rem:(p = Tast.Rem) k ~lit
|
||||||
|
| _ -> ());
|
||||||
if signed then (cqo f.b; idiv_r f.b ~src:rcx)
|
if signed then (cqo f.b; idiv_r f.b ~src:rcx)
|
||||||
else (xor_rr f.b ~dst:rdx ~src:rdx; div_r f.b ~src:rcx);
|
else (xor_rr f.b ~dst:rdx ~src:rdx; div_r f.b ~src:rcx);
|
||||||
if p = Tast.Rem then mov_rr f.b ~dst:rax ~src:rdx
|
if p = Tast.Rem then mov_rr f.b ~dst:rax ~src:rdx
|
||||||
@ -2153,6 +2299,13 @@ and cast f (a : Tast.expr) (target : Types.t) dst =
|
|||||||
fstore f.b ~src:xmm0 ~mm:(lmem f dst ~scratch:r11) ~f64:(f64_of dst_t)
|
fstore f.b ~src:xmm0 ~mm:(lmem f dst ~scratch:r11) ~f64:(f64_of dst_t)
|
||||||
| true, false ->
|
| true, false ->
|
||||||
fload f.b ~dst:xmm0 ~mm:(lmem f l ~scratch:r11) ~f64:(f64_of src_t);
|
fload f.b ~dst:xmm0 ~mm:(lmem f l ~scratch:r11) ~f64:(f64_of src_t);
|
||||||
|
(* [cvttsd2si] answers a fixed "integer indefinite" for a value out of
|
||||||
|
range, which is a number rather than an answer — and the other backend
|
||||||
|
calls the same cast undefined and will fold it to anything. So the
|
||||||
|
value is tested against the destination's range first. *)
|
||||||
|
(match src_t, dst_t with
|
||||||
|
| Types.Float sk, Types.Int k -> check_cast f a.Tast.loc sk k
|
||||||
|
| _ -> ());
|
||||||
cvttf2si f.b ~f64:(f64_of src_t) ~dst:rax ~src:xmm0;
|
cvttf2si f.b ~f64:(f64_of src_t) ~dst:rax ~src:xmm0;
|
||||||
store_loc f ~reg:rax dst dst_t
|
store_loc f ~reg:rax dst dst_t
|
||||||
|
|
||||||
|
|||||||
Loading…
x
Reference in New Issue
Block a user