585 lines
26 KiB
Plaintext
585 lines
26 KiB
Plaintext
;;;; An EDN tokenizer, in Flan, over a [const u8].
|
|
;;;;
|
|
;;;; This is the bottom layer of a reader. It answers one question — "what is
|
|
;;;; the next token, and where" — and it answers it without allocating
|
|
;;;; anything: every token's text is a `slice` of the input buffer, not a copy
|
|
;;;; of it.
|
|
;;;;
|
|
;;;; Two layers sit above it. `read.fln`, in this package, is the one that
|
|
;;;; exists: `edn/read(bytes)` walks this cursor and answers a dynamic
|
|
;;;; `Value`. The other, `read-edn(Enemy, bytes)` emitting a parser from a
|
|
;;;; compile-time walk over a struct's fields, belongs to the compiler and is
|
|
;;;; not here; until it exists a caller writes the struct reader by hand
|
|
;;;; against this cursor, and test/programs/edn.flan is a worked example of
|
|
;;;; doing exactly that.
|
|
;;;;
|
|
;;;; ── The lifetime contract, which the type system does not state ─────
|
|
;;;;
|
|
;;;; A Token's `text` is a slice INTO the buffer the Cursor was built over.
|
|
;;;; It is ptr+len and it owns nothing. Therefore:
|
|
;;;;
|
|
;;;; * the input buffer must outlive every Token taken from it, and every
|
|
;;;; Cursor over it;
|
|
;;;; * mutating the input while tokens are live changes their text under
|
|
;;;; them, because they are views and not copies;
|
|
;;;; * a Token returned out of the function that owns the buffer is a
|
|
;;;; dangling pointer, and nothing in the language will say so.
|
|
;;;;
|
|
;;;; That is the price of not allocating, and it is written here because it is
|
|
;;;; the kind of contract that otherwise gets discovered from a corrupted
|
|
;;;; string three frames later.
|
|
;;;;
|
|
;;;; **`read.fln` does not follow this rule, deliberately.** The two layers of
|
|
;;;; this package diverge on exactly this point: a Token is a view, and a Value
|
|
;;;; owns copies of every string in it. The reason is that a view is a fine
|
|
;;;; thing for a cursor a caller is driving inside the function that holds the
|
|
;;;; buffer, and a trap for a document handed back out of one. Said the other
|
|
;;;; way: the contract above is a property of the *layer*, not of the package,
|
|
;;;; and a caller who mixes them — holding a Token out of an `edn/read(...)`
|
|
;;;; that has returned — is on the tokenizer's terms and not the reader's.
|
|
;;;;
|
|
;;;; ── What is refused, and why ────────────────────────────────────────
|
|
;;;;
|
|
;;;; Every refusal below is a *named* one with a reason attached, reachable as
|
|
;;;; edn/error-message(code). A tokenizer that quietly skipped what it did not
|
|
;;;; understand would hand a caller a value that is not the one in the file.
|
|
;;;;
|
|
;;;; escaped strings "a\nb", "a\"b" — the important one. Unescaping needs
|
|
;;;; somewhere to put the unescaped copy, and there is no
|
|
;;;; allocator, so there is nowhere. Returning the raw
|
|
;;;; bytes including the backslash would be quietly wrong:
|
|
;;;; a caller comparing against "a\nb" would get a 4-byte
|
|
;;;; answer where it expected 3, and a caller printing it
|
|
;;;; would print a backslash. So a backslash inside a
|
|
;;;; string is an error at the byte it appears on.
|
|
;;;; tagged literals #foo {} — the tag decides the type, and dispatching on
|
|
;;;; a tag at run time is what a type-directed reader
|
|
;;;; exists to avoid.
|
|
;;;; #inst, #uuid named separately from tagged literals because they are
|
|
;;;; the two a real file is most likely to contain, and
|
|
;;;; "tagged literals are refused" would not tell a caller
|
|
;;;; that a timestamp is the thing to remove.
|
|
;;;; ratios 22/7 — there is no rational type.
|
|
;;;; metadata ^{:a 1} — it attaches to the value after it, and a
|
|
;;;; flat token stream has nowhere to attach anything.
|
|
;;;; characters \a — outside the requested subset; a char is not
|
|
;;;; a byte once anything is non-ASCII, and there is no
|
|
;;;; code point type.
|
|
;;;;
|
|
;;;; ── One place this is not EDN, on the record ────────────────────────
|
|
;;;;
|
|
;;;; `.5` is a float here. In EDN a number must begin with a digit and `.` is
|
|
;;;; a legal symbol-start byte, so strictly `.5` is the *symbol* `.5` — which
|
|
;;;; makes this a reinterpretation of a legal token and not an extension, and
|
|
;;;; therefore the kind of thing that gets written down rather than discovered.
|
|
;;;; It is this way because is-number-start runs before the symbol case and
|
|
;;;; parse-f64 accepts a leading dot; a caller who needs the symbol reading
|
|
;;;; should not be writing `.5` at all. `-`, by contrast, is a symbol, because
|
|
;;;; is-number-start requires a digit after the sign.
|
|
;;;;
|
|
;;;; ── Errors ──────────────────────────────────────────────────────────
|
|
;;;;
|
|
;;;; On the cursor, not in the return type. `next` answers a Token whose kind
|
|
;;;; is tok-error, and the cursor carries the code and the byte offset it was
|
|
;;;; found at; edn/error-message(code) turns the code into the sentence. The
|
|
;;;; offset is the point: an editor underlines a byte range, and an Option with
|
|
;;;; no position could not tell it where. An Option(Token) was the alternative
|
|
;;;; and it loses exactly that — None says something went wrong, and a second
|
|
;;;; out-parameter for the position is the same two fields with a worse shape.
|
|
;;;;
|
|
;;;; A failed cursor is poisoned: every later `next` answers the same error
|
|
;;;; token without advancing. That is what stops a caller's `while` loop from
|
|
;;;; spinning on a malformed file forever.
|
|
|
|
;; ── Token kinds ─────────────────────────────────────────────────────
|
|
;;
|
|
;; Plain i32 constants and not an `enum`, which is the shape that wants
|
|
;; explaining. An enum here is FFI-only: `=` on an Enum value fails in emit,
|
|
;; and a keyword is not a pattern, so `match` cannot see one either. Both fixes
|
|
;; live in check.ml and emit.ml, which this lane does not touch. An i32 loses
|
|
;; the compile-time typo check on a keyword and gains a token kind a caller can
|
|
;; actually branch on, which is the whole job.
|
|
|
|
const tok-eof = 0 ; the input is exhausted; text is empty
|
|
const tok-error = 1 ; see edn/error(c) and edn/error-message(...)
|
|
const tok-nil = 2 ; nil
|
|
const tok-bool = 3 ; true / false — text is the word
|
|
const tok-int = 4 ; text parses as i64
|
|
const tok-float = 5 ; text parses as f64
|
|
const tok-string = 6 ; text is the CONTENTS, without the quotes
|
|
const tok-keyword = 7 ; text is WITHOUT the leading colon
|
|
const tok-symbol = 8 ; text is the symbol, namespace and all
|
|
const tok-vec-open = 9 ; [
|
|
const tok-vec-close = 10 ; ]
|
|
const tok-map-open = 11 ; {
|
|
const tok-map-close = 12 ; }
|
|
const tok-list-open = 13 ; (
|
|
const tok-list-close = 14 ; )
|
|
|
|
;; #{ — and there is deliberately no tok-set-close. A set closes on `}`, the
|
|
;; same byte a map closes on, and a closer that answered a different kind
|
|
;; depending on what was open would be asking the caller to track a thing the
|
|
;; opener already told it. Appended rather than slotted in beside the other
|
|
;; openers because these are the numbers a caller branches on.
|
|
const tok-set-open = 15 ; #{
|
|
|
|
;; ── Error codes ─────────────────────────────────────────────────────
|
|
|
|
const err-none = 0
|
|
const err-unexpected-byte = 1
|
|
const err-unterminated = 2
|
|
const err-string-escape = 3 ; refusal
|
|
const err-tagged = 4 ; refusal
|
|
const err-inst = 5 ; refusal
|
|
const err-uuid = 6 ; refusal
|
|
const err-metadata = 7 ; refusal
|
|
const err-ratio = 8 ; refusal
|
|
const err-char = 9 ; refusal
|
|
const err-bad-number = 10
|
|
const err-empty-keyword = 11
|
|
const err-unbalanced = 12 ; a closer that does not match what is open
|
|
const err-too-deep = 13
|
|
const err-unexpected-token = 14 ; raised by a caller, not by the tokenizer
|
|
|
|
;; err-set was 4 and is gone rather than kept with a new message. A code that
|
|
;; nothing raises is a code a caller can still test for and never see, and
|
|
;; renumbering the rest is free: these are named constants, and the only place
|
|
;; a number appears is in this list.
|
|
|
|
;; How deep a nesting the balance check can follow. A fixed array in the
|
|
;; Cursor and not a growable stack, because there is no allocator; 32 is far
|
|
;; past anything a hand-written config file contains, and past it the answer is
|
|
;; err-too-deep rather than a silently unchecked closer.
|
|
const max-depth = 32
|
|
|
|
;; ── The types ───────────────────────────────────────────────────────
|
|
|
|
;; `text` is a slice of the Cursor's `src`. Read the lifetime contract at the
|
|
;; top of this file before storing one anywhere.
|
|
;;
|
|
;; `pos` is the offset of the token's first byte in the ORIGINAL buffer — of
|
|
;; the opening quote for a string, of the colon for a keyword — so it stays a
|
|
;; usable underline position even though `text` is narrower than the token.
|
|
struct Token(kind: i32, text: [const u8], pos: i32)
|
|
|
|
;; The cursor owns no storage either: `src` is the caller's buffer.
|
|
;;
|
|
;; `open` is the stack of delimiters still open, holding the tok-*-close kind
|
|
;; each one is waiting for. Balance is checked in `next` itself rather than
|
|
;; left to a parser, because `[1 2}` is malformed in a way only the tokenizer
|
|
;; has the position for.
|
|
struct Cursor
|
|
src: [const u8]
|
|
pos: i32
|
|
err: i32
|
|
err-pos: i32
|
|
open: [max-depth i32]
|
|
depth: i32
|
|
|
|
;; ── Construction ────────────────────────────────────────────────────
|
|
|
|
fn cursor(src: [const u8]) -> Cursor
|
|
Cursor{.src src .pos 0 .err err-none .err-pos 0 .depth 0}
|
|
|
|
fn is-ok(c: Ptr(Cursor)) -> bool = c.err == err-none
|
|
|
|
fn error(c: Ptr(Cursor)) -> i32 = c.err
|
|
|
|
fn error-pos(c: Ptr(Cursor)) -> i32 = c.err-pos
|
|
|
|
;; Each refusal names itself and says why, so a file that uses one fails with
|
|
;; the sentence explaining what to do about it rather than with a code.
|
|
fn error-message(code: i32) -> str
|
|
if code == err-none
|
|
"no error"
|
|
elif code == err-unexpected-byte
|
|
"unexpected byte: not the start of any EDN value"
|
|
elif code == err-unterminated
|
|
"unterminated string: end of input before the closing quote"
|
|
elif code == err-string-escape
|
|
"escaped strings are refused: unescaping needs a copy of the bytes, and there is no allocator to put one in"
|
|
elif code == err-tagged
|
|
"tagged literals #tag are refused: the tag would pick the type at run time, which is what a type-directed reader exists to avoid"
|
|
elif code == err-inst
|
|
"#inst is refused: it is a tagged literal, and there is no timestamp type to read it into"
|
|
elif code == err-uuid
|
|
"#uuid is refused: it is a tagged literal, and there is no uuid type to read it into"
|
|
elif code == err-metadata
|
|
"metadata ^ is refused: it attaches to the value after it, and a flat token stream has nowhere to attach it"
|
|
elif code == err-ratio
|
|
"ratios are refused: there is no rational type, and rounding one to a float would change the value"
|
|
elif code == err-char
|
|
"character literals are refused: a character is not a byte once it is not ASCII, and there is no code point type"
|
|
elif code == err-bad-number
|
|
"not a number: the token starts like one but does not parse as an integer or a float"
|
|
elif code == err-empty-keyword
|
|
"empty keyword: a colon with no name after it"
|
|
elif code == err-unbalanced
|
|
"unbalanced: this closing delimiter does not match the one that is open"
|
|
elif code == err-too-deep
|
|
"nesting is too deep: the balance stack is a fixed array and it is full"
|
|
elif code == err-unexpected-token
|
|
"unexpected token: not the kind the caller was reading"
|
|
else
|
|
"unknown error code"
|
|
|
|
;; Marks the cursor failed. Public, because a caller's own reader needs to
|
|
;; report "expected an integer here" with a position the same way this file
|
|
;; does, and there is nowhere else the position would come from.
|
|
;;
|
|
;; The first failure wins: a later one would overwrite the offset that
|
|
;; explains the file, with an offset that is merely downstream of it.
|
|
fn fail(c: Ptr(Cursor), code: i32, pos: i32) -> ()
|
|
if c.err == err-none
|
|
c.err = code
|
|
c.err-pos = pos
|
|
|
|
;; ── Byte classes ────────────────────────────────────────────────────
|
|
|
|
;; A comma is whitespace in EDN, which is the rule most hand-written readers
|
|
;; get wrong: {:a 1, :b 2} is one map and the comma is not a token.
|
|
fn- is-ws(b: u8) -> bool = is-space(b) or b == \,
|
|
|
|
;; Everything that ends an unquoted token. Note `;` is here: `[1;c` has the
|
|
;; comment start immediately after the 1, with no space, and a scanner that
|
|
;; only stopped on whitespace and brackets would read "1;c" as one number.
|
|
fn- is-delim(b: u8) -> bool
|
|
is-ws(b) or b == \( or b == \) or b == \[ or b == \] or b == \{ or b == \} or b == \" or b == \;
|
|
|
|
fn- is-alpha(b: u8) -> bool = (b >= \a and b <= \z) or (b >= \A and b <= \Z)
|
|
|
|
;; What EDN lets a symbol begin with. It matters that this is a list and not
|
|
;; "anything that is not a delimiter": without it every stray byte becomes a
|
|
;; one-character symbol, and `@` or a backtick — a Clojure reader macro, not
|
|
;; EDN — reads as a name instead of being reported at the byte it is on.
|
|
fn- is-sym-start(b: u8) -> bool
|
|
is-alpha(b) or b == \. or b == \* or b == \+ or b == \! or b == \- or b == \_ or b == \? or b == \$ or b == \% or b == \& or b == \= or b == \< or b == \> or b == \/
|
|
|
|
;; ── Internal helpers ────────────────────────────────────────────────
|
|
;;
|
|
;; "Internal" by intent and not by enforcement: a package has no visibility
|
|
;; yet, so edn/scan-atom and edn/push-open are as callable as edn/next is.
|
|
;; Nothing below is part of the API and none of it will keep its shape.
|
|
|
|
fn- is-at-end(c: Ptr(Cursor)) -> bool = c.pos >= length(c.src)
|
|
|
|
;; An empty slice of src, positioned at p. Used for the tokens that have no
|
|
;; text of their own — eof, error, and every delimiter. It is still a slice of
|
|
;; the input rather than a slice of nothing, so `text` has one meaning for all
|
|
;; token kinds.
|
|
fn- empty-at(c: Ptr(Cursor), p: i32) -> [const u8] = slice(c.src, p, p)
|
|
|
|
fn- token(c: Ptr(Cursor), kind: i32, lo: i32, hi: i32, p: i32) -> Token
|
|
Token{.kind kind .text slice(c.src, lo, hi) .pos p}
|
|
|
|
fn- error-token(c: Ptr(Cursor)) -> Token
|
|
Token{.kind tok-error .text empty-at(c, c.err-pos) .pos c.err-pos}
|
|
|
|
;; Whitespace, commas, and `;` comments, which run to the newline or to the end
|
|
;; of input — a comment on the last line of a file with no trailing newline is
|
|
;; the case that decides whether the loop tests the length before the byte.
|
|
fn- skip-trivia(c: Ptr(Cursor)) -> ()
|
|
while not is-at-end(c)
|
|
let b = c.src[c.pos]
|
|
if is-ws(b)
|
|
c.pos += 1
|
|
elif b == \;
|
|
while not is-at-end(c) and c.src[c.pos] != \newline
|
|
c.pos += 1
|
|
;; The newline itself, if there is one. If there is not, is-at-end is
|
|
;; already true and the outer loop stops.
|
|
if not is-at-end(c)
|
|
c.pos += 1
|
|
else
|
|
return
|
|
|
|
;; The end of the unquoted token starting at lo: the first delimiter, or the
|
|
;; end of input.
|
|
fn scan-atom(c: Ptr(Cursor), lo: i32) -> i32
|
|
let i = lo
|
|
while i < length(c.src) and not is-delim(c.src[i])
|
|
i += 1
|
|
i
|
|
|
|
fn push-open(c: Ptr(Cursor), closer: i32, p: i32) -> bool
|
|
if c.depth >= max-depth
|
|
fail(c, err-too-deep, p)
|
|
return false
|
|
c.open[c.depth] = closer
|
|
c.depth += 1
|
|
true
|
|
|
|
fn- pop-close(c: Ptr(Cursor), closer: i32, p: i32) -> bool
|
|
if c.depth == 0 or c.open[c.depth - 1] != closer
|
|
fail(c, err-unbalanced, p)
|
|
return false
|
|
c.depth -= 1
|
|
true
|
|
|
|
;; ── Numbers ─────────────────────────────────────────────────────────
|
|
|
|
;; A token starting with a digit, or with a sign or a dot followed by one.
|
|
;; `-` alone is a symbol in EDN and stays one here.
|
|
fn- is-number-start(c: Ptr(Cursor), i: i32) -> bool
|
|
let s = c.src
|
|
if i >= length(s)
|
|
return false
|
|
if is-digit(s[i])
|
|
return true
|
|
(s[i] == \- or s[i] == \+ or s[i] == \.) and i + 1 < length(s) and is-digit(s[i + 1])
|
|
|
|
fn- read-number(c: Ptr(Cursor), lo: i32) -> Token
|
|
let hi = scan-atom(c, lo)
|
|
c.pos = hi
|
|
let text = slice(c.src, lo, hi)
|
|
;; A ratio is caught here and not by a "contains a slash" rule over every
|
|
;; token, because a slash is perfectly ordinary in a symbol: foo/bar is a
|
|
;; namespaced name and must stay one.
|
|
if match(index-of(text, \/), Some(_), true, None, false)
|
|
fail(c, err-ratio, lo)
|
|
return error-token(c)
|
|
if match(parse-i64(text), Some(_), true, None, false)
|
|
return token(c, tok-int, lo, hi, lo)
|
|
if match(parse-f64(text), Some(_), true, None, false)
|
|
return token(c, tok-float, lo, hi, lo)
|
|
;; "12x", and also EDN's own 1N and 1M, which have no type here.
|
|
fail(c, err-bad-number, lo)
|
|
error-token(c)
|
|
|
|
;; ── Strings ─────────────────────────────────────────────────────────
|
|
|
|
;; The whole reason this is not three lines. `text` is the interior, between
|
|
;; the quotes — so the bytes are usable directly — but `pos` is the opening
|
|
;; quote, so an editor underlines the literal and not its contents.
|
|
;;
|
|
;; A backslash anywhere inside is the refusal, reported at the backslash
|
|
;; rather than at the start of the string, because the backslash is what has
|
|
;; to be removed.
|
|
fn- read-string(c: Ptr(Cursor), lo: i32) -> Token
|
|
let i = lo + 1
|
|
s = c.src
|
|
while i < length(s)
|
|
let b = s[i]
|
|
if b == \\
|
|
c.pos = i
|
|
fail(c, err-string-escape, i)
|
|
return error-token(c)
|
|
if b == \"
|
|
c.pos = i + 1
|
|
return token(c, tok-string, lo + 1, i, lo)
|
|
i += 1
|
|
;; Ran off the end with the string still open. Reported at the opening
|
|
;; quote: that is the byte a caller has to look at, not the end of the file.
|
|
c.pos = i
|
|
fail(c, err-unterminated, lo)
|
|
error-token(c)
|
|
|
|
;; ── The dispatch ────────────────────────────────────────────────────
|
|
|
|
;; The one call a caller makes. Advances the cursor past the token it returns.
|
|
;;
|
|
;; A cursor that has already failed keeps answering the same error token and
|
|
;; does not advance, so `while t.kind != tok-eof ...` terminates on a
|
|
;; malformed file instead of spinning.
|
|
fn next(c: Ptr(Cursor)) -> Token
|
|
if not is-ok(c)
|
|
return error-token(c)
|
|
skip-trivia(c)
|
|
if is-at-end(c)
|
|
;; Something still open at the end of input is malformed, and the position
|
|
;; that helps is the end — the file stopped, not the value.
|
|
if c.depth > 0
|
|
fail(c, err-unbalanced, c.pos)
|
|
return error-token(c)
|
|
return Token{.kind tok-eof .text empty-at(c, c.pos) .pos c.pos}
|
|
let s = c.src
|
|
lo = c.pos
|
|
b = s[lo]
|
|
;; ── Delimiters, each of which moves the balance stack ──────────
|
|
if b == \[
|
|
c.pos = lo + 1
|
|
if push-open(c, tok-vec-close, lo)
|
|
token(c, tok-vec-open, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \]
|
|
c.pos = lo + 1
|
|
if pop-close(c, tok-vec-close, lo)
|
|
token(c, tok-vec-close, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \{
|
|
c.pos = lo + 1
|
|
if push-open(c, tok-map-close, lo)
|
|
token(c, tok-map-open, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \}
|
|
c.pos = lo + 1
|
|
if pop-close(c, tok-map-close, lo)
|
|
token(c, tok-map-close, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \(
|
|
c.pos = lo + 1
|
|
if push-open(c, tok-list-close, lo)
|
|
token(c, tok-list-open, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \)
|
|
c.pos = lo + 1
|
|
if pop-close(c, tok-list-close, lo)
|
|
token(c, tok-list-close, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif b == \"
|
|
read-string(c, lo)
|
|
;; ── Keywords ───────────────────────────────────────────────────
|
|
elif b == \:
|
|
let hi = scan-atom(c, lo + 1)
|
|
c.pos = hi
|
|
if hi == lo + 1
|
|
fail(c, err-empty-keyword, lo)
|
|
error-token(c)
|
|
;; text drops the colon: a caller comparing against "name" should not
|
|
;; have to write ":name", and the compiler-side reader will want the
|
|
;; bare name to match a field against.
|
|
else
|
|
token(c, tok-keyword, lo + 1, hi, lo)
|
|
;; ── The refusals that have their own byte ──────────────────────
|
|
elif b == \^
|
|
c.pos = lo + 1
|
|
fail(c, err-metadata, lo)
|
|
error-token(c)
|
|
elif b == \\
|
|
c.pos = lo + 1
|
|
fail(c, err-char, lo)
|
|
error-token(c)
|
|
elif b == \#
|
|
let hi = scan-atom(c, lo + 1)
|
|
c.pos = hi
|
|
;; #{ — the brace is a delimiter, so scan-atom stopped before it and
|
|
;; hi is lo+1; the brace itself is consumed here, which is why the
|
|
;; position moves to lo+2 and not to hi.
|
|
;;
|
|
;; The closer pushed is tok-map-close, because the byte that closes a
|
|
;; set is `}`. That is not a compromise: the balance stack holds the
|
|
;; *closing kind still owed*, and a set and a map owe the same one.
|
|
if lo + 1 < length(s) and s[lo + 1] == \{
|
|
c.pos = lo + 2
|
|
if push-open(c, tok-map-close, lo)
|
|
token(c, tok-set-open, lo, lo, lo)
|
|
else
|
|
error-token(c)
|
|
elif is-bytes-equal(slice(s, lo + 1, hi), bytes-view("inst"))
|
|
fail(c, err-inst, lo)
|
|
error-token(c)
|
|
elif is-bytes-equal(slice(s, lo + 1, hi), bytes-view("uuid"))
|
|
fail(c, err-uuid, lo)
|
|
error-token(c)
|
|
else
|
|
fail(c, err-tagged, lo)
|
|
error-token(c)
|
|
;; ── Numbers, then everything else as a symbol ──────────────────
|
|
elif is-number-start(c, lo)
|
|
read-number(c, lo)
|
|
else
|
|
let hi = scan-atom(c, lo)
|
|
;; Two ways to get here without a symbol. `hi = lo` would be a
|
|
;; zero-length atom and an infinite loop; a byte that is not a symbol
|
|
;; start is `@` or a backtick, which are Clojure and not EDN. Both
|
|
;; advance one byte before failing, so the position is the offending
|
|
;; byte and the loop cannot spin on it.
|
|
if hi == lo or not is-sym-start(b)
|
|
c.pos = lo + 1
|
|
fail(c, err-unexpected-byte, lo)
|
|
return error-token(c)
|
|
c.pos = hi
|
|
let text = slice(s, lo, hi)
|
|
if is-bytes-equal(text, bytes-view("nil"))
|
|
token(c, tok-nil, lo, hi, lo)
|
|
elif is-bytes-equal(text, bytes-view("true"))
|
|
token(c, tok-bool, lo, hi, lo)
|
|
elif is-bytes-equal(text, bytes-view("false"))
|
|
token(c, tok-bool, lo, hi, lo)
|
|
else
|
|
token(c, tok-symbol, lo, hi, lo)
|
|
|
|
;; ── Reading values out of a token ───────────────────────────────────
|
|
;;
|
|
;; Each checks the kind first. None for the wrong kind rather than a parse of
|
|
;; whatever bytes happened to be there, which is the same reason parse-i64 is
|
|
;; Flan and not strtoll.
|
|
|
|
fn int-of(t: Token) -> Option(i64)
|
|
if t.kind == tok-int then parse-i64(t.text) else None
|
|
|
|
;; Accepts an integer token too: 1 and 1.0 are the same number, and a config
|
|
;; file that writes `:speed 2` for an f32 field is not making a mistake.
|
|
fn float-of(t: Token) -> Option(f64)
|
|
if t.kind == tok-float or t.kind == tok-int then parse-f64(t.text) else None
|
|
|
|
fn bool-of(t: Token) -> Option(bool)
|
|
if t.kind == tok-bool
|
|
Some(is-bytes-equal(t.text, bytes-view("true")))
|
|
else
|
|
None
|
|
|
|
fn- is-text-equal(t: Token, s: str) -> bool
|
|
is-bytes-equal(t.text, bytes-view(s))
|
|
|
|
;; A keyword whose name is s. The leading colon is not part of `text`, so this
|
|
;; is written (is-keyword-equal t "hp") and not (is-keyword-equal t ":hp").
|
|
fn is-keyword-equal(t: Token, s: str) -> bool
|
|
t.kind == tok-keyword and is-bytes-equal(t.text, bytes-view(s))
|
|
|
|
;; ── Reading past a value ────────────────────────────────────────────
|
|
|
|
;; Consumes exactly one value — a scalar, or a whole collection with everything
|
|
;; nested inside it. This is what a struct reader calls on a map key it does
|
|
;; not know, so an extra field in a data file is ignored rather than fatal.
|
|
;;
|
|
;; Iterative on the cursor's own balance depth and not recursive: the depth is
|
|
;; already tracked, and a recursive skip would put the nesting on the C stack
|
|
;; where a deep file is a crash rather than err-too-deep.
|
|
;;
|
|
;; Being written against the depth and not against the kinds is also why sets
|
|
;; cost this function nothing: `#{` pushes in `next` like every other opener,
|
|
;; so a set was already a collection here before it was one anywhere else. An
|
|
;; arm per collection kind would have been a second list to forget to add to.
|
|
fn skip-value(c: Ptr(Cursor)) -> bool
|
|
let start = c.depth
|
|
t = next(c)
|
|
if not is-ok(c)
|
|
return false
|
|
if t.kind == tok-eof
|
|
fail(c, err-unexpected-token, t.pos)
|
|
return false
|
|
;; A scalar is one token and we are done. A closer here is a value ending
|
|
;; that never began, which pop-close has already reported.
|
|
if c.depth <= start
|
|
return true
|
|
while c.depth > start
|
|
let u = next(c)
|
|
if not is-ok(c)
|
|
return false
|
|
if u.kind == tok-eof
|
|
;; next already failed on the open depth; this is belt and braces.
|
|
fail(c, err-unbalanced, u.pos)
|
|
return false
|
|
true
|
|
|
|
;; ── Expecting a kind ────────────────────────────────────────────────
|
|
|
|
;; The shape a hand-written reader is built out of: take the next token, and if
|
|
;; it is not the kind wanted, fail the cursor at that token's position with a
|
|
;; reason. The returned token is the error token in that case, so a caller that
|
|
;; forgets to test is-ok still does not read a value out of the wrong kind —
|
|
;; int-of and friends answer None for tok-error.
|
|
fn expect(c: Ptr(Cursor), kind: i32) -> Token
|
|
let t = next(c)
|
|
if is-ok(c) and t.kind != kind
|
|
fail(c, err-unexpected-token, t.pos)
|
|
return error-token(c)
|
|
t
|