Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Trellis: Design Document

Trellis is the project, the specification layer, and the IDE. Soil is the target language it lowers to. The name reflects what the tool does: vibes grow, the trellis shapes them.

Status: design phase. This document records every decision made so far, the reasoning behind each, the items explicitly deferred, and the planned build order. It is intended to be the single reference for the project until the .tr grammar and lock schema documents supersede the relevant sections.


1. Vision

1.1 The core thesis

If an AI agent writes the implementations, the human-authored layer of a program should consist of specifications, types, tests, and structural constraints rather than code. The target language those implementations are written in should be designed to be easy for a machine to write and easy for a checker to verify, rather than pleasant for a human to type. Soil is that target language; Trellis is the specification layer and the tooling that lowers specifications into Soil.

Soil is, in effect, the assembly language of vibe-coding: a small, strict, verifiable language that humans read but rarely write.

1.2 The shape of a project

Both intended users write a slice. An individual writes a small program that calls foreign libraries. A team writes a small module inside a large foreign codebase. In both cases Trellis owns a bounded region of the program and treats everything outside that region as untrusted foreign code with a tested boundary. The scope pitch is: Trellis owns the region you care about being right; the rest of your stack is FFI.

Trellis remains a vibe-coding language, with two tiers:

  • Unchecked vibing: an agent writes the .tr files and an agent lowers them. The human reads nothing. This is still strictly better than vibe-coding as commonly practiced, because every function has a type that checks and tests that pass, even if no human looked at them.
  • Disciplined vibing: the human writes the .tr files and an agent lowers them. The human’s attention goes entirely to specifying and judging; they never write a loop, a match, or a signature. They are forced to think about correctness, not about code.

Both tiers produce identical artifacts, so a project can move from unchecked to disciplined one function at a time: a human takes over a .tr file, reviews or rewrites the tests, and marks the definition accepted. This is the migration path from “prototype I vibed” to “thing I trust” without a rewrite. The “humans write tests” rule concerns what gates lowering in the disciplined tier, not what workflows are permitted.

1.3 Users

  • First user: the author, building small projects that use Python libraries.
  • Target users: both individuals writing small projects against libraries, and large teams migrating small slices of code within a larger codebase.
  • First users beyond the author: other individuals on their own small projects. This keeps team tooling deferred but pulls installation and the first-hour experience forward, since each user cannot be hand-held.

The rule adopted for resolving the tension between these two: decisions that affect format or semantics must be made now to accommodate both users; decisions that are tooling only are made in the simplest form for the first user and grown later.

1.4 Design tenet

Nothing that affects correctness exists only in an agent’s context. Signatures, human answers to agent questions, export pins, style examples, and every other input to a lowering lives in a hashed file. An agent’s “memory” of a function is its lock entry and its last lowering, nothing more.


2. The three components

2.1 Soil: the target language

A small, strict, refinement-typed ML with effect tracking, designed as an agent target and a glue language for tying other languages together.

2.2 Trellis: the specification layer

A file format and toolchain in which humans write prose, types (optionally), and tests, and an agent lowers each definition to Soil under the supervision of a type checker, a refinement checker, and the tests.

2.3 Trellis IDE and build system

A graph-oriented IDE over the definition graph, with an interactive lowering interface, and a build system that compiles a TOML declaration of dependencies down to Nix.


3. Soil language design

3.1 Surface language

AspectDecisionReasoning
StyleDirect styleCPS as a user-visible language is a liability for both agents and human readers. CPS/ANF is used as an intermediate representation only.
TypingStrict, ML-style type inference, with refinement typesML inference is kept because it reduces the surface area for agent error. Refinements add verifiable claims without requiring full dependent types.
EvaluationStrictSimpler for effects, backends, and agent reasoning about performance. Follows Koka and Idris 2.
DataSum types, product types, type aliases/typedefsStandard ML data modelling.
Pattern matchingYesStandard.
Minimalism“There is only one way to do something”Less surface for the agent to hallucinate, more for the checker to catch.
Spec sizeThe core language reference carries a CI-enforced token budgetAdopted 2026-08-23 (survey §4.4, after Mog): every language addition must pay for itself against a fixed budget (target: core reference under ~8k tokens), which operationalizes “as simple as possible” into a measurable gate and keeps the whole-spec-in-context lowering strategy viable.
Intermediate representationANF preferred over CPS for the mid-endEasier to optimize for register machines and the JVM; easier for an agent to read when debugging lowering failures.

3.2 Refinement types, not full dependent types

Full dependent types (Agda/Idris style) were considered and rejected for v1. They are hard to infer, would discard ML inference, and agents write proofs poorly. Refinement types (Liquid Haskell / F* style) keep inference, let the agent write specifications rather than proofs, and discharge obligations through an SMT solver. Full dependent types remain a possible future escape hatch for the small fraction of functions where SMT cannot help.

Where inference stops: types may mention terms only if those terms are total and fall in a decidable fragment (linear arithmetic, uninterpreted functions, lengths of lists and sets). Beyond that, the refinement is unproven and the function is demoted (see §6.4).

Refinements are erased at runtime in release builds only if they were proven (see §3.9).

Refinements may eliminate error cases. A function returning Result may carry a postcondition such as {r | is_ok r} under a precondition, allowing callers that can discharge the precondition to skip the match. This pushes obligations up the call chain, which is the intended ratchet behaviour: a caller that cannot prove the precondition simply matches on the Result as normal. The Err arm must still exist at runtime in unrefined/debug builds, because the refinement is erased.

3.3 Effects

Effects are tracked as effect rows on types, not as monads. Rows compose without the transformer-stacking problem, and row-polymorphic higher-order functions inherit the effects of their arguments automatically (so map over a total function is total).

The effect system is purely a tracking mechanism. There are no first-class algebraic effect handlers and no effect runtime. This keeps backends simple. The consequence accepted: no mock handlers for testing, which is addressed instead by the capability model (§3.5).

Effect lattice:

  • total: the empty row; the function provably terminates and has no effects.
  • div: may diverge.
  • panic: may crash at runtime (indexing, division, unproven refinements, FFI).
  • io: touches the world.
  • ffi: calls foreign code; implies panic.
  • User-declared algebraic effects are a possible future extension but not part of v1.

There is no exn effect and no exceptions. Result a e is the error channel — success type first, as in OCaml and Rust. Haskell’s error-first Either e a order exists so the partially applied constructor can be the Functor/Monad instance, a motivation that cannot arise in Soil (no type classes, no higher-kinded abstraction), so the more widely known order wins. Cases that cannot be expressed as Result and cannot be proven safe by refinement carry panic.

Function application annotations (the Trellis feature of declaring what a function may call) fall out of the effect system: f may call g iff g’s row is a subset of f’s row. The graph view of the IDE is therefore also an effect-flow diagram.

3.4 Totality

Totality is tracked as the absence of div in the effect row. Termination is established by an Idris/Agda-style termination checker: structural decrease on an argument, or a user-supplied measure (decreases n). This is preferred over Koka’s syntactic approach because an agent can usually supply the decreasing argument, and the annotation is cheap for both the agent to write and the checker to verify. If the checker fails, the function acquires div and the manifest may refuse it.

Only total functions may appear in type indices (refinements), because the type checker must normalize them. Totality is thus an effect at the term level and a gate at the type level.

Policy: total by default, div opt-in (the Idris policy), which is the right default for an agent-written language.

3.5 Capabilities

io is not ambient. A function that performs I/O receives an opaque capability value representing the permission, and the io effect in its row indicates that it uses it.

type Fs = opaque
type Net = opaque
type Clock = opaque

read_file : Fs -> Path -> io (Result String FsError)
now       : Clock -> io Time

World is the root capability, handed to main by the runtime (or by the host in embedded mode). Sub-capabilities are derived from it and cannot be constructed any other way because the types are opaque.

main : World -> io Unit
main w =
  let fs = world_fs w in
  ...

Testing: the prelude provides fake capabilities of the same type (fake_fs [{"key": "config.toml", "value": "..."}], taking a Map in the canonical value encoding), so a function under test cannot distinguish a real capability from a fake. The capability is the handler, passed by hand; no effect handlers are needed.

Capability set (confirmed): Fs, Net, Clock, Env, Proc, Rand, Py.

Effect row stays a bare io. Naming capabilities in the row (<io:fs,net>) was considered and rejected for v1 as the capability arguments already carry the same information.

FFI needs a capability too. A Python call takes Py. This makes FFI-calling functions visible, allows stubbed fakes for tests, and makes a prelude fork without a Py constructor into a Python-free sandbox.

Boilerplate accepted: the individual user writes read_file fs path rather than read_file path. A with fs sugar for implicit passing was considered and rejected, since implicit parameters are a form of type classes.

Rationale for deciding this now: the capability style must be present in every io signature in the prelude from the start. Retrofitting it would invalidate the corpus. A default “real world” capability keeps the individual user’s CLI experience unchanged.

3.6 Polymorphism

Parametric polymorphism only. No type classes, no functors, no overloading of any kind. This is accepted as painful for humans and acceptable for an agent target, and it fits “one way to do things.” Global coherence constraints of a class system would also conflict with the per-file incremental model.

Consequences accepted:

  • No overloaded numeric literals. There is no general-purpose Int or String type; see §3.12. Literals have a default type (1 is I64, 1.0 is F64, "..." is Utf8) unless annotated, as in Rust.
  • Maps take an explicit comparator (OCaml Map.Make as a plain higher-order function). This is the confirmed idiom and the prelude will show it.
  • Sort takes a key function (sort_by : (a -> k) -> List a -> List a), with structural compare on k. The Python idiom; agents know it.

3.7 Derived functions: eq, show, compare, hash

With no ad-hoc polymorphism there is exactly one possible eq per type, so all four are auto-derived for every type rather than opted into as in Rust.

Implementation: derived per type as new definitions (Foo::eq, etc. — :: is the namespace separator, . being reserved for field access), rather than as polymorphic primitives. The user’s reasoning was the ability to statically exclude closures and to give float types a specific treatment. (Both approaches are semantically equivalent given a kind restriction; per-type derivation was the chosen spelling.)

  • Functions: deriving eq on a type containing an arrow is a type error.
  • Floats: total order. NaN is equal to itself and sorts last (Rust’s total_cmp). Float therefore derives all four functions and can be a map key. Decided.
  • FFI handles and abstract types: pointer identity via the opaque strategy.
  • show is the JSON encoder and read/parse is the decoder. Expect tests compare on JSON. One value format for everything.
  • Refinements: erased at runtime; eq on {v:Int | v > 0} is eq on Int.
  • hash is FNV-1a 64-bit over the value’s canonical JSON encoding (the opaque strategy hashes the address instead). Fixed and documented because hash is language-observable and must be deterministic across platforms, runs, and toolchains; hashing the canonical bytes means “equal ⇒ same hash” follows from canonical encoding for free, and there is only one byte-form of a value in the system. Soil maps are comparator-ordered, not hash tables, so keyed/DoS-resistant hashing buys nothing. (Resolved 2026-08-22 with the soil-rt implementation plan.)
  • Large structures: structural eq is O(n); accepted.

No user override for now. A first-class override mechanism was discussed and recognized as type classes returning through the side door (coherence, hash/eq agreement, equivalence-relation guarantees). Instead, types may declare one of a small fixed menu of derivation strategies:

Strategyeqshowcompare/hash
structural (default)structuralJSONstructural
opaqueidentity"<handle>"on address
ignoredalways Trueomittedconstant

ignored is a per-field marker, not a per-type strategy, and must carry a default expression — total calls only, with the record’s other fields in scope — which refills the field wherever a value is materialized without it (JSON decode, py_to_soil, host stubs): cached_word_count : U64 ignored = word_count text. Because there is no mutation (§3.13), the default always reproduces the value the constructor stored, so omission from show is lossless. ignored thus means “excluded from derivation, reconstructible on demand.”

The cases a user override would have served are handled without it: case-insensitive strings via a newtype with normalization in the constructor; ignored cache fields via the ignored strategy; FFI handles via opaque. Real overrides, if ever needed, are understood to be the addition of a class system and are deferred indefinitely.

3.8 Recursion

Mutual recursion across Trellis definitions is forbidden. Content addressing becomes a tree, incremental checking is a topological walk, and every lowering has a well-defined “everything below me is already checked.”

What is lost: mutual algorithms (even/odd, recursive-descent parsers with expr/term, traversals of mutually recursive data). The workaround is the standard one: one function with sum-typed dispatch, or mutually recursive locals. Nothing is inexpressible; only decomposition into separately specified units is constrained. Parsers are expected to be the most painful case, and to hurt the agent more than the human because the mutual shape is the idiom it has seen most.

Within a single .soil file, let rec ... and ... is allowed freely. Mutual recursion among locals and module-private helpers is a Soil-level detail invisible to Trellis. This covers most parser cases: parse is one Trellis definition whose lowering contains mutually recursive locals.

Not a stdlib-only feature. The stdlib must be written in the same dialect it teaches, or the corpus shows idioms users cannot use.

Recursive types across definitions remain allowed. Types are definitions but are not lowered, so cycles among them do not break the lowering walk. They need a combined cycle hash in the lock; this is the one cycle the lock must represent.

3.9 Debug and release modes

First-class in Soil, not two separate lowerings. Two lowerings would mean two artifacts that can diverge and two hashes. Instead, one lowering carries refinements, invariant checks, panic guards, and test hooks as erasable annotations.

Release mode only erases what was proven. A demoted (unproven) refinement still runs its check in release, because tests are the trust root and an unproven claim must not become an unchecked one. Unverified code carries runtime checks forever, which is the correct incentive.

3.10 Runtime and embedding

The Soil runtime is a Rust crate exposing a C ABI, with a trivial main wrapper. Writing the runtime in Rust makes three things one codebase: embeddability as a C library, shared ownership with Rust batteries (Rust values are runtime-owned, so no FFI and no conversion), and the memory model (reference counting, Perceus-style reuse, cycle handling). The compiler may still emit C, native code, or JVM bytecode; this decision is about the runtime only.

Soil is embeddable as a C library. This is committed to from the first commit: no global state, explicit init/teardown, callable from C. It serves both users from the same compiled artifact:

  • Individual user: Soil is main, Python is a library it calls.
  • Migrating team: Python (or another host) is main and imports the Soil module.

Writing the runtime as a program and libraryizing it later would be a rewrite.

3.11 Backends and FFI

  • v1 backend: Cranelift. The backend proper is a pure pass emitting CLIF text (golden-testable like every compiler pass); a small Rust driver feeds Cranelift for instruction selection, register allocation, and object emission — x86-64 and arm64 from one backend, no C toolchain dependency. C emission, JVM, and direct x86 were considered and deferred: direct emission means writing the two largest, least-testable compiler phases first, and C is a semantically messy target that drags in an external toolchain. See docs/bootstrap-plan.md §4.
  • v1 FFI: both SysV/C ABI and Python. They are different kinds of work and share no duplicated code. The C ABI is a layout and calling-convention problem in the compiler backend and is close to free once the runtime is Rust (Rust extern "C" functions are the native case). Python is a marshalling and lifecycle problem: CPython embedded from the Rust runtime (via pyo3), GIL held around calls, PyObject* wrapped as an opaque refcounted handle, py_to_soil/soil_to_py defined over the same JSON-shaped value model as show, tests, and host stubs (one value model, never two). Sequencing: C ABI → Rust batteries → Python embedding → Python batteries, each usable before the next. Rust is for the stdlib; Python is for libraries. Node comes later.
  • Hand-written bindings only in v1 for both FFIs. “Hand-written” means a specific binding exists because someone asked for it, one at a time, with a spec; it does not mean a human types it. A binding is three artifacts: a .tr spec (prose, Soil signature with effects and capabilities, auto-generated contract tests per §4.5); a shim (a Rust extern "C" function, or a pyo3 function across the GIL); and a lock entry with trust level declared-only, contract, or harvested. The corpus includes one worked example of each shim kind so the agent writes them inside a normal lowering.
  • The whole-package binding generator is deferred, likely permanently. Reading a crate or package and emitting its full surface is a separate project with an unbounded difficulty profile: every foreign type system is a new translator (Rust lifetimes, traits, generics; Python’s effectively untyped .pyi stubs; C pointers and ownership); effects and capabilities are invisible in foreign signatures, so the generator either assigns everything io + ffi + panic with World (defeating the capability system) or guesses; refinements are entirely absent, so the valuable annotation is still manual; the surface is unbounded and most of it is never called, flooding the lock with untrusted entries (the same trust hole closed by rejecting helpers.tr); and subtly wrong generated bindings are memory-safety bugs surfacing far from their cause. Hand-written bindings grow the batteries layers by demand, are mostly agent-written anyway, and surface which foreign types are actually hard to map one at a time. The thing that is built, in v1, is a per-symbol binding assistant: trellis bind requests.get or trellis bind regex::Regex::new. This is vital to the first project, which is the kind of program that would otherwise have been vibed in pure Python and will call a dozen functions from a few packages; each needs a binding before any lowering that uses it can proceed, and a dozen hand-written bindings is a week of friction at the exact moment the tool’s pleasantness is being tested. The assistant is not the bulk generator and has none of its problems, because it is a lowering job with a different context bundle: the daemon fetches that one symbol’s metadata (.pyi stub, inspect.signature, docstring for Python; cargo doc JSON for Rust) into symbol.json and docstring.md, adds one corpus shim example, and runs the normal lowering loop with the normal MCP tools. The output is a .tr with prose summarized from the docstring, an agent-proposed signature with effects and capabilities, the shim, and auto-generated contract tests; the human reviews and accepts it like any other lowering. At single-symbol scale the type-mapping, effect-guessing, and refinement problems each have a human in the loop who corrects in one click, which is fine for twelve symbols and not for twelve thousand. No new agent loop, trust path, or file format. The bulk importer (trellis bind regex) is never built.
  • Embedding direction in v1 is Soil-hosts-Python. The py_module build target and generated .pyi (Python-hosts-Soil) are small but post-v1.
  • Generated bindings (reading documentation and types to produce FFI interfaces) will be constrained to sources with machine-readable types: C headers, .pyi, .d.ts. This feature is expected to be the largest bug source.
  • Primitive types: strings, integers, floats in all the variants the FFI targets need, with the stdlib providing interop. Defining these with the FFI in mind from day one is a known requirement; the specific technical decisions are deferred (§10).
  • Host stub erasure rule (confirmed): an exported Soil function with effects io, panic becomes a host-language function that may raise SoilError; capabilities become host-side objects passed in. The generated .pyi/.d.ts stub documents the effect row. Decided once, applied to every host language. In debug mode SoilError is structured and carries the trace, the demoted refinement if any, and the JSON inputs, so host test suites can locate Soil bugs.
  • Foreign values are opaque handles and carry no refinements. A Py handle stays a handle; refinements could be invalidated by foreign mutation, so they attach only to Soil values. Converting a handle to a Soil value is an explicit, visible call (py_to_soil) whose cost the caller chooses to pay. Big structures stay in Python and are manipulated by handle-in, handle-out batteries functions; small results cross the boundary and may be refined after conversion. Rust-backed batteries differ: Rust structures are owned by Soil’s runtime and are therefore ordinary Soil values, refinable and potentially total, with no conversion and no capability (capabilities are about the world, not the implementing language; a Rust function that does no io takes none).
  • The JSON value model must suffice for remote marshalling (constraint adopted 2026-08-23, survey §7). soil_to_py/py_to_soil are defined over the same JSON-shaped value model as show and tests; nothing non-serializable may sneak into the FFI boundary, because the deferred co-process isolation mode (§10) — the Python interpreter in a child process, calls marshalled over a pipe, handles as remote references — must be implementable as a deployment mode of the existing marshalling layer, not a second one.

3.12 Primitive types

There is no blessed general-purpose integer or string type. The types that exist are those with unambiguous semantics, so that backends, FFI, and refinements all agree:

  • Fixed-width integers: I64, U64, I32, U32, etc. Overflow is panic, or is refinement-checked away. Integer / and % are floor division and floor modulus (Python’s semantics, matching the reference-implementation language so differential tests agree without adjustment), not C/Rust truncation; a zero divisor is panic or refinement-checked away like overflow.
  • BigInt: arbitrary precision, implemented in the pure core prelude (not foreign-backed, so it remains total). SMT reasons about unbounded integers natively.
  • F64: total-ordered (§3.7).
  • Utf8: validated UTF-8 bytes, no O(1) indexing. Bytes for raw data.

A default literal type exists for ergonomics (1 is I64, "..." is Utf8), which is the only concession. The pressure to make I64 and Utf8 feel general-purpose falls on the prelude, not the language. Unit has no literal (there is no () in the grammar); the kernel/prelude value unit : Unit is the one way to produce it (resolved 2026-08-22 with the soil0 CLI contract, which pre-registers it as a builtin value). The Python batteries layer (§5) uses BigInt at its boundary because that is what Python numbers are.

3.13 Mutation and records

No mutation. Every function body is a term; there is no st effect and refinements are sound without an aliasing story. In-place update is recovered by the backend where possible (Perceus-style reuse analysis, as in Koka) without the language being aware.

Records only. Product types are records with named fields; there are no positional tuples and no named arguments. This is what JSON wants and what agents read best.

3.14 Module exports and types

Export lists name functions and types explicitly. If an exported function’s signature references a type that is not exported, it is an error the agent must repair, with two permitted repairs: export the type, or mark it as intentionally abstract (callers may hold values of it but not inspect them). Abstract types are therefore a deliberate feature rather than an accident.

main is an ordinary definition with effect io and a World argument; it has a .tr, tests (cram only), and a lock. soil.toml names it as the entrypoint and it receives no other special treatment.

3.15 Namespacing and name resolution

Constructors live in their type’s namespace and qualify as Type::Ctor (:: being the namespace separator, §3.7). Resolution has exactly one legal spelling per context: a bare constructor (Ok, None) is legal iff its variant name is unique among the sum types in scope — and is then the only legal form; when two types in scope share a variant name, Type::Ctor is required. Qualifying a unique constructor is an error. The rule applies identically in expressions, patterns, and the predicate language’s is tests.

Definition names resolve by the same rule (resolved 2026-08-22, surfaced by the soil0 renamer). Soil has no imports — names resolve through the manifest (§6.1) — and a bare definition name is legal iff it is unique across the visible definition set plus the builtins, making the prelude callable bare from everywhere; on a collision between modules, module::def is required, and qualifying a unique name is an error. Same-module-only bare resolution was rejected because the prelude would need an always-bare special case (a second way); Rust-style optional qualification was rejected as before (two spellings for every unique name). Private definitions are visible only within their own module and are never qualifiable.

3.16 Typed holes

Adopted 2026-08-23 (survey §4.3, after Tacit/Thermite — the Idris hole workflow §12 already cites, confirmed to work for agents specifically). ?name is an expression of any type; a definition containing holes checks, with the checker reporting each hole’s goal type, so the agent can lower a hard function outside-in — structure first, holes for the hard cases, green types at every step — and a failed attempt at hole 3 does not discard holes 1–2. A blocked lowering pauses in a principled partial(holes: n) lock state instead of all-or-nothing failure, and ask_human can point at a specific hole and its goal type rather than a prose description of being stuck.

The rule that keeps the trust model intact: a definition with holes can never be tested, accepted, or built into a release target — holes are a lowering-time state, visible in the lock, never in an artifact; run/test refuse them. Grammar in soil-syntax-spec §3.3/§5.11; soil0 CLI contract v1.1.

3.17 Literal provenance

Adopted 2026-08-23 (survey §6.2, after Vera’s <DB> literal-provenance rule that makes SQL injection a check error). Literal a is a provenance fact tracked syntactically by the checker, not the solver: string literals carry it, concatenation of literals preserves it, nothing else does. Prelude and batteries signatures for injection-shaped boundaries demand it — proc_run command text, SQL query text, a future Py::eval — with runtime values passed separately as parameters. The escape hatch is a human-only trust_literal cast, listed in the manifest with the other escape hatches. This turns the most likely dangerous agent error in glue code into a checker error at zero solver cost. Enforcement lands when the first boundary needing it does (plan 06: Proc/Py); batteries signatures are written provenance-aware from the start.

Rationale: “there is only one way to do something” (§3.1). The Rust rule (bare when unique, qualified always allowed) was rejected because it leaves every unique constructor with two legal spellings; always-qualifying was rejected as a permanent verbosity tax on the code humans read most; Type.Ctor was rejected because . is reserved for field access (§3.7). Cost accepted: adding a colliding type to a definition’s context changes the required spelling in that definition — a visible change that content addressing surfaces as an ordinary re-check. (Surfaced by impl plan 02; resolved 2026-08-22. Grammar in soil-syntax-spec §3.3–§3.4, static rule §5.9, tr-grammar §2.3.)


4. Trellis: the specification layer

4.1 Trust model

RoleOwns
HumanProse, tests, escape hatches, export pins, the “accepted” status
AgentType signatures, refinement annotations, proofs, Soil bodies, module-private helpers
TestsThe root of trust for correctness
Refinement checkerA ratchet: proves the lowering satisfies the agent’s own claims; catches internal inconsistency (off-by-one, missed cases) but does not establish intent
Reference implementationAn executable spec and differential-testing oracle

Key consequences:

  • A passing refinement check proves the body satisfies the agent’s claim, not the human’s intent. The IDE shows “typed” and “tested” as separate badges; only “tested” means correct.
  • Tests are written by humans. If someone wants to truly vibe-code, they use their agent to generate Trellis files; the Trellis layer itself does not generate the gating tests. Agent-generated tests may exist as a clearly labelled differential tier run against the reference implementation, but they do not gate lowering.
  • A cheap tightening: auto-generate property tests from refinements ({v | v > 0} becomes a QuickCheck property), so refinements are cross-checked against the trust root.
  • Escape hatches (unsafe, partial, raw FFI) are written only by the human, in the spec, and every escape hatch in the project is listed in the manifest so the audit view shows where trust is concentrated.

4.2 The unit: one file per definition

A definition is a function, a type, or a module header. Each lives in its own file.

Reasoning: the unit of lowering, checking, testing, and locking is the function; file = definition aligns every per-function artifact, makes git diffs map to semantic changes, eliminates merge conflicts between people editing different definitions, makes content addressing trivial (hash the file), and naturally bounds the agent’s context.

Costs accepted: external tooling (grep, git log, GitHub review) degrades with many small files, mitigated by a “module view” in the IDE that renders a directory as one virtual file.

Helpers are not Trellis definitions. A helpers.tr magic file with reduced requirements was considered and rejected: it would be a hole in the trust model that grows until all real logic lives there. The friction of requiring a spec and tests for anything named at the Trellis level is the correct friction; if a helper is not worth specifying, it is not Trellis’s business.

Helpers live at the Soil layer in two forms:

  1. Local let bindings inside a lowering, hashed as part of the parent, invisible to Trellis. Covers most cases.
  2. Module-private Soil definitions (_private.soil) when the same helper is needed across several lowerings in a module. Rules:
    • Cannot be called from outside the module; cannot appear in any Trellis signature or refinement.
    • Owned by the lowerings that use them; garbage-collected when no caller references them. The agent does not accumulate a private standard library.
    • The lowering skill prefers local lets and promotes to a private helper only to avoid duplication.
    • Their hashes feed into their callers’ soil_hash.
    • The global manifest lists them as nodes flagged soil-private with no spec hash; the IDE greys them out. A module with many private helpers and few definitions is a smell the IDE flags.

Promotion path: the IDE offers “promote to Trellis definition,” which stubs a spec file with the inferred signature and existing Soil as the initial lowering, and requires prose and tests before the lock entry is valid. A helper crosses the boundary only by acquiring a spec, never by exemption.

4.3 File format

  • Markdown with fenced code blocks. Prose is freeform Markdown; formal parts (signature, tests, calls) are fenced blocks with designated languages.
  • Signature: a combination of English prose and Soil types, as the human prefers. Since the agent owns types, the human may write no formal signature at all.
  • Minimal valid definition: frontmatter naming the definition, one sentence of prose, and one expect test. This is the onboarding story.
  • Filename is identity. Filenames are static; renaming a file without updating every reference is an error, and the IDE provides refactor-rename. The lock stores the filename; hashes are for invalidation, not identity. No separate name table is needed. Filenames are lowercase snake_case for every file kind, and every file carries YAML frontmatter with a required name: a function’s equals the filename stem, a type’s is PascalCase with the filename its snake_case form (the one formal record of casing, so identity survives case-insensitive filesystems), a module’s equals its directory. File kind itself stays inferred, never declared. Frontmatter also admits optional tags, drawn from a vocabulary declared in soil.toml: non-semantic, user-extensible metadata for IDE graph filtering and CI policy, hashed under prose_hash and never an input to lowering.
  • Tests are named blocks, and a file may contain any number. Names are used by the lock to report failures, by the IDE for click-to-run, and by REPL-to-test promotion to know where to append. Expect tests are call-arrow lines (("1,2") => {"tag": "Ok", "value": [1, 2]}): the function under test is implicit (file = definition), with lines bind fake capabilities with pinned seeds, panic is a legal outcome only under a panic row, and xfail is an info-string modifier.
  • Refinements are written in a prose-friendly, human-readable form that is still machine-readable. Decided: separate requires/ensures blocks of label: predicate clauses; the signature block stays plain Soil types. Labels are what the lock and checker errors pin; the shared predicate language (also used by type invariants and property tests) is restricted to total calls in the decidable fragment. ensures on a Result uses result is Ok(v) implies ….
  • There is no calls annotation. The original “function application annotation” idea is fully subsumed by effect rows and capabilities: a function without a Net argument cannot reach the network regardless of what it calls. Call edges are tracked by the lock for the graph view but are not human-written.
  • Values: JSON, with a block drag-and-drop UI in the IDE for constructing them. Every Trellis type is round-trippable through JSON; this is also the derivation mechanism for show/eq (§3.7). Sum types are internally tagged ({"tag": "Ok", "value": …}; nullary variants {"tag": "None"}): one uniform shape for every variant, self-describing for hosts and generic tooling. The verbosity is accepted because the IDE’s widgets, not humans, write and read these values. Decoding is always type-directed; show output is canonical (declaration-order fields, comparator-order maps, shortest round-trip floats) so expect tests compare on the string. Full encoding table in the grammar prototype.
  • Type definitions are definitions: prose plus a shape (or agent-inferred shape) plus optional invariants expressed as refinements on aliases, which are checked as properties every constructor must preserve. Confirmed; whether invariants are checked on every constructor call or only proven at definition sites is an open detail (§9).
  • Decision blocks (adopted 2026-08-23, survey §8.2, after Aver): decisions blocks in _module.tr and a project-level _project.tr record structured chosen/rejected entries (“all timestamps UTC”, “comparator maps keyed by user id”), hashed with the spec and included in every context bundle in scope — closing a real gap in the §1.4 tenet, where project-wide rules lived nowhere hashable and each lowering either rediscovered or violated them. The interactive lowering UI writes rule-shaped answers back to decisions rather than into one function’s prose; the IDE can query them. Grammar in tr-grammar §5.2. Invalidation resolved 2026-08-24: decisions are their own per-entry hash class, invalidated by reliance edges — every lowering cites the decisions it applied and the lock records them like call edges (lock-schema §3), so an edited entry flags only the lowerings that relied on it, while an added entry flags the scope once via a membership hash. The human may reclassify an edit as editorial (typo/wording — nothing flags; the human already owns the spec and accepted, so trusting them to say “no meaning changed” is inside the trust model), and a batched triage sweep agent-clears remaining flags by confirming the existing Soil against the entry’s diff — tokens, never a re-lowering. Rejected: prose-class semantics (scope-wide review-suggested on every edit breeds fatigue until the flags are ignored) and formal-class semantics (a typo re-lowers the world) — both scale with project size where reliance scales with actual use. Citation is self-reported, so an uncited-but-influential decision under-flags — the drift class the prose tier already accepts. Details in tr-grammar §5.2, lock-schema §3/§8.
  • The .tr grammar is prototyped in docs/tr-grammar.md with worked examples in examples/. The reserved block languages are soil-sig, requires, ensures, test, property, cram, reference, allow, soil-type, invariant, exports; fenced blocks in any other language are prose. Finalization into docs/ is pending (§11).

4.4 Module structure

  • Folder = module. _module.tr holds module-level prose and the export list.
  • Export list is an explicit list of Trellis files (functions), exactly. Easy to hash. The export list is the FFI surface; host stubs are generated from it.
  • Export pinning (proposed, unenforced in v1): because exported signatures are agent-authored, the public API of a module is agent-authored. The proposal is that exported signatures require human approval via a pinned flag in the lock, after which the agent cannot change them without a Trellis error. Same principle as “humans write the tests.” Enforcement is tooling and is deferred; the flag exists in the lock format from the start.

4.5 Tests

Tiers, expressed as a lattice the manifest can describe:

expect (cheap, always run) → property/quickcheck/fuzz → differential against reference → proof.

  • The user declares which tier each function must reach; the lowering agent escalates automatically when a cheaper tier is green.
  • Test budget is inferred from the effect row: pure functions are fuzzed hard; io functions get expect/cram tests only. Per-function override available.
  • Not every tier is equally English-describable: expect tests and properties translate well from prose; fuzz harness configuration is just code.
  • Reference implementation / validator in Python, or a CLI oracle: a JSON-in/JSON-out executable invoked as a black box and hashed like any oracle (amended for the compiler bootstrap, where soil0’s passes are the oracles — see docs/bootstrap-plan.md). Always attached explicitly by the human, never auto-detected. Differential tests call it through the Python FFI with a Py capability in the test harness. serves as a differential-testing oracle, an executable spec the agent reads when prose is ambiguous, and a migration path (an existing Python codebase is the reference spec, and Trellis becomes a verified port tool).
  • Test dependencies: tests may reference the prelude, the reference implementation, the function under test, and other user definitions only if those definitions are accepted. Acceptance is the human’s trust signal, so accepted definitions are legitimate oracles, and this creates test-level edges only to frozen work. The lock tracks the edge; if the oracle’s spec changes, dependent tests re-run. Unaccepted definitions cannot be oracles, which prevents oracle cycles among unfinished work. (v1 status, resolved 2026-08-28: the daemon records reference and CLI oracle rows only; definition-oracle rows and this acceptance rule’s enforcement wait for the prelude (plan 04), which absorbs the helper vocabulary predicates lean on today — see lock-schema §5.)
  • xfail marker: a test may be marked expected-to-fail to document a known limitation. It stays in the spec, is shown in the lock, and blocks accepted until resolved, so the disciplined alternative to deleting a test exists.
  • Non-deterministic functions: Rand and Clock fakes take seeds and timestamps; the IDE’s test widgets expose these as fields so every such test is pinned by construction.
  • Contradiction pre-flight: before spending any tokens, the daemon checks tests against each other and against the prose for mechanical contradictions (same input, different expected output). This is a distinguished check; subtler contradictions become a distinguished class of ask_human question.
  • The kind of test a definition needs depends on what it is. Every definition has tests, but not the same tests: pure functions get expect and property tests; io functions get fake-capability tests; FFI bindings get contract tests (below). The invariant “nothing in the lock is untested” holds throughout.
  • FFI bindings get auto-generated contract tests, not human-written behavioural tests. A binding’s spec is “faithfully cross the boundary,” and that is checkable without understanding the library. The daemon derives contract tests from the signature and effect row: a call with an obvious valid input yields Ok; an invalid input yields Err, not panic; returned handles are accepted by the sibling bindings for that type; a loop of calls under the debug runtime’s leak checker shows no growth; Soil values survive soil_to_py/py_to_soil unchanged. The human supplies prose and at most one example input. The lock records tests as contract rather than expect, so the trust level stays visible. Behavioural correctness of foreign code is not the binding’s job; it is caught one level up by the user’s own tested functions, which is the same place a hand-written wrapper around a C library would fail.
  • Harvested tests remain optional and strictly better when the foreign library has examples or a test suite worth translating. Trust level harvested vs contract is recorded in the lock.
  • Refined bindings need one real test per refinement. A refinement on a binding (e.g. {p | valid_regex p} -> total Regex) does work for downstream proofs that contract tests do not exercise, so each refinement requires one human-written counterexample. Most bindings carry no refinements.
  • io testing: fake capabilities (§3.5) are the primary mode. Cram tests against real side effects are the fallback for the individual user and for the FFI boundary, where fakes stop being possible. A cram transcript runs in a fresh temp dir with with file fixtures and may invoke built binary targets or trellis call <def> <json-args> (any io definition, real World-derived capabilities, canonical JSON out) — the real-mode escape for non-main functions. Cram never runs inside the lowering sandbox (§4.6). The lock tags a function’s io tests with their mode so additional modes can be added later without a format change.
  • Test strength is measured, not assumed (adopted 2026-08-23, survey §3.1, after Thermite/Vow): a post-v1 trellis mutants daemon job mutates the generated Soil (swapped comparisons, off-by-one constants, dropped match arms) and reports surviving mutants in the lock as a test-strength score. This is how tests are audited without reading implementations — exactly the position Trellis puts the human in. The IDE shows each survivor as a concrete “your tests don’t catch this” example, one click from becoming an expect test; soil.toml CI policy may require a mutation score for accepted. The lock schema reserves the fields now.
  • The contradiction pre-flight gains vacuity probes (adopted 2026-08-23, survey §3.2): a precondition no input satisfies, a postcondition implied by true, and an expect-test set that never exercises a declared variant are each flagged before tokens are spent — closing the failure mode of an agent satisfying “write a refinement” with a claim that constrains nothing.
  • The test budget gains a hostile tier (adopted 2026-08-23, survey §3.3, after Aver): property tests biased to boundary values (empty lists, integer extremes, NaN, the largest value a refinement permits) and failing capability fakes — an Fs whose reads error mid-stream, a Clock that jumps backward. Because capabilities are explicit arguments, hostile fakes are ordinary prelude values (plan 04); no mechanism needed.

4.6 Lowering

Input context per lowering is a fixed directory layout (the context bundle): spec.md, callees/ (signatures only, never bodies), tests.json, examples/ (prelude), reference.py, and previous.soil if re-lowering. The agent reads it through a tool; humans can inspect it. Unverified callees are shown as their base type plus a note that the refinement is demoted, so the agent cannot rely on an unproven claim.

Lowerings run strictly serially and in disciplined order: a definition cannot lower until all its callees have. There is no separate signature-inference step; a definition with no lowering has no checkable signature, and callers wait. The dependency tree provides the queue order, the IDE shows what is blocking what, and the IDE supports queuing lowerings while other definitions are still being written. If a human edits a definition that a queued or running lowering depends on, the daemon invalidates that job.

Fresh context per lowering. Every lowering is its own session receiving a fresh context built from the current (human-edited) files. More expensive in tokens; guaranteed to be correct and free of stale memory. A separate “memory cache” of agent observations was considered and rejected under the design tenet: anything the lowerer learns that would help next time either belongs in the prose, the tests, the module header, or the prelude fork, or it is not an input. The one exception allowed: a per-lowering log (f.log, gitignored) of what was tried and why it failed, for the human to read only; never an input to the next session.

Lowerer implementation. Each lowering is one headless agent invocation, which implements the fresh-context rule via the process boundary. The daemon (§4.9) assembles the context bundle and invokes the agent with a prompt to lower it. The daemon is an MCP server, and the agent is allow-listed to exactly its tools: read_context, check_types, check_refinements, run_tests, write_soil, read_spec (added 2026-08-25 with the daemon-contract review: diagnostics carry spec_refs, so the agent dereferences pinned spec sections directly — served from the toolchain’s embedded copies, never the working tree), ask_human. No raw shell, no filesystem outside the scratch directory. Consequences:

  • The sandbox is the tool allow-list. run_tests runs in the daemon’s sandbox with fake capabilities; real-world tests are not a tool the lowerer has.
  • The check loop lives inside one invocation. The agent writes Soil, checks, reads structured errors, revises, tests. The daemon caps turns, time, and cost rather than implementing retries.
  • Every tool call is logged by the daemon, which is f.log and the cost telemetry with no instrumentation of the agent.

The question channel ends the invocation. ask_human writes a structured question and exits. The daemon surfaces it in the IDE, the human answers, the daemon writes the answer into the .tr prose (a ## Clarifications Q/A section by default, tr-grammar §8, resolved 2026-08-24; the IDE later offers an agent-performed fold-into-prose; rule-shaped answers go to decisions blocks), and re-invokes the lowerer fresh. The agent never sees an answer that is not already in the spec, and every question/answer pair is a visible diff. A blocking in-context variant may be added later as a cost optimization.

Provider abstraction. One interface, lower(bundle, tools, budget) -> outcome, with two provider kinds: agent-CLI providers (Claude Code headless / Agent SDK, and other headless agent CLIs; nearly free to build; runs on subscription quotas) and a raw-API provider (own loop against a model API; full control; pay-as-you-go). The IDE supports multiple agents with user-supplied credentials. Agent-CLI is v1; raw-API is v2. Running the lowerer through Claude Code is a high-priority feature because it is how subscription users avoid paying twice, though subscription quota policy for headless use has shifted recently and should be re-verified. Daemon responsibilities from day one: per-invocation timeouts, turn caps, cost caps, and isolated agent home directories.

Per-function granularity, manifest/type/lint checked at each step. Failures are local and retryable.

No widening. If lowering f reveals that g’s signature is wrong, the agent does not widen g; it is a Trellis error for the human.

Agent write-back: the agent writes its inferred signature back into the .tr file. How agent-authored parts are marked, and what happens when a human edits them (presumably: they become pinned), is an open format question (§9).

Agent questions before lowering: the agent may raise an ambiguity as a blocking state (“should parse accept trailing whitespace?”). The human’s answer is written back into the prose. Ambiguity becomes spec improvement rather than silent guessing. Implied by the interactive UI decision; not separately confirmed.

Errors for the agent as a first-class audience: the LSP has two output modes, human and agent; the agent mode is structured (JSON) with concrete counterexample, violated spec clause, failing test, and a suggested repair class. Codes and repair classes are a drift-gated registry (adopted 2026-08-23, survey §4.1, after Vera/Zero): every diagnostic carries a stable code from a registry living in the toolchain source, each code mapping to a typed repair class the agent acts on mechanically (add-decreases, widen-match, insert-guard, export-type-or-mark-abstract, …) plus a spec_ref into the sectioned Soil spec, and CI fails if registry, docs, and emitted diagnostics disagree. The soil0 CLI contract’s error-code registry is the first instance; the repair-class layer is the daemon’s (plan 03).

The context-bundle assembler is a budgeted, prioritized packer (adopted 2026-08-23, survey §4.4, after Tacit/Aver): trellis context <def> --budget <n> packs spec > tests > callee signatures > nearest corpus examples > module prose, with the budget per model recorded in the provider config — resolving the bundle-sizing question by making the budget explicit and the priority fixed. The Soil spec itself is written in pinned, individually addressable sections that spec_ref points into and a daemon tool serves; the whole spec is never shipped blind.

The lowering skill is generated, never hand-maintained (adopted 2026-08-23, survey §4.5, after Vow’s compiler-emitted skill): trellis skill assembles the lowering skill from the compiler’s own registries — error codes, repair classes, the effect lattice, derivation strategies, primitive types — plus the pinned prelude corpus, and CI fails if a committed copy drifts from the toolchain that emitted it. The skill is a build artifact versioned with the toolchain, which also serves non-Claude harnesses with one bundle.

Interactive lowering UI: the user interacts with lowering errors and reports through a prompt or UI. Rule: anything the human says to the lowerer that changes the outcome must be persisted to the .tr file, or the lowerer refuses to act on it. An answer that is really a project rule goes into a decisions block (§4.3), not into one function’s prose.

Model selection: user-customizable, with “auto” functionality that chooses when no model is selected. Auto keys on spec size, effect row, presence of refinements, and number of past lowering attempts; escalates on retry. A per-project cost ceiling is planned because auto-with-escalation is exactly the setting where one pathological function burns a budget. All of this is tooling and ships minimal for v1 (one model, fixed retry count).

Style and the prelude corpus: see §5.

4.7 “Done”

Done is the user’s decision. The lowering’s current status (tests, proofs, demotions) is presented transparently and the user decides when they are happy. The lock status vocabulary is typed, tested, verified, accepted; only accepted is set by a human. CI policy (“every exported definition must be accepted”) is a line in soil.toml, not a semantic rule.

4.8 Generated Soil

  • Checked in. Agent output is not reproducible, so the Soil is the artifact and the agent is a code generator like protoc; regeneration is a deliberate step.
  • Editable by humans. People will do it anyway. The lock records hand-edited and skips re-lowering until the spec changes.
  • Round-trip check: handled by the spec rather than by diffing prose summaries. The checked type of the generated function must entail the spec type from the manifest (a mechanical subtyping/entailment check). Prose drift is handled separately via hashing (§6.2).
  • Exactly one canonical text form (adopted 2026-08-23, survey §5.1, after Vow/Tacit): the printer is a compiler pass with parse → print → parse idempotence as a conformance test, and the lowerer’s write_soil canonicalizes on write — the agent never controls formatting, so every checked-in diff is semantic and soil_hash is meaningful. Rule pinned now; the formatting spec and soil0 print land as soil0 impl step 11, before the first hash ships (plan 03).

4.9 The daemon

Because the IDE is the primary product (§7.1), the lowerer, checker, REPL, and runtime are services rather than batch commands. A long-running trellis daemon holds incremental compiler state, exposes the LSP, runs lowerings as jobs with the question channel, serves the REPL, and exposes the MCP tool surface used by the lowering agent. The CLI and the IDE are both thin clients; CI support later is “run the daemon headless.”

REPL: nearly free given incremental compilation, since the daemon already holds every definition compiled. It works at the Trellis level (call a definition with JSON) as well as the Soil level; the former feeds REPL-to-expect-test promotion.

Debug mode: debug builds instrument every Trellis definition boundary (not every Soil function) with entry/exit, JSON arguments and return, timing, and which refinement checks fired. One run yields a trace convertible to expect tests by clicking a call, a flame graph at definition granularity, and a repro for any panic with exact inputs. Slowness is treated as a bug; the spec-granularity flame graph lets the human see it without reading Soil. Run logs are labelled sandboxed (from lowering) or real (from the human), and the IDE shows which one a proposed test came from.


5. The prelude as a trusted corpus

The standard library is a Trellis project whose lock file is trusted. Each entry is a spec plus a blessed, human-verified lowering. This serves simultaneously as:

  • The stdlib.
  • The few-shot example corpus that defines the style of Soil the agent writes. Examples are real checked code and cannot drift from the language.
  • The mechanism for forks with different capabilities: a no-io fork is a sandboxed profile; a fork without Py is Python-free. Forking the stdlib is forking a repo; no new mechanism.
  • The governance mechanism: users contribute by promoting their own definitions; style is a PR-review question rather than a prompt-engineering one.
  • The home for shared oracles (§4.5).

Structure: a pure core plus a foreign-backed batteries layer. The core prelude is written in Soil and is pure and total where possible: Option, Result, List, Map (with comparator), BigInt, Utf8, Bytes, JSON encode/decode, the derived-function primitives, the capabilities and their fakes, and Py. This is small (a few thousand lines) and is what teaches the agent what Soil looks like. A separate batteries layer provides refinement-typed Soil signatures over foreign stdlib functions. soil-rs-std is built first: Rust-backed functions are runtime-owned, can be total, and are refinable, so the batteries layer teaches the agent sound idioms. soil-py-std follows in v1 as the library layer, with handle-in/handle-out style, ffi/panic, and the Py capability. Python-backed functions must not be the core or the first batteries, or nothing would be total and the corpus would teach “call out” as the idiom. BigInt may be Rust-backed (num-bigint) and still count as core, since runtime-owned values are ordinary.

Trust is by full hash per package, pinned in soil.toml: the core prelude and each batteries package are separate hashes, and the lock holds the list of trusted packages. A fork is a different hash; partial trust of a package is not possible. Batteries functions come in two visible flavours: handle-in/handle-out (cheap, unrefined) and handle-in/Soil-out (converts, refinable on the result).

Corpus retrieval: for v1 the whole prelude fits in context. At scale, retrieval by signature similarity and effect row is the obvious mechanism, and the lock file is already most of the index. Deferred.

The prelude must be written in the same dialect users are permitted to use (no stdlib-only features), and must use capability-style io signatures from the start.

First prelude definition to write: read_file, since it exercises capabilities, effects, Result, and the FFI boundary at once.


6. Hashing, locking, and incrementality

6.1 Content addressing (Unison-style)

Every definition is identified by the hash of its syntax tree with free variables replaced by the hashes of what they refer to. Names are a lookup table on the side. Consequences:

  • A function’s identity includes the identities of everything it calls. Change g and f gets a new hash automatically; unchanged functions do not.
  • No name-based conflicts; renames are free.
  • Check results, test results, and lowering results are cached by hash.
  • Composes with Nix, which is also content-addressed.
  • soil_hash is computed over the canonical text of the alpha-normalized AST (adopted 2026-08-23, survey §5.2, after Tacit/Unison; the hash form is specified in soil-syntax-spec §9.1, added 2026-08-24): local binders hash as indices with display names excluded, so renaming a local in a hand-edit or re-lowering never invalidates verification or caching. Definition-level names stay load-bearing (filename is identity, §4.3); local names are display metadata for hashing purposes. The identifier-leakage research (Wang et al.) also motivates a later misleading-name lint (§10) — wrong names damage the next agent to read the code — but not nameless surface syntax (§12).

Because cross-definition mutual recursion is forbidden (§3.8), the definition graph is a tree for functions. Recursive types need a combined cycle hash.

6.2 Three-part hashing per definition

HashCoversOn change
formal_hashSignature, effect row, refinements — the formal .tr blocks per tr-grammar §1 (spec-side only, resolved 2026-08-24: the computed import set lives in the lowering’s calls[].hash edges, which is how caller invalidation routes — folding it in here would make the spec hash depend on the artifact it gates)Must re-lower or re-verify
test_hashHuman-written tests; the reference attachmentMust re-lower or re-verify (a reference change re-runs differential tests only — the reference is an oracle, not an input to lowering)
prose_hashEverything elseFlag review-suggested; existing lowering stays valid

A prose-stale state may be auto-cleared when an agent re-reads the prose and confirms the existing Soil still matches. This gives a cheap round-trip check without forced regeneration. Prose is thus “somewhere between hashed exactly and allowed to drift.”

decisions blocks are in none of the three: they form a fourth, per-entry hash class invalidated by reliance edges rather than file-level hashing (resolved 2026-08-24, §4.3; tr-grammar §1/§5.2; lock-schema §3/§8).

6.3 Lock file

  • One lock file per definition: f.tr → f.lock. Sidecar.
  • Global manifest is derived, gitignored, regenerated from the sidecars. It is what Nix consumes. It must be a merge of the sidecars, never separately maintained.
  • Merge conflicts can only arise when two people edit the same definition, which is a real conflict anyway.
  • Verbosity is fine. Fields expected per entry: formal_hash, test_hash, prose_hash, soil_hash (covering the Soil body plus transitively referenced private helpers), check status, test status and test mode tags, provenance (agent, human-verified, hand-edited, prelude-fork), trust level for FFI bindings (harvested, generated, declared-only), language/format version, pinned flag, accepted flag, escape hatch list, cycle hash for recursive types, and for FFI bindings the symbol hash plus the Nix store path of the package.
  • Language versioning: the lock records which Soil and Trellis versions a lowering targeted, so upgrades do not invalidate silently.
  • The lock schema is prototyped in docs/lock-schema.md with example sidecars in examples/. Decisions: JSON in the canonical value form (one format, one parser, diff-stable key order); component statuses (checks facts plus per-test results) with the typed/tested/verified/accepted ladder derived by the IDE, never stored; the lowering record carries provider and model for audit while costs, retries, and timings stay in the gitignored f.log; per-block provenance under spec.blocks implements the @agent write-back scheme; oracles records test-level edges by hash; accepted requires no xfail/xpass results.

6.4 Incremental refinement checking and demotion

Refinement types are modular: each function is checked against its own signature plus the signatures of callees. Checking is per-function; changing a body without changing its signature invalidates nothing downstream. Changing a signature invalidates callers via content addressing.

Fallback on proof failure: the function is demoted to its base ML type (refinements erased from the checker’s view), marked unverified, and tests remain required. Callers relying on the refinement are checked against the weaker type. Demotion is explicit and local; the agent never silently widens a signature. The IDE shows the unproven chain. This is effectively gradual refinement typing (Lehmann & Tanter).

The solver outcome is three-way, and a counterexample is a failure, never a downgrade (adopted 2026-08-23, survey §2.2, after Thermite): unsat proves the clause; unknown or timeout demotes it — the only demotion path; sat with a model fails the lowering, because the checker has found a concrete input on which the claim is wrong and demoting would ship code known-wrong on a known input. The counterexample is handed to the agent as a structured repair input (it is a failing test the solver wrote) and offered to the human as a one-click expect test — the solver just found an input the human’s tests missed, strengthening the trust root. runtime status therefore means “undecided”, never “known wrong”.

Assurance is recorded per clause, not per definition (adopted 2026-08-23, survey §2.1): each refinement clause carries proven (SMT, erased in release) | runtime (demoted, guard active) | trusted (human escape hatch) in the lock (§6.3), so a definition with three clauses can honestly be two-proven-one-runtime instead of a single flag. Module and project badges aggregate by minimum over exported definitions’ clauses (§7.1), so one runtime-only boundary is never hidden behind a proof elsewhere.

Violations carry blame (adopted 2026-08-23, survey §2.3, after Vow): a requires violation faults the caller, an ensures/invariant violation faults the callee, as a structured field in every runtime check and SoilError (§3.11’s debug payload). Blame tells the daemon which definition to queue for re-lowering and which lock entry to mark suspect — the units of repair are definitions with separate specs, so routing the repair automatically is worth a field in every guard, and retrofitting it into emitted guards later would be tedious. Lands with the runtime guards (plan 05).

Blast radius is kept small by design: the checker is for extra safety; tests are what correctness is based on.


7. Product and IDE

7.1 IDE-first

The IDE experience is the primary goal; the first milestone is a tool the author wants to use. Industry adoption within larger teams is the secondary goal (such teams will want a trellis derive command that stubs .tr files from an existing codebase). CI and headless operation are later concerns, served by running the daemon headless.

  • Platform: a web app served by Electron, for portability and simplicity, and so the same UI can later be served remotely. The Electron tax is accepted.
  • Editing model: a text editor with widgets. Widgets are the JSON test blocks (drag-and-drop), REPL-to-test promotion, and the question/answer panel. Everything else is Markdown.
  • Lowering UX: clicking “lower” starts an interactive session in which the agent can prompt back with questions and error reports. The human answers in the IDE; answers are persisted to the .tr.
  • Fix mode: a debug-mode run produces a trace, flame graph, and proposed test cases; the IDE supports turning any of these into tests.
  • Git: the IDE never commits on its own. At accept, pin, and rename it suggests a commit with a message; one click to commit, one to decline.
  • Lints: definition-size and split-suggestion lints (e.g. a 400-line lowering for a one-paragraph spec) come from the Soil checker and surface in the IDE as suggestions, never blocking.
  • Badges aggregate by minimum (adopted 2026-08-23, survey §2.1): a module or project badge is the minimum over its exported definitions’ per-clause assurance and status — a green project means every export’s every clause is at least runtime-checked and every export is accepted. One aggregation rule; prevents the dashboard lie.
  • Status: the lock is ugly and never read directly; the IDE renders it. Code review is supported by the IDE rendering “specs changed, tests changed, N re-lowered, M newly accepted.”
  • Telemetry: the lowerer tracks tokens, cost, retries, and provider per lowering from day one; this is what later makes automatic model selection possible.
  • Build targets: trellis build builds whatever soil.toml specifies: binary, py_module, shared_lib, jvm_jar, etc. One tool, one flag surface.
  • Upgrades: lowerings record the language version they targeted and are left alone on upgrade until their spec changes; the Soil compiler therefore needs a compatibility policy.
  • Licensing (resolved 2026-08-22): the open parts are GPL-3.0-or-later, with the GCC Runtime Library Exception 3.1 additionally applied to everything that ends up inside compiled user programs (soil-rt, the prelude, later batteries) — the GCC model, so forks of the toolchain must stay open while user binaries carry no obligations. Docs, specs, and examples are CC BY-SA 4.0 (copyleft for prose without GPL’s ill-fitting source-form mechanics). Contributions are DCO-only, no CLA; consequently the closed-source IDE shares no code with the open repo and talks to the daemon only over its API — which is already the architecture (§4.9). AGPL for the daemon was considered (hosted-lowering loophole) and rejected in favor of one uniform license; MPL was rejected as too weak (closed files around the open core); a CLA was rejected as the wrong asymmetry for a copyleft project. License texts: LICENSE, LICENSE.exception, LICENSE.docs at the repo root.
  • Hosted lowering: eventually possible by design (the daemon is a service), but not a v1 or v2 concern.
  • Naming: the project, the specification layer, and the IDE are Trellis; the target language is Soil. The CLI is trellis (e.g. trellis lower, trellis bind). “Trellis” was the placeholder (declarative vibe-coding).

7.2 On-disk layout

Soil lives in its own directory, any directory with a soil.toml at its root. Zip files were rejected as opaque to git; inline embedding in foreign source was rejected as worse. A repo may contain multiple independent Soil roots with no cross-root imports; slices that need to share are one root.

Spec, generated Soil, and lock sit side by side per definition. A split into spec/ and gen/ was rejected as losing the locality that justified one-file-per-definition.

soil/
  soil.toml            -- deps, prelude fork, build targets, CI policy, tag vocabulary
  soil.lock            -- derived global manifest, gitignored
  parser/
    _module.tr         -- module prose, export list
    parse.tr
    parse.soil
    parse.lock
    parse.log          -- gitignored, human-readable lowering log
    tokenize.tr
    tokenize.soil
    tokenize.lock
    _private.soil      -- module-private helpers
  soil/__init__.pyi    -- generated host stub

Every definition is three files; the module header and private helpers are the two exceptions; the only non-derived global file is soil.toml.


8. Build system

  • Declare dependencies in TOML; Trellis compiles it to a Nix build script. Nix is the right model (hermetic, content-addressed) but adopting it wholesale couples users to its ecosystem. The approach mirrors dream2nix, crate2nix, poetry2nix; Trellis is the polyglot roof over them.
  • Generate flake.lock-style pinned inputs.
  • Toolchain pin, refuse-on-mismatch (adopted 2026-08-23, survey §5.3, after Tacit): soil.toml pins the compiler, the skill bundle, and the prelude/batteries hashes, and the daemon refuses to lower or verify under a mismatched toolchain with a structured diagnostic, rather than silently writing lock entries that claim more than they should. Upgrades are explicit trellis toolchain update events — the natural trigger for the §7.1 re-verification sweep.
  • No “escape to raw Nix” field in the TOML, or every project will use it and the tool becomes Nix with extra steps.
  • A dev mode that shells out to native toolchains without Nix is planned for contributor onboarding.
  • Foreign file locking: Nix hashes lock provenance (which bytes); Trellis’s lock records interface (which shape) via symbol hashes of .pyi, .d.ts, or C header declarations. v1 accepts Nix’s coarse invalidation (any upstream commit invalidates all bindings); the lock format is designed so finer symbol-level invalidation slots in later.
  • Buck2 was noted as an alternative if Nix proves too heavy. Deferred.
  • FFI boundary tests: harvesting tests from the dependency’s own suite is the default; generated boundary property tests from declared types are the fallback; each binding records its trust level. A foreign call has effect ffi (implying panic) and “tests are the only guarantee here”; no attempt to shrink effect rows for foreign code.

9. Open questions

  1. .tr grammar — prototyped in docs/tr-grammar.md (§4.3) with its follow-up questions resolved: type files carry YAML frontmatter declaring the type’s cased name (filenames are snake_case everywhere); property where filters use constrained generation, not rejection sampling; the cram block is a minimal cram subset (with file fixtures, fresh temp dir, literal output, [n] exit codes, trellis call for real-capability invocation); property-only and cram-only files are valid (main and other toplevel functions are typically of that shape); every file carries frontmatter with a required name and optional non-semantic tags whose vocabulary is declared in soil.toml, while file kind stays inferred. Finalization into docs/ pending.
  2. JSON encoding of Soil values — resolved: internally tagged sums, type-directed decode, canonical show output, opaque one-way "<handle>", functions a hard error (grammar prototype §7). BigInt is hybrid by range: a JSON number within ±(2^53−1), a string beyond, and decode accepts either — small values stay readable while big ones survive float-only host JSON parsers.
  3. Lock entry schema — prototyped in docs/lock-schema.md (§6.3), including .tr provenance for the vibing tiers and per-block provenance for the write-back scheme (5). Newly open from the prototype: whether module entries participate in accepted; a fixed naming scheme for derived tests; whether entries pin the prelude-fork hash or leave it global in soil.toml.
  4. Prose-friendly refinement syntax — resolved: labelled requires/ensures clauses over a shared predicate language (§4.3; grammar prototype §2.3).
  5. Agent write-back markers — tentative proposal (grammar prototype §8): agent-authored blocks carry @agent in the info string; a human edit removes the marker, and an unmarked formal block is pinned — the agent may not change it, only ask_human.
  6. Type invariants: checked on every constructor call, or only proven at definition sites.
  7. Naming conventions for agent-created private helpers.

10. Deferred decisions and their triggers

ItemCurrent stanceRevisit when
Memory model details (cycle collection strategy, regions)Runtime is Rust; reference counting with Perceus-style reuse; cycle collection strategy openWhen the Rust runtime is built
Primitive type details for FFI (UTF-8/16, bigints, floats)Language supports all variants; stdlib provides interop; agent handles polyglot stringsStrict technical decision when writing the FFI layer
Additional io testing modes beyond fake capabilitiesCram tests available as fallback; lock tags modeFakes prove insufficient
Export pin enforcementFlag exists in lock; nothing enforcesSecond user
Retry policy, model routing, cost ceilingMinimal fixed versionsSecond user / cost pain
Corpus retrievalWhole prelude in contextPrelude outgrows context
Merge tooling for locksNone needed with per-definition sidecarsSecond user
IDE (graph view, click-to-generate, Q/A buttons, JSON block editor, module view)Not builtAfter the CLI loop proves pleasant
Nix backendNot builtAfter the CLI loop
Whole-package binding generatorNot built; per-symbol trellis bind is v1Likely never; see §3.11
Backends beyond the firstNot builtAfter the loop works
User-declared algebraic effects, effect handlers, concurrencyNot in language; concurrency via FFI to host librariesOnly if glue-language positioning changes
User overrides of derived functionsNot allowed; would be a class systemIndefinitely
Full dependent types beyond refinementsNot in languageSMT proves insufficient for a meaningful fraction of functions
Buck2 as build alternativeNot pursuedNix proves too heavy
Capability names in effect rowsBare ioNever, unless capability arguments prove insufficient
Symbol-level invalidation of foreign bindingsCoarse Nix invalidationAfter v1
py_module build target and Python-hosts-Soil embeddingNot builtAfter v1
Blocking in-context ask_humanExit-and-reinvoke onlyIf question round-trips prove too costly
Raw-API lowering providerAgent-CLI providers onlyv2
Hosted loweringNot builtPost-v2
trellis derive from an existing codebaseNot builtSecond user / team adoption
Concurrent loweringsStrictly serial with a queueIf serial throughput hurts
Separate signature-inference step before loweringNot allowed; disciplined order enforcedIf waiting on callees proves too annoying
seccomp/Landlock policy emitted from the capability set; trellis run --denyDesign adopted 2026-08-23 (survey §6.1); not builtWith trellis build (plan 06+) — “the sandbox is derived from the types”
Per-resource capability confinement (scoped Fs via Landlock paths)Not built; kind-level caveat documentedAfter the kind-level sandbox
TrellisBench (spec+tests problems through the real daemon, per release, published results)Adopted 2026-08-23 (survey §8.1); not builtBefore the prelude grows (plan 04 kickoff) — corpus changes must be measured
Speculative proof-delta daemon tools (speculative_check, spec-edit blast radius)Adopted 2026-08-23 (survey §4.2); not builtOnce incremental checking is warm (plan 03+)
Co-process Python isolation (coproc per-binding mode)JSON-marshalling constraint adopted (§3.11); mode not builtPost-v1; lowering sandbox runs Python coproc first
Mutation-testing job (trellis mutants)Lock fields reserved 2026-08-23 (§4.5)Post-v1 daemon job
Bounded model checking as an assurance rung between proven and runtimeNot adopted (§12)If SMT coverage proves insufficient
Misleading-name lint on Soil bindersNot built (§6.1)After soilc
Solver-backed constrained generation for property where filters and invariant-derived propertiesv1 runner: random generation from the type; where filters refused with a structured error (never rejection sampling — tr-grammar §3.4; resolved 2026-08-24, impl plan 03 §9.8). Invariant-derived properties refused the same way (unsupported-invariant-property, resolved 2026-08-28): generating self from the bare structure tests exactly the values the invariant excludes, so the honest generator is the constrained one; type acceptance is ungated meanwhile (lock-schema §8)Z3 lands (plan 05). Caveat recorded now: raw solver models cluster — sampling an SMT solution space uniformly is unsolved, so distribution quality must be checked before adoption

11. Build order

The build order follows the bootstrap plan (docs/bootstrap-plan.md): the compiler itself is the first Trellis project, self-hosted on a minimal Rust implementation that is kept forever as a differential oracle. Each milestone has an implementation guide in docs/plans/.

  1. soil0 + soil-rt. A Rust workspace: the runtime crate (values, reference counting, JSON bridge, C ABI with a trivial main wrapper from the first commit) and a minimal Soil implementation — parser, ML + effect-row inference, exhaustiveness, tree-walking interpreter, test runner. No refinements, no termination checker, no codegen, no FFI yet. Every pass is exposed as a JSON-in/JSON-out CLI command (soil0 parse, soil0 infer, soil0 run) — the future differential oracles. Deliberately small; interpreted execution is the engine for the whole bootstrap, and slow is accepted.
  2. The daemon. Incremental compiler state, LSP, context-bundle assembler, MCP tool surface, lowering jobs with the exit-and-reinvoke question channel, serial queue, agent-CLI provider (Claude Code headless first), REPL endpoint, debug-mode instrumentation, cost telemetry, per-definition locks and the derived manifest. This is where the effort goes; it is smaller than it sounds because the agent CLI supplies the loop.
  3. The pure-core prelude as the first Trellis code, interpreted on soil0.
  4. soilc: the compiler as the first Trellis project. Passes in oracle-ready order — lexer, parser (the mutual-recursion stress test, deliberately early), renamer, type + effect inference, exhaustiveness/patterns, ANF, termination checker, refinement checker (SMT-LIB out, Z3 behind a Solver capability), CLIF backend — each differentially tested against the matching soil0 command and swapped into the daemon at accepted (strangler pattern). Refinements and demotion therefore arrive here, as compiler passes, not as a later milestone. Closure: soilc interpreted compiles the prelude and itself via the Cranelift driver → soilc₁; soilc₁ compiles the same sources → soilc₂; the build requires soilc₁ ≡ soilc₂ byte-identical. soil0 is retained permanently as oracle and debug-mode engine.
  5. FFI, trellis bind, minimal IDE. C-ABI FFI to Rust and Python FFI via embedded CPython, hand-written bindings, sequenced C ABI → Rust batteries → Python embedding → Python batteries; one corpus shim example per FFI; the per-symbol bind assistant; the Electron-served IDE over the daemon (Markdown editor with test-block widgets, graph view, lower button with streaming output and the question panel, REPL, lock rendering), as thin as possible.
  6. The Python-glue project as the second Trellis project. A program the author would otherwise have had an agent write in pure Python: Rust crates through soil-rs-std, a dozen Python functions through trellis bind, real work in Soil. The compiler validates the pure core; this validates the FFI/capability/bind half of the pitch and the pleasantness test — writing .tr files and reading generated Soil.
  7. Then trace-to-test fix mode, profiling, build targets, TOML→Nix, raw-API provider, trellis derive, headless CI mode.

Immediate next artifacts:

  • The .tr grammar specification — prototyped (docs/tr-grammar.md plus examples/, including the JSON value encoding and the refinement prose syntax); to be finalized into docs/ once the prototype has been exercised.
  • The lock entry schema — prototyped (docs/lock-schema.md plus example .lock sidecars); to be finalized into docs/ with the grammar.
  • The prelude’s read_file as the first real definition (drafted as examples/read_file.tr with read_file.lock and read_file.soil).
  • A high-level Soil surface syntax: prototyped in docs/soil-syntax.md (Haskell-style signatures and inline refinements, OCaml-style terms, decreases lines, comparison operators as derived-function notation, no imports; guards/;/let? deliberately absent or deferred), elaborated into a lexical spec, EBNF, and static rules in docs/soil-syntax-spec.md (OCaml-style match with parenthesized nesting; one connective spelling and/or/not shared by terms and predicates, implies predicate-only; shadowing forbidden; :: namespacing; .. required in partial record patterns; parameterless let bindings with explicit lambdas; a fixed scalar-value-only string escape set; arithmetic operators as notation with overflow/zero-divisor as refinement obligations; Bool encoding as JSON booleans). Result a e is success-first (§3.3). Full semantics arrive with the Soil core milestone.

12. Prior art to consult

  • Idris 2 / Agda / Lean: type-as-spec, hole-driven workflow (Trellis without the LLM); termination checking; total-by-default policy.
  • Hazel: live holes, typed structure editing, for the IDE.
  • Unison: content-addressed definitions, incremental typechecking, codebase-as-database; cycle hashing; closest existing thing to the manifest idea.
  • Dafny / Verus: spec-then-implementation with a checker in the loop.
  • Liquid Haskell / F*: refinement types with SMT discharge.
  • Lehmann & Tanter: gradual refinement types, for the demotion model.
  • Koka: direct-style surface with effect rows, div as an effect, Perceus reference counting; its effect types subsume function-application annotations.
  • OCaml: polymorphic compare, Map.Make comparator idiom.
  • Rust: derive, total_cmp for floats, debug/release split.
  • dream2nix / crate2nix / poetry2nix: per-ecosystem TOML-to-Nix precedent.
  • Buck2: polyglot build alternative.
  • Inform 7, AppleScript, COBOL, Wolfram: history of natural-language programming. The consistent lesson: prose as syntax fails; prose as spec alongside formal structure works. Trellis is on the right side of that line.

The 2026-08 agent-language survey (agentlanguages.dev, 38 entries; adoption record in docs/plans/extra/agentlanguages-adoptions.md) supplied every decision marked “adopted 2026-08-23” above — chiefly from Thermite (assurance ladder, counterexample rule, holes, seccomp), Vera (drift-gated codes, proof-delta LSP, literal provenance), Vow (blame, generated skill, canonical printer, mutation testing), Tacit (canonical form, toolchain pin, sectioned primer), and Aver (decision blocks, hostile profiles, budgeted context). Deliberately declined, with reasons: De Bruijn surface syntax (humans review this code and filename-identity depends on names; alpha-normalized hashing captures the benefit); AST-as-source / JSON programs (canonical text keeps greppability); mandatory contracts with no opt-out (the one-sentence-one-test minimal definition is the point — the human’s attention is the scarce resource in disciplined vibing); bounded model checking as the primary engine (verification-artifact bounds leak into contracts; at most a future assurance rung); first-person compiler personas (not the product).

The .tr File Format — Prototype Grammar

Status: prototype. Tentatively resolves §9.1 (grammar), §9.2 (JSON value encoding), and §9.4 (refinement prose syntax) of docs/design.md, and proposes an answer to §9.5 (write-back markers). Worked examples live in examples/.


1. Container format

A .tr file is CommonMark. Formal content lives in fenced code blocks whose info string begins with a reserved word. Everything else — including fenced blocks in unreserved languages such as python or text — is prose: it is hashed under prose_hash and never parsed.

There are four file kinds, distinguished by content, with filename as identity (design §4.3):

KindFilenameDefines
Function<name>.trone function; name = filename stem verbatim
Type<name>.trone type; PascalCase name declared in frontmatter, filename is its snake_case form
Module header_module.trmodule prose, the export list, module decisions
Project header_project.trproject prose and project-wide decisions (at the Soil root; added 2026-08-23, design §4.3)

Filenames are lowercase snake_case for every kind, so identity survives case-insensitive filesystems.

Every .tr file begins with YAML frontmatter. name is required and is the definition’s canonical name: a function’s equals the filename stem; a type’s is PascalCase and the filename is its snake_case form; a module’s equals its directory name; a project header’s equals the Soil root directory’s name. There is no kind field — kind is inferred (_module.tr and _project.tr by filename; a soil-type block makes a type file; otherwise the file is a function).

tags is optional: a list drawn from a per-project vocabulary declared in a [tags] table in soil.toml; an undeclared tag is an error. Tags are non-semantic metadata — IDE graph filtering and colouring, CI policy in soil.toml (e.g. “every definition tagged api must be accepted”) — and are never part of the lowering context bundle. Unknown frontmatter keys are errors.

---
name: ParseError
tags: [parser]
---

Reserved block languages

Info stringFile kindCountPurpose
soil-sigfunction≤ 1Soil type signature
requiresfunction≤ 1preconditions
ensuresfunction≤ 1postconditions
test <name> [xfail]functionanyexpect test
property <name> [xfail]functionanyproperty test
cram <name> [xfail]functionanyshell-transcript test (io fallback, main)
referencefunction≤ 1reference-implementation attachment
allowfunction≤ 1escape hatches (human-only)
soil-typetype= 1the type’s shape
invarianttype≤ 1type invariants
exportsmodule= 1export list
decisionsmodule, project≤ 1structured chosen/rejected decisions (design §4.3, adopted 2026-08-23)

Mapping to the three-part hash (design §6.2)

  • formal_hash: the frontmatter name, soil-sig, requires, ensures, soil-type, invariant, exports, allow.
  • test_hash: test, property, cram, reference (the reference is an oracle; changing it re-runs differential tests, not the lowering).
  • prose_hash: everything else in the file, including tags.
  • decisions blocks are in none of the three (resolved 2026-08-24, §9): each entry is hashed individually (label → hash of the entry’s text, rejected lines included) into the lock, and invalidation follows the reliance edges of §5.2 rather than any whole-file hash.

2. Shared mini-languages

2.1 JSON values

json below means an RFC 8259 JSON value, interpreted type-directedly under the encoding of §7.

2.2 Identifiers

ident is lowercase snake_case (functions, parameters, fields). Ctor and TypeName are PascalCase.

2.3 The predicate language

Shared by requires, ensures, invariant, and property. It is a restricted Soil boolean expression: calls may target only total definitions, and predicates outside the decidable fragment (linear arithmetic, uninterpreted functions, lengths — design §3.2) still parse but demote the function to unverified.

predicate  ::= disj [ "implies" predicate ]            (right-assoc)
disj       ::= conj { "or" conj }
conj       ::= neg { "and" neg }
neg        ::= [ "not" ] atom
atom       ::= comparison | is-test | call | "(" predicate ")"
comparison ::= expr relop expr
relop      ::= "==" | "!=" | "<" | "<=" | ">" | ">="
is-test    ::= expr "is" [ TypeName "::" ] Ctor [ "(" ident ")" ]
                                                       (binds the payload)
expr       ::= mul { ("+" | "-") mul }
mul        ::= app { "*" app }
app        ::= call | aexpr
call       ::= ident aexpr { aexpr }                   (juxtaposition; the head is a bare ident)
aexpr      ::= literal | path | "(" expr ")"
path       ::= ident { "." ident }                     (record field access)
literal    ::= JSON literal

Calls are juxtaposed exactly as in Soil terms — len v > 0, not len(v) > 0 — so there is one application spelling everywhere, by the same argument that gave the connectives one spelling (resolved 2026-08-22; the original parenthesized-comma call form was dropped, and examples/csvstats/median.soil is normative). Consequently is and implies are reserved words within predicates and cannot name paths there.

Names in scope: the signature’s named parameters; result (in ensures only); self (in invariant only); forall binders (in property only); and total definitions visible to the file.

Constructor references in is tests follow the exactly-one-spelling rule (soil-syntax-spec §5.9, design §3.15): bare when the variant name is unique among the types in scope, Type::Ctor required when it collides — qualifying a unique constructor is an error.


3. Function-definition blocks

3.1 soil-sig

sig    ::= defname ":" arrow
arrow  ::= param "->" arrow | ret
param  ::= "(" ident ":" type ")" | type
ret    ::= [ row ] type
row    ::= effect { effect }
effect ::= "div" | "panic" | "io" | "ffi"

An empty row means total. type is Soil type syntax, specified separately; this grammar treats it as opaque. The signature may be absent — the agent infers it and writes it back (§8). If requires/ensures refer to a parameter by name, the signature must exist and use the named-parameter form.

3.2 requires / ensures

block  ::= clause { clause }
clause ::= label ":" predicate                          (one per line)
label  ::= free text not containing ":"

Labels are the pinnable names: the lock, the IDE, and checker errors refer to clauses by label. In ensures, result is bound to the return value; for Result-typed functions the idiom is result is Ok(v) implies … (design §3.2: refinements may eliminate error cases).

3.3 test

The function under test is implicit — it is the file’s definition. A block holds any number of cases; case k of block name is reported as name#k.

block     ::= { with-line } case-line { case-line }
with-line ::= "with" ident "=" fake-call
fake-call ::= ident { json }
case-line ::= "(" [ arg { "," arg } ] ")" "=>" outcome
arg       ::= json | ident                              (a with-binding)
outcome   ::= json | "panic"
  • with lines construct fake capabilities from the prelude (design §3.5); their arguments are JSON, so seeds and timestamps are pinned by construction (with clock = fake_clock 1700000000).
  • panic as an outcome is only legal if the signature’s row carries panic.
  • xfail in the info string marks the whole block expected-to-fail; it blocks accepted until resolved (design §4.5).

3.4 property

block       ::= forall-line { forall-line } predicate
forall-line ::= "forall" ident ":" type [ "where" predicate ]

Generators are derived from the binder’s type; a where filter is a generator constraint, satisfied by constrained generation rather than rejection sampling. Properties may call the function under test, the prelude, the reference implementation, and accepted definitions (design §4.5). (Toolchain status, resolved 2026-08-24: the v1 runner generates randomly from the type and refuses where filters with a structured error — constrained generation arrives with the solver, plan 05; design §10. The format is unchanged.)

3.5 cram

For io functions where fakes stop being possible, for FFI bindings, and for main. The dialect is a minimal subset of classic cram: unindented, no (re)/(glob) matchers (extensible later).

block       ::= { with-line } step { step }
with-line   ::= "with" "file" string "=" string      (JSON strings)
step        ::= command { output-line } [ exit ]
command     ::= "$ " rest-of-line                    (a shell command)
output-line ::= any line not beginning "$ " or "["
exit        ::= "[" integer "]"
  • Each block runs in a fresh temp dir. with file lines materialize fixtures before the transcript runs — path and contents are JSON strings, so escapes are pinned.
  • Commands run sequentially in one shell session in that dir. Expected output is combined stdout+stderr, matched literally. An omitted exit line means 0.
  • The transcript may invoke built binary targets from soil.toml (placed on PATH), and trellis call <def> <json-arg>…, which runs an io definition with real World-derived capabilities and prints its result as canonical JSON — the real-mode escape for non-main io functions and FFI bindings. Capability parameters are injected from World; the JSON arguments fill the remaining parameters in order.
  • The lock tags these tests mode real (design §4.5). Cram never runs inside the lowering sandbox: the lowerer’s run_tests tool exposes only fake-capability tests (design §4.6).

3.6 reference

One line, in one of two forms:

line ::= relpath "::" symbol          (Python reference)
       | "cli" command-line          (CLI oracle)

ref/stats.py::median attaches a Python reference implementation. cli soil0 parse attaches a JSON-in/JSON-out executable as a black-box oracle: the daemon passes the test’s JSON arguments on stdin and expects canonical JSON on stdout (added for the compiler bootstrap; design §4.5). Attached explicitly by the human, never auto-detected.

3.7 allow

One escape hatch per line, from the fixed set partial, unsafe, ffi-raw. Written only by the human; every allow block in the project is surfaced in the manifest’s audit view (design §4.1).


4. Type-definition blocks

4.1 soil-type

typedef ::= "type" TypeName { tyvar } "=" body
body    ::= "opaque" | record | sum | type              (last = alias)
record  ::= "{" field { "," field } "}"
field   ::= ident ":" type [ "ignored" "=" expr ]
sum     ::= [ "|" ] ctor { "|" ctor }
ctor    ::= Ctor [ type ]

A variant carries at most one payload of any type; multi-field payloads are inline records, since there are no positional products (design §3.13) — Ok a, but BadCell { index : U64, text : Utf8 }.

Sum vs alias (resolved 2026-08-28, surfaced by the daemon’s parser): the grammar above makes type Path = Utf8 ambiguous — a one-variant sum or an alias — and type W = MkW Utf8 ambiguous between a one-variant sum and an applied-type alias. The rule: a body is a sum iff it begins with | or contains a top-level |; otherwise it is an alias (or record/opaque by its leading token). A single-variant sum therefore must write the leading | (type W = | MkW Utf8) — one way to write each thing, and the kernel’s type Path = Utf8 keeps its alias reading. Derivation strategies (design §3.7): structural is the default; an opaque body selects the opaque strategy; the ignored field marker selects the ignored strategy for that field.

An ignored field must carry a default: an expr from the predicate language (§2.3) — so calls target only total definitions — with the record’s non-ignored fields in scope. The default materializes the field wherever a value is built without it: JSON decode, py_to_soil, host stubs. ignored thus means “excluded from derivation, reconstructible on demand”:

type Doc = { text : Utf8, cached_word_count : U64 ignored = word_count text }

4.2 invariant

Same clause grammar as ensures, with self bound to a value of the type. Invariants are properties every constructor must preserve (whether checked at every construction or proven at definition sites remains open, design §9.6). Each invariant clause auto-generates a property test (design §4.1), which is why type files carry no hand-written test blocks. (Toolchain status, resolved 2026-08-28: the v1 runner refuses invariant-derived properties with unsupported-invariant-property — generating self from the bare structure would test values the invariant is precisely meant to exclude, i.e. the honest generator is constrained generation, which arrives with the plan-05 solver alongside where filters (§3.4). The invariant remains a runtime assurance in the lock’s checks, and type acceptance is ungated in v1 (lock-schema §8). The format is unchanged.)


5. Module headers

5.1 exports

line ::= ident
       | "type" TypeName
       | "abstract" "type" TypeName

An exact list of definitions (design §4.4). If an exported signature references an unexported type, the two permitted repairs are exporting it or marking it abstract here (design §3.14).

5.2 decisions

Structured project or module rules — chosen, rejected, rationale — queryable by the IDE and included in every context bundle in scope (the module’s lowerings for _module.tr, every lowering for _project.tr), closing the tenet gap where such rules lived only in an agent’s context (design §4.3, adopted 2026-08-23 after Aver).

block    ::= entry { entry }
entry    ::= label ":" text                    (the decision, one line)
             { "rejected:" text }              (zero or more alternatives, with reasons)
label    ::= free text not containing ":"

Labels are addressable: the interactive lowering UI’s write-back targets them (an answer that is a project rule lands here, design §4.6), and the IDE answers “why Result here?” by label. Example:

```decisions
timestamps: all timestamps are UTC seconds (I64)
  rejected: local time (DST bugs); F64 epoch (precision loss)
errors: Result everywhere; panic only via refinement-checked prelude calls
```

Hashing and invalidation (resolved 2026-08-24; design §4.3). Decisions are neither prose nor formal — forcing them into that binary meant a one-word edit either re-lowered the world (formal semantics) or flagged the whole scope review-suggested (prose semantics), and scope-wide flags that are usually noise train the user to ignore the one that matters. Decisions are therefore their own hash class, invalidated by reliance — the same shape the lock already uses for callee edges:

  • Each entry is hashed individually; labels are identity (renaming a label is remove + add).
  • Every lowering cites the decisions it applied: the lowerer’s completion output carries the list of labels (possibly empty), recorded in the lock as {scope, label, hash} edges beside the call edges, together with a hash over the entry labels in scope (membership, not text). Lock shapes in lock-schema §3.
  • An edited or removed entry flags only the lowerings whose edges cite it; an added entry flags everything in scope once, via the membership hash — a new rule is new information to every existing lowering. Staleness is derived by hash comparison, never stored (lock-schema §8).
  • The human may reclassify a detected change as editorial (typo, wording): the daemon re-stamps the recorded hashes and nothing is flagged. The human already owns the spec, the tests, and accepted; trusting them to say “this edit changed no meaning” is inside the trust model.
  • Whatever still flags is cleared by the triage sweep: a batched daemon job that hands an agent the entry’s diff and asks, per flagged lowering, whether the existing Soil still conforms — the prose-stale auto-clear of design §6.2, batched. It costs tokens, never a re-lowering.

Citation is self-reported by the lowerer, so an uncited-but-influential decision under-flags — the same drift class the prose tier already accepts (flag-and-confirm; tests remain the trust root). Rejected: pure prose semantics (review-suggested fatigue at project scale) and formal semantics (a typo re-lowers the world) — both scale with project size, where reliance edges scale with actual use.


6. Validity rules

  1. Every file begins with frontmatter carrying a name. Unknown frontmatter keys, and tags not declared in soil.toml, are errors.
  2. Function file: name equals the filename stem; at least one prose paragraph and at least one test, property, or cram block. The minimal valid definition is the frontmatter, one sentence of prose, and one expect test (design §4.3); property-only and cram-only files are valid — main and other toplevel functions are typically of that shape.
  3. Type file: the snake_case form of name equals the filename stem; at least one prose paragraph and exactly one soil-type block. No test blocks.
  4. Module header: name equals the containing directory’s name; exactly one exports block; at most one decisions block.
  5. Project header: lives at the Soil root beside soil.toml; name equals the root directory’s name; no exports; at most one decisions block.
  6. Block multiplicities per the table in §1; test/property/cram names unique within a file.
  7. The names declared in soil-sig and soil-type, when present, must equal the frontmatter name. Renaming is refactor-rename (design §4.3).
  8. Predicates may reference parameters only via a named-parameter soil-sig.

7. JSON value encoding (resolves design §9.2)

One encoding serves show/parse, tests, the REPL, and host stubs. Encoding and decoding are always type-directed. Sum types are internally tagged: one uniform shape for every variant, self-describing for hosts and generic tooling. The verbosity is accepted because JSON values are primarily written and read through the IDE’s block widgets (design §4.3), not typed by hand.

Soil typeJSON
I64, U64, I32, …number (integer)
BigIntnumber within ±(2^53−1), string beyond; decode accepts either
F64number; "NaN", "Inf", "-Inf" as strings
Utf8string
Bytesstring, base64
Unitnull
Booltrue / false — a prelude sum type, but the one special case in the sum encoding
recordobject; every field present; ignored fields omitted by show, refilled from their default on decode
sum, nullary variant{"tag": "Name"}
sum, payload variant{"tag": "Name", "value": <payload>}
List aarray
Map k varray of {"key": k, "value": v} in comparator order
opaqueshow emits "<handle>"; decoding is an error
functionhard error in both directions

Canonical output: show emits record fields in declaration order, map entries in comparator order, and floats in shortest round-trip form, so equal values produce byte-equal JSON and expect tests can compare on the string (design §3.7).

The keys "tag" and "value" are produced only by the sum encoding; since decoding is type-directed, a record field named tag is not ambiguous, but the linter warns on it.


8. Provenance and write-back (tentative, design §9.5)

Agent-authored formal blocks carry an @agent marker at the end of the info string:

```soil-sig @agent
mean : (xs : List F64) -> F64
```

The agent may freely rewrite blocks marked @agent. When a human edits such a block they remove the marker; an unmarked formal block is human-authored and therefore pinned — the agent may not change it, only raise ask_human. The lock records provenance per block alongside the hashes.

The marker is part of the block’s hashed info string (impl plan 03 §8.5; resolved 2026-08-28): removing @agent — even without touching the content — moves the block’s class hash and re-verifies the lowering. Taking ownership of a block is a formal event, not a metadata flip; the conservative cost of one re-verification per adoption was chosen over content-only hashing.

Answer write-back (resolved 2026-08-24; design §4.6, impl plan 03 §9.7). A function-scoped ask_human answer is written into the .tr under a ## Clarifications heading (created at the end of the file on first use), one Q/A pair per entry:

## Clarifications

- **Q:** should `parse` accept trailing whitespace?
  **A:** yes — trim before parsing.

This is ordinary prose — hashed under prose_hash, visible in the diff the human already reviews, and included in the bundle like the rest of the prose. Rule-shaped answers go to a decisions block instead (§5.2), never here. The IDE later offers a fold-into-prose action: an agent rewrites the relevant prose to absorb a Q/A pair and deletes it — an ordinary prose edit under the same hashing. Rejected: human-placed inline answers as the default (the round-trip stops being automatic) and a reserved qa block (a new mini-language for what prose already expresses).


9. Open questions raised by this prototype

All questions raised by the first draft — type-name casing (frontmatter, §1), ignored-field decoding (defaults, §4.1), BigInt interop (hybrid by range, §7), generator strategy (constrained, §3.4), the cram grammar (§3.5), and property-only files (valid, §6) — have been resolved and folded into the sections above.

The decisions-hash question (newly open 2026-08-23 with the decisions adoption: prose-class flagging vs formal-class invalidation) was resolved 2026-08-24 as neither: decisions are their own per-entry hash class, invalidated by reliance edges, with editorial reclassification and the triage sweep. Both binary options scaled with project size (scope-wide review-suggested fatigue, or a typo re-lowering the world); reliance scales with actual use. Rules and rationale in §5.2; lock shapes in lock-schema §3/§8; design §4.3.

The Lock Entry Schema — Prototype

Status: prototype. Tentatively resolves §9.3 of docs/design.md. Example sidecars live next to the .tr examples in examples/. Hashes in examples are abbreviated.


1. Shape

One lock file per definition, sidecar: f.tr → f.lock (design §6.3). A lock file is a single JSON object in the canonical form of the value encoding (grammar prototype §7): sorted-stable key order, so diffs are minimal and semantic. The global soil.lock manifest is derived by merging sidecars (§7 below) and is gitignored.

Top-level keys, in order:

KeyKindPurpose
lock_formatallschema version of the lock itself
namealldefinition name, matching the .tr frontmatter
kindallfunction | type | module (inferred from the .tr, recorded for the manifest)
versionsall{trellis, soil} versions the current artifacts target (design §6.3: upgrades must not invalidate silently)
specallhashes and provenance of the .tr (§2)
loweringfunctionthe generated Soil (§3); null before first lowering
checksfunction, typechecker facts (§4)
testsfunction, typeper-case results (§5)
oraclesfunctiontest-level edges (§5)
ffifunctionbinding-only trust record (§6)
cycle_hashtypecombined hash for recursive type groups (design §3.8); null if acyclic
acceptedallthe human’s trust flag (design §4.7) — the only field a human sets directly

2. spec

"spec": {
  "provenance": "human",
  "hashes": {
    "formal": "sha256:9f2c41aa",
    "test": "sha256:41aa73c0",
    "prose": "sha256:c8172d99"
  },
  "prose_state": "fresh",
  "blocks": [
    { "block": "soil-sig", "author": "agent" },
    { "block": "test odd-length", "author": "human" }
  ],
  "pinned": false,
  "escape_hatches": []
}
  • provenance: human | agent — who authored the .tr overall, distinguishing the two vibing tiers (design §1.2, §9.3).
  • hashes: the three-part hash (design §6.2). The block-to-hash mapping is grammar prototype §1. test is null for type and module files.
  • prose_state: fresh | review-suggested. Set to review-suggested when prose changes under an unchanged lowering; auto-cleared when an agent re-reads the prose and confirms the Soil still matches (design §6.2).
  • blocks: per-block provenance for the write-back scheme (grammar prototype §8). block is the info string minus modifiers (xfail is not identity). author: human means unmarked, therefore pinned — the agent may not edit it. Frontmatter is implicitly human.
  • pinned: the export-pin flag (design §4.4) — human approval of the public signature. Present from the start, unenforced in v1.
  • escape_hatches: the contents of the allow block, aggregated by the manifest into the audit view (design §4.1).
  • decisions (module and project entries only; adopted 2026-08-24, tr-grammar §5.2): a map from entry label to the hash of that entry’s text (rejected lines included). Decisions are their own hash class — decisions blocks are excluded from prose — and labels are identity (a rename is remove + add). Function entries record their reliance on these under lowering.decisions (§3).

3. lowering

"lowering": {
  "soil_hash": "sha256:77b04e12",
  "provenance": "agent",
  "provider": "claude-code",
  "model": "claude-opus-4-7",
  "private_helpers": [ { "name": "_parse_cell", "hash": "sha256:3fe210bb" } ],
  "calls": [ { "name": "sort_by", "hash": "sha256:aa90b1f3" } ],
  "decisions": {
    "scope_hash": "sha256:5b21aa04",
    "applied": [ { "scope": "project", "label": "timestamps", "hash": "sha256:c91d20fe" } ]
  }
}
  • soil_hash: content-address of the Soil body plus transitively referenced private helpers (design §6.3), free variables replaced by callee hashes (design §6.1). Adopted 2026-08-23 (design §6.1, §4.8): the hash is computed over the canonical text of the alpha-normalized AST — local binders as indices, display names excluded, exactly one printed form per AST (the printer is a compiler pass; soil0 impl step 11) — so local renames and formatting can never invalidate caching or verification.
  • provenance: agent | human-verified | hand-edited | prelude-fork. hand-edited skips re-lowering until the spec changes (design §4.8).
  • provider / model: which agent produced the current Soil, for audit and future model routing. Costs, retries, and timings live in the gitignored f.log, not here (design §4.6). Present only when an agent ran: hand-written entries (human-verified, hand-edited without a prior agent run) carry plain null — the lock’s own null convention, as with an absent lowering or test hash (resolved 2026-08-28 with the examples-regeneration policy — a lock must not assert a lowering event that never happened).
  • private_helpers: the _private.soil definitions this lowering owns (design §4.2); the manifest derives its soil-private nodes from these, and a helper with no remaining owner is garbage-collected.
  • calls: callee edges with the hashes they were checked against — the graph view and the invalidation record (design §4.3, §6.1).
  • decisions (adopted 2026-08-24, design §4.3; tr-grammar §5.2): reliance edges, the decisions analog of calls. applied lists the decision entries the lowerer cited as applied — the citation is a required part of the lowering’s completion output, possibly empty — each pinned at the entry hash it was applied under (scope is project | module). scope_hash covers the sorted entry labels in scope (membership, not text), so additions and removals are visible without text edits flagging everyone. Staleness is derived, never stored (§8): an applied hash that no longer matches the current entry — or whose entry is gone — flags this lowering; a scope_hash mismatch flags everything in scope. An editorial reclassification by the human, or a triage-sweep confirmation that the Soil still conforms, re-stamps these hashes without re-lowering.

4. checks

"checks": {
  "types": "ok",
  "termination": "verified",
  "refinements": { "upper bound": "proven", "non-empty": "runtime" },
  "holes": 0
}
  • types: ok | error.
  • termination: verified | unverified | n/a — whether the claimed row’s absence of div is established. Until the termination checker exists (plans 04–05), recursive definitions claiming totality carry unverified: the demotion philosophy applied to div — unproven, visible, tests still gate.
  • refinements: assurance per clause (adopted 2026-08-23, design §6.4), a map from clause label (grammar prototype §3.2) to proven (SMT unsat; erased in release) | runtime (solver unknown/timeout — demoted, guard active in every build) | trusted (human escape hatch). {} when the definition has no refinements. A definition can honestly be two-proven-one-runtime; the IDE renders the unproven chain from the runtime entries. A counterexample is never a state: sat with a model fails the lowering, and the counterexample becomes a structured repair input and an offered expect test (design §6.4) — runtime means “undecided”, never “known wrong”.
  • holes: the count of unfilled typed holes (design §3.16). Nonzero is the partial state: the lowering paused with goals open.
  • Type entries use invariants in place of refinements (same per-clause map).

5. tests and oracles

"tests": [
  { "name": "odd-length#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
  { "name": "ensures lower bound", "tier": "property", "mode": "sandboxed", "origin": "derived", "result": "pass" },
  { "name": "real-read#1", "tier": "cram", "mode": "real", "origin": "spec", "result": "pass" }
],
"oracles": [
  { "kind": "reference", "path": "ref/stats.py::median", "hash": "sha256:5e8f0b2a" },
  { "kind": "definition", "name": "sort_by", "formal_hash": "sha256:aa90b1f3" }
]
  • name: block-name#k for expect and cram rows (k 1-based, even when the block has a single case — single#1 — so a case added to a block never renames an existing row); the bare block name for property rows, which are one row per block by construction. (Clarified 2026-08-28, step 7: earlier prose said single-case blocks drop the #k, but the checked-in examples — this section’s own example included — never did, and the always-#k form is the one with stable names under case addition.) Derived tests are namespaced (resolved 2026-08-24 with impl plan 03 §9.5): derived:ensures:<label> / derived:requires:<label> / derived:invariant:<label> for clause-derived properties, derived:differential:<reference-line> for the reference tier, derived:contract:<n> reserved for FFI contract tests (plan 06). The derived: prefix can never collide with a spec block name; chosen over bare clause labels for exactly that reason. (The pre-daemon example locks used ad hoc names; the regeneration pass rewrites them.)
  • tier: the test lattice (design §4.5): expect | property | cram | contract | differential | proof.
  • mode: sandboxed (fakes, runnable by the lowerer’s run_tests) | real (cram; never available inside the lowering sandbox). New modes can be added without a format change (design §4.5).
  • origin: spec (a block in the .tr) | derived (auto-generated: properties from refinement clauses per design §4.1, contract tests for bindings per design §4.5, invariant properties for types).
  • result: pass | fail | xfail (expected failure, failed) | xpass (expected failure, passed — needs spec attention). Failure details (counterexample, diff, and the blame field — caller for a requires violation, callee for ensures/invariant, design §6.4) are not stored here; they live in f.log and the IDE. Results are cache entries keyed by the hashes (design §6.1), so they churn only when inputs do.
  • mutation (reserved 2026-08-23, design §4.5; written by the post-v1 trellis mutants job): { "score": 0.87, "survivors": n } — the test-strength record, reservable now so its arrival is not a format change. Survivor details live in f.log and the IDE’s mutant-to-test flow.
  • oracles: what the tests depend on beyond the definition itself — the reference implementation (by file hash), accepted definitions used as oracles (by formal_hash), and CLI oracles (by executable hash): { "kind": "cli", "command": "soil0 parse", "hash": "sha256:…" }. A changed oracle re-runs dependent tests (design §4.5). v1 gap (resolved 2026-08-28, impl plan 03 step 7): the daemon records reference and cli rows only. Definition-oracle rows for user definitions referenced by clause and property predicates — and the acceptance rule design §4.5 attaches to them — wait for the prelude (plan 04), which will absorb the helper vocabulary such predicates lean on today (len, sort, the csvstats predicate suite) and shrink the honest edge set to something worth enforcing. The schema row is unchanged; only the writer is deferred.

6. FFI bindings

A binding is an ordinary function entry plus an ffi section (design §3.11, §6.3, §8):

"ffi": {
  "trust": "contract",
  "symbol": "requests.get",
  "symbol_hash": "sha256:be77a0c1",
  "package": "/nix/store/a1b2…-python3.12-requests-2.32.3"
}
  • trust: declared-only | contract | harvested.
  • symbol_hash: hash of the symbol’s machine-readable declaration (.pyi stub, cargo doc JSON, C header decl) — the interface record, designed so symbol-level invalidation can slot in later (design §8).
  • package: Nix store path — the provenance record.

7. The derived manifest (soil.lock)

A pure merge of the sidecars — never separately maintained (design §6.3) — plus nodes and aggregates that only exist globally:

{
  "lock_format": 1,
  "definitions": { "csvstats/median": { …sidecar contents… } },
  "soil_private": [
    { "name": "csvstats/_parse_cell", "hash": "sha256:3fe210bb", "owners": ["csvstats/parse_row"] }
  ],
  "escape_hatches": {},
  "trusted_packages": { "prelude": "sha256:e0a1b2c3", "soil-rs-std": "sha256:f1b2c3d4" }
}

soil_private nodes are derived from every lowering.private_helpers list (greyed out in the IDE; many helpers + few definitions is a flagged smell, design §4.2). escape_hatches is the audit view. trusted_packages pins the prelude and batteries by full package hash (design §5).

8. Derived status and invariants

The typed / tested / verified / accepted ladder (design §4.7) is derived, never stored:

  • typed — checks.types == "ok" and checks.holes == 0.
  • partial — checks.holes > 0 (design §3.16): the lowering paused with open goals; never tested, never accepted, excluded from release targets.
  • tested — typed, and every test result is pass or xfail.
  • verified — tested, and every checks.refinements clause is proven (unreachable when the map is empty; only tested means correct, design §4.1).
  • accepted — the stored flag.

Module and project badges are the minimum over exported definitions’ statuses and clause assurances (design §7.1) — derived by the IDE, never stored.

Decision staleness (2026-08-24) is likewise derived, never stored: compare lowering.decisions against the module/project entries’ current hashes (§3). It surfaces alongside review-suggested and does not block the ladder — tests remain the trust root; the triage sweep and the editorial reclassification (tr-grammar §5.2) are the clearing paths.

Invariants the daemon enforces:

  1. accepted may be true only if tested holds and no result is xfail or xpass (design §4.5) — resolving an xfail is a spec change, which clears accepted via the hash rules below. This gate applies to function entries; module and project entries carry accepted ungated (they have no tests — §9, resolved 2026-08-24), and type entries are accepted-ungated in v1 (resolved 2026-08-28): their only possible tests are invariant-derived properties, which the v1 runner refuses (tr-grammar §4.2 toolchain note — they are constrained generation in disguise), so the invariant stays visible as a runtime assurance in checks and the gate tightens automatically when plan 05 makes invariant properties runnable.
  2. A human-authored block may not be rewritten by the agent; the lowerer can only ask_human.
  3. Only accepted definitions may appear in another entry’s oracles.
  4. A definition with checks.holes > 0 cannot be tested or accepted and never reaches a build target (design §3.16).

Invalidation on change:

ChangedEffect
spec.hashes.formallowering invalid; re-lower or re-verify; callers follow via calls[].hash
spec.hashes.testre-run tests; re-lower if failing
spec.hashes.proseprose_state: "review-suggested"; lowering stays valid
a decision entry’s textciting lowerings become decision-stale (derived from lowering.decisions.applied); cleared by editorial reclassification, triage confirmation, or re-lowering
a decision entry added or removedevery in-scope lowering’s scope_hash mismatches — flagged once; a removal may be reclassified editorial
a decision edit reclassified editorial by the humanrecorded hashes re-stamped; nothing flagged
an oracle’s hashre-run the dependent tests only
Soil hand-editsoil_hash updated, provenance: "hand-edited", re-lowering skipped until spec changes
versionsnothing, until the spec changes (design §7.1 upgrade policy)

9. Open questions raised by this prototype

All three questions raised by the first draft were resolved 2026-08-24 with impl plan 03 (§9.5, §9.6):

  • Module and project entries carry accepted — a human can meaningfully accept an export list and a decisions block — but CI policy and badges continue to key on exported function definitions, so nothing changes downstream. Module/project entries have no tests, so their accepted has no derived gate (invariant 1 in §8 applies to function entries).
  • The derived naming scheme is the derived: namespace (§5).
  • The prelude-fork hash stays global in soil.toml’s trusted_packages: per-entry pins would repeat one hash everywhere until per-entry forks exist, and the toolchain pin (design §8) already refuses to lower or verify under a changed prelude.

Soil Surface Syntax — Prototype Highlights

Status: prototype, highlights only — the full grammar comes with the Soil core (design §11, milestone 1). Haskell for the type-level look, OCaml for the term-level look, minimal everywhere the design tenets demand it. First sample: examples/read_file.soil.


1. Definitions

Haskell-style signature line, then one equation. Curried, strict. The signature is mandatory in .soil — the file must check standalone, and the daemon verifies it entails the .tr spec type (design §4.8).

read_file : (fs : Fs) -> (path : Path) -> io (Result Utf8 FsError)
read_file fs path = ...

Named parameters (x : T) are optional except where refinements refer to them.

2. Effect rows

Space-separated, before the return type; the empty row is total.

io (Result Utf8 FsError)
ffi panic io PyObject

3. Terms

OCaml: let … in, let rec … and … (free within a file, design §3.8), match … with, fun x -> e, if/then/else. No ; — sequencing an effect is let _ = log clock msg in …. Let bindings are parameterless: a let-bound function is an explicit lambda (let go = fun x -> … in). Boolean connectives are the words and/or/not, the same spelling as the predicate language. One way to do things.

4. Pattern matching

Variant patterns bind the payload; record payloads destructure by name with punning. Exhaustiveness is enforced. _ is the wildcard. No guards — nested if/match instead.

match parse_cell text with
| Ok row                  -> ...
| Err { index, text = t } -> ...

5. Records

Construct with =, access with ., functional (non-mutating) update with with:

let r = { cells = xs } in
let r2 = { r with cells = ys } in
r2.cells

6. Sums

Nullary variants are bare; a variant carries at most one payload of any type, and multi-field payloads are inline records (no tuples): None, Ok bytes, BadCell { index = 1, text = "x" }.

Result a e puts the success type first (OCaml/Rust order). Haskell’s error-first Either e a exists so the partially applied constructor can be a Functor instance — impossible in Soil (no type classes, no higher-kinded abstraction), so the widely known order wins.

7. Refinements

Inline in .soil signatures, Liquid-style — the agent-facing spelling that the .tr’s requires/ensures clauses desugar into:

median : (xs : {v : List F64 | len v > 0}) -> {r : F64 | min xs <= r and r <= max xs}

8. Termination

A decreases line between signature and equation when structural decrease is not inferable (Idris-style measure, design §3.4):

gcd : U64 -> U64 -> U64
decreases b
gcd a b = if b == 0 then a else gcd b (a % b)

9. Comparison operators are notation, not overloading

Every type has exactly one derived eq/compare (design §3.7), so the elaborator rewrites x == y to T::eq x y at the inferred monomorphic type. In polymorphic code the operators are unavailable — take the function as a parameter (sort_by, map comparators), the confirmed idiom (design §3.6).

10. Names and modules

No import statements; the daemon resolves names through the manifest, and the lock’s import set is computed, never written. Same-module definitions and the prelude are bare; cross-module exports are qualified (csvstats::median); derived functions are Row::eq; private helpers are underscore-prefixed and live only in _private.soil. :: is the namespace separator, keeping . exclusively for record field access.

11. Literals and comments

1 is I64, 1.0 is F64, "…" is Utf8; other widths by annotation (42 : U32), no suffixes. Comments are --.

12. Deliberately absent

Tuples, guards, ;, do-notation, exceptions, mutation, type classes, operator sections, user-defined operators, parameterized let bindings.

Deferred, not rejected: a let? x = e in … sugar for Result propagation. v1 writes the match explicitly; if the corpus shows it is the dominant noise, the sugar is one desugaring rule later.

Soil Surface Syntax — Elaborated Specification

Status: prototype. Elaborates soil-syntax.md into a lexical spec, EBNF, and static rules. Decisions folded in from review: OCaml-style match (no terminator, parenthesize non-tail nested matches), word connectives (and/or/not) shared with the predicate language, shadowing forbidden, :: for namespaces with . reserved for field access, .. required in partial record patterns, parameterless let bindings (functions are explicit lambdas), and (2026-08-22, surfaced by impl plan 02) constructor qualification Type::Ctor with the exactly-one-spelling rule (§5.9). Semantics (typing, effect, and refinement rules) arrive with the Soil core milestone (design §11); this document is the parser’s contract.


1. Lexical structure

  • Source is UTF-8. Whitespace separates tokens and is otherwise insignificant — there is no layout rule.
  • Comments run from -- to end of line. No block comments.

Identifiers

ClassFormUsed for
ident[a-z][a-z0-9_]*values, parameters, fields, type variables, row variables, module names
private-ident_ identmodule-private definitions (only in _private.soil)
TypeName[A-Z][A-Za-z0-9]*types and constructors (one class; constructors live in the type’s namespace)

Keywords

let  rec  and  or  not  in  fun  match  with  if  then  else  decreases

Effect names are reserved in type position only: div, panic, io, ffi. The boolean connectives are the keywords and/or/not — the same spelling as the predicate language, so specs and code read identically (implies remains predicate-exclusive). There are no boolean literals: Bool is the prelude sum True | False (see §6 for its JSON special case).

and serves both as the mutual-binding connector and the boolean connective; see §3.3 for the disambiguation rule.

Operators and punctuation

->  =  |  :  ::  .  ,  ..  ( )  { }
==  !=  <  <=  >  >=  +  -  *  /  %  _

? appears only as the head of a typed hole: ? immediately followed by an ident lexes as one hole token (?rest); a bare ? is an error (§5.11, added 2026-08-23).

Literals

  • Integer: [0-9][0-9_]*, default type I64. Other widths by annotation ((42 : U32)) or by context — an integer literal adopts the type inference demands (“as in Rust”, design §3.6), so b == 0 checks with b : U64; the default applies only when unconstrained. No suffixes. A literal is range-checked against its resolved width at check time (impl plan 02 §8.14, amended 2026-08-23 — the fully-fixed-at-lex reading would have broken the normative gcd example); since negation is an operator, I64::MIN is not writable as a literal (the C/Rust wart, accepted — a prelude constant covers it).
  • Float: digits . digits, with an optional exponent — lowercase e, an optional - (no +, no E), digits — as in 1.5e3, 1.5e-3. Default type F64. Underscores are permitted in every digit run of numeric literals and are not part of the value. Lexing extends a numeric literal only when what follows can continue it: 1..2 lexes as 1 .. 2, and 1.5e as 1.5 followed by the identifier e. (Edges resolved 2026-08-22 with the soil0 lexer.)
  • String: "…", default Utf8. The escape set is exactly \", \\, \n, \r, \t, and \u{hex} with one to six hex digits denoting a Unicode scalar value (surrogates U+D800–U+DFFF and values above U+10FFFF are errors — a Utf8 value can never hold them). Any other character after \ is an error; there are no octal/hex byte escapes (Bytes are built by prelude functions, not literals) and no line-continuation escapes. Raw characters are unrestricted: any character other than " and \ — including newlines and non-ASCII — stands for itself, so strings may span lines; a string is unterminated only at end of file. (Resolved 2026-08-22 with the soil0 lexer.)
  • No character literals; no Bytes literals (construct via prelude functions).
  • Negation is the unary operator, not part of the literal.

2. Files

  • <name>.soil holds exactly one definition; its name equals the filename stem and the .tr frontmatter name.
  • _private.soil holds any number of private-ident definitions (design §4.2).
  • .soil files contain no type declarations — type shapes live in the .tr soil-type blocks and the compiler materializes them — and no imports: names resolve through the manifest (design §6.1), and the lock’s import set is computed.

3. Grammar

Notation: { x } is zero-or-more, [ x ] optional, | alternation, terminals quoted.

3.1 Definitions

soil-file      ::= definition
private-file   ::= { definition }

definition     ::= signature [ decreases-line ] equation
signature      ::= defname ":" type
decreases-line ::= "decreases" expr
equation       ::= defname { ident } "=" expr
defname        ::= ident | private-ident

Equation parameters are bare identifiers — destructuring happens in the body. The defname of the signature and equation must agree.

3.2 Types

type       ::= [ dom "->" ] cod                      (right-assoc arrows)
dom        ::= "(" ident ":" type ")" | app-type | refinement
cod        ::= [ row ] type
row        ::= effect { effect } [ ident ]           (trailing ident = row variable)
             | ident                                  (row variable alone — only before a type)
effect     ::= "div" | "panic" | "io" | "ffi"
app-type   ::= TypeName { atype } | atype
atype      ::= TypeName | ident | refinement | "(" type ")"
refinement ::= "{" ident ":" type "|" predicate "}"
  • Quantification is implicit and prenex: free lowercase type variables are universally quantified. There is no forall.
  • Row parsing is unambiguous without HKT: type variables have kind * and are never applied, so in a -> e b and a -> e (List b) the leading lowercase ident followed by another type can only be a row variable. A lone ident after -> is the return type. An absent row is the empty row (total).
  • predicate is the shared predicate language of the .tr grammar (tr-grammar §2.3): refinements are the spec language embedded in Soil, desugaring one-to-one from requires/ensures clauses. Its connectives (and/or/not) are the same words the term layer uses, so there is exactly one spelling everywhere; implies exists only in predicates.

3.3 Expressions

expr      ::= "let" [ "rec" ] binding { "and" binding } "in" expr
            | "fun" ident { ident } "->" expr
            | "if" expr "then" expr "else" expr
            | "match" expr "with" arms
            | or-expr

binding   ::= ( defname | "_" ) "=" expr
arms      ::= "|" arm { "|" arm }
arm       ::= pattern "->" expr

or-expr   ::= and-expr { "or" and-expr }
and-expr  ::= not-expr { "and" not-expr }
not-expr  ::= "not" not-expr | cmp-expr
cmp-expr  ::= add-expr [ cmpop add-expr ]            (non-associative)
cmpop     ::= "==" | "!=" | "<" | "<=" | ">" | ">="
add-expr  ::= mul-expr { ("+" | "-") mul-expr }
mul-expr  ::= unary { ("*" | "/" | "%") unary }
unary     ::= "-" unary | app
app       ::= atom { atom }                          (left-assoc application)

atom      ::= literal
            | path
            | qualified
            | TypeName                               (constructor)
            | record
            | hole
            | "(" expr [ ":" type ] ")"

hole      ::= "?" ident                              (typed hole, §5.11; design §3.16)

path      ::= (ident | private-ident) { "." ident }  (variable + field projections)
qualified ::= ident "::" ident                       (module::def)
            | TypeName "::" ident                    (Type::derived)
            | TypeName "::" TypeName                 (Type::Ctor — §5.9)
record    ::= "{" [ path "with" ] field { "," field } "}"
field     ::= ident "=" expr
  • Bindings carry no parameters — a let-bound function is an explicit lambda: let go = fun x -> … in …. One way to write a function.
  • let _ = e in … discards the result — the sequencing idiom for effects. _ binds nothing and is exempt from the no-shadowing rule; it is not permitted in let rec.
  • The and disambiguation: after and, the two-token sequence defname "=" begins a new binding of the enclosing let rec; anything else makes and the boolean connective. This is unambiguous because = never occurs in expressions (equality is ==) and bindings are parameterless.
  • An arm’s body extends as far as possible; a subsequent | belongs to the innermost open match. A nested match in non-tail position must be parenthesized (the OCaml rule, chosen deliberately).
  • Constructor application is ordinary application with a TypeName head: Ok bytes, BadCell { index = 1, text = "x" }, nullary None — or a qualified head Result::Ok bytes exactly when the variant name collides in scope (§5.9).
  • Record update bases are paths, not arbitrary expressions: { r with cells = ys }.
  • (e : type) is a local annotation, the only way to give a literal a non-default type.

3.4 Patterns

pattern    ::= "_" | ident | literal
             | ctor-name [ pat-atom ]
             | record-pat
             | "(" pattern ")"
pat-atom   ::= "_" | ident | literal | ctor-name | record-pat | "(" pattern ")"
ctor-name  ::= [ TypeName "::" ] TypeName            (qualification per §5.9)
record-pat ::= "{" fieldpat { "," fieldpat } [ "," ".." ] "}"
fieldpat   ::= ident [ "=" pattern ]                 (bare ident = punning)

A record pattern must name every field of the record type unless it ends with .. — omitting fields silently is an error, so adding a field to a type breaks exactly the patterns that need reviewing.

Patterns nest arbitrarily. There are no or-patterns, no guards, and no as bindings (nested match and fresh lets instead).

4. Precedence

Tightest to loosest:

  1. field access .
  2. application (juxtaposition), ::
  3. unary -
  4. * / %
  5. + -
  6. == != < <= > >= (non-associative — a < b < c is a parse error)
  7. not
  8. and
  9. or
  10. if / fun / let / match bodies

This ladder is the predicate language’s ladder (tr-grammar §2.3) with the arithmetic tiers inserted below the comparisons and implies absent.

5. Static rules

  1. No shadowing. Binding a name already in scope — by let, fun, an equation parameter, or a pattern — is an error. Fresh names only.

  2. Exhaustive matches, and redundant arms are errors.

  3. Effect subsumption: f may call g iff g’s row ⊆ f’s row; ffi implies panic (design §3.3).

  4. Operators are notation, not overloading (design §3.6, §3.7). The elaborator rewrites at the inferred monomorphic type; in polymorphic position the operators are unavailable and the function is taken as a parameter.

    SurfaceElaborates to
    == !=T::eq (negated for !=)
    < <= > >=T::compare
    + - * / %, unary -per-type numeric primitives (I64::add, F64::div, …)
    and orshort-circuit builtins on Bool
    notbuiltin on Bool
  5. Arithmetic obligations: overflow, and a zero divisor for / and % on integers, are refinement obligations (design §3.2, §3.12). If SMT discharges the obligation the operation is total; otherwise the enclosing function’s row acquires panic. There is no panic syntax; explicit panics are a prelude function.

  6. Integer division is floor division. On integer types / rounds toward negative infinity and % is the matching floor modulus (the result carries the divisor’s sign), preserving (a / b) * b + a % b == a. These are Python’s semantics — the reference-implementation language — not C/Rust truncation, so differential tests agree without adjustment. I64::MIN / -1 is an overflow obligation like any other. On F64 the operators are IEEE 754.

  7. Termination: a recursive definition needs structural decrease or a decreases measure; failing both, its row acquires div (design §3.4).

  8. Derived functions are reached by qualification: Row::eq, Row::show, Row::compare, Row::hash (design §3.7).

  9. Constructor resolution has exactly one legal spelling. A bare constructor resolves iff its variant name is unique among the sum types in scope, and the bare form is then the only legal form; when two types in scope share the variant name, the qualified Type::Ctor form is required. Qualifying a unique constructor is an error. One spelling per context (design §3.15); the error on a collision names the candidate types. The rule applies identically in expressions, patterns, and the predicate language’s is tests (tr-grammar §2.3).

  10. Definition names follow the same exactly-one-spelling rule (resolved 2026-08-22, design §3.15). A bare name resolves iff it is unique across the visible definition set (the manifest, design §6.1) plus the builtins — so the prelude is called bare — and bare is then the only legal form; when two modules define the name, the qualified module::def form is required, and qualifying a unique name is an error. Private definitions are visible only within their own module and cannot be qualified (:: takes a plain ident on the right; private-idents never cross modules).

  11. Typed holes (adopted 2026-08-23, design §3.16). ?name is an expression of any type; hole names are unique within a definition (a repeated hole name is an error). A definition containing holes checks, with each hole’s goal type reported; it is partial — never testable, never accepted, never built, and run/test refuse it (unfilled-hole). Holes are a lowering-time state.

6. Interaction with the value encoding

Bool is a prelude sum type but encodes as JSON true/false, not {"tag": "True"} — the one special case in the sum encoding, matching what every host expects (recorded in tr-grammar §7).

7. Worked examples

The checked-in sample (examples/read_file.soil):

read_file : (fs : Fs) -> (path : Path) -> io (Result Utf8 FsError)
read_file fs path =
  match fs_read_bytes fs path with
  | Err e -> Err e
  | Ok bytes ->
    match utf8_decode bytes with
    | Ok text -> Ok text
    | Err _ -> Err (NotUtf8 { path = path })

A measure, symbolic operators, and notation elaboration:

gcd : U64 -> U64 -> U64
decreases b
gcd a b = if b == 0 then a else gcd b (a % b)

Row polymorphism, a lambda, and qualification:

sum_lengths : (rows : List Row) -> I64
sum_lengths rows =
  fold (fun acc r -> acc + len r.cells) 0 rows

8. Open questions raised by this spec

All questions raised by the first draft have been resolved and folded in: partial record patterns require .. (§3.4); the term layer uses the word connectives, unifying with the predicate language (§1, §3.3); the string escape set is fixed and scalar-value-only (§1); let bindings are parameterless and functions are explicit lambdas (§3.3).

Constructor-name ambiguity, surfaced by impl plan 02 (its §9.1), was resolved 2026-08-22: Type::Ctor qualification with the exactly-one-spelling rule (§3.3, §3.4, §5.9); design §3.15 records the rationale and the rejected alternatives.

Canonical text form: adopted 2026-08-23 and specified in §9 below (soil0 impl step 11); soil_hash is computed over the canonical text of the alpha-normalized AST (design §6.1).

9. Canonical form

Added 2026-08-23 (impl plan 02 step 11; design §4.8). Soil has exactly one printed form per AST: the printer is a compiler pass, soil0 print emits it, parse → print is a fixpoint, and the lowerer’s write_soil canonicalizes — the agent never controls formatting. Layout is structural and width-independent: no rule consults line length. The checked-in examples are the style oracle — printing them is byte-identity, which is a conformance test.

  • Definitions: the signature on one line (name : type); the decreases line; the equation head name p1 … = with the body inline on the same line when it is not a block, else on the next line at indent 1. Definitions in _private.soil are separated by one blank line. Indentation is two spaces per level; no trailing whitespace; the file ends with one newline.
  • Blocks are let, if, and match; they lay out multi-line in statement positions (definition bodies, let bodies, arm bodies, then/else operands) and single-line inline everywhere else (parenthesized argument positions).
  • let at indent k: let [rec] name = value in on one line when the value is inline; the body follows at indent k. A block-valued or block-bodied-lambda binding prints its header (let f = fun a b ->), the value block at k+1, then in alone at k, then the body at k.
  • if at k: if cond, then X, else Y on three lines at k; a block operand moves to k+1 on the following line.
  • match at k: match scrutinee with then one | pattern -> body line per arm at k; a block arm body moves to k+1.
  • Parentheses are minimal by precedence (§4): emitted only where reparsing would change the tree — plus one style rule from the examples: an applied type constructor after an effect row is parenthesized (io (Result Utf8 FsError)).
  • Spacing: single spaces around binary operators, : in signatures and refinements, = in bindings and record fields, and | in refinements; { a = 1, b = 2 } record spacing; x.f and m::f unspaced; unary - attached — except that a negated negation prints -(-x), since attached -- would lex as a comment (a consequence of the minimal-parens rule; found by the printer’s fixpoint property, 2026-08-26).
  • Literals: integers and floats print their underscore-free source text; strings escape exactly \", \\, \n, \r, \t, and lowercase \u{…} for remaining control characters, all other characters raw.

9.1 Hash form

Added 2026-08-24 (impl plan 03 §9.2; design §6.1, lock-schema §3): the text soil_hash is computed over. The canonical form of §9 is the display form; the hash form is the same text with local names alpha-normalized, so renaming a local binder can never move a hash.

  • The hash form is the §9 canonical text with every local binder — equation parameters, let/let rec binders, fun parameters, and pattern binders — replaced by %N, where N numbers binding sites 0, 1, 2, … in pre-order of appearance within the definition; every occurrence of a bound name prints its binder’s %N. _ binds nothing and stays _.
  • Everything at definition level or above is untouched: the definition’s own name, callee and private-helper names, type names, constructors, field names, builtin names, and hole names (?name is a named goal; renaming it is a visible change to a lowering-time state that never reaches tested).
  • The signature line keeps its declared parameter names — they are spec surface: the .tr signature and its requires/ensures predicates name them. Consequently renaming a parameter is a visible spec change that moves formal_hash and soil_hash together; only let, fun, and pattern binder renames are invisible (clarified 2026-08-28, with the step-4 hash tests).
  • The hash form is not reparsable (%N is not in the grammar) and is not a CLI surface: it is produced by a soil0 library function the daemon calls. The CLI contract’s print remains the display form; soilc’s printer (plan 05) is differentially tested against print, with the hash transform one shared implementation above it. Exposing a print --hash-form flag would be a future additive contract amendment; none is planned.

soil.toml — Prototype (v1 subset)

Status: prototype — approved 2026-08-25 (drafted as impl plan 03 step 1 with docs/contracts/trellis-daemon.md; §4 questions resolved and the whole reviewed with the user; prototype status like tr-grammar and lock-schema — exercised by plan 03, finalized with them). soil.toml is the only non-derived global file in a Soil root (design §7.2), but no document specified it; this prototype pins the subset the daemon needs — the toolchain pin, the tag vocabulary, the entrypoint — and reserves the rest of design §8’s surface as named future sections. Prototype status like tr-grammar and lock-schema: exercised by plan 03, finalized with them. References: design §7.2 (layout), §8 (build system, toolchain pin), §4.3 (tags), §3.14 (main), §5 (trusted packages), tr-grammar §1 (tag validation), impl plan 03 §4.


1. Role and location

A Soil root is any directory with a soil.toml at its top; every .tr/.soil/.lock under it belongs to that root, and roots never import across each other (design §7.2). The CLI discovers its root by walking up from the working directory, unless TRELLIS_ROOT is set — then that directory is the root (and must hold a soil.toml, else config-no-root). The env override exists for processes running outside the tree that still belong to it: the cram runner (impl plan 03 step 8, resolved 2026-08-28) exports it so a transcript in its fresh temp dir can invoke trellis call back into the root. The file is TOML, written by the human (and by trellis toolchain update, which rewrites exactly one table). The derived soil.lock manifest sits beside it, gitignored.

Parsing is strict, matching the .tr frontmatter rule (tr-grammar §6): an unknown table or key is an error (config-unknown-key), so typos fail loudly instead of silently configuring nothing.

2. v1 surface

entrypoint = "main"            # optional

[toolchain]
trellis = "0.1.0"
soil0_cli = 1
# skill = "sha256:…"           # written by `trellis toolchain update`
                               # once `trellis skill` exists (§2.1)

[tags]
api = "public surface; CI policy keys on this"
parser = "the parsing subsystem"

2.1 [toolchain] — the pin

Refuse-on-mismatch (design §8, adopted 2026-08-23): any command that lowers or verifies compares this table against the running binary and refuses with the structured toolchain-mismatch diagnostic when they disagree — never silently writing lock entries that claim more than they should. trellis toolchain update is the explicit upgrade event: it rewrites the table to the running toolchain’s identity and is the natural trigger for a re-verification sweep (design §7.1).

  • trellis — the required trellis crate version (the daemon and CLI are one binary, so one version covers both).
  • soil0_cli — the soil0 contract major (soil0 --version’s soil0_cli); builtins and error codes key on it.
  • skill — the hash of the generated skill bundle (trellis skill, design §4.6). Optional only while no generator exists (impl plan 03 steps 2–11, which run on a hand-carried interim skill): during that window a missing pin warns. Once trellis skill ships (step 12), a missing skill under a generator-capable toolchain refuses like any other mismatch (resolved 2026-08-25, §4.1): the skill is a lowering input under the design §1.4 tenet, so an unpinned one is exactly the “exists only in an agent’s context” hole. trellis toolchain update writes and maintains it — the fix is one command.
  • The prelude and batteries hashes land in [trusted_packages] with plan 04 (§3); the pin table stays about the toolchain.

Read-only commands (status, check of the static pipeline, context) run under a mismatch — inspecting a foreign-pinned root must be possible — but anything that writes a lock refuses.

2.2 [tags] — the vocabulary

The per-project tag vocabulary of tr-grammar §1: keys are the legal tags: values in .tr frontmatter (a .tr tag not declared here is the tr-undeclared-tag error); values are one-line descriptions the IDE shows. Tags are non-semantic — graph filtering and CI policy — and never enter a context bundle.

2.3 entrypoint

Optional; names the definition trellis build will treat as main (an ordinary io definition with a World argument, design §3.14). Parsed and validated (the definition must exist when set); unused until build targets (plan 06+). Recorded now so the key’s shape is settled before anything depends on it.

3. Reserved sections (named, unparsed until their milestone)

Declaring one of these before its milestone is an error (strictness rule, §1) — the names are reserved so their arrival is additive:

TableArrivesContent (design ref)
[trusted_packages]plan 04prelude and batteries pinned by full package hash (design §5); merged into soil.lock
[deps]plan 06+dependency declarations compiled to Nix (design §8); no raw-Nix escape field, ever
[build]plan 06+build targets: binary, py_module, shared_lib, … (design §7.1)
[ci]plan 06+policy lines (“every definition tagged api must be accepted”, mutation-score floors — design §4.7, §4.5)
[providers]never (see below)—

Provider configuration (models, budgets, caps, credentials) is deliberately not in soil.toml: it is user-level, not project-level, and credentials must never sit in a committed file. It lives in ~/.config/trellis/providers.toml (impl plan 03 §8.9, non-contractual).

4. Open questions raised by this prototype — resolved 2026-08-25

  1. The skill bootstrap leniency (§2.1). Resolved: required once the generator exists — from step 12 a missing skill pin is toolchain-mismatch (an unpinned skill is a lowering input living outside a hashed file, the design §1.4 hole); warning-only remains solely for the pre-generator window. Recorded in §2.1; the step-12 drift gate enforces it.
  2. [ci] shape. Resolved: deferred to its milestone (plan 06+): the table stays reserved-by-name — which already prevents ad hoc invention — and its schema is drafted when headless CI and build targets arrive, informed by real badge rendering and the mutants job’s actual fields.

Bootstrapping Plan: Rust Interpreter → Self-Hosted Compiler

Status: prototype plan. The Soil compiler (soilc) is written as a Trellis project — the first Trellis project — bootstrapped on a minimal Rust implementation (soil0) that is never thrown away. Decisions folded in from review: slow interpreted bootstrapping is accepted; the reference-oracle rule is amended to allow CLI oracles (design §4.5); the compiler is the first project with the Python-glue project second; the backend emits Cranelift CLIF, not C and not direct x86.


0. Premise

A compiler is the ideal Trellis dogfood: every pass is a pure, total function over trees, which is exactly what the JSON value model, expect tables, property tests, and differential testing handle best. The AST is a family of Trellis type definitions, so serialization (show/parse) and golden tests come for free. And the stage-0 implementation is not scaffold but permanent infrastructure: an independent implementation to differentially test the real compiler against, forever.

The bias to keep in view: compiler code is the easiest kind of Trellis code (pure, total, no FFI, no capabilities). It validates the language core and the lowering workflow; it validates nothing about bindings, capabilities, or trellis bind. Hence the Python-glue project stays as the second project.

1. Stage 0 — soil0 + soil-rt (hand-written Rust)

One workspace, two crates:

  • soil-rt — the runtime, never bootstrapped away (design §3.10): values, reference counting, the JSON bridge, C ABI from the first commit, later pyo3.
  • soil0 — parser, ML type + effect-row inference, exhaustiveness, tree-walking interpreter, test runner. No refinements, no termination checker, no codegen. Deliberately the size of a class project.

The load-bearing requirement: every pass is a JSON-in/JSON-out CLI command — soil0 lex, soil0 parse, soil0 infer, soil0 run. These are the future differential oracles. Slowness is explicitly fine; the whole bootstrap runs interpreted.

2. Stage 1 — Trellis toolchain on the interpreter

Design §11 milestones 2–3 unchanged in substance: daemon (incremental state, hashing/locks, context bundles, MCP tool surface, lowering jobs, agent-CLI provider), then the pure-core prelude as the first Trellis code. Soil execution is soil0 interpretation throughout. At the end of stage 1, Trellis is a working language whose execution engine happens to be an interpreter.

3. Stage 2 — soilc: the compiler as the first Trellis project

Each pass is a module of .tr specs, agent-lowered, running interpreted. Pass order, chosen so each has its oracle before it is written:

  1. Lexer — warm-up; calibrates the spec-and-lower workflow.
  2. Parser — deliberately early: the predicted worst case for the mutual-recursion ban (design §3.8), written as one definition with let rec … and … locals. Stress-tests the language design while it is still cheap to change.
  3. Renamer / scope checker — no-shadowing, :: qualification.
  4. Type + effect inference.
  5. Exhaustiveness + pattern compilation.
  6. ANF lowering (the mid-end, design §3.1).
  7. Termination checker — new functionality, no soil0 counterpart; tested by spec only.
  8. Refinement checker — emits SMT-LIB text as a pure function; Z3 runs behind a new Solver capability (the Py pattern). Refinements and demotion therefore arrive as compiler passes, not a separate milestone.
  9. CLIF backend — see stage 3.

Strangler pattern: as each pass reaches accepted, the daemon swaps its soil0 pass for the Trellis one (shelling to soil0 run while interpreted). The soil0 passes are demoted to oracles, never deleted.

Testing: every pass gets JSON→JSON expect tables, properties (e.g. parse after print is identity), and differential tests against the matching soil0 command via CLI oracle. Passes 7–8 are spec-only.

4. Stage 3 — closing the loop

The backend is a pure Trellis pass ANF → CLIF text (golden-testable like every other pass), plus a small Rust driver in the workspace that feeds Cranelift for isel/regalloc/object emission — x86-64 and arm64, no C toolchain.

Why not direct x86: instruction selection, register allocation, and ELF emission are the largest and least testable chunk of a native backend, platform-locked, with zero free optimization. ANF is already shaped like portable three-address code; emitting a textual SSA IR outsources exactly the bad part. Direct emission remains a deferred independence move, not a foreclosed one.

Closure:

  1. soilc (interpreted on soil0) compiles the prelude and itself → native soilc₁.
  2. soilc₁ compiles the same sources → soilc₂.
  3. Fixed point: soilc₁ and soilc₂ are byte-identical. Soil is unusually well-positioned for this — no mutation, canonical JSON, comparator-ordered maps, content addressing — so determinism is the default, but the fixed-point test is what enforces it.

5. Permanent division of labor

Stays Rust foreverBecomes Trellis
soil-rt (runtime, C ABI, pyo3)all compiler passes
the Cranelift driverSMT-LIB generation
the daemon’s process/IO shellpure daemon logic later (hashing, lock manipulation)
Z3 itself
soil0 — permanent differential oracle and debug-mode engine

6. Mapping to the build order

Design §11 is reordered (recorded there): soil0+soil-rt, daemon, prelude, soilc as first project (through self-hosting closure), then trellis bind + FFI batteries + minimal IDE, then the Python-glue project as second project — it validates the FFI/capability/bind half of the pitch that the compiler cannot touch.

7. Remaining risks

  • Biased dogfood: the compiler proves the pure core, not the FFI story; mitigated only by actually doing the second project.
  • Fixed-point determinism: any iteration-order or fresh-name nondeterminism in lowered passes breaks stage 3; the renamer must allocate names canonically.
  • Type-inference pass size: the hardest single lowering target in the plan; if agent lowering strains anywhere, it is there — split the module aggressively.

Trellis-prose: Design Document

Trellis-prose is a sibling of Trellis that lowers a human-written sentence plan into finished prose. The human writes directives — the content and the rhetorical moves — and an agent writes the sentences.

Status: draft, design phase. This document records the decisions made so far with their reasoning, in the style of docs/design.md. It is self-contained: nothing here changes Trellis. Worked examples live in examples/prose/. Decisions below marked (provisional) were drafted to complete the format and await user confirmation; they are listed in §8.


1. Vision

1.1 The core thesis

Trellis applies one idea to code: the human authors the specification, an agent authors the artifact, and a checker stands between them. Trellis-prose applies the same idea to writing. The human writes a sentence plan — a sequence of directives, each naming a rhetorical move and its payload — and an agent lowers the plan into polished full-sentence prose. The human’s attention goes entirely to content and structure: what each sentence says, where the paragraph breaks fall, which sentence is an example and which is a pivot. The agent’s job is wording.

The plan is to prose what a .tr file is to Soil: the layer where intent lives, hashed and diffable, so that “what the document says” can be reviewed and re-lowered independently of “how it is phrased.”

1.2 What is checkable about prose

Code lowerings are gated by a type checker and tests. Prose has no type checker, but the directive language is designed so that one structural property is mechanically checkable: the strict 1:1 mapping (§4). Each prose directive produces exactly one output sentence, in plan order, inside the paragraph the plan declares. The checker can verify paragraph count, per-paragraph sentence count, order, and verbatim headings without any judgment. Style conformance, which cannot be checked mechanically, is judged against a shared prose style guide (§5) and recorded in the lock as a judgment, not a proof — the analog of an unproven refinement.

1.3 Design tenets inherited from Trellis

  • Nothing that affects the lowering exists only in an agent’s context. The plan, the surrounding notes, and the style guide are all hashed into the lock.
  • There is only one way to do something. The directive set is closed and small; a new move must pay for itself (§8.1). Verbosity for the human is preferred over ambiguity for the checker.

2. The .trp file (provisional container)

A .trp file is CommonMark with YAML frontmatter, mirroring the .tr container: formal content lives in one fenced block whose info string is the reserved word plan; everything else is prose notes — context for the agent (audience, purpose, background facts), hashed but never parsed.

Frontmatter keys (unknown keys are errors, as in .tr):

  • name (required): the document’s canonical name; equals the filename stem, lowercase snake_case.
  • style (required): relative path to the style guide (§5). Required rather than defaulted so the lowering’s full input set is visible in the file.

Exactly one plan block is required. Its contents are directive lines only: blank lines are permitted for human readability but carry no meaning (paragraphing is explicit, §3.2), and any non-blank line that does not parse as a known directive is an error.

The three-part hash, mirroring .tr:

  • plan_hash: the plan block. Changing it invalidates the lowering.
  • notes_hash: frontmatter and all surrounding prose. Changing it marks the lowering review-suggested rather than invalidating it.
  • style_hash: the resolved style guide file. Changing it marks every document that references it review-suggested (§5.2).

3. The directive language

A directive is one line: <move>: <payload>. The move names what kind of sentence this is; the payload carries the content in the human’s shorthand — sentence fragments, abbreviations, and bullet-point telegraphese are all fine, because wording is the agent’s job.

3.1 Prose moves

Each prose move renders exactly one sentence (§4).

MovePayloadRenders
sentence:the sentence’s contentone sentence asserting the payload
bridge to:an ideaone transition sentence connecting the preceding content to the idea
example:a concrete instanceone sentence presenting the payload as an example of the preceding claim
contrast:a counterpointone sentence pivoting against the preceding content; the agent supplies the “but / whereas / by contrast” framing
restate:an idea already made in this documentone sentence re-expressing that idea in fresh words
define:a term, optionally — glossone sentence defining the term

Reasoning for this particular set: sentence: and bridge to: are the two primitives — content the human dictates, and connective tissue the human cannot fully dictate because it depends on the wording the agent chose for the neighbours. The other four are the moves that a sentence: payload expresses badly: an example needs “for instance” framing relative to the previous sentence; a contrast needs the pivot word; a restatement must repeat earlier content, which is otherwise forbidden (the agent may not introduce claims absent from the plan, and restate: is the one move licensed to echo); a definition has a fixed rhetorical shape worth naming. Candidate moves that did not make v1 are listed in §8.1.

3.2 Structural moves

MovePayloadEffect
paragraph:optional topicopens a new paragraph; renders zero sentences. A topic payload is a cohesion hint to the agent and a target the judge can hold the paragraph to; an empty payload is allowed.
heading <level>:verbatim texta heading at the given level. The payload passes through verbatim — headings are structure, not wording, so the agent may not rewrite them. The level is always explicit (heading 1:, heading 2:); a defaulting rule would be a second way to say the same thing.

Structure is entirely explicit: blank lines in the plan are ignored, and a prose directive outside an open paragraph is an error (a heading closes any open paragraph). “Blank line = paragraph break” was considered and rejected: it reintroduces significant whitespace alongside the directive language, i.e. a second syntax. Agent-chosen paragraphing was rejected because it is uncheckable and would churn between lowerings.

3.3 What the agent may and may not do

The agent chooses wording, sentence-internal structure, connectives, pronouns, and rhythm. It may not add claims absent from the plan, drop or merge directives, reorder, or rewrite heading text. The output is plain paragraphs and headings — no lists, block quotes, or emphasis in v1 (§8.4).


4. The 1:1 mapping — the type check of prose

Decision: strict 1:1, order-preserving. Each prose directive renders exactly one output sentence; sentences appear in directive order inside the declared paragraphs.

Reasoning: this is the property that makes a prose lowering checkable rather than plausible. The checker verifies, mechanically:

  1. every plan line parses as a known directive (closed set; unknown move is an error);
  2. the output’s headings match the heading payloads verbatim, at the declared levels, in order;
  3. the output’s paragraph count equals the plan’s paragraph: count;
  4. each output paragraph’s sentence count equals its prose-directive count.

With counts and order verified, output sentence N of a paragraph is known to correspond to prose directive N — the lock needs no explicit map, the way a Trellis lock needs no line table.

Alternatives rejected:

  • Free counts (order only): loses the count gate; the only mechanical check left is heading text.
  • 1:1 with opt-in spans (sentence*:): keeps the gate but adds a second way to write a sentence. If one directive genuinely needs two sentences, the human writes two directives; that verbosity is the tenet working as intended.

Cost accepted: the checker needs a sentence segmenter, and segmentation of arbitrary prose is not exact. The rule is that the agent carries the burden: it must write prose the conservative segmenter parses unambiguously (e.g. avoiding sentence-medial abbreviations the segmenter would split on). A lowering the segmenter cannot parse cleanly is a failed lowering, exactly as Soil that does not parse is.


5. Style

5.1 A prose style guide, not a rule file

Decision: the shared style input is a free-prose style guide — a house style document (voice, register, preferences, banned habits) that every document’s frontmatter references. There are no mechanical style rules in v1.

Reasoning: mechanical rules (max words per sentence, forbidden-word lists, declared tense/person) were considered, both alone and as a hybrid rules-plus-exemplar file. They were rejected for v1 because the rules that are cheap to check are not the ones that make prose good, and a rule file grows into a second style guide that fights the first. One document, one voice, one place to edit it. The consequence is accepted openly: style conformance is judged, not proven. A judge agent reads the style guide and the lowering and returns pass/fail with cited violations; the lock records the verdict as a judgment. If judging proves unreliable, mechanical rules can be added later as a floor under the guide (§8.2) without invalidating the format.

V1 staging (decided 2026-09-14, plan P1): the first tool ships with the human as the judge — the mechanical checks gate, and setting accepted doubles as the style verdict, recorded in the lock as judge: human. An agent judge (in-session first, tool-invoked later) remains the intended end state; it is deferred so that v1 carries no judging machinery to trust before the workflow has been exercised.

5.2 Freshness

The style guide is hashed into every lock that references it. Editing it marks those lowerings review-suggested rather than stale: the plans still say what they said, the prose may merely be off-voice — the same treatment .tr gives a changed decisions block.


6. Unit of lowering and freshness

Decision: the whole document. One .trp lowers to one output document; the lock covers the pair.

Reasoning: prose has document-global cohesion that code does not — anaphora, transitions, the “we said this above” texture. Per-paragraph locks would pin wording that a neighbouring edit needs to shift. Per-paragraph and per-section units were rejected for v1.

Cost accepted: any plan edit invalidates the entire lowering, and re-lowering may re-word paragraphs whose directives did not change. This is tolerable because re-lowering a document is cheap relative to code, but the churn is real and a stability mechanism is an open question (§8.5).

7. The lock (provisional schema)

A <name>.lock sidecar in the spirit of docs/lock-schema.md, holding: the three spec hashes (§2); the output hash; lowering provenance (provider, model); the mechanical check results (parse, structure); the style judgment (judge provenance and verdict); and the accepted flag, which only a human sets. See examples/prose/small_languages.lock for the worked shape.


8. Open questions

  1. Additional moves. The user asked for candidates beyond the six. Proposed, each pending a hand-written example where a sentence: rendering is genuinely tortured (the adoption rule; a move that fails it stays out):
    • question: $idea — a rhetorical question; the one move ending in ?.
    • concede: $point — a concession (“admittedly …”), the mirror of contrast:.
    • therefore: $conclusion — draws the consequence of the preceding sentences; differs from bridge to: in closing a line of argument rather than opening one.
    • quote: $text — $source — verbatim pass-through, like headings; the only prose move the agent may not reword.
  2. Mechanical style floor. Revisit rule-based checks if the style judge proves unreliable (§5.1).
  3. Sentence segmentation. Specify the conservative segmenter and the prose forms the agent must avoid (abbreviations, quoted periods, ellipses).
  4. Richer output. Lists, block quotes, and emphasis are absent from v1; each would need a directive and a checkable rendering rule.
  5. Re-lowering churn. Whether accepted documents need paragraph-level wording stability (e.g. the agent is shown the previous lowering and told to preserve wording for unchanged directives), given whole-document invalidation (§6).
  6. Provisional decisions to confirm: the CommonMark container with a single plan block (§2), required style frontmatter (§2), explicit heading levels (§3.2), and the lock field shapes (§7).
  7. Beyond one document: multi-document projects, cross-references, and whether restate: may reference another document.

The soil0 CLI — Oracle Contract

Status: frozen — contract v1.3 (v1 approved 2026-08-22; amended to v1.1 on 2026-08-23 with the agentlanguages-survey adoptions — typed holes: the Hole expression, checks.holes, the unfilled-hole code, run/test refusal; and the reserved print command for the canonical printer, arriving with impl step 11; amended to v1.2 on 2026-08-24, resolving §13.5 for the daemon’s env generator — VariantD.fields for inline-record variant payloads, §6; amended to v1.3 on 2026-08-28 with three Utf8 builtins, §11 — the list-primitives argument applied to strings: nothing string-shaped was writable, and the examples’ parse_row needs to exist. The amendments are additive: v1 outputs are byte-identical for hole-free programs, v1.1 env files decode unchanged, programs not using the new builtins are unaffected, and soil0_cli stays 1 for additive amendments.) This document is the compatibility contract of the soil0 binary, the format every differential oracle for soilc (plan 05) compares against, and the surface the daemon (plan 03) drives. Changing anything here is a contract change requiring a user decision. The soil0 library API is explicitly not covered (§12). References: impl plan 02 (§1 decisions, §8 micro-pins), plan 02 decision points, docs/soil-syntax-spec.md, tr-grammar §2.3 and §7, lock-schema §4–§5, impl plan 01 §4.7 (descriptor JSON).


1. Conventions

  • Invocation: soil0 <command> [flags] [input], commands lex | parse | rename | infer | check | run | test (v1.1 additionally reserves print, the canonical printer — impl step 11). A global --version flag prints {"soil0_cli":1,"soil0":"<crate version>"} and exits 0; soil0_cli is this document’s major version and does not bump for additive amendments.
  • lex and parse take one .soil file; rename, infer, check, run take a program manifest (§7); test takes a bundle (§10). All inputs are UTF-8 files named by path — nothing is read from stdin.
  • stdout carries exactly one JSON document on success: compact, no whitespace, single line, object keys in the field order this document declares. Golden tests byte-compare stdout.
  • stderr carries exactly one diagnostics document (§2) on failure, same compactness rules.
  • Exit codes: 0 success; 1 the input program violates the spec (any diagnostic), and for run/test the runtime outcomes noted in §8.6–§8.7; 2 usage or environment error — malformed manifest, env, bundle, or AST JSON, unreadable file, manifest order violation, or an internal bug.
  • JSON documents embedded in inputs and outputs follow tr-grammar §7 conventions wherever the data is Soil-shaped: internally tagged sums ({"tag": …, "value": …}, nullary tags bare), records with every field present, optionality via the Option encoding ({"tag":"None"} / {"tag":"Some","value":…}), booleans as JSON true/false. Soil values (run results, expected test values) are produced only by soil-rt’s canonical encoder.

2. Diagnostics

{"errors":[{"code":"shadowing","message":"…","file":"median.soil",
  "span":{…},"notes":[{"message":"…","file":"…","span":{…}}]}]}
  • code is a stable kebab-case string from the registry below — the rejection corpus keys on it; renaming a code is a contract change, adding one (for a new rule) is not.
  • file and span are Option-encoded (a usage error has neither).
  • notes attach secondary locations and facts: the prior binding site for shadowing, candidate types for ambiguous-constructor, the witness pattern for non-exhaustive-match.

Error-code registry

PhaseCodes
lexunterminated-string, bad-escape, escape-not-scalar, stray-character
parseparse-expected (generic expected/found), nonassoc-comparison, name-mismatch (signature vs equation defname), wildcard-in-let-rec, misplaced-private-def (private def outside _private.soil, or a non-private def inside one), multiple-defs (a non-private file with more than one definition)
renameshadowing, unknown-name, unknown-type, unknown-qualified, unknown-constructor, ambiguous-constructor, ambiguous-name (definition names follow the same exactly-one-spelling rule, syntax spec §5.10), needless-qualification, private-cross-module
infertype-mismatch, effect-violation (io/ffi deficits only — §8.5), operator-polymorphic, non-derivable, annotation-needed (unresolved record type at field access or constructor), unknown-field, literal-out-of-range
matchnon-exhaustive-match, redundant-arm, missing-record-rest, duplicate-field
runtimeruntime-panic (§8.6), unfilled-hole (v1.1: run/test on a definition — or transitive callee — containing typed holes, §8.6)
class 2usage, io, malformed-input, forward-reference, internal

3. Spans

{"start":42,"end":47,"line":3,"col":5}

start/end are UTF-8 byte offsets into the file, end-exclusive — the canonical span (plan 02 resolved). line/col are 1-based, derived from start, display-only: oracle comparisons may not depend on them beyond the derivation rule.

4. The AST JSON schema

The AST is specified as Soil type declarations; its JSON encoding follows mechanically from tr-grammar §7 (this is the schema soilc’s Trellis type definitions must reproduce, plan 05). Every tree node is wrapped:

type Span      = { start : U64, end : U64, line : U64, col : U64 }
type Spanned a = { span : Span, item : a }
type Binder    = { span : Span, name : Utf8 }

4.1 Files and definitions

type File = { defs : List (Spanned Def) }
type Def  = { name : Utf8
            , sig : Spanned Type
            , decreases : Option (Spanned Expr)
            , params : List Binder
            , body : Spanned Expr }

parse emits one File; a <name>.soil file has exactly one def, a _private.soil any number (syntax spec §2).

4.2 Types, rows, predicates

type Effect = Div | Panic | Io | Ffi
type Row    = { effects : List Effect, var : Option Utf8 }

type Type =
  | Arrow   { param : Option Binder, dom : Spanned Type
            , row : Row, cod : Spanned Type }
  | Con     { name : Utf8, args : List (Spanned Type) }
  | TVar    { name : Utf8 }
  | Refined { binder : Binder, base : Spanned Type, pred : Spanned Pred }
  • effects always in the fixed order Div, Panic, Io, Ffi; the empty row (total) is {"effects":[],"var":{"tag":"None"}}. A row attaches only to an arrow’s codomain position, per the grammar.
  • Con covers both bare (I64, args []) and applied (Result Utf8 FsError) type constructors — one node.
  • Refinement predicates are parsed and retained (plan 02 scope 4):
type Pred =
  | Implies { lhs : Spanned Pred, rhs : Spanned Pred }
  | POr     { lhs : Spanned Pred, rhs : Spanned Pred }
  | PAnd    { lhs : Spanned Pred, rhs : Spanned Pred }
  | PNot    { pred : Spanned Pred }
  | PCmp    { op : CmpOp, lhs : Spanned PExpr, rhs : Spanned PExpr }
  | Is      { scrutinee : Spanned PExpr, type_name : Option Utf8
            , ctor : Utf8, payload : Option Binder }
  | PCall   { name : Utf8, args : List (Spanned PExpr) }

type PExpr =
  | PLiteral { lit : Lit }
  | PPath    { root : Utf8, fields : List Utf8 }
  | PArith   { op : ArithOp, lhs : Spanned PExpr, rhs : Spanned PExpr }
  | PCallE   { name : Utf8, args : List (Spanned PExpr) }

PArith admits only Add/Sub/Mul (tr-grammar §2.3); the parser enforces it. Is.type_name is the optional Type:: qualification (exactly-one-spelling rule, syntax spec §5.9).

4.3 Expressions

type Lit = LInt { digits : Utf8 }     -- underscores stripped; sign never included
         | LFloat { text : Utf8 }     -- source text, underscore-free
         | LStr { value : Utf8 }      -- escapes decoded

type CmpOp   = Eq | Ne | Lt | Le | Gt | Ge
type ArithOp = Add | Sub | Mul | Div | Mod

type Expr =
  | Let       { is_rec : Bool, bindings : List LetBinding, body : Spanned Expr }
  | Fun       { params : List Binder, body : Spanned Expr }
  | If        { cond : Spanned Expr, then_branch : Spanned Expr
              , else_branch : Spanned Expr }
  | Match     { scrutinee : Spanned Expr, arms : List (Spanned Arm) }
  | OrE       { lhs : Spanned Expr, rhs : Spanned Expr }
  | AndE      { lhs : Spanned Expr, rhs : Spanned Expr }
  | NotE      { operand : Spanned Expr }
  | Cmp       { op : CmpOp, lhs : Spanned Expr, rhs : Spanned Expr }
  | Arith     { op : ArithOp, lhs : Spanned Expr, rhs : Spanned Expr }
  | Neg       { operand : Spanned Expr }
  | App       { fn : Spanned Expr, arg : Spanned Expr }
  | Path      { root : Utf8, fields : List Utf8 }
  | Qualified { space : Utf8, name : Utf8 }
  | CtorE     { name : Utf8 }
  | RecordE   { update : Option PathBase, fields : List FieldInit }
  | Literal   { lit : Lit }
  | Annot     { expr : Spanned Expr, ty : Spanned Type }
  | Hole      { name : Utf8 }                          -- typed hole `?name` (v1.1, design §3.16)

type LetBinding = { binder : Option Binder, value : Spanned Expr }  -- None = "_"
type Arm        = { pattern : Spanned Pattern, body : Spanned Expr }
type PathBase   = { root : Utf8, fields : List Utf8 }
type FieldInit  = { name : Utf8, value : Spanned Expr }
  • A bare variable is Path with fields = [] — there is no separate Var node (one way). root may be a private-ident.
  • App is binary and left-nested, faithful to juxtaposition.
  • Qualified covers module::def, Type::derived, and Type::Ctor; the distinction is semantic (rename), not syntactic.
  • Constructor application is App with a CtorE or Qualified head.
  • Parenthesization is not represented; precedence is already resolved.

4.4 Patterns

type Pattern =
  | PWild
  | PBind   { binder : Binder }
  | PLit    { lit : Lit }
  | PCtor   { type_name : Option Utf8, name : Utf8
            , arg : Option (Spanned Pattern) }
  | PRecord { fields : List FieldPat, open : Bool }

type FieldPat = { name : Utf8, pattern : Option (Spanned Pattern) }  -- None = punning

PRecord.open is the trailing ..; PCtor.type_name is the Type:: qualification (§5.9 rule applies in patterns identically).

4.5 Tokens (lex output)

type Token = TIdent { name : Utf8 } | TPrivate { name : Utf8 }
           | TTypeName { name : Utf8 } | TKeyword { word : Utf8 }
           | TOp { op : Utf8 }
           | TInt { digits : Utf8 } | TFloat { text : Utf8 }
           | TStr { value : Utf8 }

lex emits {"tokens": List (Spanned Token)}. Comments and whitespace are dropped; there is no EOF token. Keywords are the §1 keyword list of the syntax spec (effect names are not keywords — they lex as TIdent).

5. Semantic types (signature output)

infer/check report signatures span-free, refinements erased:

type SigType = SArrow { param : Option Utf8, dom : SigType
                      , row : Row, cod : SigType }
             | SCon   { name : Utf8, args : List SigType }
             | SVar   { name : Utf8 }

Canonical renaming: type variables and row variables form one sequence a, b, c, … (…, z, a1, b1, …) assigned in order of first appearance in a pre-order walk of the signature. Named parameters ((fs : Fs)) keep their declared names.

6. env.json — the type environment

type EnvFile  = { types : List TypeDef }
type TypeDef  = { name : Utf8, params : List Utf8
                , strategy : Strategy, body : TypeBody }
type Strategy = Structural | Opaque
type TypeBody = Record { fields : List FieldD }
              | Sum { variants : List VariantD }
              | OpaqueBody
              | Alias { ty : SigType }
type FieldD   = { name : Utf8, shape : SigType
                , ignored : Option IgnoredDefault }   -- Const | CopyField, impl plan 01 §8.14
type VariantD = { name : Utf8, payload : Option SigType
                , fields : Option (List FieldD) }     -- v1.2 (§13.5)

payload and fields are mutually exclusive (both Some is malformed-input); fields declares an inline-record payload (tr-grammar §4.1’s NotFound { path : Path } form), registered as a record payload exactly as the kernel already models its own — the value encoding is unchanged ({"tag": …, "value": {…fields}}, tr-grammar §7). v1.1 env files, which omit the key, decode with fields = None (v1.2, 2026-08-24).

Field, payload, and alias types use the semantic type encoding (§5) — one type encoding for signatures and environments (resolved 2026-08-22, §13.3). SCon covers the scalars (I64, Utf8, Unit, …), List/Map, and named types with arguments; SVar references a params entry (params is required, [] for ground types). SArrow shapes are rejected in v1 (class-2 malformed-input): nothing needs function-typed fields yet, and lifting the rejection later is a behavior change, not a format change. soil0 lowers SigType to soil-rt descriptor shapes internally when registering ground instances on demand (an arrow would lower to the runtime’s shapeless Closure, which poisons derivation); the runtime’s own descriptor JSON (impl plan 01 §4.7) keeps its roles — fixtures and the C ABI — unchanged. Aliases are checker-level and erased before registration.

6.1 The kernel

soil0 pre-registers the kernel at startup; an env.json entry reusing a kernel name is malformed-input. Normative declarations (shapes provisional where marked — plan 04’s prelude .tr specs must adopt or revise them with the user, like the builtin names):

type Bool     = True | False            -- pre-registered by soil-rt
type Option a = None | Some a
type Result a e = Ok a | Err e          -- success-first (design §3.3)
type Path     = Utf8                    -- alias
type FsError  =                         -- provisional shape
  | NotFound   { path : Path }
  | ReadFailed { path : Path, message : Utf8 }
  | NotUtf8    { path : Path }
type Utf8Error = InvalidUtf8 { at : U64 }   -- provisional shape
type World = opaque
type Fs = opaque      type Net = opaque    type Clock = opaque
type Env = opaque     type Proc = opaque   type Rand = opaque
type Py = opaque

All seven capability types exist (design §3.5); only Fs, Clock, Rand have primitives in v1 (§11). Unit and the numeric/string/bytes scalars are built-in shapes, not declarations.

7. program.json — the manifest

{"types":"env.json","defs":["stubs/list_wrappers.soil","sort_by.soil","median.soil"]}
  • types: path to an env.json, relative to the manifest’s directory (an empty environment is {"types":[]} in that file).
  • defs: .soil paths relative to the manifest’s directory, ordered callee-before-caller; a forward reference is class-2 forward-reference (the caller owns the dependency tree — the daemon in production, the test harness before that; handwritten manifests are test fixtures only).
  • A definition’s module, for _private visibility and module::def resolution, is its file’s parent directory name.

8. Commands

8.1 soil0 lex <file.soil>

Tokens per §4.5. Fails only with lex-phase diagnostics.

8.2 soil0 parse <file.soil>

File per §4.1–§4.4. File role (private or not) comes from the filename.

8.3 soil0 rename <program.json>

Success: per-definition reference sets — the computed import set (design §6.1) and the lock’s call-edge data:

{"defs":[{"name":"median","refs":{"defs":["len","nth","sort"],
  "types":["F64","List"],"builtins":[],"privates":[]}}]}

Definitions in manifest order; each list sorted lexicographically. References are surface references — operator elaboration targets (F64::add, …) appear only after infer and are not listed.

8.4 soil0 infer <program.json>

Success, per definition in manifest order:

{"defs":[{"name":"median","type":{…SigType…},"row":{"effects":[],"var":{"tag":"None"}},
  "checks":{"termination":"unverified","panic_obligations":[
    {"kind":{"tag":"DivZero"},"span":{…}}]}}]}

type/row per §5 (the declared signature, checked). checks per §8.5. The --dump-ast flag additionally writes the elaborated, per-node-typed AST to stdout instead — non-contractual, free to change without notice (impl plan 02 §1).

8.5 The check-facts model

soil0 has no SMT solver and no termination checker, but the normative examples claim total. Failing them would contradict the exit criteria, and widening their signatures would contradict design §6.4 (“the agent never silently widens”). The resolution is the demotion philosophy plans 04–05 already adopted for the lock (lock-schema §4):

  • io and ffi subsumption is enforced. A callee’s declared io/ffi not covered by the caller’s declared row is effect-violation — these effects are about capabilities and reality, not proof.
  • div and panic deficits are recorded, never errors. They are proof obligations a later stage discharges (termination checker, refinement checker — plan 05); until then they are unproven claims: visible, runtime-checked, tests still gating.
    • termination: "verified" (no recursion and no div deficit) | "unverified" (self-recursion, let rec, or a call to a declared-div callee, while the row claims no div) | "n/a" (the declared row carries div) — lock-schema §4 vocabulary.
    • panic_obligations: the sites whose potential panic is not covered by the declared row, each {"kind": Overflow | DivZero | CalleePanic { callee : Utf8 }, "span": …}; empty when the row carries panic or no such site exists. Arithmetic sites are always-on runtime checks regardless (plan 02 scope 6).
    • holes (v1.1): the definition’s typed holes in source order, each {"name": …, "span": …, "ty": <SigType>} with the goal type’s unresolved variables canonically renamed. Nonzero holes is the partial state (lock-schema §4): the definition checks but can never be tested, accepted, or built (design §3.16).

Declared signatures remain the modular truth for callers; facts do not propagate (the daemon tracks them per definition in the lock).

8.6 soil0 run <program.json> --entry <name> --args <json-array>

Static pipeline first (any diagnostic aborts, exit 1), then: capability-typed entry parameters are injected from a real World in order; the remaining parameters decode type-directedly from the --args array positionally (entry parameter types must be ground). Real capabilities: Fs is the process filesystem with paths resolved against the working directory, Clock the system clock, Rand OS entropy. Mirrors trellis call (tr-grammar §3.5).

stdout on success: the result value’s canonical JSON, nothing else. A runtime SoilError exits 1 with a single runtime-panic diagnostic whose message is the panic kind and message and whose notes carry the definition-level trace, outermost first. If the entry — or any transitive callee — contains typed holes, run (and test, for the definition under test) refuses with unfilled-hole (v1.1, design §3.16).

8.7 soil0 check <program.json>

The full static pipeline — parse, rename, infer, exhaustiveness and all static rules. Success output is identical to infer’s (§8.4); failure is the combined diagnostics of every phase that ran.

8.8 soil0 test <bundle.json>

See §10.

8.9 soil0 print <file.soil> (v1.1)

The canonical printer (syntax-spec §9): stdout is the file’s canonical Soil text — the one documented exception to the JSON-stdout rule (§1), since the output is source code. parse → print is a fixpoint and the checked-in examples are byte-identical under it; soil_hash (plan 03) is computed over this text after alpha-normalization (lock-schema §3).

9. Exhaustiveness diagnostics

non-exhaustive-match carries one note per missing case with a witness pattern rendered in surface syntax (Err _, { kind = NotFound, .. }); redundant-arm spans the useless arm. Literal patterns (LInt, LStr, LFloat) never exhaust their type; a match over them requires a default arm.

10. Test bundles

Assembled by the caller (daemon; harness; fixtures), one bundle per invocation:

{"program":"program.json",
 "cases":[
   {"name":"found-and-missing#1","def":"read_file",
    "binds":[{"bind":"fs","ctor":"fake_fs",
      "args":[[{"key":"config.toml","value":"port = 8080"}]]}],
    "args":[{"tag":"Binding","value":{"name":"fs"}},
            {"tag":"Json","value":"config.toml"}],
    "expect":{"tag":"Value","value":{"tag":"Ok","value":"port = 8080"}},
    "xfail":false}]}
  • binds construct fake capabilities: ctor names a builtin fake constructor, its args are JSON decoded against the constructor’s parameter types (canonical §7 forms — a map is a key/value array).
  • args fill the definition’s parameters in order: Binding references a bind, Json decodes against the parameter’s type.
  • expect is Value (compared by canonical-JSON byte equality after decode/re-encode normalization, impl plan 02 §8.10) or Panic (matches any SoilError raised by the call).
  • Output, cases in bundle order:
{"cases":[{"name":"found-and-missing#1","result":"pass",
  "details":{"tag":"None"}}]}

result ∈ pass | fail | xfail | xpass (lock-schema §5); details is Some {expected, actual} (both strings: canonical JSON, or "panic" / the panic rendering) exactly when the result is fail or xpass. Exit code: 0 iff no case is fail or xpass; the report prints regardless. Bundle-level problems (unknown def, undecodable args) are class-2, never case results.

11. Builtins

The native definitions visible to every program. Names and signatures are provisional prelude surface (impl plan 02 §8.7, resolved 2026-08-22): plan 04’s prelude .tr specs adopt these names, or revise them with the user and update this table. Strict-minimum inventory (2026-08-22): Env/Proc/Net/Py have no primitives yet; adding one is a new decision point.

NameSignatureSemantics
world_fsWorld -> Fsderive the real filesystem capability
world_clockWorld -> Clockderive the real clock
world_randWorld -> Randderive real entropy
fs_read_bytesFs -> Path -> io (Result Bytes FsError)read a whole file; NotFound / ReadFailed on error
clock_nowClock -> io I64seconds since the Unix epoch (provisional — design §3.5 sketches now : Clock -> io Time; plan 04 reconciles)
rand_u64Rand -> io U64next random value
fake_fsMap Utf8 Utf8 -> Fsin-memory fs: hit returns Ok (utf8_encode contents), miss Err (NotFound { path })
fake_clockI64 -> Clockclock_now returns the given constant on every call
fake_randU64 -> Randdeterministic SplitMix64 stream from the seed (§11.1)
utf8_decodeBytes -> Result Utf8 Utf8Errorvalidate; InvalidUtf8 { at } gives the first bad byte offset
utf8_encodeUtf8 -> Bytesthe underlying bytes
utf8_split(sep : Utf8) -> Utf8 -> List Utf8v1.3: split on every occurrence of sep (separator-first for partial application); Python semantics — n separators yield n+1 cells, so the empty string yields [""]; an empty separator yields the whole string as one cell (pinned; Python errors here, but a total corpus workhorse beats a panic row)
utf8_trimUtf8 -> Utf8v1.3: strip ASCII whitespace (space, \t, \r, \n) from both ends — ASCII only, pinned for determinism
utf8_parse_f64Utf8 -> Option F64v1.3: parse exactly -? digits [ "." digits ] (underscores not accepted — this is data, not source); anything else — including exponents — is None. Grammar provisional prelude surface like the rest of §11; extending it (plan 04) is a behavior change to record
unitUnitthe unit value — there is no () literal; this is the one way to write Unit (§13.2)
list_lenList a -> I64length
list_nthList a -> I64 -> panic azero-based index; out of range panics
list_emptyList athe empty list (a polymorphic value)
list_appendList a -> a -> List aappend one element

Numeric primitives (the §5.4 elaboration targets, also reachable as Type::op): for each fixed-width integer type T — T::add, T::sub, T::mul, T::neg (T -> T -> panic T / T -> panic T, overflow panics) and T::div, T::mod (T -> T -> panic T, floor division and floor modulus — Python semantics, syntax spec §5.6 — zero divisor and MIN / -1 panic). All panics here are obligation sites under §8.5, runtime-checked always. F64 ops are total IEEE 754; BigInt::add/sub/mul/neg are total, BigInt::div/mod panic on zero. Derived functions come from soil-rt, are total (closure poisoning is rejected statically as non-derivable), and have the signatures T::eq : T -> T -> Bool, T::show : T -> Utf8, T::hash : T -> U64, and T::compare : T -> T -> I64 returning −1/0/1 (provisional like the rest of the kernel surface — an Ordering sum is a plan-04 option, scope 6 there). and/or/not elaborate to built-in short-circuit forms on Bool and are not callable by name.

11.1 SplitMix64 (pinned)

fake_rand outcomes appear in expect tests, so the algorithm is observable and must never drift. State starts at the seed; each rand_u64 call advances state by 0x9E3779B97F4A7C15 and returns mix(state) where mix(z): z ^= z >> 30; z *= 0xBF58476D1CE4E5B9; z ^= z >> 27; z *= 0x94D049BB133111EB; z ^= z >> 31.

12. Non-contractual surfaces

Free to change without notice: the soil0 Rust library API (the daemon links it, but conformance is defined by the CLI alone — plan 02 resolved); --dump-ast output; diagnostic message and notes wording (codes, spans, and note structure are contractual); performance of everything.

13. Open questions raised by this draft

  1. The check-facts model (§8.5). Approved 2026-08-22. It extends the impl-plan-02 §8.3 output (adds checks) and interprets plan 02’s “conservatively acquires div” as a recorded fact rather than a hard failure; impl plan 02 §8.3/step 6 and plan 02 scope 4 note the reconciliation.
  2. There is no way to write a Unit value. Resolved 2026-08-22: a kernel value unit : Unit (§11) — the one way to produce Unit, like list_empty a builtin value. Chosen over a () literal or a grammar change: no new syntax, and the grammar stays as specified. Plan 04’s prelude adopts it with the rest of the kernel surface.
  3. Function-typed fields in env.json. Resolved 2026-08-22: unify on the semantic type encoding now — field/payload/alias types are SigType (§6), with SArrow rejected until a prelude type first needs a function field; lifting the rejection is behavior, not format, so the frozen contract survives. Rejected alternatives: a shapeless Closure shape (cannot type a call through the field — every access would need a local annotation) and adding an Fn variant to a separate Shape grammar later (a third type grammar duplicating SigType, and extending a frozen sum is a breaking change for consumers that match on shapes).
  4. FsError/Utf8Error shapes and clock_now are provisional kernel surface for plan 04 to adopt or revise (§6.1, §11). Resolved 2026-08-22: deferred to plan 04, whose scope now carries an explicit kernel-reconciliation item.
  5. Inline-record variant payloads are not expressible in env.json v1. The soil-type grammar (tr-grammar §4.1) gives variants inline record payloads (NotFound { path : Path }), but VariantD.payload is a SigType shape, which cannot carry fields. The kernel models its own record payloads internally; a user env.json cannot declare one yet. The encoding decision belongs to the daemon’s env generator (plan 03), which is the first thing that will need it — raise it there rather than inventing a shape now. Resolved 2026-08-24 (v1.2, with impl plan 03 §9.1): VariantD gains fields : Option (List FieldD) (§6) — the honest representation, additive like v1.1. Rejected: generator-synthesized hidden record types (reserved names leaking into diagnostics and show).
  6. v1.1 amendments (2026-08-23, from the agentlanguages-survey adoptions; approved with them). Typed holes: Expr.Hole, checks.holes, unfilled-hole, run/test refusal — design §3.16; syntax-spec §5.11. The print command is reserved for the canonical printer (one byte-exact form per AST, parse → print → parse idempotence), specified and implemented as impl step 11 before the daemon computes the first soil_hash.

The Trellis Daemon — Agent-Facing Contract

Status: frozen — contract v1 (drafted as impl plan 03 step 1; §10 open questions resolved with the user on 2026-08-25, adding read_spec to the v1 surface; approved in full 2026-08-25). Changing anything here is a contract change requiring a user decision; additive amendments follow the soil0-cli precedent. This is the contract of the daemon’s agent- and user-facing surfaces: the MCP tool surface the lowering agent is allow-listed to, the context-bundle layout it reads, the diagnostic-registry format and spec_ref scheme, and the question/answer schema. These are the surfaces plan 05’s strangler swap and every future provider must keep stable. Explicitly not contractual: the CLI↔daemon JSON-RPC (both ends live in one binary and ship together), the trellis library API, providers.toml, f.log beyond its documented event kinds (§9), and diagnostic message wording (codes, repair classes, and spec_refs are contractual; prose is not — the soil0-cli §12 rule). References: design §4.6 (lowering, tools, packer), tr-grammar (blocks, decisions, write-back §8), lock-schema (§3 edges, §5 test names), docs/contracts/soil0-cli.md (§8.4–§8.6 outputs, §10 bundles, whose shapes are reused verbatim), impl plan 03 §8 (approved micro-pins).


1. Conventions

  • JSON documents follow tr-grammar §7 conventions wherever the data is Soil-shaped: internally tagged sums, records with every field present, Option encoding for optionality, booleans as JSON true/false. Soil values are produced only by soil-rt’s canonical encoder.
  • Schemas below are written as Soil type declarations, encoding per tr-grammar §7 (the soil0-cli §4 device).
  • Hashes are sha256: + 64 lowercase hex. Paths inside bundles are bundle-relative, /-separated, and never contain ...
  • This contract has a version, surfaced to the agent in the MCP initialize response (serverInfo.version carries the trellis crate version; the skill states the contract version). Additive amendments follow the soil0-cli precedent: recorded here, dated, old inputs unaffected.

2. The MCP surface

The daemon’s MCP face is the stdio shim: the agent CLI spawns trellis _mcp --job <id> per its MCP config, and the shim forwards to the daemon’s socket (impl plan 03 §8.8). The shim implements the MCP subset a tools-only server needs — initialize, tools/list, tools/call, and notifications — as JSON-RPC 2.0 over stdio.

  • tools/list returns exactly the seven tools of §3, with the input schemas below as JSON Schema. The tool allow-list is the sandbox (design §4.6): there is no shell tool, no filesystem tool, and no real-capability path; run_tests runs sandboxed tests only.
  • Every tools/call result carries exactly one text content item containing exactly one compact JSON document (the soil0 stdout discipline, applied to tool results). Structured failures set isError and the text item is a diagnostics document (§6 shape) — the agent always parses one JSON document either way.
  • Every tool call is logged by the daemon to f.log (§9) with duration and a result digest; the daemon, not the agent CLI, enforces the job’s turn, wall-clock, and cost caps.

3. The seven tools

3.1 read_context

input:  { path : Utf8 }
output: File { content : Utf8 } | Dir { entries : List Utf8 }

Reads a file or lists a directory inside the job’s context bundle (§4). entries are sorted names, directories suffixed /. A path outside the bundle or not found is an isError result (bundle-path, bundle-not-found).

3.2 check_types

input:  {}   -- operates on the current draft (§3.5)

Runs the full static pipeline — the daemon-assembled environment and callee-first program, then the soil0 library equivalent of soil0 check — over the current draft. Success output is the soil0-cli §8.4 document for the definition under lowering (declared signature, row, checks facts including holes with goal types — a draft with holes checks; design §3.16). Failure output is the combined diagnostics, each enriched with repair_class and spec_ref (§6). Calling with no draft written yet is isError (no-draft).

3.3 check_refinements

input:  {}
output: {"tag":"None"}

A stub until soilc (plan 05): the Option encoding’s None, meaning “no refinement checker is available — clauses are recorded runtime”. When soilc arrives this returns Some { clauses : Map Utf8 Assurance } in the lock-schema §4 vocabulary; the stub shape is chosen so that arrival is additive.

3.4 run_tests

input:  {}
output: { cases : List CaseResult }    -- soil0-cli §10 output, verbatim

Runs every sandboxed case in tests.json (§4.5) — spec and derived — plus the property tiers, against the current draft, with fakes constructed per the bundle cases. Property inputs are generated daemon-side under the pinned seed (impl plan 03 §8.12), so the agent sees the same cases the final gate will run; their results report under the property’s name (spec properties) or derived: name (clause-derived). Real-mode (cram) tests are not present in the bundle and not runnable here, by construction. A draft with unfilled holes is refused with the soil0 unfilled-hole diagnostic (contract §8.6); no draft is no-draft.

3.5 write_soil

input:  { soil : Utf8
        , decisions_applied : List DecisionRef }
type DecisionRef = { scope : Scope, label : Utf8 }
type Scope      = Project | Module
output: { canonical : Utf8 }

Parses and canonicalizes the draft through the printer (soil-syntax-spec §9) — the agent never controls formatting (design §4.8) — and stores it as the job’s current draft. The echoed canonical text is what every later check, test, and hash sees. Parse failures are isError with diagnostics; nothing is stored.

decisions_applied is the required citation (design §4.3, resolved 2026-08-24): the decision entries, by scope and label, the agent applied in this draft — [] when none. Each call restates the full list; the last successful write_soil before a green outcome is what the lock records as lowering.decisions.applied (lock-schema §3). An unknown label is isError (unknown-decision).

3.6 read_spec

input:  { ref : Utf8 }            -- a §7 anchor, e.g. "soil-syntax-spec#5.9"
output: { title : Utf8, text : Utf8 }

Serves one pinned spec section by anchor (added to v1 on 2026-08-25, resolving §10.2: diagnostics carry spec_refs, so the agent must be able to dereference them — “the whole spec is never shipped blind”, design §4.6). Sections come from the toolchain’s embedded copies of the §7 documents, never the repo working tree: the spec an agent reads must be the one the toolchain that checks it was built from (the design §8 pin rationale). An unresolvable anchor is isError (unknown-spec-ref). The skill still carries the budgeted core reference; this tool is the lookup path for everything beyond it.

3.7 ask_human

input:  { kind : Ambiguity | Contradiction | Rule
        , question : Utf8
        , hole : Option Utf8          -- a hole name this blocks on
        , options : List Utf8 }       -- proposed answers, may be empty

Ends the invocation (design §4.6): the daemon acknowledges the call, persists the question (§8), and terminates the job. Contradiction is the distinguished class for spec-internal conflict the pre-flight could not catch mechanically; Rule marks an answer expected to be a project/module rule — its answer is written to a decisions block rather than the definition’s prose (tr-grammar §8). The agent never sees an answer that is not already in the spec: the answered question re-runs the job fresh.

4. The context bundle

The fixed directory layout (design §4.6) the job’s read_context serves. Assembled by the budgeted packer (§5); also written to disk by trellis context <def> --budget <n> for human inspection.

bundle/
  bundle.json          -- the manifest (§4.1)
  spec.md              -- §4.2
  decisions.md         -- §4.3
  callees/<name>.md    -- §4.4, one per direct callee
  tests.json           -- §4.5
  examples/            -- §4.6
  reference.py         -- §4.7, when a Python reference is attached
  previous.soil        -- §4.8, when re-lowering

4.1 bundle.json

type BundleManifest =
  { definition : Utf8                  -- e.g. "csvstats/median"
  , budget : U64                       -- tokens (estimate: bytes/4)
  , estimate : U64                     -- packed-size estimate
  , packed : List Utf8                 -- bundle-relative paths
  , dropped : List Dropped             -- §5; never silent
  , oracle : Option Oracle             -- §4.7
  , spec_hashes : SpecHashes }         -- what this bundle was built from
type Dropped    = { item : Utf8, reason : Utf8 }
type Oracle     = Reference { path : Utf8, symbol : Utf8 }
                | Cli { command : Utf8 }
type SpecHashes = { formal : Utf8, test : Utf8, prose : Utf8 }

spec_hashes is the staleness anchor: mid-flight invalidation and question staleness (§8) compare against it.

4.2 spec.md

The definition’s .tr file, verbatim (frontmatter included). Never dropped by the packer.

4.3 decisions.md

The decisions in scope, project section first, then the module’s, each entry rendered with its scope and label:

# Decisions

## project

- **timestamps** — all timestamps are UTC seconds (I64)
  rejected: local time (DST bugs); F64 epoch (precision loss)

## module csvstats

- **cell-errors** — BadCell carries index and raw text

These are the labels write_soil.decisions_applied cites. Absent when no decisions exist in scope; travels at spec priority (§5).

4.4 callees/<name>.md

One file per definition on the callable surface — every lowered function in the root visible to the target, nearest-first (the target’s module before others), the target itself excluded. (Amended 2026-09-14, step 9: the draft said “one per direct callee”, which is circular on a first lowering — no body, no known edges — and the first lowering is the bundle’s main audience. The serial lowering order, design §4.6, makes exactly the lowered definitions available, so the surface is the menu; the packer’s callee tier prunes farthest-first at scale.) Signatures only, never bodies (design §4.6). Content: the callee’s soil-sig block; its requires/ensures clauses, each annotated with its lock-schema §4 assurance; a runtime (demoted) clause is shown as the base type plus a demotion note — the agent must not rely on an unproven claim:

# sort_by

```soil-sig
sort_by : (key : a -> k) -> (xs : List a) -> List a
```

- ensures `stable` (proven): …
- ensures `sorted` (runtime — demoted, guard active; do not rely on
  this claim when reasoning about your own obligations): …

4.5 tests.json

{ "cases": [ …soil0-cli §10 case shape… ] }

Exactly the soil0 bundle-case shape — one format — minus the program key (the daemon owns program assembly). Contains every sandboxed case: the spec’s expect tests, property cases the runner will generate are not enumerated (properties appear as their forall text inside spec.md; generated inputs are not bundle content), and derived expect-shaped cases under their lock-schema §5 derived: names. Never dropped by the packer.

4.6 examples/

The corpus (design §5): .tr/.soil pairs. Until the prelude exists (plan 04) this is the repo’s examples/; afterwards, the pinned prelude corpus. Dropped last among droppables, whole pairs at a time.

4.7 reference.py / the CLI stanza

A Python reference attachment (tr-grammar §3.6) copies the referenced file verbatim as reference.py, with the symbol named in bundle.json.oracle. A CLI oracle contributes no file — the stanza in bundle.json.oracle is the whole record (the agent cannot run either; they are context, and the differential tier runs daemon-side).

4.8 previous.soil

On re-lowering: the current checked-in canonical Soil. Absent on a first lowering.

5. The packer

trellis context <def> --budget <n>; the same assembler feeds every job. Fixed priority (design §4.6, adopted 2026-08-23):

spec (spec.md, decisions.md, bundle.json)  >  tests (tests.json)
  >  callee signatures  >  corpus examples  >  module prose
  • Budget is in tokens, estimated at bytes/4 (impl plan 03 §8.10); the per-model budget lives in the provider config.
  • spec.md, decisions.md, bundle.json, and tests.json are never dropped: a budget too small for them is a structured error (budget-too-small), not a silent truncation.
  • Dropping is coarsest-first within a tier (a whole callee file, a whole example pair) and every drop is listed in bundle.json.dropped and logged — no silent caps.
  • previous.soil and reference.py sit between tests and callee signatures (step 9, 2026-09-14): the body under revision and the executable spec outrank the callable menu, but unlike the §4.5 set they are droppable under extreme budgets.
  • The packer keeps the longest priority-order prefix that fits: the first over-budget item and everything after it drop (step 9, 2026-09-14). This makes the packed set and the estimate monotone in the budget and the drop order predictable for the agent; greedy skip-and-continue was rejected because a grown budget could then produce a different (not larger) bundle. estimate counts the packed bytes plus the rendered manifest, so it can exceed a tight budget by the manifest’s own size — the bytes/4 estimate is approximate by declaration (§8.10).
  • Module prose (_module.tr’s prose, lowest tier) is packed as module.md when it fits.

6. Diagnostics, the registry, and repair classes

Every diagnostic the daemon emits — its own and soil0’s passed through — carries:

type Diagnostic = { code : Utf8            -- kebab-case, from the registry
                  , message : Utf8         -- non-contractual wording
                  , file : Option Utf8
                  , span : Option Span     -- soil0-cli §3 shape
                  , notes : List Note
                  , repair_class : Utf8    -- from the registry
                  , spec_ref : Utf8 }      -- §7 anchor

An error document is { "errors": List Diagnostic } (the soil0-cli §2 shape, extended by the last two fields).

6.1 The registry file

registry/diagnostics.json in the trellis crate is the single source; the skill and this contract’s rendered tables are generated from it, and CI fails when registry, docs, and emitted codes disagree (design §4.6). The format is contractual; the contents grow additively and are governed by the drift gate, not frozen here:

type Registry     = { registry_format : U64        -- 1
                    , repair_classes : List RepairClass
                    , codes : List CodeEntry }
type RepairClass  = { name : Utf8, summary : Utf8 }
type CodeEntry    = { code : Utf8
                    , source : Utf8                -- "soil0" | "trellis"
                    , repair_class : Utf8
                    , spec_ref : Utf8 }
  • Every soil0-cli §2 code appears with source: "soil0", unchanged; the registry adds only the mapping (soil0’s own outputs stay byte-identical — the enrichment happens daemon-side).
  • Daemon codes are namespaced by prefix: tr-… (.tr validity), config-… (soil.toml / provider config), lock-…, bundle-…, job-…, preflight-… (the design §4.5 contradiction check and vacuity probes — step 7), plus toolchain-mismatch, budget-too-small, no-draft, unknown-decision, unknown-spec-ref, unsupported-where-filter, unsupported-invariant-property, test-ungenerable-type, and oracle-error.

6.2 Initial repair classes

The starting vocabulary (typed actions an agent takes mechanically, design §4.6); extension is additive and drift-gated:

ClassActs on
rename-bindingshadowing and collision errors
qualify-name / unqualify-namethe §5.9/§5.10 exactly-one-spelling errors
widen-matchnon-exhaustive-match (the witness note names the arm to add)
remove-redundant-armredundant-arm
add-record-restmissing-record-rest
add-annotationannotation-needed, operator-polymorphic
fix-typetype-mismatch, unknown-field, arity errors
thread-capabilityeffect-violation (an io/ffi callee needs its capability and row threaded)
fix-literalliteral-out-of-range
insert-guardpanic-obligation repairs (match the zero/overflow case away)
add-decreasestermination deficits (parsed now, discharged plan 05)
export-type-or-mark-abstractthe design §3.14 export error
fill-holeunfilled-hole at run/test time
fix-speccontradictions and .tr validity errors only a spec edit can fix — the class whose repair is ask_human when the block is pinned

7. spec_ref anchors

<doc-slug>#<section>: slugs design, tr-grammar, lock-schema, soil-syntax-spec, soil0-cli, trellis-daemon; <section> is the printed section number (soil-syntax-spec#5.9, tr-grammar#3.3). The CI drift gate resolves every registry spec_ref against the target document’s headers, so renumbering a referenced section fails the build — which is the point: the Soil spec is pinned, individually addressable sections a daemon tool serves (design §4.6); read_spec (§3.6) is that tool.

8. Questions and answers

ask_human persists the question as <def>.question.json beside the .tr (gitignored, like f.log):

type Question = { def : Utf8
                , kind : Ambiguity | Contradiction | Rule
                , question : Utf8
                , hole : Option Utf8
                , options : List Utf8
                , asked_against : SpecHashes    -- §4.1
                , bundle : Utf8                 -- bundle content hash
                , created : Utf8 }              -- RFC 3339, display-only

trellis status surfaces open questions; trellis answer <def> (interactive, or --text <answer>) applies the write-back and re-queues the job fresh:

  • kind Ambiguity/Contradiction → a Q/A pair in the .tr’s ## Clarifications section (tr-grammar §8);
  • kind Rule → an entry appended to the module’s or project’s decisions block (the CLI asks which scope when ambiguous);
  • if the .tr changed since asked_against, the answer is refused with a diff (job-stale-question) — the invalidation rule already requeued or killed the job, and the question may no longer apply.

Every question/answer round-trip is therefore a visible spec diff, and the re-invoked agent reads the answer from the spec like any other input (design §4.6).

9. f.log (documented, non-contractual)

JSONL beside the definition, gitignored, human-readable, never an input (design §4.6). One event per line, event plus fields: job_started (definition, spec hashes, provider, model, caps), preflight (contradiction/vacuity findings), bundle_packed (the §4.1 manifest), tool_call (name, duration_ms, result digest), question, agent_result (turns, tokens in/out, cost, provider session id), outcome (green | partial holes | failed | invalidated, plus per-test failure details — the lock stores results only, lock-schema §5), triage (entry, verdict). Field sets may grow; nothing may come to depend on them.

10. Open questions raised by this draft — resolved 2026-08-25

All four were resolved with the user on 2026-08-25 during the step-1 review:

  1. Registry completeness. Resolved: freeze the format, contents evolve. The §6.2 repair classes and the daemon code list are a starting vocabulary; step 6 wires them, gaps land additively under the drift gate with no per-addition review gate.
  2. A read_spec tool. Resolved: ship it in v1 — §3.6. The diagnostics carry spec_refs, so the agent can dereference them immediately; the skill still carries the budgeted core reference, and read_spec is the lookup path beyond it. Sections are served from the toolchain’s embedded doc copies, matching the pin.
  3. MCP protocol revision. Resolved: pin at implementation time — steps 10/11 pin the revision the Claude Code headless client then speaks, recorded here when known; the tool surface is small enough that revision drift is absorbed in the shim.
  4. run_tests case filtering. Resolved: all cases, always — the agent always sees the same picture as the final gate. A { cases : List Utf8 } filter input remains a possible additive amendment if long suites make the loop slow.

Milestone Plans — Overview

These plans are implementation guides for future agents. Each corresponds to a milestone of the build order (design §11, bootstrap plan) and states its goal, scope, work breakdown, testing strategy, exit criteria, and open decision points.

How to use these plans

  • Read CLAUDE.md, then docs/design.md for the decisions and reasoning, then the plan for your milestone. The plan tells you what to build and in what order; the specs (docs/tr-grammar.md, docs/lock-schema.md, docs/soil-syntax-spec.md) tell you what the artifacts must look like.
  • Decision points listed in a plan are not yours to make. The initial sets were all resolved with the user on 2026-08-22 and are recorded in each plan; if implementation surfaces a new open choice, present options to the user and record the outcome in the plan and in docs/design.md the same way.
  • Exit criteria are the definition of done. Do not start the next milestone’s work early; the ordering is load-bearing (each stage is the oracle or substrate for the next).
  • When implementation reveals a spec contradiction or gap, stop and raise it — spec fixes propagate to docs/ and examples/ before code works around them.

Sequence

PlanMilestoneDepends on
01-soil-rt.mdThe runtime crate (values, JSON, C ABI)—
02-soil0.mdMinimal Rust Soil: parser, checker, interpreter, CLI oracles01
03-daemon.mdThe Trellis daemon: hashing, locks, lowering jobs, MCP tools01, 02
04-prelude.mdThe pure-core prelude, first Trellis code03
05-soilc.mdThe compiler as the first Trellis project; self-hosting04
06-ffi-bind-ide.mdFFI (Rust + Python), trellis bind, minimal IDE05
07-python-glue.mdThe second project: validate the FFI half of the pitch06

Sibling track: trellis-prose

PlanMilestoneDepends on
prose-01-trp.mdThe trp checker crate (parse, verify, locks)— (shares the Rust workspace; independent of the Soil sequence)

Rust code accumulates in one workspace (soil-rt, soil0, later the daemon and the Cranelift driver). Trellis code (prelude, soilc) lives in Soil roots with soil.toml. Nothing in these plans exists yet; the repo is design-only until plan 01 begins.

Plan 01 — soil-rt: the runtime crate

References: design §3.7 (derived functions), §3.10 (runtime and embedding), §3.12 (primitives), tr-grammar §7 (JSON value encoding).

Goal

A Rust crate owning the Soil value model, memory management, derived operations, the canonical JSON bridge, and a C ABI for embedding. This crate is permanent — it is never bootstrapped away — and everything else (interpreter, compiled code, hosts) manipulates values only through it.

Scope

  1. Workspace skeleton. Cargo workspace at the repo root (or a rust/ subdirectory — decision point) containing soil-rt; later crates (soil0, daemon, Cranelift driver) join the same workspace.
  2. Value model. A Value representation covering: fixed-width integers (I64, U64, I32, U32, I16, U16, I8, U8), BigInt, F64 with total order (total_cmp; NaN equal to itself, sorts last), Utf8 (immutable, validated), Bytes, Unit, records (fields in declaration order), sums (tag + at most one payload), List, Map (stored in comparator order — the comparator is a Soil closure carried by the map), closures, and opaque handles. Bool is the prelude sum True | False; the runtime may represent it natively as an optimization but must present it as a sum.
  3. Type descriptors. Runtime-registered metadata per type: field names and order, variant names, derivation strategy (structural / opaque / per-field ignored with default thunks). Descriptors drive derived ops and type-directed JSON decode. Registration API used by the interpreter now and compiled code later.
  4. Memory. Reference counting. No Perceus reuse yet (that is compiler-side, stage 3+); no cycle collector (strategy explicitly deferred, design §10 — document that cycles leak for now).
  5. Derived operations. One implementation each of structural eq/compare/hash/show over Value + descriptor, honoring strategies: opaque → identity/"<handle>"/address; ignored → skipped. Deriving eq over a closure is rejected at descriptor registration.
  6. Canonical JSON. Encode and type-directed decode per tr-grammar §7, exactly: internally tagged sums, Bool as JSON booleans, BigInt hybrid by range (±(2^53−1)), F64 "NaN"/"Inf"/"-Inf" strings, Bytes base64, Unit null, maps as comparator-ordered {"key","value"} arrays, ignored fields omitted on encode and refilled from default thunks on decode, opaque decode error, closures a hard error both ways. Encoding is canonical: declaration-order fields, shortest round-trip floats — byte-equal output for equal values.
  7. Errors. A structured SoilError (panic kind, message, trace hook, offending JSON inputs when available) — the debug-mode payload of design §3.11’s host-stub rule.
  8. C ABI. No global state; explicit soil_init/soil_teardown; opaque SoilValue* handles with constructors/accessors/refcount ops; error out-parameters; a trivial main wrapper entry. Header via cbindgen. A minimal C demo program proves embeddability.

Non-goals

Execution (plan 02), Perceus reuse, cycle collection, pyo3 (plan 06), refinements (checker-side only, and later).

Testing

  • Rust unit tests per module; property tests (proptest) for: JSON round-trip identity on random well-typed values, eq/compare/hash agreement (equal ⇒ same hash; compare total order laws incl. NaN), canonical encoding determinism (encode twice, byte-equal).
  • A fixtures/ corpus of (type descriptor, value, canonical JSON) triples, checked in — these become shared goldens for soilc’s passes later.
  • The C demo compiled and run in CI fashion (a script for now).

Exit criteria

  • The C demo constructs values through the ABI, round-trips them through canonical JSON, and tears down cleanly (no leaks under a debug counter).
  • All tr-grammar §7 rows demonstrably implemented, including the Bool, BigInt-range, and ignored-field-default cases, backed by fixtures.
  • Descriptor API documented well enough that plan 02 needs no runtime changes to interpret the prelude.

Decision points — resolved 2026-08-22

  • Workspace location: rust/ subdirectory holding the Cargo workspace (soil-rt, later soil0, the daemon, the Cranelift driver); the repo root stays docs/examples/Trellis-roots.
  • Integer widths: all eight (I64…U8) from the start — adding widths later ripples through JSON, the C ABI, and descriptors, and together they are mostly a macro.
  • BigInt: num-bigint.
  • Map representation: sorted vec of pairs with the carried comparator closure — trivially correct and canonically ordered for show/JSON; swap for a tree behind the same API only if profiling demands it.

Implementation plan

A detailed implementation guide exists at impls/01-soil-rt-impl.md. It records a second round of decisions resolved 2026-08-22 — Value as a Rust enum with nonatomic refcounted boxes, closures as boxed Rust callables, hand-rolled canonical encoder with serde_json decode, interned per-instance TypeIds, and hash as FNV-1a 64 over canonical JSON bytes (also recorded in design §3.7, being language-observable) — plus the module map, build order, fixture format, C ABI surface, and a set of micro-pins (its §8, approved 2026-08-22, including: all dev dependencies managed through a repo-root shell.nix, which is also the single pin for the Rust toolchain).

Implementation Plan 01 — soil-rt

Detailed implementation guide for ../01-soil-rt.md. That plan states the scope and exit criteria; this one states how the crate is actually structured and built, records the implementation decisions resolved with the user on 2026-08-22, and proposes the remaining micro-details (§8) for review before code is written. References: design §3.7, §3.10, §3.12, tr-grammar §7, soil-syntax-spec §1 (string escapes), plan 02 (the first consumer).


1. Resolved implementation decisions (2026-08-22)

DecisionChoiceRationale
Value representationRust enum, heap variants behind refcounted boxesIdiomatic and mostly safe; layout stays internal to the crate, so the C ABI and later compiled code see only opaque SoilValue* handles and the representation can change without ABI breakage. A uniform tagged word was rejected as unsafe-heavy before any profiling justifies it — “slow is accepted” (bootstrap plan).
Reference countsNon-atomic, single-threaded instancesRc-style counts; a runtime instance and all its values belong to one thread (documented C ABI rule). The daemon parallelizes with one instance per thread. Atomic counts were rejected as a permanent tax for a concurrency story the design defers (concurrency is FFI to host libraries).
Ref<T>Thin newtype over std::rc::Rc<T> plus a debug-build live-allocation counterMinimal unsafe, satisfies the leak-check exit criterion. A custom refcount header was rejected as premature: Perceus reuse is compiler-side and far away, and nothing inspects headers yet.
Canonical JSONHand-rolled encoder; decode parses via serde_json into a generic tree, then a type-directed walkThe encoder is the canonicality contract (byte-equal output), so it must be owned code, using ryu/itoa for shortest-round-trip numerals. Parsing is borrowed from a battle-tested crate; outsourcing the encoder to serde_json was rejected because the core spec guarantee would depend on a third-party crate’s formatting stability.
Closure invocationValue::Closure wraps a boxed Fn(&[Value]) -> Result<Value, SoilError>One invocation path for everyone: plan-01 tests pass plain Rust closures; soil0’s interpreter (which links this crate) captures AST+env in a Rust closure; compiled code later wraps an extern "C" fn + env value in the same shape. A registered-invoker callback was rejected: it needs a second native mechanism for tests anyway.
Derived hashFNV-1a 64-bit, fixed and documentedhash is language-observable (design §3.7), so determinism across OS/arch/runs/toolchains is the requirement. Soil maps are comparator-ordered, not hash tables, so there is no DoS surface and SipHash’s machinery buys nothing; std::DefaultHasher is explicitly unstable across Rust releases.
hash input bytesThe value’s canonical JSON encoding; the opaque strategy hashes the addressThere is exactly one byte-form of a value in the whole system. Equal ⇒ byte-equal JSON is already the canonical-encoding guarantee, so equal ⇒ same hash follows for free, and ignored fields contribute a constant by construction (they are omitted from the encoding). Cost is an encode per hash call — acceptable; only the algorithm, not the definition, would change if it ever binds. A parallel structural byte-feed was rejected as a second definition of a value’s byte form to keep in agreement.
Type identityInterned per-instance TypeId(u32), dense index into a per-runtime registryCheap comparisons and lookups; the C ABI passes the u32. Ids are not stable across runs — anything persistent uses the type name, which is fine because filename/name is identity in Trellis (design §4.3). String names everywhere was rejected: every lookup becomes map-by-string and ids get retrofitted later anyway.

The one language-observable item (the hash definition) is recorded in design §3.7; the rest are runtime-internal and live here.


2. Workspace and crate skeleton

Per plan 01’s resolved decision points: the Cargo workspace lives in rust/; the repo root stays docs/examples/Trellis-roots.

shell.nix              -- repo root; the dev environment, single source of truth (§8.12)
rust/
  Cargo.toml           -- workspace; members = ["soil-rt"] (later soil0, daemon, clif driver)
  check.sh             -- runs inside nix-shell: fmt --check, clippy -D warnings, test, C demo build+run
  soil-rt/
    Cargo.toml
    cbindgen.toml
    src/               -- module map in §3
    fixtures/          -- (descriptor, value, canonical JSON) triples, §6
    cdemo/
      main.c
      build.sh         -- cc against the staticlib + generated header

All dev dependencies are managed through a repo-root shell.nix: the Rust toolchain (rustc/cargo — no rust-toolchain.toml; the nixpkgs pin in shell.nix is the single source of the compiler version), cbindgen, the C compiler for the demo, and whatever later milestones add (Z3, Python). Building outside nix-shell is off the supported path. Rust crate dependencies remain in Cargo.toml/Cargo.lock as usual — shell.nix manages tools, Cargo manages crates.

Crate dependencies, kept deliberately short:

CrateWhyWhere
num-bigintBigInt (plan 01 resolved decision)runtime
ryu, itoashortest-round-trip float / integer formatting in the canonical encoderruntime
serde_jsondecode-side parsing to a generic tree only; never encodesruntime
base64Bytes encodingruntime
proptestproperty testsdev-only
cbindgenheader generationtool from shell.nix, invoked from check.sh, not a build.rs dependency

soil-rt builds as both rlib (for soil0 and the daemon) and staticlib (for the C demo and embedding).


3. Module map

src/
  lib.rs         -- Runtime (owns Registry + debug counters), re-exports, crate docs
  error.rs       -- SoilError, PanicKind, TraceFrame
  value.rs       -- Value, Ref<T>, heap payloads (Record, SumVal, MapVal, Closure, OpaqueVal)
  descriptor.rs  -- TypeShape, TypeDesc, FieldDesc, Registry, TypeId; descriptor JSON (§5)
  ops.rs         -- derived eq / compare / hash over Value + descriptor
  json/
    mod.rs
    encode.rs    -- the canonical encoder; owns every canonicality guarantee
    decode.rs    -- type-directed decode over the serde_json tree
  capi.rs        -- the C ABI surface (§7), the only module containing `extern "C"`

Dependency direction is strictly downward: capi → {json, ops} → {descriptor, value} → error. No module reaches back up; no global state anywhere (Runtime is a value the embedder holds).


4. Build order

Each step names its deliverable, the API it stabilizes, and its tests. Steps are sequential; a step is done when its tests pass and check.sh is green.

Step 1 — skeleton and SoilError

Workspace files, empty modules, check.sh, toolchain pin.

#![allow(unused)]
fn main() {
pub struct SoilError {
    pub kind: PanicKind,       // Overflow, DivideByZero, DecodeError,
                               // TypeError, DerivationError, CapiMisuse, …
    pub message: String,
    pub trace: Vec<TraceFrame>, // hook only; filled by soil0/daemon later
    pub inputs: Option<String>, // offending canonical-JSON inputs when available
}
}

This is the debug-mode payload of design §3.11’s host-stub rule; plan 02 raises interpreter panics as SoilError, so the shape is API from day one. Tests: construction and Display formatting only.

Step 2 — Value and Ref<T>

#![allow(unused)]
fn main() {
pub enum Value {
    I64(i64), U64(u64), I32(i32), U32(u32),
    I16(i16), U16(u16), I8(i8), U8(u8),
    F64(f64),
    BigInt(Ref<BigInt>),
    Utf8(Ref<str>),
    Bytes(Ref<[u8]>),
    Unit,
    Record(Ref<Record>),     // TypeId + field values in declaration order
    Sum(Ref<SumVal>),        // TypeId + variant index + optional payload
    List(Ref<Vec<Value>>),
    Map(Ref<MapVal>),        // sorted Vec<(Value, Value)> + comparator Closure
    Closure(Ref<Closure>),   // Box<dyn Fn(&[Value]) -> Result<Value, SoilError>>
    Opaque(Ref<OpaqueVal>),  // TypeId + Box<dyn Any>; identity = address
}
}
  • Value is small and passed by value; clone is a refcount bump on heap variants.
  • Ref<T> wraps Rc<T>; in debug builds every allocation increments and every final drop decrements a thread-local live counter (consistent with instances being single-threaded), exposed as soil_debug_live_values() for the leak-check exit criterion. Release builds compile the counter out.
  • Bool is not a Value variant: it is the prelude sum True | False (plan 01 scope item 2). The registry pre-registers it (step 3) and the JSON layer special-cases its TypeId.
  • Cycles leak; documented on Ref (design §10 defers the strategy).
  • MapVal maintains the sorted-vec invariant internally: insertion is binary search via the carried comparator (a Closure — fallible, so every map operation is fallible). Plan 01’s resolved decision: swap for a tree behind the same API only if profiling demands it.

Tests: refcount behavior (clone/drop, counter returns to zero), map insert/lookup/remove with a native comparator closure, closure invocation, Value size assertion (fits two words + discriminant).

Step 3 — descriptors and the registry

#![allow(unused)]
fn main() {
pub enum TypeShape {
    I64, U64, I32, U32, I16, U16, I8, U8,
    BigInt, F64, Utf8, Bytes, Unit,
    List(Box<TypeShape>),
    Map(Box<TypeShape>, Box<TypeShape>),
    Closure,                  // may appear in shapes; poisons derivation
    Named(TypeId),            // records, sums, opaques — including recursion
}

pub struct TypeDesc {
    pub name: String,               // cased name, e.g. "SummaryRow"
    pub strategy: Strategy,         // Structural | Opaque
    pub body: TypeBody,             // Record(Vec<FieldDesc>) | Sum(Vec<VariantDesc>) | Opaque
}
pub struct FieldDesc {
    pub name: String,
    pub shape: TypeShape,
    pub ignored: Option<Closure>,   // default thunk; presence marks the field ignored
}

impl Registry {
    pub fn declare(&mut self, name: &str) -> Result<TypeId, SoilError>;
    pub fn define(&mut self, id: TypeId, desc: TypeDesc) -> Result<(), SoilError>;
    pub fn register(&mut self, desc: TypeDesc) -> Result<TypeId, SoilError>; // declare+define
    pub fn get(&self, id: TypeId) -> &TypeDesc;
    pub fn lookup(&self, name: &str) -> Option<TypeId>;
}
}
  • Two-phase registration (declare then define) exists because recursive types across definitions are legal (design §3.8); a self-referential shape names its own declared id. Using an undefined id in any runtime operation is CapiMisuse.
  • Derivability is computed at define time: a type whose shape transitively contains Closure (through fields, payloads, list/map elements — map comparators excluded, they are structure, not content) is marked non-derivable, per “deriving eq on a type containing an arrow is a type error” (design §3.7). Derived ops on such a type return DerivationError; the static rejection happens in soil0’s checker.
  • The registry pre-registers Bool (sum True | False, in that declaration order) at construction and exposes Registry::BOOL.
  • Strategies per design §3.7: Structural (default), Opaque (identity/"<handle>"/address); ignored is per-field with a mandatory default thunk.

Tests: registration round-trips, recursive type via declare/define, closure-poisoning marks non-derivable, duplicate names rejected, Bool present.

Step 4 — derived operations (ops.rs)

One implementation each of eq, compare, hash over (&Runtime, &Value); show is the canonical encoder (step 5), per “show is the JSON encoder” (design §3.7).

  • compare is a total order per type. Numeric types compare numerically within their own type; comparing values of different types is TypeError (the checker prevents it; the runtime is defensive).
  • F64: all NaN bit patterns are one logical value — equal to each other, sorting after +Inf (design §3.7: “NaN is equal to itself and sorts last”). This is total_cmp semantics with the NaN payload/sign distinctions collapsed, because canonical JSON has a single "NaN" and round-tripping must preserve eq. -0.0 < 0.0 stays distinct (canonical JSON distinguishes -0.0 from 0.0). Pinned in §8.
  • Structural equality: records field-wise in declaration order; sums by variant index then payload; lists element-wise then by length; maps by sorted entry sequence (the comparator closure is excluded from eq — it is structure, not content); Utf8/Bytes byte-wise; BigInt numerically.
  • opaque strategy: eq/compare/hash on the payload address.
  • ignored fields: skipped by eq/compare; constant hash contribution by construction (omitted from the canonical encoding).
  • hash: fnv1a64(canonical_json_bytes(v)) with the standard FNV-1a offset basis and prime, documented in the module; opaque hashes the address bytes instead.
  • eq/compare/hash reaching a Closure value is DerivationError (defense in depth behind the descriptor-level rejection).

Tests: unit tests per type; property tests deferred to step 6 where generators exist.

Step 5 — the canonical encoder (json/encode.rs)

Implements tr-grammar §7 exactly, writing into a Vec<u8>. This module owns every canonicality guarantee; nothing else in the system may produce value JSON.

CaseEncoding
fixed-width intsitoa
BigIntJSON number within ±(2^53−1); decimal string beyond
F64ryu shortest round-trip; "NaN", "Inf", "-Inf" as strings
Utf8JSON string, escape set in §8
Bytesbase64 string (standard alphabet, padded, unwrapped — §8)
Unitnull
Booltrue / false (TypeId special case)
recordobject, fields in declaration order, ignored fields omitted
sum{"tag": "Name"} / {"tag": "Name", "value": …}
Listarray
Maparray of {"key": …, "value": …} in comparator order (storage order)
opaque"<handle>"
closurehard error

No whitespace anywhere (element separators are , and : alone — §8). Encoding is value-directed: scalars self-describe via their variant, records/sums carry their TypeId, and the registry supplies field and variant names.

Tests: golden strings per row of the table, including BigInt at ±(2^53−1)±1, negative zero, integral floats (1.0 not 1), and the determinism test (encode twice, byte-equal).

Step 6 — type-directed decode (json/decode.rs) and fixtures

decode(&Runtime, &TypeShape, &str) -> Result<Value, SoilError>: parse with serde_json into serde_json::Value, then walk the shape. Strict: every deviation is a DecodeError naming the JSON path.

  • Numbers must be integral and in range for the target width; BigInt accepts number-or-string (hybrid); F64 accepts numbers and the three strings.
  • Records: missing non-ignored field, unknown field, or a present ignored field are errors (§8); after the other fields decode, each ignored field is refilled by invoking its default thunk with the non-ignored fields as arguments in declaration order (§8).
  • Sums: exactly the keys tag (+ value iff the variant has a payload); unknown tag is an error. Bool accepts only true/false.
  • Opaque shapes and closure shapes are errors (tr-grammar §7).
  • Recursion depth is bounded by serde_json’s parser limit; the walk itself is iterative or depth-checked to keep panic out of the crate.

Fixtures (fixtures/*.json, the shared goldens plan 01 requires, reused later by soilc’s passes):

{
  "types": [ …descriptor JSON, §5-encoded… ],
  "type": "SummaryRow",
  "canonical": "{\"count\":3,\"mean\":1.5}",
  "accepts": ["{\"mean\":1.5,\"count\":3}"],
  "rejects": ["{\"count\":3}", "{\"count\":3,\"mean\":1.5,\"x\":0}"]
}

The harness decodes canonical, re-encodes, requires byte equality; decodes each accepts entry and requires eq with the canonical value; requires each rejects entry to fail. Fixture files cover every row of the tr-grammar §7 table, including the Bool, BigInt-range, and ignored-field-default cases named in the exit criteria. Default thunks in fixtures are limited to a tiny built-in vocabulary the harness provides (e.g. a constant, a field copy), since fixtures are data, not code.

Property tests (proptest), closing plan 01’s testing section:

  • generator of random well-formed descriptors + well-typed values;
  • decode(encode(v)) eq v; encode determinism (byte-equal);
  • eq ⇒ same hash; compare total-order laws (reflexive, antisymmetric, transitive, total) including NaN and -0.0;
  • compare == Equal ⇔ eq.

Step 7 — descriptor JSON (descriptor.rs, serialization half)

The fixtures and the C ABI both need descriptors as data. The encoding follows tr-grammar §7’s own conventions, as if descriptors were Soil values (internally tagged sums, records):

{ "name": "SummaryRow",
  "strategy": {"tag": "Structural"},
  "body": {"tag": "Record", "value": [
    {"name": "count", "shape": {"tag": "U64"}, "ignored": null},
    {"name": "mean",  "shape": {"tag": "F64"}, "ignored": null} ] } }

TypeShape::Named serializes by name, not id (ids are per-instance); loading a descriptor list resolves names in two passes (declare all, then define all), which handles recursion for free.

This format is the seed of plan 02’s --types env.json contract; plan 02 freezes it in docs/contracts/soil0-cli.md, so it should be reviewed with that in mind, but it is not frozen by this plan.

Step 8 — the C ABI (capi.rs) and the demo

Surface (prefix soil_, verbatim rules: no global state, explicit init/teardown, no unwinding across the boundary):

SoilRuntime *soil_init(void);
void         soil_teardown(SoilRuntime *);

/* types: registered as descriptor JSON — one format everywhere */
int  soil_register_types(SoilRuntime *, const char *desc_json, SoilError **err);

/* values: opaque handles; constructors, accessors, refcounting */
SoilValue *soil_i64_new(int64_t);                 /* …one per scalar kind */
SoilValue *soil_record_new(SoilRuntime *, uint32_t type_id,
                           SoilValue *const *fields, size_t n, SoilError **err);
/* …sum_new, list_new, accessors (checked, error out-param)… */
SoilValue *soil_value_clone(const SoilValue *);
void       soil_value_free(SoilValue *);

/* derived ops and the JSON bridge */
bool   soil_eq(SoilRuntime *, const SoilValue *, const SoilValue *, SoilError **);
char  *soil_show(SoilRuntime *, const SoilValue *, SoilError **);   /* canonical JSON */
SoilValue *soil_decode(SoilRuntime *, const char *shape_json,
                       const char *value_json, SoilError **);

/* errors */
const char *soil_error_message(const SoilError *);
int         soil_error_kind(const SoilError *);
void        soil_error_free(SoilError *);

/* debug builds only */
size_t soil_debug_live_values(void);

int soil_main(int argc, char **argv, SoilMainFn);   /* trivial main wrapper */
  • SoilValue* is a leaked Box<Value>; clone/free manage it. Handles and the runtime are single-threaded (decision §1); documented in the header.
  • Every entry point wraps its body in catch_unwind; a caught panic becomes a SoilError of kind Internal — Rust panics never cross the boundary.
  • Registration goes through descriptor JSON rather than a C struct surface: one format for fixtures, plan 02’s env.json, and embedding, and the C API stays five functions instead of thirty (§8).
  • Header generated by cbindgen into soil_rt.h; cdemo/main.c registers a record type, constructs a value through the ABI, shows it, decodes it back, checks eq, frees everything, and asserts soil_debug_live_values() == 0. check.sh builds and runs it — the embeddability exit criterion.

Closures are not constructible over the C ABI in this plan (soil_closure_new arrives with compiled code, plan 05); the demo therefore uses map-free, ignored-free types, and Rust tests cover the rest.

Step 9 — docs pass

Rustdoc on Runtime, Value, Ref, Registry, TypeShape, the JSON modules (stating the canonicality guarantees and the FNV-1a definition), and capi (threading rule, error contract). Exit criterion: plan 02 can interpret the prelude against this API without runtime changes, judged by walking plan 02’s scope list against the rustdoc.


5. Exit-criteria traceability

Plan 01 exit criterionWhere it lands here
C demo constructs, round-trips, tears down leak-freeStep 8 (cdemo + debug counter)
Every tr-grammar §7 row implemented, incl. Bool, BigInt range, ignored defaults, backed by fixturesSteps 5–6 (encoder table, fixtures corpus)
Descriptor API documented for plan 02Steps 3, 7, 9

6. Non-goals (restated from plan 01)

Execution of any kind, Perceus reuse, cycle collection, pyo3, refinements. Additionally out of scope here: C-ABI closure construction (plan 05), stable TypeIds across runs (names are the stable identity), and any performance work beyond the size assertion on Value.


7. Risks and checks

  • Canonicality regressions are spec violations, not bugs of degree; the determinism property test and fixture byte-comparisons run in every check.sh.
  • ryu output drift (crate update changing formatting) would break canonical bytes: the fixtures pin the expected strings, so an update that changes output fails loudly; the lockfile pins the version.
  • serde_json float parsing is imprecise without the float_roundtrip feature — found by the property tests during implementation (§8.15). The feature is on; the round-trip property guards against regression.
  • Map comparator misbehavior (a comparator that is not a total order) silently corrupts the sorted-vec invariant. The runtime does not detect it (that is the refinement checker’s future job); documented on MapVal.
  • Descriptor/value mismatch through the C ABI (wrong field count, wrong scalar kind) must be a checked SoilError, never UB: record_new and friends validate against the descriptor.

8. Micro-pins — approved 2026-08-22

Per the overview’s rule that new choices go to the user, these were presented as proposals and approved by the user on 2026-08-22 (item 12 added at approval time). They are pinned; reopening one is a new decision point.

  1. String escape set (canonical JSON): escape exactly " , \, and control characters U+0000–U+001F; use the short forms \n \r \t \b \f where they exist and lowercase \u00xx otherwise; all other characters (including non-ASCII) are raw UTF-8; no \/. (RFC 8785’s choices; deliberately not Soil’s source escape set, which is a different layer — soil-syntax-spec §1.)
  2. Whitespace: none. Separators are , and : only.
  3. NaN and zero: all NaN bit patterns are one logical value (equal, sorts after +Inf); -0.0 and 0.0 are distinct with -0.0 < 0.0, and ryu renders them -0.0 / 0.0, which round-trip. NaN collapses because canonical JSON has a single "NaN"; zeroes stay distinct because the encoding distinguishes them.
  4. Base64 for Bytes: standard alphabet, with padding, no line wrapping; decode rejects non-canonical padding/alphabet.
  5. Strict decode: unknown record fields, missing non-ignored fields, and present ignored fields are all DecodeErrors. Rationale for the last: encode omits them, so accepting them would admit a second, unverifiable source for a field whose value is defined to be reconstructed (design §3.7); one canonical form in both directions.
  6. Default-thunk arity: an ignored field’s default closure receives the record’s non-ignored fields, in declaration order, as its arguments (“the record’s other fields in scope”, design §3.7, restricted to non-ignored to avoid ordering dependencies among ignored fields).
  7. Duplicate keys in decoded JSON: rejected (DecodeError), not last-wins. Requires walking with a duplicate check since serde_json’s default map is last-wins — use its preserve_order/raw-value facilities or a custom visitor; whichever is chosen, the observable rule is “duplicates reject”.
  8. Debug leak counter: thread-local (instances are single-threaded), debug builds only, exposed as soil_debug_live_values().
  9. Panic policy: the crate itself never panics on valid API use; catch_unwind at the C ABI converts bugs to SoilError::Internal. Rust-side consumers (soil0) get Result everywhere.
  10. C-ABI type registration via descriptor JSON (not a C struct builder API): one descriptor format across fixtures, env.json, and embedding.
  11. Descriptor JSON is reviewable but not frozen here; plan 02 freezes it inside docs/contracts/soil0-cli.md as the --types contract.
  12. All dev dependencies are managed through a repo-root shell.nix (§2): it is the single pin for the Rust toolchain (no rust-toolchain.toml) and provides every tool (cbindgen, the C compiler, later Z3/Python); building outside nix-shell is unsupported. Cargo continues to manage Rust crate dependencies. Location and single-pin choice resolved with the user 2026-08-22.

Items 13–15 were added during implementation (2026-08-22, autonomous session) — recorded here and flagged for user review:

  1. Decoded maps carry the derived structural order. A map arriving through JSON has no program-supplied comparator closure to carry, so MapVal orders are Structural (the derived compare) or Custom(closure); decode always builds Structural, map operations take the runtime (structural order consults the registry), and the order is structure, not content — eq/compare see only the entries. Map decode accepts entries in any order (re-sorted) but rejects duplicate keys.
  2. Ignored-default vocabulary. FieldDesc.ignored holds an IgnoredDefault: Const (a constant of the field’s own shape, scalar-only), CopyField (a non-ignored, same-shaped field), or Native (an arbitrary embedder thunk — what a Soil default expression eventually compiles to). Const/CopyField are the serializable subset used by descriptor JSON and fixtures; Native does not serialize. The C ABI’s soil_record_new takes non-ignored fields and refills ignored ones, mirroring decode.
  3. serde_json needs its float_roundtrip feature. The default float parse is not correctly rounded (1.8821735589659427e48 re-parses to different bits), which breaks canonical round-trips; the round-trip property test caught it. The feature is enabled and pinned in Cargo.toml with a comment.

9. Open questions

  • None currently. §8 was approved 2026-08-22; reopening any of its items is a new decision point for the user.

Plan 02 — soil0: the minimal Soil implementation

References: docs/soil-syntax-spec.md (the parser’s contract), design §3 (language), bootstrap plan §1.

Goal

A deliberately small Rust implementation of Soil — parser, static checks, tree-walking interpreter over soil-rt values — whose every pass is a JSON-in/JSON-out CLI command. It is the execution engine for the entire bootstrap and the permanent differential oracle for soilc. Resist every temptation to make it good; make it correct and small.

Scope

  1. The AST JSON schema — a spec artifact, written first. The exact JSON encoding of the surface AST and of check results. This is the CLI oracle contract that soilc’s Trellis type definitions must later reproduce, so it follows the tr-grammar §7 value-encoding conventions (internally tagged sums, records) as if the AST were already Soil data. Deliverable: docs/contracts/soil0-cli.md, reviewed before code.
  2. Lexer + parser. Hand-written recursive descent implementing docs/soil-syntax-spec.md exactly: the and binding/connective disambiguation, non-associative comparisons, parenthesized non-tail match, .. record patterns, the fixed escape set, decreases lines. Errors carry source spans.
  3. Renamer. Scope resolution, the no-shadowing rule, :: qualification, _private visibility.
  4. Type + effect inference. ML inference (HM with records and sums from registered type shapes — no row-polymorphic records needed), effect rows as sets with row variables, subsumption at calls, operator-notation elaboration (== → T::eq, arithmetic → per-type primitives at monomorphic types only). No refinements (parsed, retained in the AST, otherwise ignored). No termination checking: every self-recursive or let rec definition conservatively acquires div (consequence handled in plan 04) — recorded by check as an unverified-termination fact rather than a hard error, so total-signed examples still pass (the demotion philosophy; soil0-cli §8.5, approved 2026-08-22).
  5. Exhaustiveness + redundancy checking for matches.
  6. Interpreter. Strict tree-walk over soil-rt values; closures; capability primitives implemented natively (fs_read_bytes, clock, rand, env, proc) plus the prelude’s fake capability constructors (fake_fs, seeded fake_rand, fake_clock); arithmetic follows floor division/modulus and panics (as SoilError) on overflow and zero divisors — obligations don’t exist yet, so these are always-on runtime checks, matching debug-mode semantics.
  7. CLI. soil0 lex|parse|rename|infer|check|run|test, each reading source (or AST JSON) and emitting schema-conformant JSON on stdout, errors as structured JSON on stderr, nonzero exit. run takes --entry name --args <json-array> and prints the canonical JSON result; test executes a JSON test bundle (assembled by the daemon) and reports per-case results in the lock’s result vocabulary.

Non-goals

Refinement checking, termination checking, codegen, optimization of any kind, .tr parsing (daemon’s job, plan 03), Python FFI.

Testing

  • Golden tests: a corpus of .soil fragments → expected AST/type/effect JSON, including every syntax-spec static rule as a rejection test (one test per rule: shadowing, redundant arm, a < b < c, missing .., parameterless let violation, bad escape…).
  • examples/read_file.soil and examples/csvstats/median.soil must parse, check, and (with stub callees) run.
  • Interpreter: expect-style tests mirroring the .tr examples’ test blocks, run through soil0 test.
  • Fuzz the parser (cargo-fuzz or a simple generator) for panic-freedom.

Exit criteria

  • docs/contracts/soil0-cli.md exists and every command conforms to it.
  • Both checked-in .soil examples check and run with correct results.
  • The rejection-test corpus covers every static rule in docs/soil-syntax-spec.md §5.
  • A downstream consumer (plan 03) can drive lex→check→test entirely through the CLI without linking soil0 as a library.

Decision points — resolved 2026-08-22

  • Library + CLI: the daemon links soil0 as a crate for in-process checking, but the CLI is the frozen compatibility contract — oracle and conformance tests always go through the CLI, never the library API.
  • Spans: UTF-8 byte offsets are canonical (start/end), with derived line/col included alongside for display.
  • Type environment: commands take --types env.json — type descriptors in the schema’s own encoding. Early tests hand-write it; the daemon generates it from .tr files later. soil0 never parses .tr.

Implementation plan

A detailed implementation guide exists at impls/02-soil0-impl.md. It records a second round of decisions resolved 2026-08-22 — union-find inference with naive env-scan generalization, Spanned wrapper records in the AST JSON, a signature-only frozen contract for soil0 infer (with a non-contractual --dump-ast), and a machine-written program.json manifest referencing env.json — plus the module map, build order, the required contents of docs/contracts/soil0-cli.md, a set of micro-pins (its §8, approved 2026-08-22), and three spec gaps it surfaced, all resolved 2026-08-22 (its §9): constructor qualification Type::Ctor with the exactly-one-spelling rule (now soil-syntax-spec §5.9, design §3.15), the fake_fs example corrected to the canonical map encoding, and the list primitives adopted as native-backed prelude surface (plan 04). Step 1’s deliverable, docs/contracts/soil0-cli.md, was reviewed and frozen as contract v1 on 2026-08-22 (amended v1.1 on 2026-08-23: typed holes, print). All implementation steps (2–11) completed 2026-08-23: every contract command is live, both normative examples check and run with correct results (median through a real insertion sort; read_file through fake and real capabilities), the rejection corpus is audited by an executable checklist, and the canonical printer is byte-identical on the examples with parse → print a fixpoint. The exit criteria are met; plan 03 is unblocked.

Implementation Plan 02 — soil0

Detailed implementation guide for ../02-soil0.md. That plan states the scope and exit criteria; this one states how the crate is structured and built, records the implementation decisions resolved with the user on 2026-08-22, and proposes the remaining micro-details (§8) and newly surfaced spec gaps (§9) for review before code is written. References: docs/soil-syntax-spec.md (the parser’s contract), tr-grammar §2.3 (predicate language) and §7 (value encoding), lock-schema §5 (test-result vocabulary), design §3, impl plan 01 (the runtime this builds on).


1. Resolved implementation decisions (2026-08-22)

DecisionChoiceRationale
Inference engineMutable union-find for unification; generalization by scanning the environment’s free variables (no levels)No substitution composition to get wrong (the textbook-W error source), and no level bookkeeping (the classic subtle-bug source in the OCaml approach). Env scanning is quadratic — “slow is accepted” (plan 02). Signatures are mandatory on every definition (syntax spec §3.1), so inference is mostly checking against declared types; the machinery can stay small.
AST spansEvery node is a Spanned wrapper record: {"span": …, "item": <tagged sum>}Expressible as one generic Soil record type Spanned a, so soilc’s Trellis AST types (plan 05) reproduce one wrapper rather than a span field repeated per variant or a node-id side table soilc would have to renumber identically. Follows “as if the AST were already Soil data” (plan 02 scope 1).
soil0 infer frozen contractSignature only: inferred type + effect row per definition, type variables canonically renamed. A --dump-ast flag additionally emits the elaborated, per-node-typed AST, explicitly non-contractualBehavioral divergence in local types or operator elaboration surfaces in run/test results, which are already differential oracles; freezing per-node annotations would constrain soilc’s inference internals and drag normalization rules into the contract. The dump exists because localizing a soilc divergence without it means bisecting by hand. Signature write-back into .tr files is the daemon/agent’s job (design §4.6) — soil0 never sees .tr.
Program inputA program manifest program.json: {"types": "env.json", "defs": [paths…]} — types by reference, definitions as an ordered file listOne growable place for future options instead of a widening flag surface. The manifest is machine-written: the daemon (plan 03) generates it per invocation; before the daemon exists, soil0’s own test harness generates it; handwritten manifests exist only as checked-in test fixtures. Types stay a reference so the descriptor format remains exactly the one plan 02 froze (--types env.json), usable alone by single-file commands. Inline embedding was rejected as duplicating the contract; a stdin stream was rejected as inventing a multi-definition container the language deliberately lacks.

None of these are language-observable (the CLI contract itself is recorded in docs/contracts/soil0-cli.md, step 1), so per the plan-01 precedent they live here and not in docs/design.md.


2. Crate skeleton

A new workspace member beside soil-rt; per plan 02’s resolved decision, the crate builds a library (linked by the daemon for in-process checking; API unstable) and a binary (the frozen compatibility contract; oracle and conformance tests always go through it).

rust/
  Cargo.toml           -- members = ["soil-rt", "soil0"]
  check.sh             -- extended: soil0 fmt/clippy/test + CLI conformance run
  soil0/
    Cargo.toml         -- lib + [[bin]] soil0
    src/               -- module map in §3
    tests/             -- golden harness + corpus (§6 of ../02-soil0.md)
      golden/          -- accept corpus: <case>/{*.soil, program.json, env.json, expected/*}
      reject/          -- rejection corpus: one case per static rule, keyed by error code
      bundles/         -- test-bundle fixtures for `soil0 test`

Dependencies, kept deliberately short (§8.11):

CrateWhyWhere
soil-rt (path)values, registry, canonical JSON, SoilErrorruntime
serde, serde_jsonCLI input/output JSON (manifests, bundles, AST, errors) — Soil values embedded in outputs still go through soil-rt’s canonical encoder, never serde_jsonruntime
proptestparser panic-freedom + determinism properties (§8.12)dev-only

Argument parsing is hand-rolled (§8.11): seven subcommands and a handful of flags do not justify a dependency. All tools come from the repo-root shell.nix (impl plan 01 §8.12); the toolchain stays the single pinned stable Rust.


3. Module map

src/
  lib.rs         -- pipeline entry points (parse_file, check_program, …); unstable API
  span.rs        -- Span {start, end, line, col}; byte offsets canonical (plan 02)
  diag.rs        -- Diagnostic {code, message, span, file, notes}; JSON rendering (§8.1)
  token.rs       -- token set
  lexer.rs       -- hand-written; owns the escape-set and literal rules
  ast.rs         -- surface AST: Spanned<T>, defs, types, rows, exprs, patterns, predicates
  ast_json.rs    -- AST <-> JSON exactly per docs/contracts/soil0-cli.md
  parser.rs      -- recursive descent per soil-syntax-spec §3
  kernel.rs      -- kernel typedefs (contract §6.1), builtin/derived/numeric name tables (§11)
  manifest.rs    -- program.json + env.json loading (type descriptors with params, §8.6)
  rename.rs      -- scopes, no-shadowing, `::` resolution, `_private` visibility, reference sets
  types.rs       -- semantic types: union-find vars, rows (effect set + optional tail var)
  infer.rs       -- checking/inference, row constraint solving, operator elaboration
  elab.rs        -- the elaborated AST the interpreter consumes; --dump-ast rendering
  exhaust.rs     -- usefulness-based exhaustiveness + redundancy (Maranget)
  interp/
    mod.rs       -- tree-walk evaluator over soil-rt values; trace frames
    env.rs       -- cons-list environments; let-rec knot-tying
    builtins.rs  -- the provisional builtin table (§8.7): numerics, conversions, lists
    caps.rs      -- real capability primitives + fakes (fake_fs, fake_rand, fake_clock)
  testrun.rs     -- test-bundle execution, result vocabulary per lock-schema §5
  cli.rs         -- subcommand dispatch, exit codes, stdout/stderr discipline
  main.rs        -- thin wrapper over cli.rs

Dependency direction is a pipeline: cli → testrun → interp → {elab, exhaust} → infer → rename → parser → lexer, everything using span, diag, ast, manifest. Nothing reaches back; soil-rt is below all of it.


4. The CLI contract: docs/contracts/soil0-cli.md

Written and reviewed by the user before any code (plan 02 scope 1). It is the freeze point for everything plan 05’s soilc must reproduce and plan 03’s daemon consumes. Required sections:

  1. Conventions. Invocation shapes, which commands take a source file vs a manifest, stdout is exclusively the command’s JSON result, structured errors on stderr, exit codes (§8.1), and the compact single-line output rule (§8.13).
  2. Span and diagnostic schemas (§8.1, §8.2).
  3. The AST JSON schema, the largest section: every node kind for definitions, types, rows, expressions, patterns, and predicates (parsed and retained per plan 02), in tr-grammar §7 conventions — internally tagged sums, records, the Spanned wrapper. Predicates use the tr-grammar §2.3 grammar verbatim.
  4. env.json: the type-environment schema — soil-rt’s descriptor JSON generalized with type parameters and aliases (§8.6). This is where plan 02’s “the daemon generates it from .tr files later” contract is frozen.
  5. program.json (§8.5).
  6. Per-command output schemas: token list (lex), AST (parse), reference sets (rename, §8.4), signatures (infer/check, §8.3), canonical-JSON result (run), per-case results (test, §8.10).
  7. The builtin table (§8.7): every native definition’s name, Soil signature (effects and capabilities included), and semantics — including the fakes’ determinism guarantees (§8.8).
  8. Non-contractual surfaces, listed explicitly: --dump-ast output, the library API, and anything else free to change.

5. Build order

Steps are sequential; a step is done when its tests pass and check.sh is green. Golden and rejection tests shell out to the built soil0 binary rather than calling the library — the harness thereby proves continuously that a downstream consumer can drive everything through the CLI, which is an exit criterion.

Step 1 — docs/contracts/soil0-cli.md

The contract document (§4), reviewed before code. Micro-pins §8.1–§8.10 land in it; user approval of this plan plus that document unblocks everything below.

Step 2 — skeleton, spans, diagnostics, CLI shell

Workspace member, span.rs, diag.rs, cli.rs with all seven subcommands wired to stubs, exit-code discipline, check.sh extension. Tests: diagnostic JSON golden, exit codes.

Step 3 — lexer and soil0 lex

The token set of syntax-spec §1: keywords (effect names are not keywords — reserved in type position only, handled by the parser), identifier classes (ident, private-ident, TypeName), operators including .. and :: under maximal munch, -- comments, integer literals with underscores, float literals, and the string escape set — exactly \" \\ \n \r \t \u{1–6 hex}, rejecting surrogates, values above U+10FFFF, and any other escape. Every token carries a span.

Tests: golden token streams; one rejection per lexical rule (bad escape, surrogate, unterminated string, stray character).

Step 4 — parser, AST JSON, and soil0 parse

Hand-written recursive descent per syntax-spec §3, with the named disambiguations:

  • and: after and, the two-token lookahead defname "=" continues the enclosing let’s bindings; anything else parses and as the boolean connective (spec §3.3).
  • Non-associative comparisons: a < b < c is a parse error with its own code (spec §4).
  • Match arm extent: an arm body extends maximally; | attaches to the innermost open match. The “parenthesize non-tail nested matches” rule is a consequence of this attachment, not a separate check.
  • Row vs return type: after ->, a lowercase ident followed by another type is a row variable; alone it is the return type (spec §3.2).
  • Named domain: ( then ident : begins a named parameter type; otherwise a parenthesized type.
  • Refinements: { in type position opens a refinement; its predicate is parsed with the tr-grammar §2.3 grammar and retained in the AST.
  • decreases lines between signature and equation; signature/equation defname agreement is a parse-level error.

ast_json.rs implements §4.3’s schema in both directions (parse emits; rename/infer accept AST JSON per plan 02’s CLI list).

Tests: golden ASTs for the syntax-spec §7 worked examples and both checked-in .soil examples; parse-level rejection cases (nonassoc comparison, missing .. is not here — it is a checker rule — but malformed record patterns, bad decreases placement, name mismatch are); proptest properties — arbitrary byte input never panics, parsing is deterministic, spans are well-formed and properly nested (§8.12).

Step 5 — renamer and soil0 rename

Operates on a manifest (all defs). Scope stack per definition: equation parameters, let/fun binders, pattern binders, over an outer scope of manifest definition names and builtins.

  • No shadowing (spec §5.1): any binding of a name already in scope — including a parameter colliding with a manifest definition or builtin — is an error. _ binds nothing, is exempt, and is rejected in let rec.
  • :: resolution: module::def resolves iff a manifest definition def lives in directory module; Type::derived resolves against the env’s types and the fixed derived-function set (eq, show, compare, hash); derivability itself is checked later (infer), existence here.
  • _private visibility: private-ident definitions are legal only in files named _private.soil, and callable only from definitions whose file shares that directory (design §4.2). Module identity is the definition file’s parent directory (§8.5).
  • Success output: per-definition reference sets — the definitions, types, builtins, and private helpers it references (§8.4). This is the computed import set design §6.1 and the lock’s call-edge data need, so it is load-bearing for plan 03, not debug output.

Tests: golden reference sets; rejections for every scoping rule (shadowing in each binder position, _ in let rec, private call across modules, private def outside _private.soil, unknown name, unknown qualification).

Step 6 — types, effect rows, inference, elaboration; soil0 infer

The largest step. Semantic types in types.rs: scalars and List/Map as builtin constructors, named records/sums/aliases from the env (instantiated at their parameters), arrows carrying a row, unification variables as union-find indices. Rows are an effect set (div, panic, io, ffi) plus at most one tail variable; ffi ⇒ panic is normalized at construction.

Checking discipline: every definition carries a declared signature, so the checker verifies the body against it, using callees’ declared signatures (already checked — manifest order is callee-first, §8.5). Within a body: HM inference with union-find; generalization only at let (scan env free vars, no levels); no polymorphic recursion.

  • Effect rows: collect subset constraints (callee row ⊆ context row) during inference; solve by fixpoint propagation after unification. Row-polymorphic higher-order signatures (map : (a -> e b) -> List a -> e (List b) style) instantiate their tail variables per call site, so map over a total function stays total (design §3.3).
  • Recursion: any self-recursive definition or let rec group conservatively acquires div (plan 02 scope 4); a declared row lacking div yields termination: "unverified" in the check facts rather than an error (soil0-cli §8.5, approved 2026-08-22; io/ffi deficits remain hard errors). decreases lines are parsed and ignored.
  • Operator elaboration (spec §5.4): at the zonked monomorphic type, ==/!= → T::eq, comparisons → T::compare, arithmetic and unary minus → per-type numeric primitives, and/or → short-circuit Bool builtins, not → Bool builtin. An operator whose operand type is still a variable or is polymorphic is an error (“operators are notation, not overloading”). Derivability (no closures in the type) is checked here for ==/!=/comparisons and explicit Type::derived uses.
  • Literals: type = the immediately enclosing annotation if present, else the default (I64, F64, Utf8); integer literals are range-checked against that width at check time (§8.14).
  • Field access: resolved against the zonked record type; if still a variable, an “annotation needed” error (no row-polymorphic records, plan 02 scope 4).
  • Constructors: exactly-one-spelling resolution (syntax-spec §5.9): a bare constructor resolves iff its variant name is unique among the env’s sum types; on a collision the qualified Type::Ctor form is required, and qualifying a unique constructor is an error. Rejection tests cover all three failure modes (ambiguous bare, unknown qualified, needlessly qualified).
  • Refinements: erased to their base types before checking (parsed, retained, otherwise ignored — plan 02 scope 4).

Output (§8.3): per-definition inferred signature — structured type JSON plus effect row, type variables canonically renamed in order of first appearance. --dump-ast additionally emits the elaborated AST with per-node types, marked non-contractual.

Tests: golden signatures and check-facts for all examples (undeclared-div recursion and panic obligations are golden facts, not rejections — soil0-cli §8.5); rejections for io/ffi subsumption violations, operator-at-polymorphic-type, literal out of range, arity/type mismatches, non-derivable ==; a dedicated HOF row-propagation suite (sum_lengths, map-over-total, map-over-panicking).

Step 7 — exhaustiveness, redundancy; soil0 check

Maranget-style usefulness over the pattern matrix: constructor sets from the env’s sum types, records with the mandatory-.. rule (a partial record pattern without .. is an error, spec §3.4), literal patterns useful only under a wildcard/binder default. Non-exhaustive matches report a witness pattern in the diagnostic; redundant arms are errors per arm.

soil0 check is the full static pipeline — parse → rename → infer → exhaustiveness — over a manifest; its success output is infer’s (§8.13), its failure output the combined diagnostics.

Tests: exhaustiveness/redundancy golden + rejection cases (missing variant, missing .., redundant arm, literal match without default); this completes the “every static rule in syntax-spec §5” corpus obligation together with steps 5–6 (checklist in step 10).

Step 8 — interpreter

Strict tree-walk over the elaborated AST, producing soil-rt values.

  • Environments are persistent cons-lists (env.rs); let rec closures tie the knot via Rc<RefCell<…>> backpatching.
  • Closures are soil_rt::Closure::native wrapping Rc’d elaborated AST plus captured env — the single invocation path impl plan 01 §1 committed to.
  • Arithmetic: hand-rolled floor division and floor modulus on integers (Python semantics; note Rust’s div_euclid is not floor for negative divisors — §7 risks); checked overflow everywhere; overflow and zero divisors raise SoilError (Overflow, DivideByZero) — always-on runtime checks matching debug-mode semantics (plan 02 scope 6). F64 operators are IEEE 754. BigInt primitives are total.
  • Ground-instance registration: env types are parameterized, but soil-rt’s registry holds ground descriptors; the interpreter instantiates and registers ground instances on demand behind a cache keyed by the instantiated shape (§8.6). Every runtime value thereby carries a real TypeId, so show/eq/decode work unchanged.
  • Trace frames: entry to every definition pushes a TraceFrame; raised SoilErrors carry the stack (the design §3.11 debug payload).
  • Capabilities (caps.rs): opaque values wrapping native handles. Real primitives (fs_read_bytes, clock, rand, env, proc) and the fake constructors (fake_fs, seeded fake_rand, fake_clock) per the builtin table (§8.7, §8.8).

Tests: expect-style unit tests over the interpreter for each expression form; the arithmetic golden table (all sign combinations of / and %, overflow edges, I64::MIN / -1); fake-capability behavior (seeded rand reproducibility, fake_fs hit/miss).

Step 9 — soil0 run and soil0 test

  • run --entry name --args <json-array>: decodes each non-capability parameter from the array in order (type-directed, via soil-rt decode; entry parameters must be ground types), injects real World-derived capabilities for capability-typed parameters (mirroring trellis call, tr-grammar §3.5 — §8.9), prints the result’s canonical JSON on stdout; a runtime SoilError is reported as a structured error on stderr with nonzero exit.
  • test <bundle.json>: executes a daemon-assembled bundle — per case: construct with-bound fakes, decode args, call the definition under test, compare outcomes (§8.10) — and reports per-case results in the lock’s vocabulary: pass | fail | xfail | xpass (lock-schema §5). panic expectations are legal outcomes matching any raised SoilError.

Tests: bundle fixtures mirroring the .tr examples’ test blocks (read_file.tr’s fake-fs cases, median.tr’s expect cases), xfail/xpass reporting, capability injection.

Step 6b — typed holes (contract v1.1, adopted 2026-08-23)

?name per syntax-spec §3.3/§5.11 and design §3.16: lexer hole token, Expr::Hole (contract §4.3), checker support (a hole is a fresh variable; its zonked goal type reported in checks.holes with canonical renaming; duplicate hole names error), and the unfilled-hole refusal rule for run/test (lands with step 9). Tests: goal-type reporting, duplicate names, hole-free outputs byte-identical to v1.

Step 10 — corpus completion, examples, conformance pass

  • The rejection corpus is audited against a checklist enumerating every static rule: syntax-spec §5.1–§5.8 plus every lexical/parse rule from steps 3–4 — one case per rule, keyed by error code.
  • examples/read_file.soil checks and runs against an env defining Path, FsError, and a fake fs; examples/csvstats/median.soil checks and runs with stub callees — len/nth as thin Soil wrappers over the list builtins and sort_by as a real Soil insertion sort (~10 lines), which doubles as the integration test for recursion, closures, and comparison elaboration.
  • Full proptest run; docs/contracts/soil0-cli.md conformance re-review command by command; rustdoc pass on the library surface (marked unstable) and the builtin table.

Step 11 — canonical printer (soil0 print, adopted 2026-08-23)

Before plan 03 computes the first soil_hash (design §4.8/§6.1): write the canonical-form section into soil-syntax-spec (user-reviewed — it is spec), implement the printer as a pass with parse → print → parse idempotence as a conformance test (proptest over the golden corpus), and wire soil0 print <file.soil> (contract v1.1’s reserved command). The alpha-normalization for hashing (local binders as indices) is specified with it; the daemon’s write_soil canonicalizes through this printer.


6. Exit-criteria traceability

Plan 02 exit criterionWhere it lands here
docs/contracts/soil0-cli.md exists; every command conformsStep 1 (written, reviewed), step 10 (conformance pass)
Both checked-in .soil examples check and run correctlySteps 8–10 (interpreter, run, stubs)
Rejection corpus covers every syntax-spec §5 static ruleSteps 3–7 accumulate; step 10 audits against the checklist
Plan 03 can drive lex→check→test through the CLI aloneThe golden/rejection/bundle harness shells out to the binary exclusively (§5 preamble)

7. Risks and checks

  • Floor vs Euclidean vs truncating division. Rust’s / truncates and div_euclid differs from floor for negative divisors (7.div_euclid(-2) == -3 but Python 7 // -2 == -4). The arithmetic golden table (step 8) covers all sign combinations and is generated once from the reference semantics (Python) rather than written from memory.
  • Row-constraint solving subtleties. HOF effect propagation is where hand-rolled row systems typically go wrong (a dropped tail variable silently widens or narrows a row). The dedicated HOF suite in step 6 exists for this; any fix adds a case.
  • The frozen CLI drifting from the evolving library. Prevented structurally: no test calls the library except the interpreter/infer unit tests; everything observable goes through the binary.
  • No-shadowing strictness surprising the corpus. Parameters colliding with builtin or definition names will be the most common agent error; the diagnostic must name the colliding binding site. Rejection tests pin the message shape.
  • Ground-instance cache correctness. Two structurally equal instantiations must map to one TypeId (or eq/show split behavior); keyed by fully-zonked shape, with a property test.
  • serde_json recursion limits on deep AST JSON: acceptable — depth-limited inputs fail loudly with a structured error, and the parser side (source text) is iterative-or-depth-checked like plan 01’s decode walk.

8. Micro-pins — approved 2026-08-22

Per the overview’s rule that new choices go to the user, these were presented as proposals and approved by the user on 2026-08-22. They are pinned; each lands in docs/contracts/soil0-cli.md (step 1), and reopening one is a new decision point. The one language-observable item (§8.14, literal range checking) is also recorded in soil-syntax-spec §1; the rest are CLI- or crate-internal and live here.

  1. Diagnostics and exit codes. stderr carries {"errors": [{"code", "message", "file", "span", "notes": […]}]}; code is a stable kebab-case string (shadowing, redundant-arm, nonassoc-comparison, missing-record-rest, …), one per rule — the rejection corpus keys on it. Exit codes: 0 success, 1 the input violates the spec (any diagnostic), 2 usage or internal error (malformed manifest/bundle, I/O failure, bug).
  2. Span JSON: {"start", "end", "line", "col"} — UTF-8 byte offsets canonical (plan 02 resolved), line/col 1-based and derived from start, display-only.
  3. Signature output (infer/check): structured type JSON (the AST type encoding), not a pretty string; effect rows as {"effects": […], "var": name|null} with effects in the fixed order div, panic, io, ffi; type and row variables renamed a, b, c, … in one sequence by order of first appearance. A pretty-printed string may accompany it but is non-contractual. Approved extension (2026-08-22, with docs/contracts/soil0-cli.md §8.5): the output also carries per-definition checks facts — termination in the lock-schema §4 vocabulary plus panic_obligations — because div/panic deficits are recorded facts, not errors, while io/ffi subsumption remains enforced.
  4. rename output: per-definition reference sets {"defs", "types", "builtins", "privates"}, each sorted lexicographically — the computed import set for the lock and the graph view (design §6.1).
  5. Manifest semantics: defs must be ordered callee-before-caller (the daemon owns the dependency tree; disciplined order, design §4.6); a forward reference is a 2-class error, not a checker diagnostic. Module identity for _private visibility and module::def is the definition file’s parent directory name.
  6. env.json generalizes descriptor JSON. soil-rt descriptors are ground, but Soil type definitions are parameterized (Result a e) and include aliases (Path = Utf8), which TypeBody cannot express. env.json therefore extends the impl-plan-01 §4.7 format with "params": [tyvars…], a Var shape, and an Alias body. The checker consumes it directly; the interpreter registers ground instances on demand (cache keyed by instantiated shape) so runtime values carry real TypeIds and soil-rt stays unchanged. Aliases are checker-level and erased before registration. Revised 2026-08-22 (with docs/contracts/soil0-cli.md §6/§13.3): field, payload, and alias types use the semantic type encoding (SigType) rather than a generalized descriptor Shape — one type encoding for signatures and environments; SArrow rejected until a prelude type needs a function field; soil0 lowers SigType to runtime descriptor shapes at ground-instance registration.
  7. A provisional builtin table, enumerated in full in docs/contracts/soil0-cli.md: capability primitives and their fakes; utf8_decode and sibling conversions; per-type numeric primitives (the elaboration targets, including total BigInt arithmetic); and a minimal set of list primitives (list_len, list_nth, list_empty, list_append, …) — necessary because the grammar has no list literals or list patterns, so nothing list-shaped is writable in Soil without them. Resolved 2026-08-22 (was §9.3): these are prelude surface — native-backed prelude definitions like the fakes (plan 04 scope 2–3); the names frozen in docs/contracts/soil0-cli.md are the prelude names.
  8. Fake determinism: fake_rand seed is SplitMix64 (fixed, documented — outcomes are pinned in expect tests, so the algorithm is observable and must never drift); fake_clock t returns the constant t on every call; fake_fs files holds an in-memory Map Utf8 Utf8, fs_read_bytes returning the UTF-8 bytes of the stored string, misses returning the env’s FsError not-found variant.
  9. run capability injection: parameters whose type is a capability are injected from a real World in order; remaining parameters decode from --args positionally; entry parameters must be ground types. Mirrors trellis call (tr-grammar §3.5) so plan 03 reuses the semantics.
  10. Test-outcome comparison: the actual result is canonically encoded and byte-compared against the expected value’s canonical form (expected JSON is decoded type-directedly and re-encoded, so accepted non-canonical spellings normalize — same normalization as plan 01’s fixtures). panic expectations match any SoilError raised by the call; bundle-level problems (undecodable args, unknown definition) are 2-class invocation errors, never fail results.
  11. Dependency floor: soil-rt, serde, serde_json, dev-only proptest; hand-rolled argument parsing (no clap).
  12. Fuzzing via proptest generators on the pinned stable toolchain, not cargo-fuzz: cargo-fuzz wants a nightly toolchain and sanitizer plumbing, and shell.nix pins exactly one stable Rust (impl plan 01 §8.12). Properties: arbitrary bytes never panic the lexer/parser; generated token sequences never panic the parser; parse is deterministic. Revisit if coverage-guided fuzzing proves necessary.
  13. Output discipline: every command’s stdout is one compact (no-whitespace) single-line JSON document with field order fixed by the schema doc; golden tests byte-compare. Success output of check equals infer’s.
  14. Integer literals are range-checked against their annotated-or- default width at check time. Consequence: I64::MIN is not writable as a literal (9223372036854775808 overflows before negation is applied); accepted — a prelude constant covers it later, and the same wart exists in C and Rust’s literal grammars. Amended 2026-08-23: “or-default” is contextual, as design §3.6’s “as in Rust” intends — a literal adopts the type inference demands and defaults to I64 only when unconstrained; the fixed-at-lex reading broke the normative gcd example (b == 0 with b : U64). Recorded in syntax-spec §1.
  15. lex output: a JSON array of Spanned tokens, each an internally tagged sum ({"tag": "Ident", "value": {"name": …}}, literal tokens carrying their decoded value).

9. Spec gaps surfaced by this plan — all resolved 2026-08-22

Raised per the overview’s rule (“when implementation reveals a spec contradiction or gap, stop and raise it”); all three were resolved with the user on 2026-08-22 and propagated to the spec docs as noted below.

  1. Constructor-name ambiguity. Resolved 2026-08-22: Type::Ctor qualification with the exactly-one-spelling rule — bare iff the variant name is unique among the types in scope (bare then being the only legal form); qualified required on a collision; qualifying a unique constructor is an error. Recorded in soil-syntax-spec §3.3–§3.4 and §5.9, tr-grammar §2.3 (is tests), and design §3.15. Affects steps 4 (grammar: ctor-name, qualified production), 5–6 (resolution), and the rejection corpus.
  2. fake_fs argument shape. Resolved 2026-08-22: the example was wrong. examples/read_file.tr wrote a JSON object for a Map Utf8 Utf8; tr-grammar §7 encodes maps as arrays of {"key", "value"} pairs, and there is no object convenience form — one canonical shape. The example (and the design §3.5 illustration) now use the array form; soil0’s test bundles do the same.
  3. The provisional list primitives (§8.7). Resolved 2026-08-22: they are prelude surface — native-backed prelude definitions, names frozen in docs/contracts/soil0-cli.md and adopted by plan 04 (its scope 2 now records this).

Plan 03 — the Trellis daemon

References: design §4 (the specification layer, lowering, the daemon), §6 (hashing and locks), docs/tr-grammar.md, docs/lock-schema.md.

Goal

The long-running service that makes Trellis a language: parses .tr files, computes the hashes, maintains locks and the derived manifest, assembles context bundles, runs lowering jobs against an agent CLI with the MCP tool surface, and serves the CLI (and later the IDE and LSP). Execution is delegated to soil0 throughout.

Scope

  1. .tr parsing. Frontmatter, reserved fenced blocks, block mini-languages (signature, requires/ensures/invariant predicates, test call-arrow lines with with bindings, property, cram, reference, exports, allow), validity rules — all per docs/tr-grammar.md. Unknown blocks pass through as prose.
  2. Hashing. The three-part hash (formal/test/prose per tr-grammar §1), soil_hash with content addressing (free variables replaced by callee hashes, private helpers folded in), the recursive-type cycle hash. Canonicalization rules documented; hashes must be reproducible across machines.
  3. Locks and manifest. Read/write .lock sidecars per docs/lock-schema.md in canonical JSON; derive soil.lock by merge (definitions, soil_private nodes, escape-hatch audit, trusted packages); enforce the schema invariants (accepted gating, oracle acceptance rule, block-author pinning).
  4. Incremental state. Watch the Soil root; on change, recompute hashes, apply the invalidation table (lock-schema §8), update statuses.
  5. Context bundles. The fixed directory layout per design §4.6: spec.md, callees/ (signatures only; demoted refinements shown as base type + note), tests.json, examples/ (prelude corpus), reference.py or CLI-oracle stanza, previous.soil.
  6. MCP server + lowering jobs. The seven tools (read_context, check_types, check_refinements — a stub returning none until soilc, run_tests (sandboxed, fakes only), write_soil, read_spec — pinned spec sections by anchor, added 2026-08-25 with the daemon contract, ask_human); one headless agent invocation per lowering (Claude Code provider first) with turn/time/cost caps, isolated agent home, every tool call logged to f.log; ask_human ends the invocation, the answer is written into the .tr prose, and the job re-runs fresh; serial queue with invalidation when a dependency is edited mid-flight.
  7. Test running. Expect/property via soil0 test with fakes; contradiction pre-flight (mechanical same-input/different-output check before any tokens are spent); cram runner (temp dir, with file fixtures, literal output, [n] exit codes) and trellis call (real capabilities) — real mode never exposed to the lowering sandbox.
  8. Thin CLI. trellis check|status|lower|test|call|repl speaking to the daemon. REPL: call any definition with JSON args (Trellis level), feeding later REPL-to-test promotion.
  9. Telemetry. Tokens, cost, retries, provider, model per lowering, to f.log; provider/model into the lock’s lowering record.

Adoptions (2026-08-23, agentlanguages survey)

The survey adoptions (design doc, entries marked “adopted 2026-08-23”; record in docs/plans/extra/agentlanguages-adoptions.md) land here:

  • Repair-class registry, drift-gated (design §4.6): every daemon/LSP diagnostic carries a stable code mapping to a typed repair class plus a spec_ref into the sectioned Soil spec; CI fails when registry, docs, and emissions disagree. soil0’s contract §2 registry is the substrate.
  • Budgeted context packer (design §4.6): the bundle assembler is trellis context <def> --budget <n> with fixed priority (spec > tests

    callee signatures > corpus examples > module prose).

  • Decision blocks in bundles (design §4.3, tr-grammar §5.2): _project.tr decisions in every bundle, module decisions in the module’s; rule-shaped ask_human answers write back to decisions. tr-grammar §9’s open question (flag vs invalidate) was resolved 2026-08-24: per-entry hashes with reliance edges, editorial reclassification, and the triage sweep (design §4.3; impl plan §9.4).
  • Generated skill (design §4.6): trellis skill assembles the lowering skill from the toolchain’s registries + the pinned prelude; drift-gated in CI.
  • Toolchain pin, refuse-on-mismatch (design §8): the daemon refuses to lower/verify under a mismatched pin; trellis toolchain update is the explicit upgrade event.
  • partial(holes) lock state (design §3.16, lock-schema §4): the lowering loop treats the hole as the retry unit; budget exhaustion pauses with goals open instead of failing whole.
  • Vacuity probes in the contradiction pre-flight (design §4.5).
  • Canonical soil_hash (design §6.1): hashing consumes soil0’s step-11 printer (canonical text of the alpha-normalized AST); the lowerer’s write_soil canonicalizes.
  • Deferred but shaped here: speculative proof-delta tools, trellis mutants (lock fields reserved) — design §10.

Non-goals

The IDE (plan 06), LSP beyond bare diagnostics, refinement checking, concurrent lowerings, raw-API provider, hosted anything.

Testing

  • Hash stability: golden hashes over examples/; mutation tests (edit prose → only prose_hash moves, etc. — one test per invalidation row).
  • Lock round-trip: examples/*.lock re-emitted byte-identical.
  • A scripted fake agent provider (plays back canned tool-call sequences) to test the job loop, question channel, caps, and log without spending tokens; one live smoke test against real Claude Code headless.
  • End-to-end: a fixture project lowers mean.tr-sized definitions to green through the fake provider.

Exit criteria

  • examples/ fully round-trips: parse → hash → lock regeneration matches the checked-in sidecars (update examples if the daemon exposes spec drift — spec first, then code).
  • One real headless lowering of a trivial definition completes: bundle → agent → write_soil → checks → tests → lock, with f.log populated.
  • A question round-trip works: ask_human → IDEless CLI answer → prose diff → re-invoke → green.

Decision points — resolved 2026-08-22

  • IPC: JSON-RPC over a Unix socket ($XDG_RUNTIME_DIR/trellis/<root-hash>.sock); LSP speaks its own stdio transport; TCP is the future remote-serving path.
  • Language: Rust, in the rust/ workspace as trellis-daemon, linking soil0 and soil-rt (the soil0 CLI remains the oracle contract regardless).
  • Sandboxing: all three layers. Tool allow-list (the scope-6 MCP tools, no shell) + the agent CLI’s own sandbox and isolated home + a chroot-style OS jail around the agent process (unprivileged via user namespaces / bubblewrap in practice). Defense in depth from v1.
  • Incremental state: explicit refresh — hashes re-checked at request boundaries plus trellis refresh; the watcher arrives with the IDE milestone on the same invalidation code path.

Implementation Plan 03 — the Trellis daemon

Detailed implementation guide for ../03-daemon.md. That plan states the scope and exit criteria; this one states how the crate is structured and built, restates the decisions resolved with the user on 2026-08-22, and records the micro-details (§8, approved 2026-08-24) and the newly surfaced spec gaps (§9, all resolved — 2026-08-24, plus the step-7 batch on 2026-08-28 — and propagated to the spec docs). References: design §4 (specification layer, lowering, the daemon), §6 (hashing and locks), §8 (toolchain pin), docs/tr-grammar.md, docs/lock-schema.md, docs/contracts/soil0-cli.md (frozen v1.1 — the surface this milestone drives), soil-syntax-spec §9 (canonical form), impl plan 02 (the library this links), and docs/plans/extra/agentlanguages-adoptions.md (the survey items that land here).


1. Resolved decisions (2026-08-22, from plan 03)

Restated from the plan’s decision points; none are reopened here.

DecisionChoiceRationale (recorded in plan 03)
IPCJSON-RPC over a Unix socket at $XDG_RUNTIME_DIR/trellis/<root-hash>.sockOne transport for CLI and (later) IDE; the LSP speaks its own stdio transport when it arrives; TCP is the future remote-serving path.
LanguageRust, workspace member beside soil-rt/soil0, linking bothThe daemon does in-process checking through the soil0 library; the soil0 CLI remains the oracle contract regardless (contract §12).
SandboxingAll three layers: MCP tool allow-list + the agent CLI’s own sandbox and isolated home + an OS jail (unprivileged user namespaces / bubblewrap) around the agent processDefense in depth from v1.
Incremental stateExplicit refresh — hashes re-checked at request boundaries plus trellis refresh; no watcher until the IDE milestoneThe watcher later reuses the same invalidation code path.

Also load-bearing and already settled elsewhere: soil_hash consumes the canonical printer (soil0 step 11 is done; parse → print is a fixpoint on the examples); the lowering loop treats typed holes as the retry unit with partial(holes) in the lock (design §3.16, contract v1.1); the MCP tools (seven since 2026-08-25 — read_spec added with the step-1 contract review) and the exit-and-reinvoke question channel (design §4.6); Claude Code headless as the first provider.

Boundary notes. No LSP server in this milestone — the plan’s goal line reads “serves the CLI (and later the IDE and LSP)”, and the repair-class registry (§4.4) is built now so the LSP inherits it later. No refinement checking (check_refinements is a stub returning none), no concurrent lowerings, no raw-API provider, no watcher, no trellis mutants (lock fields stay reserved), no speculative proof-delta tools (shaped by the registry and RPC design only).


2. Crate skeleton

One new workspace member. Per §8.1 the CLI and the daemon are one binary (trellis), so there is exactly one artifact to version, pin, and jail-mount.

rust/
  Cargo.toml           -- members += ["trellis"]
  check.sh             -- extended: trellis fmt/clippy/test + drift gates (§5 step 12)
  trellis/
    Cargo.toml         -- [[bin]] trellis; lib for tests
    registry/
      diagnostics.json -- error code -> repair class + spec_ref (§4.4); drift-gated
    src/               -- module map in §3
    tests/
      golden/          -- .tr parse trees, hashes, bundles, packed contexts
      reject/          -- one case per .tr validity rule, keyed by error code
      fixture_root/    -- a small Soil root (csvstats-shaped) for e2e through the fake provider
      scripts/         -- scripted-provider playbacks (§8.9)

Dependencies, kept deliberately short:

CrateWhyWhere
soil-rt (path)canonical value JSON for everything Soil-shaped in bundles and resultsruntime
soil0 (path)parser (types, predicates), checker, interpreter, test runner, canonical printer — in-process (§8.11)runtime
serde, serde_jsonRPC, locks, bundles, registry, telemetryruntime
sha2SHA-256 for every hash the lock storesruntime
pulldown-cmarkCommonMark fenced-block extraction for .tr (§8.3)runtime
tomlsoil.toml parsing (added at step 2, 2026-08-25: the table above missed that a TOML file needs a TOML parser; hand-rolling TOML repeats exactly the subtle-divergence risk the pulldown-cmark pin rejected)runtime
proptesthash canonicalization + packer propertiesdev-only

No async runtime, no clap, no YAML crate, no MCP SDK (§8.2, §8.3, §8.7). Tools added to the repo-root shell.nix: bubblewrap (the OS jail) and python3 (the reference-oracle runner, §9.9). The claude CLI is deliberately not managed by shell.nix — it is a user-supplied ambient tool, exercised only by the opt-in live smoke test (§5 step 11, §7).


3. Module map

src/
  main.rs        -- thin dispatch: client subcommands vs `daemon run` vs `_mcp`
  cli.rs         -- the thin client: parse args, ensure daemon, RPC, render
  rpc.rs         -- JSON-RPC 2.0 over ndjson frames; request/response types (§8.2)
  daemon.rs      -- socket server, connection threads, shared state, lifecycle
  config.rs      -- soil.toml (toolchain pin, tags, entrypoint) + providers.toml
  trfile/
    mod.rs       -- TrFile: frontmatter, block list, kind inference, validity rules
    frontmatter.rs -- the two-key YAML subset (§8.4)
    blocks.rs    -- fence extraction via pulldown-cmark; info-string parsing
    sig.rs       -- soil-sig via soil0's type parser
    predicate.rs -- requires/ensures/invariant/property via soil0's predicate parser
    tests_blk.rs -- test call-arrow lines, with-bindings; property; cram; reference
    exports.rs   -- exports + decisions block grammars
  diag.rs        -- daemon diagnostics: code, message, span, repair_class, spec_ref
  registry.rs    -- loads registry/diagnostics.json; the drift-gate helpers
  hash/
    mod.rs       -- the three-part hash (§8.5), soil_hash (§8.6), cycle hash
  lock.rs        -- lock model, canonical reader/writer (§8.7), schema invariants
  manifest.rs    -- soil.lock derivation by merge (lock-schema §7)
  state.rs       -- the definition graph, statuses, invalidation table, refresh
  envgen.rs      -- .tr types -> env.json (§9.1); program.json assembly
  testrun.rs     -- bundle assembly, expect via soil0, derived tests, results -> lock
  propgen.rs     -- property generators, where-filter constraints, seeds (§8.12)
  preflight.rs   -- contradiction check + vacuity probes
  oracle.rs      -- reference (python3 subprocess) and CLI oracles; differential tier
  cram.rs        -- cram runner: temp dir, fixtures, transcript match; `trellis call`
  bundle.rs      -- context-bundle assembly; the budgeted packer (§8.10)
  mcp.rs         -- the `_mcp` stdio shim + the seven tool implementations (§8.8)
  jail.rs        -- bubblewrap invocation, isolated home, env scrubbing (§8.9)
  provider.rs    -- Provider trait; claude-code provider; scripted fake provider
  job.rs         -- the lowering job loop: queue, caps, invalidation, outcomes
  question.rs    -- ask_human persistence, `trellis answer`, prose write-back
  telemetry.rs   -- f.log JSONL events (§8.14)
  repl.rs        -- `trellis repl` and `trellis call` evaluation path
  skill.rs       -- `trellis skill` assembly from registries + corpus

Dependency direction: cli/daemon → job → {provider, mcp, jail, bundle, question} → {testrun, preflight, oracle, cram} → {state, envgen} → {lock, manifest, hash} → trfile → soil0/soil-rt, everything using diag, registry, config, telemetry. Nothing reaches back up.


4. The contract surface: docs/contracts/trellis-daemon.md and friends

Written and reviewed by the user before code (step 1), because this milestone freezes three agent- or user-facing surfaces. The internal CLI↔daemon RPC is explicitly not contractual (both ends live in one binary and ship together).

  1. The MCP tool surface — the seven tools’ names, input schemas, output schemas, and error shapes: read_context, check_types, check_refinements (returns the Option encoding’s None until soilc), run_tests, write_soil, read_spec (pinned spec sections by anchor, served from the toolchain’s embedded doc copies — added 2026-08-25 at review), ask_human. This is what the generated skill documents and what plan 05’s strangler swap must keep stable.
  2. The context-bundle layout (design §4.6): the fixed directory tree — spec.md, decisions.md, callees/<name>.md (signature + refinement clauses with per-clause assurance; a runtime clause shown as base type plus a demotion note), tests.json (the soil0 bundle-case shapes, reused verbatim), examples/, reference.py or the CLI-oracle stanza, previous.soil — plus the packer’s priority order and drop rules (§8.10).
  3. The diagnostic registry format (§4.4 adoption): every daemon diagnostic carries code, repair_class, and spec_ref; registry/diagnostics.json is the single source; the doc renders the table from it. The soil0 contract §2 registry is the substrate: soil0 diagnostics pass through with their codes unchanged and gain repair_class/spec_ref by registry lookup.
  4. spec_ref anchors: the scheme <doc>#<section> pinned against the numbered sections of docs/soil-syntax-spec.md and docs/tr-grammar.md (e.g. soil-syntax-spec#5.9). Renumbering a referenced section becomes a drift-gate failure, which is the point.
  5. The question and answer schema: what ask_human writes, what trellis answer consumes, and the prose write-back format (§9.7).
  6. f.log event schema (§8.14) — documented because humans read it, marked non-contractual.

Separately, docs/soil-toml.md (status: prototype, like the grammar docs) specifies the v1 subset the daemon needs: [toolchain] (the pin — trellis, soil0_cli, skill hash; prelude hash arrives in plan 04), [tags] (the vocabulary .tr validation reads), entrypoint (parsed, unused until build targets). Everything else in design §8 stays future sections. trellis toolchain update rewrites [toolchain] to the running binary’s identity; any command that lowers or verifies under a mismatched pin refuses with the toolchain-mismatch diagnostic (design §8).


5. Build order

Steps are sequential; a step is done when its tests pass and check.sh is green. Everything through step 9 is testable hermetically; step 10 adds the scripted fake provider so the whole loop is CI-testable without tokens; only step 11 touches a real agent.

Step 1 — contracts and spec resolutions

docs/contracts/trellis-daemon.md and docs/soil-toml.md (§4), reviewed by the user. The §9 spec gaps were all resolved with the user on 2026-08-24 and are recorded in the spec docs; this step also lands their document artifacts that belong to plan 03 — syntax-spec §9.1 (the hash form) is already written, the contract v1.2 amendment is recorded, and the two new docs above are drafted here. Approval of those two documents unblocks everything below.

Step 2 — skeleton, lifecycle, RPC, tooling

Workspace member, main.rs dispatch, the ndjson JSON-RPC layer, the socket server with one thread per connection, trellis daemon run|stop|status, client auto-spawn (§8.1), shell.nix additions (bubblewrap, python3), .gitignore entries (.trellis/, *.log, soil.lock), check.sh extension. config.rs reads soil.toml and enforces the toolchain pin from the first command.

Tests: socket round-trip, concurrent clients, stale-socket recovery, auto-spawn, pin mismatch refusal.

Step 3 — .tr parsing

The full tr-grammar: frontmatter (name required, tags against the soil.toml vocabulary, unknown keys are errors), fenced-block extraction with reserved-language recognition (everything else is prose, including unreserved fences), the block mini-languages — soil-sig and the predicate blocks parsed through the linked soil0 parsers so there is exactly one grammar implementation for types and predicates; test (with-lines, call-arrow cases, JSON args via canonical decode), property (forall/where), cram (the §3.5 subset), reference (both forms), allow, exports, decisions, soil-type (via soil0’s type declarations). Kind inference and every §6 validity rule. @agent info-string markers recorded per block (tr-grammar §8).

Tests: golden parse trees for every examples/*.tr; a rejection corpus with one case per validity rule, keyed by daemon error code (unknown-frontmatter-key, undeclared-tag, name-mismatch, duplicate-test-name, missing-test-block, panic-without-row, …); proptest: arbitrary bytes never panic the parser, block spans partition the file.

Step 4 — hashing

The three-part hash over the §8.5 canonicalization; per-entry decision hashes and the scope-membership hash (§9.4); soil_hash per §8.6 — canonical printer output in hash form (§9.2), free references replaced by referent hashes, private helpers folded in; the recursive-type cycle hash. All hashes are sha256: + 64 hex (§9.10).

Tests: golden hashes over examples/; one mutation test per invalidation row (lock-schema §8): edit prose → only prose_hash moves; edit a test block → only test_hash; edit the signature → formal_hash; rename a local binder in the .soil → no hash moves (the alpha-normalization test); reformat the .soil → no hash moves (the canonical-form test); edit a callee body → the caller’s soil_hash moves; edit one decision entry → only that entry’s hash moves, the three-part hashes and the scope hash untouched; add an entry → only the scope hash moves. Reproducibility: hashes byte-identical across two runs and under a copied tree at a different absolute path.

Step 5 — locks and the manifest

The lock model exactly per lock-schema, the canonical writer (§8.7), the reader with schema validation, the §8 invariants enforced on write (accepted gating, oracle acceptance, block-author pinning, holes never tested); soil.lock derived by pure merge — definitions, soil_private nodes from private_helpers with ownership GC, the escape-hatch audit view, trusted_packages.

This step includes the examples regeneration pass (§9.10; policy resolved 2026-08-28): staged — step 5 regenerates the spec hashes, decisions, lowering records, and checks facts, carrying the existing test-result rows forward verbatim; step 7 re-runs every test and writes the final rows, and the byte-identity exit criterion is judged there. Scope: every .tr gets a real lock (mean.tr and parse_error.tr backfilled; mean.tr stays the not-yet-lowered exemplar with lowering: null), and parse_row.soil plus csvstats/_private.soil (_parse_cell) are authored so parse_row.lock’s private-helper and call edges describe real files — reviewed like any normative example. Provenance goes honest: human-verified with provider/model absent (the Soil is hand-written; no agent ran). Drift is reviewed with the user at each stage.

Tests: every regenerated examples/*.lock re-emitted byte-identical through read → write; invariant rejection cases; manifest merge golden including helper GC.

Step 6 — incremental state and the check pipeline

The definition graph (references from soil0 rename reference sets; tests’ oracle edges; type cycles), the status ladder derived per lock-schema §8, the full invalidation table applied on refresh, and the static pipeline: .tr validity → env/program generation → soil0 check in-process → check facts into the lock. Every diagnostic leaves through diag.rs and therefore carries repair_class and spec_ref (step 1’s registry).

Decision staleness is part of the derived pass (§9.4): compare lowering.decisions edges and scope hashes against current entry hashes; when a change is detected the CLI reports which lowerings it flags and names the reclassification escape.

CLI: trellis check [def], trellis status (human table; --json), trellis refresh, trellis decisions editorial <label> (re-stamp a detected change the human declares meaning-preserving; --module <m> when the label exists in both scopes).

Tests: invalidation-table cases end-to-end (edit file on disk → refresh → expected statuses); a review-suggested prose flow; the decision flows — edit flags only citers, addition flags the scope once, editorial re-stamp flags nothing; status goldens over the fixture root; registry coverage test — every code the daemon can emit is in the registry (the drift gate’s first half).

Step 7 — test running

  • env/program generation: .tr type files → env.json (§9.1 resolution), dependency-ordered program.json (callee-first — the daemon owns the tree; contract §7).
  • Expect tests → soil0 bundles, verbatim shapes (contract §10); results into the lock in the §5 vocabulary.
  • Derived tests (design §4.1, §4.5): properties from requires/ensures/invariant clauses; the differential tier from the reference attachment; named per §9.5.
  • Property runner (§8.12): random generators from types (where filters refused with unsupported-where-filter — §9.8; constrained generation arrives with Z3, plan 05), pinned seeds, predicates evaluated by synthesizing a Bool-returning wrapper definition and running it through the interpreter.
  • Differential oracles: Python references via a python3 subprocess driver, CLI oracles by pipe; both hashed into oracles and compared through canonical re-encode (§9.9).
  • Contradiction pre-flight + vacuity probes (design §4.5): the mechanical same-input/different-output check over decoded canonical args; a requires no test input satisfies; a syntactically trivial ensures; an expect set that never exercises a declared result variant. All run before any tokens are spent.

CLI: trellis test [def] (sandboxed tiers only).

Tests: bundle-assembly goldens against examples/ test blocks; property determinism under a fixed seed; each vacuity probe has a firing and a non-firing fixture; a deliberately contradictory pair of expect cases.

Step 8 — cram and trellis call

The cram runner: fresh temp dir per block, with file fixtures, sequential shell session, literal combined-output match, [n] exit codes; built targets and trellis call on PATH. trellis call <def> <json-args> runs the definition with real World-derived capabilities via the soil0 run injection semantics (contract §8.6), canonical JSON out. Cram results enter the lock with mode real; real mode is never reachable from the lowering sandbox (the MCP run_tests tool simply has no path to it).

Tests: read_file.tr’s cram block green against a real temp file; exit-code and output-mismatch failures; a fixture proving trellis call rejects non-ground or capability-mismatched args cleanly.

Step 9 — context bundles and the packer

Bundle assembly per the step-1 contract; the budgeted packer with the fixed priority spec > tests > callee signatures > corpus examples > module prose, decisions always included per §9.4’s resolution; demoted-refinement rendering in callees/; previous.soil on re-lower; the corpus is examples/ until plan 04 replaces it with the prelude. CLI: trellis context <def> --budget <n> writes the bundle to a directory and prints the manifest of what was packed and what was dropped.

Tests: packed-bundle goldens for fixture definitions at generous and starvation budgets (drop order is observable and pinned); the never-drop rule for spec; token-estimate monotonicity property.

Step 10 — the lowering loop against the fake provider

The MCP stdio shim (trellis _mcp, §8.8) and the seven tools; the bubblewrap jail and isolated agent home (§8.9); the Provider trait with the scripted fake provider (§8.9) that plays back canned tool-call sequences; the job loop — serial queue in disciplined callee-first order, pre-flight, bundle, invoke, per-tool logging to f.log, turn/time caps enforced by the daemon, outcome handling: green → canonicalized .soil moved into place + lock updated; ask_human → question persisted, job ends, trellis answer writes the answer into the .tr (§9.7) and re-queues fresh; budget exhausted with holes → partial(holes) committed honestly (design §3.16); dependency edited mid-flight → job invalidated and requeued. write_soil canonicalizes through the printer and rejects non-parsing text with structured diagnostics; the job’s completion payload requires the decisions_applied citation list (§9.4), recorded as lowering.decisions edges. The triage sweep is a second job kind on the same queue and provider machinery: bundle = a decision entry’s diff plus the flagged definition’s spec and Soil; outcome = conforms (re-stamp the edge) or not (leave flagged, report in f.log and status).

Tests: the whole plan-03 testing bullet — scripted playbacks for the happy path, the question path, cap exhaustion, invalidation mid-flight, a write_soil of hole-bearing Soil, a citation-recording playback plus a scripted triage run (one conforming, one not), and the e2e fixture: the fixture root’s mean-sized definitions lower to green entirely through the fake provider, locks and hashes regenerating deterministically.

Step 11 — the live provider

The claude-code provider: headless invocation with the MCP config pointing at the shim, tool allow-list, --max-turns, stream-json parsed for token/cost telemetry into f.log and provider/model into the lock. Re-verify the headless-subscription quota policy noted in design §4.6 before relying on it, and record the finding in the provider doc. One opt-in live smoke test (TRELLIS_LIVE=1): a trivial fixture definition lowers bundle → agent → write_soil → checks → tests → lock with f.log populated — the plan’s second exit criterion. CLI: trellis lower <def>.

Step 12 — REPL, skill, drift gates, conformance

trellis repl (call any definition with JSON args at the Trellis level, trellis call semantics, the future REPL-to-test promotion hook); trellis skill assembling the lowering skill from the diagnostic registry, the effect lattice, the derivation strategies, the tool contract, and the corpus, with the committed copy under a CI drift gate; the second half of the registry drift gate (registry ↔ contract doc ↔ emitted codes); the skill pin flips from warning to required — a missing pin under a generator-capable toolchain is toolchain-mismatch (soil-toml §2.1, resolved 2026-08-25); a final conformance pass over docs/contracts/trellis-daemon.md; the exit-criteria audit (§6).


6. Exit-criteria traceability

Plan 03 exit criterionWhere it lands here
examples/ fully round-trips: parse → hash → lock regeneration matches checked-in sidecarsSteps 3–5 (parse, hash, lock goldens; the regeneration pass makes the sidecars real, §9.10); step 6 keeps them stable under refresh
One real headless lowering completes with f.log populatedStep 11 (live smoke test; telemetry from step 10)
Question round-trip: ask_human → CLI answer → prose diff → re-invoke → greenStep 10 (fake provider) proves the machinery; step 11 exercises it live if the smoke lowering asks
Plan’s testing bullets (golden hashes, mutation tests, lock round-trip, scripted fake agent, e2e fixture)Steps 4, 5, 10 respectively

7. Risks and checks

  • Hash reproducibility. Everything feeding a hash must be bytes the repo controls: .tr/.soil are hashed as raw bytes with .gitattributes pinning eol=lf, paths never enter any hash (§8.5), and the reproducibility test runs the tree from two locations. The alpha-normalized hash form (§9.2) is specified before step 4 writes a single hash.
  • bubblewrap availability. Unprivileged user namespaces are off on some kernels/CI images. jail.rs probes at startup; a failed probe refuses to lower (the sandbox is not best-effort) with a diagnostic naming the missing capability. CI runs the jail tests only where the probe passes; the fake-provider loop is also exercised jail-less so the queue logic stays covered everywhere.
  • The claude CLI is unpinned. It is an ambient tool with a moving flag surface, and headless quota policy has shifted recently (design §4.6). Containment: the provider is one module behind the trait, the live test is opt-in, and step 11 starts by re-verifying invocation flags and quota policy.
  • Python reference oracles and float text. Python’s repr is not the canonical float form. The driver never compares strings from Python: its JSON is parsed and re-encoded through soil-rt’s canonical encoder before comparison, same normalization as test expectations (impl plan 02 §8.10).
  • Serial queue vs a stuck agent. The wall-clock cap is enforced by the daemon (kill the process group), not trusted to the agent CLI, or one hung lowering blocks the queue forever.
  • Lock canonical-writer drift vs examples. Prevented structurally the plan-02 way: the byte-identity test over regenerated examples runs in check.sh from step 5 onward.
  • Packer token estimates are approximate (§8.10). Accepted: the budget is a packing target, not a hard API limit; the provider’s own context limit is the backstop and the estimate constant is one number in providers.toml.
  • prose write-back merge conflicts. trellis answer re-parses and re-hashes the .tr before writing; if the file changed since the question was asked, the answer is refused with a diff rather than blindly appended (the invalidation rule already requeued the job).

8. Micro-pins — approved 2026-08-24

Per the overview’s rule, these were presented as proposals and approved by the user on 2026-08-24 (all sixteen, as written); each lands in the step-1 contract docs where user-visible, and reopening one is a new decision point. None are language-observable.

  1. One binary, auto-spawned daemon. trellis is both client and server: trellis daemon run serves; every other subcommand connects to $XDG_RUNTIME_DIR/trellis/<root-hash>.sock (root-hash = first 16 hex of SHA-256 of the canonicalized absolute root path) and spawns trellis daemon run for the root if absent. Explicit trellis daemon stop|status; no idle shutdown in v1. A version mismatch between client and daemon restarts the daemon (one binary, so mismatch means an upgrade happened).
  2. RPC framing: JSON-RPC 2.0, one compact JSON document per line (ndjson) both directions. Not contractual. Threads + Mutex state, a dedicated job-runner thread for the serial queue; no async runtime (dependency-floor precedent, impl plans 01–02).
  3. CommonMark via pulldown-cmark. Fenced-block extraction has real corner cases (tildes, indented fences, fences inside quoted or listed prose) and the crate is small and pure; hand-rolling was rejected as a corpus of subtle divergences from the CommonMark .tr files are defined to be. Only block extraction is used; prose is never rendered.
  4. Frontmatter is a hand-rolled two-key YAML subset: name: <scalar> and tags: [a, b] / block-list form, nothing else — tr-grammar makes unknown keys errors, so the subset is the format. A YAML crate would be a large dependency for two keys.
  5. Three-part hash canonicalization. The parsed file partitions into spans; each hash is SHA-256 over a length-prefixed concatenation, in document order, of its class’s parts: formal = the frontmatter name value + every formal-class block’s (info string, content bytes); test = every test-class block’s (info string, content bytes) — the info string includes xfail, so toggling it re-runs what depends on test_hash; prose = the remaining file bytes verbatim (frontmatter minus name, prose, unreserved fences) with formal/test and decisions block spans (fences included) elided — decisions are their own class (tr-grammar §1, resolved 2026-08-24): each entry hashes as SHA-256 over its bytes from the label line through its rejected: lines, and the scope hash covers the sorted labels, NUL-separated. Length-prefixing prevents concatenation ambiguity; no whitespace normalization inside blocks (the bytes are the spec); class membership per tr-grammar §1. The info string includes the @agent marker (resolved 2026-08-28): removing it — a human adopting a block unchanged — is a formal event that moves the class hash and re-verifies, chosen over content-only hashing; recorded in tr-grammar §8.
  6. soil_hash recipe. SHA-256 over: a format tag (soil-hash/1); the definition’s canonical hash-form text (§9.2); then, sorted by name, one name NUL referent-hash line per free reference — callee definitions by their soil_hash, private helpers by the same recipe applied recursively (this is the “folded in”: a helper’s change moves every owner), types by their formal_hash (cycle members by the cycle hash), builtins by soil0-builtin/<soil0_cli major>. The cycle hash is SHA-256 over the members’ canonical soil-type texts sorted by name. Rationale: substitution-by-edge-list gives the Unison property (change propagates hash-by-hash) without rewriting names inside printed text.
  7. Lock file surface form: UTF-8, 2-space-indented pretty JSON, key order exactly the schema tables’ order, elements of blocks, tests, oracles, calls, private_helpers printed one object per line (the checked-in examples’ shape — diff-per-row), one trailing newline. soil.lock same style. Recorded in lock-schema §1 as the normative surface form.
  8. MCP is hand-rolled over a stdio shim. The agent CLI spawns MCP servers as processes, so the daemon’s MCP face is trellis _mcp --job <id>: a shim speaking MCP (JSON-RPC 2.0: initialize, tools/list, tools/call) on stdio and forwarding to the daemon’s socket. The protocol subset for seven tools does not justify an SDK and its async runtime. The shim binary inside the jail is the same pinned trellis.
  9. Jail profile and providers. bubblewrap with: fresh tmpfs HOME, the bundle’s scratch dir bound read-write, the trellis binary and the agent CLI’s install (plus CA certs, resolv.conf) read-only, network allowed (the agent must reach its API), the project tree not mounted, environment scrubbed to an allow-list. Provider config lives in ~/.config/trellis/providers.toml (user-level, never in the repo): provider kind, model, turn/time/cost caps, context budget per model. The scripted provider replays a JSONL playbook of {tool, args} calls through the real shim path (so the loop, logging, and caps are exercised), asserting each result against an expectation; playbooks live in tests/scripts/.
  10. Packer accounting: tokens ≈ bytes/4, the constant recorded in providers.toml per model. Priority is fixed (design §4.6): spec > tests > callee signatures > corpus examples > module prose; decisions travel with the spec tier (§9.4); spec.md and tests.json are never dropped — a budget too small for them is an error, not a silent truncation; every drop is listed in the bundle manifest and f.log.
  11. soil0 in-process. All checking, running, and test execution call the soil0 library (the resolved plan-03 decision); one integration test per command class asserts library and shelled binary agree byte-for-byte on the examples, keeping the CLI contract honest as the daemon’s oracle.
  12. Property execution. For each property the daemon synthesizes a private wrapper definition whose params are the forall binders and whose body is the predicate lowered to Soil (the predicate grammar is a Soil-expression subset; implies becomes not a or b, is becomes a match), checked and run in-process; a case passes iff the wrapper returns True. Case counts by effect row (design §4.5): 128 for pure signatures, 32 when fakes are involved. Seeds are pinned by derivation: SplitMix64 seeded from the first 8 bytes of test_hash, so runs are reproducible and re-seed exactly when the tests change. Generators: uniform-with-boundary-bias for integer widths, finite F64 plus the specials, length-geometric lists and Utf8, field/variant-recursive for records and sums (depth-capped), comparator-consistent Maps. where filters per §9.8.
  13. trellis call and the REPL reuse contract §8.6 injection semantics exactly: capability params from a real World in order, remaining params decoded positionally, canonical JSON out, structured runtime-panic rendering on error. The REPL is a readline loop over the same RPC; no state between lines in v1.
  14. f.log is JSONL, one event per line: job_started (def, hashes, provider, model, budget), preflight, bundle_packed (contents, drops, estimate), tool_call (name, duration, result digest), question, agent_result (turns, tokens in/out, cost), outcome (green / partial holes / failed / invalidated), plus per-test failure details (the lock stores results only, lock-schema §5). Human-readable, gitignored, never an input (design §4.6).
  15. Question persistence: an open question is <def>.question.json beside the .tr (gitignored, like f.log): the structured question, the bundle hash, and the spec hashes it was asked against. trellis status surfaces it; trellis answer <def> (interactive or --text) validates staleness (§7), applies the write-back (§9.7), deletes the file, and re-queues.
  16. Daemon-owned diagnostics get their own code namespace in the same registry (tr- prefix for .tr validity, lock-, job-, toolchain-mismatch, …), kebab-case like soil0’s; the registry file is the union and the drift gate covers both halves.

9. Spec gaps and new decision points — all resolved

Raised per the overview’s rule; all were resolved with the user (§9.4 on 2026-08-24 in its own session; the rest on 2026-08-24 in the review pass) and propagated to the spec docs as noted.

  1. env.json inline-record variant payloads (contract §13.5 had deferred this to the daemon’s env generator; bites at step 7). Resolved: contract v1.2 — VariantD gains fields : Option (List FieldD) (payload and fields mutually exclusive), soil0 registers the record payload directly; additive like v1.1 (v1.1 env files decode with fields = None). The soil0 side (env decoding + registration) is implemented as part of step 7. Rejected: generator-synthesized hidden record types (reserved names leaking into diagnostics and show). Recorded in docs/contracts/soil0-cli.md §6/§13.5.

  2. The alpha-normalized hash form was never actually specified. Resolved: syntax-spec §9.1 “Hash form” — the canonical text with local binding sites numbered %N in pre-order and occurrences printing their binder’s number; definition names, callees, types, constructors, fields, and hole names untouched; emitted by a soil0 library function, not a CLI command — the CLI contract stays the display-form oracle, soilc’s printer is differentially tested against print, and the hash transform is one shared implementation above it. A print --hash-form flag remains a possible future additive amendment; none planned.

  3. formal_hash and the “import set”. Resolved: tr-grammar is the authority — formal_hash is spec-side only; callee identity flows through lowering.calls[].hash edges (already how lock-schema §8 routes caller invalidation; folding the computed import set into the spec hash would make it depend on the artifact it gates). Design §6.2’s wording amended.

  4. Decisions: flag or invalidate — resolved 2026-08-24 as neither: reliance edges + editorial reclassification + triage sweep, recorded in design §4.3, tr-grammar §1/§5.2/§9, lock-schema §2/§3/§8. Both binary options scaled with project size (scope-wide review-suggested fatigue, or a typo re-lowering the world); reliance scales with actual use. Consequences for this plan: decision entries are hashed individually and stored per label in the module/project locks (step 4/5); the lowering’s completion payload gains a required decisions_applied label list (possibly empty), recorded as lowering.decisions edges plus the scope-membership hash (steps 5, 10; the citation requirement goes into the step-1 tool contract and the generated skill); decision staleness is derived by hash comparison in the invalidation pass (step 6); trellis decisions editorial <label> re-stamps a detected change the human declares meaning-preserving; and the triage sweep is a job kind on the same queue and provider machinery as lowerings — bundle: the entry’s diff plus the flagged definition’s spec and Soil; outcome: conforms (re-stamp) or not (leave flagged, report) (step 10).

  5. Derived-test naming. Resolved: the namespaced scheme — derived:ensures:<label> / derived:requires:<label> / derived:invariant:<label>, derived:differential:<reference-line>, derived:contract:<n> reserved for plan 06. Chosen over the bare clause-label shapes in the pre-daemon example locks because the derived: prefix can never collide with a spec block name; the regeneration pass (step 5) rewrites the examples. Recorded in lock-schema §5/§9.

  6. Lock-schema §9 leftovers. Resolved: module/project entries carry accepted (a human can accept exports and decisions; ungated — they have no tests) while CI policy and badges keep keying on exported function definitions; the prelude-fork hash stays global in soil.toml trusted_packages (per-entry pins would repeat one hash until per-entry forks exist, and the toolchain pin already refuses under a changed prelude). Recorded in lock-schema §8/§9.

  7. The ask_human write-back format. Resolved: function-scoped answers land in a ## Clarifications prose section (created once, at the end of the file) as - **Q:** … / **A:** … pairs — plain prose under prose_hash, visible in the review diff; rule-shaped answers (the CLI asks which kind) go to the module’s or project’s decisions block, the tr-grammar §5.2 write-back target. The IDE later adds a fold-into-prose action: an agent absorbs a Q/A pair into the prose proper and deletes it, as an ordinary prose edit. Recorded in tr-grammar §8 and design §4.6; rejected alternatives noted there.

  8. where-filter support (tr-grammar §3.4 mandates constrained generation, not rejection sampling — full generality is a constraint solver). Resolved: the trivial v1 — generators are random from the binder’s type; a where filter is refused with a structured unsupported-where-filter error (never silently ignored — that would test outside the intended domain — and never rejection sampling, which tr-grammar forbids). Constrained generation is deferred to plan 05, when Z3 lands for the refinement checker: solver-backed generation for in-fragment predicates with model-blocking randomization — with the caveat recorded that raw solver models cluster (uniform SMT-space sampling is unsolved), so distribution quality gates adoption. Coverage-guided fuzzing (libFuzzer-style) was considered and rejected: coverage feedback would reflect the soil0 interpreter’s branches rather than the Soil program’s, it needs the nightly/sanitizer plumbing impl plan 02 §8.12 already rejected, and corpus evolution conflicts with pinned-seed reproducible results. Recorded in tr-grammar §3.4 (toolchain note) and design §10 (deferred row).

  9. Differential tests before the Python FFI exists. Resolved: the interim subprocess runner — python3 (from shell.nix) with a small driver importing relpath::symbol, feeding the test’s JSON args, comparing through canonical re-encode; CLI oracles likewise by pipe. The lock is transport-agnostic (tier + oracle hashes), so plan 06 swaps in the real FFI without a schema change. Rejected: stripping differential rows from the examples until plan 06 (weakens the round-trip exit criterion).

  10. Examples stop being fakes. Resolved: regenerate. The step-5 pass replaces abbreviated placeholder hashes with real 64-hex digests and updates the stale checks shapes (the current median.lock predates per-clause refinements, termination, and holes) and the old derived-test names (§9.5). The repo convention is amended in CLAUDE.md (2026-08-24): abbreviated fakes remain the style for docs prose; examples/*.lock carry real digests and are regenerated, never hand-edited. Missing sidecars are backfilled or documented during the pass (mean.tr stays the not-yet-lowered example), reviewed with the user. Policy resolved 2026-08-28 (see step 5): staged regeneration (step 5 partial, step 7 final byte-identity), locks for every .tr, parse_row.soil + _private.soil authored, provenance human-verified with provider/model absent.

  11. Step-7 resolutions (2026-08-28, with the user). Five decision points surfaced while building the test runner:

    • Invariant-derived properties are refused in v1 (unsupported-invariant-property): generating self from the bare structure tests exactly the values the invariant excludes — the honest generator is constrained generation, deferred with where filters to plan 05. Consequence: type entries are accepted-ungated in v1 (chosen over flipping Row to unaccepted). Recorded in tr-grammar §4.2, lock-schema §8, design §10. Rejected: checking the invariant over expect-test values (a different, weaker check wearing a property’s name) and honest-failure rows (breaks the flagship example for a toolchain gap, not a spec bug).
    • Definition-oracle rows are deferred to plan 04: the daemon records reference/cli oracle rows only. Recording (and enforcing acceptance for) every helper a predicate mentions would today flag len/min/max/sort — vocabulary the prelude will absorb. Recorded in lock-schema §5, design §4.5.
    • Differential inputs are the expect cases’ args (§9.9’s “feeding the test’s JSON args” read literally); generated-input replay was rejected as seed-coupled and generation-quality-bound.
    • Expect/cram row naming is always block#k (property rows stay bare); lock-schema §5’s stray “bare name for single-case blocks” sentence was corrected to match its own example and the corpus.
    • The F64 generator keeps §8.12 as pinned (finite plus the specials). The flagship specs were improved instead of the generator weakened: testing exposed that Soil’s F64 comparisons are the derived total order (design §3.7 — one logical NaN sorting last, -0.0 < 0.0), so sort is fully specified with no domain guard (sorted + is_permutation ensures), median keeps an all_finite requires because the mean of ±Inf middles is NaN (which sorts past +Inf, escaping the bounds), median.soil’s even-length mean became overflow-safe, and examples/csvstats/ gained a predicate vocabulary suite (finite, all_finite, non_empty, sorted, contains, count_of, is_permutation) — authored like parse_row.soil, provenance human-verified.
  12. Step-8 resolutions (2026-08-28, with the user).

    • trellis call executes in the client, not over the daemon socket: it links soil0 anyway, and client-local execution makes the real Fs’s relative paths resolve against the caller’s working directory — the cram temp dir — for free. The daemon route was rejected as needing either a carried-cwd/rooted-Fs mechanism or a per-call chdir in a threaded process. Contract §8.6 injection semantics are shared with soil0 run through the library (interp::run_json), so the two cannot drift; micro-pin §8.13’s REPL decision is untouched (step 12 revisits transport).
    • Root discovery honors TRELLIS_ROOT before the soil.toml walk (recorded in soil-toml §1): the cram runner exports it so transcripts in temp dirs can call back into their root. Rejected: temp dirs under <root>/.trellis/ (transcripts writing inside the project tree).
    • trellis test runs cram (mode real) by default — the real-mode exclusion is about the lowering sandbox’s run_tests tool, not the human CLI; transcripts are hermetic temp-dir sessions the human wrote. The step-7 carried-rows mechanism is removed.
    • read_file.tr’s transcript expectation corrected to canonical compact JSON ({"tag":"Ok",…}) — the hand-written spaced form predated the real runner; cram matches literally and canonical JSON is the §7 byte-equal form.
  13. Step-9 resolutions (2026-09-14, with the user).

    • callees/ is the callable surface: one signature file per lowered function in the root, nearest-first (target’s module before others), target excluded. Contract §4.4’s “one per direct callee” was circular on a first lowering (no body, no edges) — the case the bundle chiefly serves; the serial lowering order makes the lowered set exactly the callable menu. Rejected: known-edges-only (an empty callees/ for mean hides the module’s own helpers) and same-module-only (exported cross-module helpers vanish). Amended in contract §4.4.
    • trellis context writes .trellis/context/<def>/ by default (already gitignored; wiped per invocation); --out <dir> must be empty or absent and is resolved client-side (the daemon’s cwd is not the caller’s).
    • Prefix packing: the packer keeps the longest priority-order prefix that fits; the first over-budget item and everything after it drop. Chosen over greedy skip-and-continue so the packed set and estimate are monotone in the budget (the pinned property) and drops read predictably. previous.soil and reference.py sit between tests and callees — droppable, but only under extreme budgets. Both recorded in contract §5.
    • The corpus is the root’s own .tr/.soil pairs until plan 04 (the examples root is the v1 corpus and is also the only v1 root); module.md is the target module’s _module.tr verbatim.

Plan 04 — the pure-core prelude

References: design §5 (the prelude as a trusted corpus), §3.5 (capabilities), examples/read_file.tr.

Goal

The first Trellis code: the pure core prelude, written as .tr specs and agent-lowered through the real daemon pipeline, human-reviewed to accepted. It is simultaneously the stdlib, the few-shot corpus that defines the agent’s Soil style, and the first honest test of the whole loop. Interpreted on soil0; small (a few thousand lines of Soil).

Scope

  1. read_file first (design §5): capabilities, effects, Result, and the runtime boundary in one definition. Promote examples/read_file.tr into the real prelude root and lower it for real; reconcile any drift back into examples/.
  2. Core types and functions, roughly in dependency order: Bool (with its JSON special case), Option, Result, List (map, filter, fold, len, nth, append, reverse, sort_by…), Utf8 (split, trim, parse-number…), Bytes, BigInt, Map with explicit comparator (the Map.Make-as-function idiom, design §3.6), JSON encode/decode surface (thin wrappers over the runtime). The list primitives (list_len, list_nth, list_empty, list_append, …) are native-backed prelude definitions like the fakes — Soil has no list literals or patterns, so nothing list-shaped is writable without them — with names frozen in docs/contracts/soil0-cli.md (impl plan 02 §8.7, resolved 2026-08-22); the rest of the List functions are written in Soil on top of them.
  3. Capabilities and fakes. The capability types (Fs, Net, Clock, Env, Proc, Rand) as opaque types with their World derivations, and the fake constructors with pinned seeds/timestamps — signatures in Trellis, backed by soil0/soil-rt native primitives.
  4. Corpus duty. Every lowering is reviewed as a style exemplar, not just for correctness: idiomatic match shapes, naming, use of local lets vs private helpers. Style disagreements are settled by PR-style review with the user and become the corpus.
  5. Trust and packaging. The prelude is a Soil root with soil.toml; on completion, pin its package hash as trusted (design §5); all exported definitions accepted and pinned.
  6. Kernel reconciliation (deferred here from the soil0 CLI review, 2026-08-22). docs/contracts/soil0-cli.md §6.1 and §11 pin provisional prelude surface: the kernel type shapes (FsError, Utf8Error, Path), the builtin names and signatures (clock_now : Clock -> io I64 vs design §3.5’s sketched now : Clock -> io Time; the list primitives; unit; T::compare : T -> T -> I64 returning −1/0/1 vs an Ordering sum; the fake constructors and their determinism guarantees). The prelude’s first .tr specs must adopt these exactly, or revise them with the user and update the contract — before the corpus teaches them.

Adoptions (2026-08-23, agentlanguages survey)

  • Hostile capability fakes (design §4.5): failing fakes — an Fs that errors mid-stream, a backwards-jumping Clock — are ordinary prelude values in the capability modules from the start.
  • TrellisBench before the corpus grows (design §10): even ~10 spec+tests problems wired through the real daemon per release, so prelude/corpus changes are measured, not vibed. Stand it up at this plan’s kickoff.

The totality problem (known, planned for)

soil0 has no termination checker, so every recursive prelude function conservatively carries div — but the prelude’s signatures claim total, and those claims matter for the corpus and for callers. Interim policy (confirm with user at kickoff): the .tr signatures state the intended row; the daemon records a per-definition div-unverified flag in the lock (like a demoted refinement) rather than widening signatures; the stage-2 termination checker (plan 05) later discharges them in bulk. This mirrors the demotion philosophy: unproven, visible, tests still gate.

Non-goals

Batteries layers (soil-rs-std, soil-py-std — plan 06), Py capability, retrieval (whole prelude fits in context), performance.

Testing

  • Every definition: expect tests + properties per the effect-row budget (pure functions fuzzed hard); capability functions get fake-capability tests; read_file keeps its real-mode cram test.
  • Cross-cutting properties: sort_by stability and order laws, parse ∘ show identity on prelude types, Map comparator-order invariants.
  • Differential where cheap: CLI oracles against Python equivalents (statistics, str methods) for the numeric/string corners.

Exit criteria

  • Every exported definition accepted, pinned, tests green, package hash pinned in soil.toml.
  • The corpus test: a fresh lowering of a new small function, given the prelude as examples, produces Soil the user judges idiomatic without style corrections.
  • examples/ and the real prelude agree wherever they overlap.

Decision points — resolved 2026-08-22

  • Totality gap: signatures claim the intended row; the lock records checks.termination: "unverified" (mirroring refinement demotion — unproven, visible, tests still gate); soilc’s termination checker (plan 05) discharges the flags in bulk. Recorded in docs/lock-schema.md §4.
  • Location: in this repo, as a prelude/ Soil root; extraction into its own forkable repo waits for a second user.
  • Inventory: a concrete reviewed list before lowering begins (read_file + the §5 core types with ~6–12 functions each), then additions strictly by consumer need — the corpus stays curated.

Plan 05 — soilc: the compiler as the first Trellis project

References: docs/bootstrap-plan.md §3–4, design §3.1 (ANF), §3.4 (termination), §6.4 (demotion), §3.11 (Cranelift backend).

Goal

The Soil compiler written as a Trellis project — specs, agent lowerings, locks — running interpreted on soil0, differentially tested against it, and finally compiling itself to native code through Cranelift with a byte-identical fixed point. Refinement checking and termination checking enter the system here, as passes.

Scope

Passes in order; each is a Trellis module of pure functions with the AST as Trellis type definitions (the JSON schema from docs/contracts/soil0-cli.md is the conformance target — soilc’s AST types must round-trip it).

  1. Lexer. Warm-up; oracle soil0 lex.
  2. Parser. One definition with let rec … and … locals — the mutual-recursion-ban stress test, taken early on purpose. Oracle soil0 parse. If the ban genuinely fails here, that is a design finding to raise, not to code around.
  3. Renamer. No-shadowing, ::, _private visibility. Canonical fresh-name allocation (deterministic counters, no iteration-order dependence) — this is where fixed-point determinism is won or lost.
  4. Type + effect inference. The hardest lowering target in the whole plan; split the module aggressively (unify, generalize, rows, operator elaboration as separate definitions). Oracle soil0 infer.
  5. Exhaustiveness + pattern compilation (to decision trees). Oracle soil0 check for the boolean verdicts; pattern compilation is new but testable by semantics (compiled and source matches agree — property tests through the interpreter).
  6. ANF transformation. New; tested by properties (well-formedness of the output IR; evaluation equivalence via soil0 run on both forms).
  7. Termination checker. New functionality: structural decrease + decreases measures. Tested by spec (accept/reject corpus). On completion, run over the prelude to discharge the interim div flags from plan 04.
  8. Refinement checker. Desugars requires/ensures/inline refinements to SMT-LIB text (a pure function, golden-testable); Z3 runs behind a new Solver capability added to the prelude (the Py pattern: opaque type, fake for tests). Implements demotion (design §6.4) and arithmetic obligations (syntax spec §5). The daemon’s check_refinements stub goes live here.
  9. CLIF backend. Pure pass ANF → CLIF text, plus a small Rust Cranelift driver crate in the workspace (CLIF in, object files out, links soil-rt, x86-64 + arm64). Golden CLIF tests plus execution equivalence: compiled output vs soil0 run on the test corpus.

Strangler integration: as each pass reaches accepted, the daemon swaps its soil0 counterpart for the Trellis pass (invoked via soil0 run while interpreted). soil0 passes are demoted to oracles, never deleted.

Self-hosting closure: interpreted soilc compiles the prelude and itself → soilc₁; soilc₁ compiles the same sources → soilc₂; the build fails unless soilc₁ ≡ soilc₂ byte-identical. Then the daemon uses soilc₁ for execution, keeping soil0 for differential runs.

Adoptions (2026-08-23, agentlanguages survey)

For the refinement-checker pass (and the runtime guards it emits):

  • Three-way solver outcome (design §6.4): unsat = proven; unknown/timeout = runtime (the only demotion); sat with a model fails the lowering, the counterexample becoming a structured repair input and an offered expect test.
  • Assurance per clause (design §6.4, lock-schema §4): the checker emits proven/runtime/trusted per clause label.
  • Blame in every emitted guard (design §6.4): requires violations fault the caller, ensures/invariant the callee; carried in SoilError so the daemon routes repairs.

Non-goals

Optimization (beyond what Cranelift gives), Perceus reuse analysis, FFI codegen (plan 06 extends the backend), JVM/C/direct-x86 backends, concurrent lowering.

Testing

  • Per pass: differential against the soil0 CLI oracle over (a) the golden corpus from plan 02, (b) the prelude, (c) soilc’s own sources — the compiler is its own largest test input.
  • Property tests per pass (round-trips, well-formedness, evaluation equivalence through the interpreter).
  • Determinism harness: compile the corpus twice from clean state, byte-compare all outputs — run continuously from pass 3 onward, not discovered at stage 3.

Exit criteria

  • All passes accepted; daemon runs with soilc passes strangled in.
  • Prelude totality flags discharged by the termination checker; refinement demotion live end-to-end (a deliberately unprovable example demotes, is visible in the lock, and still runs its check).
  • The fixed point holds: soilc₁ ≡ soilc₂.
  • examples/csvstats/median.soil’s refinements actually prove.

Decision points — resolved 2026-08-22

  • Solver surface: one-shot — solve : (s : Solver) -> (script : SmtScript) -> io SolveResult with SolveResult = Sat CounterModel | Unsat | Unknown { reason }. Each obligation is an independent script: trivially fakeable, cacheable by script hash. Incremental sessions only if solve time ever hurts.
  • Decision trees are internal. The public schema covers surface AST and ANF; pattern-compilation output is free to change and is tested by semantic equivalence against soil0, not by goldens.
  • Fixed-point scope: all emitted artifacts — per-definition CLIF text, object files, and the linked binary must be byte-identical, so nondeterminism is caught at the layer that caused it. Artifacts may contain no timestamps or logs by construction.

Plan 06 — FFI, trellis bind, and the minimal IDE

References: design §3.11 (backends and FFI), §5 (batteries), §7.1 (IDE), tr-grammar §3.5–3.6 (cram, contract tests via design §4.5).

Goal

Open the foreign world in the committed sequence — C ABI → Rust batteries → Python embedding → Python batteries — with hand-written (agent-written, per-symbol) bindings, the trellis bind assistant, and the thinnest IDE that makes the lowering loop pleasant.

Scope

FFI (sequenced; each step usable before the next)

  1. C ABI codegen. The Cranelift backend learns extern calls against soil-rt’s C ABI; a worked shim example (a Rust extern "C" function wrapped as a Soil binding) joins the corpus.
  2. soil-rs-std batteries. Refinement-typed Soil signatures over Rust-backed functions — runtime-owned values, so ordinary, refinable, capability-free when pure (design §3.11). Start from demand: what the compiler and prelude wished they had.
  3. Python embedding. CPython via pyo3 inside soil-rt: GIL held around calls, PyObject* as opaque refcounted handles, py_to_soil / soil_to_py over the same JSON-shaped value model (one value model, never two), the Py capability + fake in the prelude. One worked pyo3 shim example joins the corpus.
  4. soil-py-std batteries. Handle-in/handle-out style, ffi panic io rows, Py capability; the visible two flavours (cheap handles vs converting/refinable).
  5. Binding trust plumbing. Contract tests auto-generated from signature + effect row (design §4.5: valid input → Ok, invalid → Err not panic, handle compatibility, leak-check loop, round-trip); lock ffi records (trust level, symbol hash, Nix store path — store path may be a plain path until the Nix milestone).

trellis bind <symbol>

  1. Symbol metadata fetch (Python: .pyi/inspect/docstring; Rust: cargo doc JSON) into symbol.json + docstring.md; the binding context bundle (one shim corpus example included); the normal lowering loop producing .tr + shim + contract tests + lock entry for human review. No bulk importer — ever (design §3.11).

Minimal IDE

  1. Electron-served web app over the daemon, as thin as possible: Markdown editor with test-block widgets (JSON drag-and-drop can start as guided JSON editing), lower button with streaming agent output, the question/answer panel (persisting answers to prose), graph view over the manifest (effect-flow coloring, soil-private greying), lock rendering as status badges (typed/tested/verified/accepted derived), REPL pane with REPL-to-expect-test promotion, accept/pin buttons with suggested git commits (never auto-commit).

Adoptions (2026-08-23, agentlanguages survey)

  • Literal provenance enforced (design §3.17): the checker’s Literal fact ships with the first boundary that needs it — Proc command text, Py eval, SQL batteries — and every batteries signature is written provenance-aware. trust_literal is a human-only escape hatch in the manifest.
  • Capability sandbox (design §10, deferred here): trellis build emits a seccomp/Landlock policy from main’s transitive capability set; trellis run --deny net shrinks World. Kind-level, not per-resource (documented caveat).
  • JSON marshalling constraint (design §3.11): nothing non-serializable crosses the FFI boundary, keeping the deferred co-process Python mode implementable as a deployment mode.

Non-goals

py_module/Python-hosts-Soil, .pyi stub generation, Node, TOML→Nix (later milestone), whole-package binding generation (never), IDE polish.

Testing

  • FFI: contract-test generator exercised against deliberately broken shims (each failure mode caught); leak checks under the debug runtime; py_to_soil/soil_to_py round-trip properties.
  • trellis bind: end-to-end against a fixed known symbol set (e.g. json.dumps, re.compile, one Rust crate fn) with recorded metadata so tests don’t depend on the network.
  • IDE: the golden path exercised in-browser against a fixture project — open, edit a test, lower, answer a question, accept — before calling any feature done.

Exit criteria

  • Both shim kinds exist as accepted corpus examples; a Soil program calls one Rust and one Python function through real bindings with contract tests green.
  • trellis bind requests.get (offline-recorded metadata) produces a reviewable binding end-to-end.
  • The IDE golden path works against the real daemon; question round-trip and accept flow usable without touching the CLI.

Decision points — resolved 2026-08-22

  • IDE delivery: browser first. The daemon grows an HTTP/WebSocket facade (trellis daemon --serve) and the web app is developed in a normal browser; Electron becomes a thin packaging wrapper later (design §7.1 unchanged — this sequences the wrapper last).
  • IDE stack: Svelte + CodeMirror 6 (widget decorations for test blocks and the question panel); graph view via SVG or cytoscape.js.
  • soil-rs-std starts with the Rust standard library only — a curated set drawn from std (math on floats, hashing, path/string utilities, whatever plans 04–05 wished for), no external crates initially. External crates (regex, chrono, …) arrive by demand through trellis bind, one symbol at a time.
  • Contract-test generation lives in the daemon (Rust), next to the test runner it feeds; rewriting it in Trellis is possible dogfood later, not v1.

Plan 07 — the Python-glue project: the second Trellis project

References: design §11 step 6, §1.2–1.3 (users and the pitch), bootstrap plan §0 (the bias this milestone exists to correct).

Goal

Write a real program the author would otherwise have vibed in pure Python: Rust crates through soil-rs-std, a dozen Python functions through trellis bind, real work done in Soil. The compiler validated the pure core; this validates FFI, capabilities, bindings, and — the actual product question — whether writing .tr files and reading generated Soil is pleasant. This milestone’s deliverable is as much a verdict as a program.

Scope

  1. Pick the program with the user. Criteria: genuinely wanted (not a demo), touches files/network/clock (exercises three capabilities), needs ~a dozen foreign symbols across 2–3 Python packages plus at least one Rust crate, small enough to finish (order of 30–60 definitions).
  2. Work disciplined-tier by the book. Human-written .tr prose and tests, agent lowerings, accepted gates, export pins, no hand-edited Soil unless the escape is genuinely needed (and then noted). The point is to feel the friction a real user feels; do not use insider shortcuts.
  3. Bind on demand. Every foreign symbol through trellis bind as encountered, never batched up front — this tests the assistant’s real cadence (design §3.11’s “a dozen bindings is a week of friction” claim, now measured).
  4. Keep a friction log. A running document (not memory, not code): every point where the format, the tooling, the corpus, or the agent made the wrong thing easy or the right thing hard, with enough context to act on. This log is the primary input to the next round of design changes.
  5. Ship it. The program builds via trellis build to a native binary, runs cram-tested against the real world, and gets used.

Non-goals

New toolchain features mid-flight (log them instead — resist fixing the tool from inside the project except for outright blockers); team features; trellis derive; performance work beyond what debug-mode flame graphs reveal for free.

Testing

The project’s own tests are the milestone’s tests: expect/property on pure logic, fake-capability tests on io, contract tests on every binding, cram on the entry point. CI policy line in soil.toml: every exported definition accepted.

Exit criteria

  • The program works and is actually used by the author.
  • Every binding came through trellis bind; every definition is accepted; the lock audit view shows exactly which trust levels and escape hatches exist.
  • The friction log is reviewed with the user and triaged into design changes, tooling issues, and corpus fixes — closing the loop that the compiler-first ordering deliberately left open.

Decision points — resolved 2026-08-22

  • The program is chosen at milestone kickoff, not now — the right project is whatever the author genuinely wants built when the tooling is real. The criteria in §1 are the filter; a stale pre-commitment would defeat the “genuinely wanted” requirement.
  • Friction cadence: fix after shipping, except outright blockers. The project is a measurement of real-user friction; mid-flight fixes with insider knowledge would contaminate it. Log and work around during; review and triage (design changes / tooling issues / corpus fixes) after.

Plan P1 — trp: the trellis-prose checker

References: docs/trellis-prose.md (the whole design), docs/lock-schema.md (for the spirit of lock conventions), examples/prose/ (normative fixtures). This is the first milestone of the trellis-prose sibling track; it shares the repo’s Rust workspace but depends on nothing in the Soil build order.

Decisions resolved with the user (2026-09-14)

These four were decided before this plan was written; they are recorded here with reasoning, per the working rules.

  1. Deliverable: plan doc first, build later. Building starts in a separate session from this plan, the way plans 01–07 work.
  2. Code home: a trp crate in the existing rust/ workspace. Shares the toolchain, check.sh, and testing conventions; parser/checker/lock code benefits from Rust’s strictness. A separate repo was rejected while trellis-prose remains a sibling design inside this one; Python was rejected as a weaker fit with the workspace (and the segmentation rule below removes the need for NLP libraries).
  3. Lowering: checker-only tool, in-session agent. trp is a pure oracle: it parses, verifies, and manages locks, and never calls a model. A Claude Code session performs the lowering and iterates against trp check until green — mirroring how Trellis lowering works pre-daemon. A trp lower API wrapper was rejected for v1 (it would own prompting, keys, and retry policy before the workflow is understood); it may return as a later milestone.
  4. Style judge: the human, in v1. The mechanical checks gate; setting accepted doubles as the style verdict, recorded in the lock as judge: human. Agent judging (in-session or tool-invoked) is deferred until the workflow has been exercised. Recorded in docs/trellis-prose.md §5.1.

Goal

A Rust crate trp providing the trellis-prose oracle: parse .trp files, mechanically verify a candidate lowering against its plan (the strict 1:1 mapping, §4 of the design), compute the three-part hashes, and read, write, and report on .lock sidecars. With trp green and a human accept, a document has the full trellis-prose discipline: hashed spec, verified structure, tracked freshness.

Scope

  1. Spec finalization gate. Before code: resolve design §8.3 (segmentation — proposal below) and §8.6 (the provisional container, frontmatter, heading-level, and lock decisions) with the user, and freeze the lock’s v1 field set. Outcomes are recorded in docs/trellis-prose.md in place, and examples/prose/ is updated to conform before the parser is written against it.
  2. Crate and CLI skeleton. rust/trp joins the workspace. Binary with three subcommands:
    • trp check <name>.trp — parse and validate the spec; if the sibling <name>.md exists, verify the lowering against the plan; --write-lock updates the lock on green (with the lowering provenance passed by flag).
    • trp status [dir] — freshness table across documents: fresh, stale (plan hash mismatch), review-suggested (notes or style hash mismatch), unlowered, unlocked.
    • trp accept <name> — the human-only action: sets accepted, records the style verdict as judge: human. Refuses if checks are not green or the lock is stale.
  3. Parser. Frontmatter (name, style; unknown keys are errors), the single required plan fence, directive lines against the closed move set, with positioned diagnostics (file, line, what was expected). Structural rules from design §3.2: prose directive outside an open paragraph is an error; heading closes the open paragraph.
  4. Segmenter and structural verifier. Split the output .md into headings and paragraphs; segment paragraphs into sentences; verify the four mechanical checks of design §4 (known directives, verbatim headings in order and level, paragraph count, per-paragraph sentence counts). Diagnostics name the first offending directive/sentence pair.
  5. Hashing, locks, freshness. Compute plan/notes/style hashes and the output hash; serialize and parse the lock; implement the freshness states and the review-suggested semantics for notes and style edits.
  6. Fixtures. examples/prose/ is the normative positive fixture; add a negative corpus under rust/trp/tests/fixtures/ (unknown move, sentence outside a paragraph, count mismatch, reordered/reworded heading, ambiguous segmentation, two plan blocks, unknown frontmatter key).
  7. Workflow validation. Lower at least two real documents end-to-end in a Claude Code session using trp check as the oracle, through accept. Friction observed here (especially re-lowering churn, design §8.5) is recorded back into the design’s open questions, not fixed ad hoc.

Segmentation proposal (decision point 2)

Resolve design §8.3 as: sentence terminators are ., !, ?; a sentence ends at the first terminator; terminators may not appear sentence-internally in v1 output. No abbreviations (“e.g.”, “Dr.”), no decimal numerals, no ellipses. Segmentation becomes trivial and exact — the checker needs no heuristics — and the burden falls on the agent’s wording, which is the tenet working as intended (“verbosity for the agent is acceptable”). A lowering that needs an internal period must re-word.

Non-goals

trp lower (API-driven lowering), the agent style judge, a mechanical style floor (design §8.2), richer output than headings and paragraphs (§8.4), wording-stability across re-lowerings (§8.5), multi-document projects (§8.7). Each stays in the design’s open questions until the workflow validation produces evidence.

Testing

  • Rust unit tests per module; golden-file tests for diagnostics (match soil0’s conventions).
  • trp check runs green over examples/prose/ in rust/check.sh; the negative corpus asserts each documented diagnostic.
  • Lock round-trip: parse → serialize is byte-identical for the example lock; freshness states are exercised by mutating fixture copies (edit plan → stale; edit notes → review-suggested; edit style → review-suggested for every referencing document).

Exit criteria

  1. trp check examples/prose/small_languages.trp is green, and every negative fixture produces its documented diagnostic.
  2. trp status and trp accept demonstrate all freshness states and the accept gate on the fixtures.
  3. Two new real documents lowered in-session to green, accepted, with locks committed.
  4. Design §8.3 and §8.6 are resolved in place in docs/trellis-prose.md, examples/prose/ conforms, and this plan is updated with the decision outcomes.
  5. rust/check.sh covers trp.

Decision points (for the user, at milestone start)

  1. §8.6 confirmations: the CommonMark container with a single plan fence, required style frontmatter, explicit heading levels, lock field shapes.
  2. Segmentation: the no-internal-terminators rule proposed above.
  3. Hash canonicalization: hash raw bytes of each region, or newline-normalize first (affects cross-platform lock stability).
  4. Example hash policy: repo convention is abbreviated fake hashes in examples, but trp check in CI cannot verify those. Options: a --no-hash-verify mode for the docs examples; or real (still abbreviated?) hashes in examples/prose/ regenerated by CI.
  5. CLI contract: whether to freeze a docs/contracts/trp-cli.md (as soil0 did) before agents start relying on the interface, or let the CLI settle through the workflow validation first.

read_file

The prelude’s read_file, slated to be the first real definition: a .tr spec (prose, capability-style signature, fake-capability test, real-mode cram test), its .soil lowering, and its .lock entry. The canonical files live in examples/; these are included verbatim.

read_file.tr

---
name: read_file
---

# read_file

Reads an entire file into a UTF-8 string. Returns `Err` if the file does not
exist, cannot be read, or is not valid UTF-8. Never panics.

```soil-sig
read_file : (fs : Fs) -> (path : Path) -> io (Result Utf8 FsError)
```

```test found-and-missing
with fs = fake_fs [{"key": "config.toml", "value": "port = 8080"}]
(fs, "config.toml") => {"tag": "Ok", "value": "port = 8080"}
(fs, "missing.toml") => {"tag": "Err", "value": {"tag": "NotFound", "value": {"path": "missing.toml"}}}
```

One real-mode test against the actual filesystem:

```cram real-read
with file "config.toml" = "port = 8080"
$ trellis call read_file '"config.toml"'
{"tag":"Ok","value":"port = 8080"}
```

read_file.soil

read_file : (fs : Fs) -> (path : Path) -> io (Result Utf8 FsError)
read_file fs path =
  match fs_read_bytes fs path with
  | Err e -> Err e
  | Ok bytes ->
    match utf8_decode bytes with
    | Ok text -> Ok text
    | Err _ -> Err (NotUtf8 { path = path })

read_file.lock

{
  "lock_format": 1,
  "name": "read_file",
  "kind": "function",
  "versions": { "trellis": "0.1.0", "soil": "0.1" },
  "spec": {
    "provenance": "human",
    "hashes": {
      "formal": "sha256:c14546862714dbf1a62bf8d6ea285ced0cf0af805af683145d0d672d8da93740",
      "test": "sha256:426b96cc18ce0e5d007181c432390fe0c430141b8d3c149d4097bf6764d4a419",
      "prose": "sha256:0e747d1eb101ab805a69b0a529b9467a6ea2060b1999d77ef35ec35a44589bbf"
    },
    "prose_state": "fresh",
    "blocks": [
      { "block": "soil-sig", "author": "human" },
      { "block": "test found-and-missing", "author": "human" },
      { "block": "cram real-read", "author": "human" }
    ],
    "pinned": true,
    "escape_hatches": []
  },
  "lowering": {
    "soil_hash": "sha256:ad309685e561013a54b436d32c13e64659d04c9857b77e7687b416f60635ee91",
    "provenance": "human-verified",
    "provider": null,
    "model": null,
    "private_helpers": [],
    "calls": [],
    "decisions": {
      "scope_hash": "sha256:af5570f5a1810b7af78caf4bc70a660f0df51e42baf91d4de5b2328de0e83dfc",
      "applied": []
    }
  },
  "checks": {
    "types": "ok",
    "termination": "verified",
    "refinements": {},
    "holes": 0
  },
  "tests": [
    { "name": "found-and-missing#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "found-and-missing#2", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "real-read#1", "tier": "cram", "mode": "real", "origin": "spec", "result": "pass" }
  ],
  "oracles": [],
  "accepted": true
}

csvstats

A hand-written module exercising the formats: a module header, two types (Row, ParseError), and three functions (parse_row, mean, median). Not every definition has every layer — median is the only one carried through all three (.tr → .soil → .lock); _module, row, and parse_row have .tr + .lock; parse_error and mean are spec-only. The canonical files live in examples/csvstats/; these are included verbatim.

_module (module header: .tr + .lock)

_module.tr

---
name: csvstats
---

# csvstats

Reads rows of decimal numbers from CSV lines and computes summary
statistics. The parsing functions own input validation; the statistics
functions assume validated `Row` values and stay total.

```exports
parse_row
median
mean
type Row
type ParseError
```

_module.lock

{
  "lock_format": 1,
  "name": "csvstats",
  "kind": "module",
  "versions": { "trellis": "0.1.0", "soil": "0.1" },
  "spec": {
    "provenance": "human",
    "hashes": {
      "formal": "sha256:4f523045c3a5c366ee4d455053150edda848a78e66028785b13f9a7dedfda790",
      "test": null,
      "prose": "sha256:ab30ca95085f61fa4f5af771b61e0a36cd316e93356b50ba6e54fcbceaf69730"
    },
    "prose_state": "fresh",
    "blocks": [
      { "block": "exports", "author": "human" }
    ],
    "pinned": false,
    "escape_hatches": []
  },
  "checks": {
    "exports": "ok"
  },
  "accepted": false
}

row (type: .tr + .lock)

row.tr

---
name: Row
---

# Row

One parsed CSV row. Construction goes through `parse_row` or row literals;
every row has at least one cell.

```soil-type
type Row = { cells : List F64 }
```

```invariant
non-empty: len(self.cells) > 0
```

row.lock

{
  "lock_format": 1,
  "name": "Row",
  "kind": "type",
  "versions": { "trellis": "0.1.0", "soil": "0.1" },
  "spec": {
    "provenance": "human",
    "hashes": {
      "formal": "sha256:36af0f58e2bd320599cdcced097bf362d82842a7000b7fcf878ca04005cdd5f8",
      "test": null,
      "prose": "sha256:1334d065c2a08b423f5a99a2c417fda1915967977e42b7c79c16e0fb2ec8e282"
    },
    "prose_state": "fresh",
    "blocks": [
      { "block": "soil-type", "author": "human" },
      { "block": "invariant", "author": "human" }
    ],
    "pinned": false,
    "escape_hatches": []
  },
  "checks": {
    "types": "ok",
    "invariants": {
      "non-empty": "runtime"
    }
  },
  "tests": [],
  "cycle_hash": null,
  "accepted": true
}

parse_error (type: .tr only)

parse_error.tr

---
name: ParseError
---

# ParseError

Why a CSV line failed to parse. `BadCell` carries the zero-based index of
the offending cell and its raw text.

```soil-type
type ParseError =
  | EmptyLine
  | BadCell { index : I64, text : Utf8 }
```

parse_row (function: .tr + .lock)

parse_row.tr

---
name: parse_row
tags: [parser]
---

# parse_row

Parses one CSV line of decimal numbers into a `Row`. Cells are separated by
commas; surrounding whitespace in a cell is ignored. An empty line, or any
cell that is not a decimal number, is an error.

```soil-sig
parse_row : (line : Utf8) -> Result Row ParseError
```

```test happy-path
("1.0,2.5,3.0") => {"tag": "Ok", "value": {"cells": [1.0, 2.5, 3.0]}}
("  4.0 , 5.0") => {"tag": "Ok", "value": {"cells": [4.0, 5.0]}}
```

```test errors
("") => {"tag": "Err", "value": {"tag": "EmptyLine"}}
("1.0,x,3.0") => {"tag": "Err", "value": {"tag": "BadCell", "value": {"index": 1, "text": "x"}}}
```

Scientific notation is a known gap, blocked on deciding the cell grammar:

```test scientific-notation xfail
("1e3") => {"tag": "Ok", "value": {"cells": [1000.0]}}
```

parse_row.lock

{
  "lock_format": 1,
  "name": "parse_row",
  "kind": "function",
  "versions": { "trellis": "0.1.0", "soil": "0.1" },
  "spec": {
    "provenance": "human",
    "hashes": {
      "formal": "sha256:bf4c9b19e1ca0e7050768a3cf2b707acb7c3b30bbce5eebd34a0fe57955e584d",
      "test": "sha256:f234b5385d86dc133a6e40652bf95789ce50a98e832822865f26733500c2edc2",
      "prose": "sha256:5ba26adaea0a1d2f2e5a5042570f2c83f167041d289cc5cae6dbd8ad228149bd"
    },
    "prose_state": "fresh",
    "blocks": [
      { "block": "soil-sig", "author": "human" },
      { "block": "test happy-path", "author": "human" },
      { "block": "test errors", "author": "human" },
      { "block": "test scientific-notation", "author": "human" }
    ],
    "pinned": false,
    "escape_hatches": []
  },
  "lowering": {
    "soil_hash": "sha256:c3f99d99d11fa4352ff1f9f252f5933ae28cd4f278ac0a5b6ce971c411407391",
    "provenance": "human-verified",
    "provider": null,
    "model": null,
    "private_helpers": [
      { "name": "_parse_cell", "hash": "sha256:fbbc9e37d86ec7298fd1ac46213aa6fc7cb19c208963fcad8381c110f1eb1a24" }
    ],
    "calls": [],
    "decisions": {
      "scope_hash": "sha256:af5570f5a1810b7af78caf4bc70a660f0df51e42baf91d4de5b2328de0e83dfc",
      "applied": []
    }
  },
  "checks": {
    "types": "ok",
    "termination": "unverified",
    "refinements": {},
    "holes": 0
  },
  "tests": [
    { "name": "happy-path#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "happy-path#2", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "errors#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "errors#2", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "scientific-notation#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "xfail" }
  ],
  "oracles": [],
  "accepted": false
}

mean (function: .tr only)

mean.tr

---
name: mean
---

The arithmetic mean of a non-empty list of floats.

```test simple
([1.0, 2.0, 3.0]) => 2.0
```

median (function: .tr + .soil + .lock)

median.tr

---
name: median
---

# median

Returns the median of a non-empty list of finite floats. For an even
number of elements, returns the mean of the two middle elements. The
bounds below need the all-finite domain: the mean of middles `-Inf`
and `+Inf` is `NaN`, which sorts *after* `+Inf` in Soil's total order
on `F64` (design §3.7) and so escapes `result <= max(xs)`.

```soil-sig
median : (xs : List F64) -> F64
```

```requires
non-empty: len(xs) > 0
all-finite: all_finite xs
```

```ensures
lower bound: min(xs) <= result
upper bound: result <= max(xs)
```

The reference implementation wraps `statistics.median` from the Python
standard library.

```reference
ref/stats.py::median
```

```test odd-length
([1.0, 3.0, 2.0]) => 2.0
```

```test even-length
([1.0, 2.0, 3.0, 4.0]) => 2.5
```

```test single
([42.0]) => 42.0
```

Sorting first never changes the answer, and under the total order this
holds for *every* non-empty list — even `NaN`-bearing ones, since both
sides reduce to the same middles of the same sorted list.

```property sort-invariant
forall xs : List F64
len(xs) > 0 implies median (sort xs) == median xs
```

median.soil

median : (xs : {v : List F64 | len v > 0 and all_finite v}) -> {r : F64 | min xs <= r and r <= max xs}
median xs =
  let ordered = sort xs in
  let n = len ordered in
  let mid = n / 2 in
  if n % 2 == 1
  then nth ordered mid
  else
    let a = nth ordered (mid - 1) in
    let b = nth ordered mid in
    let s = a + b in
    if s - s == 0.0
    then s / 2.0
    else a / 2.0 + b / 2.0

median.lock

{
  "lock_format": 1,
  "name": "median",
  "kind": "function",
  "versions": { "trellis": "0.1.0", "soil": "0.1" },
  "spec": {
    "provenance": "human",
    "hashes": {
      "formal": "sha256:8efa23f9ee5e01193b92be29c015233544916d41e0957938a721d5ee9d706a43",
      "test": "sha256:63929b6a73266fabd7ec20a7c54389b5a463e91780d40facdd3bc7b84f836800",
      "prose": "sha256:51f17924a3f698e3d6c289ee0fb8946d8cad30d5ee8ecc603c34fe736d3a75d3"
    },
    "prose_state": "fresh",
    "blocks": [
      { "block": "soil-sig", "author": "human" },
      { "block": "requires", "author": "human" },
      { "block": "ensures", "author": "human" },
      { "block": "reference", "author": "human" },
      { "block": "test odd-length", "author": "human" },
      { "block": "test even-length", "author": "human" },
      { "block": "test single", "author": "human" },
      { "block": "property sort-invariant", "author": "human" }
    ],
    "pinned": true,
    "escape_hatches": []
  },
  "lowering": {
    "soil_hash": "sha256:2ab06f664670b3982b5a1dbbb2213cabc4d4f61aa1174f685aa2f6fe5fd901a7",
    "provenance": "human-verified",
    "provider": null,
    "model": null,
    "private_helpers": [],
    "calls": [
      { "name": "len", "hash": "sha256:1d0db7042147887e451444fd609243b27aabb1fd18bfe3fb678823dc2f1e1d8e" },
      { "name": "nth", "hash": "sha256:d32fd45ca339a7f7cc735e362a3e2621dbd451c2f732fc06f0de82e63e4a9c7c" },
      { "name": "sort", "hash": "sha256:ecb7c98a9c4fbcc3a802c3c224067136fc92df2c39bb7501277ee82c759a8188" }
    ],
    "decisions": {
      "scope_hash": "sha256:af5570f5a1810b7af78caf4bc70a660f0df51e42baf91d4de5b2328de0e83dfc",
      "applied": []
    }
  },
  "checks": {
    "types": "ok",
    "termination": "verified",
    "refinements": {
      "all-finite": "runtime",
      "lower bound": "runtime",
      "non-empty": "runtime",
      "upper bound": "runtime"
    },
    "holes": 0
  },
  "tests": [
    { "name": "odd-length#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "even-length#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "single#1", "tier": "expect", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "sort-invariant", "tier": "property", "mode": "sandboxed", "origin": "spec", "result": "pass" },
    { "name": "derived:ensures:lower bound", "tier": "property", "mode": "sandboxed", "origin": "derived", "result": "pass" },
    { "name": "derived:ensures:upper bound", "tier": "property", "mode": "sandboxed", "origin": "derived", "result": "pass" },
    { "name": "derived:differential:ref/stats.py::median", "tier": "differential", "mode": "sandboxed", "origin": "derived", "result": "pass" }
  ],
  "oracles": [
    { "kind": "reference", "path": "ref/stats.py::median", "hash": "sha256:4c6a37d371476edb8bb5beb89290908c19a7f8a8c10adf5d4fa8bffd32617242" }
  ],
  "accepted": true
}