Research

Public summaries of the pre-registered program

Our research stance, restated: we pre-register hypotheses and experimental designs before measuring; no numeric claim below is a result — none exist yet, and we say so. Negative results are first-class outcomes and will be published here alongside positive ones.

Executor context custody & harness economics

Motivation
Agent harnesses spend heavily on coordination bookkeeping executed as model tokens, and the context that would make each call good is too heavy for users to supply.
Hypotheses
A verified obligation can replace transcript reconstruction when it is a sufficient statistic for the call; deterministically projected, role-scoped context improves first-pass acceptance; a versioned presentation profile can improve phrasing without weakening semantic quality; the combined system lowers coordination tokens/retries enough to justify its fixed cost.
Pre-registered design
Four matched arms — a frontier model with a conventional transcript loop vs a small model given (i) minimal verified-obligation context, (ii) deterministically projected dependency- and role-scoped context, (iii) the same plus a hand-engineered, versioned presentation profile (no training/optimization loop in this round; learning is deferred). Blinded task-quality scoring independent of the acceptance gate; separated token accounting; anti-gaming controls (frozen gate, held-out scoring, immutable context manifests).
Status
Pre-registered; execution awaits two internal prerequisites — the kernel-realization milestone and the frozen execution-binding anchor/reveal protocol. No measurements yet.
Falsified by
No material token/retry reduction at equal blinded quality; projected context failing to beat minimal context; presentation gains that coincide with held-out quality loss (gate exploitation).

Neuro-symbolic bottleneck & agent economics

Motivation
If a closed, verified coordination vocabulary lowers the entropy of what a model must produce, smaller specialized models may suffice for authoring — changing the economics of the whole path.
Hypotheses
Closing the output vocabulary makes constrained decoding useful; a small specialized model can reach a fixed intent-fidelity threshold at competitive end-to-end cost; constrained decoding's effect on semantic (not just structural) accuracy must be measured and may be negative; there exists a traffic volume above which the specialized path wins on total cost of ownership.
Pre-registered design
Matched arms over the same typed-interpretation → compiler path; a total-cost model including amortized training, ontology maintenance, clarification, and repair; the break-even volume is a reported output, not an assumption.
Status
Proposed research plan; not ratified; no measurements.
Falsified by
Small-model arms missing the fidelity threshold at competitive cost. (A break-even viability ceiling will be pre-registered with the plan's ratification.)

Kernel adequacy (GO/NO-GO)

Motivation
The coordination kernel must be adequate for coordination beyond value settlement before it is generalized — and that question deserves a pre-registered answer, not an accumulation of intuitions.
Hypotheses
The combined kernel (typed operations + verified plan + semantic layer) can represent a defined non-settlement coordination corpus within pre-registered thresholds.
Pre-registered design (drafted)
An open classifier over a held-out corpus, baselines against the current kernel, numeric GO/NO-GO thresholds proposed for advance commitment; both outcomes first-class — only a later GO decision would authorize generalizing the kernel.
Status
Draft pre-registration prepared; not ratified; execution parked. No measurements.
Falsified by
The corpus exceeding the kernel's representational envelope beyond the proposed thresholds once ratified — a NO-GO, which would be published as such.

An orthogonal coordination algebra

Motivation
Recurring coordination semantics — undertakings, conformance, non-interference, temporal validity and forfeiture — were being flattened into guards and ledger entries, losing distinctions that matter across domains.
Hypotheses
A small orthogonal operator basis (obligation · conformance · independence · validity window · deadline with scoped consequence) covers these semantics across independent domains without domain- or backend-specific exceptions.
Pre-registered design
Candidate operators nominated from evidence in independent domain families; remove-one tests (an operator that reduces to compositions of the others is dropped — one already was); behavioral witnesses with positive/negative contrasts; a reference evaluator whose expected traces are frozen and reproducible; held-out validation on a genuinely untouched family before any promotion.
Status
A provisional five-operator proposition is frozen with its reference evaluator and traces; its first compiler realization is in progress. Held-out validation has not run.
Falsified by
The held-out family requiring a domain-specific exception; an operator failing remove-one; realization forcing semantics the proposition cannot express.

Ontology-first domain research (documentary-credit pilot)

Motivation
Domain semantics should be derived from primary-source evidence, not intuition — and the language corrected from what domains actually require.
Hypotheses
Statute-grounded behavioral witnesses (each an exact source excerpt, an explicit interpretation, a discriminating behavioral test, and a modeling-pressure hypothesis) are a sound, honest unit for domain-to-language pressure.
Pre-registered design
A bounded witness packet on one domain (documentary letters of credit under one enacted statute), each witness carrying provenance digests and positive/boundary/negative contrasts; explicit epistemic labels (direct vs interpretive; open vs settled); coverage claims limited to what the sources actually support.
Status
The pilot packet is complete and merged after bounded review; it fed the coordination-algebra nomination above. Broader, practice-source coverage is scoped as follow-up work with independent re-sealing.
Falsified by
Independently authored and re-audited witnesses repeatedly failing to preserve source-to-interpretation traceability, or failing to yield stable discriminating outcomes — a single witness failing re-audit invalidates that witness, not the method.

Execution-binding adequacy

Motivation
Binding verified semantics to real execution — libraries, processes, remote APIs, durable jobs — must not smuggle semantic decisions into adapters; whether the current binding contract suffices is an empirical question.
Hypotheses
A small, technology-neutral set of binding properties (identity, authority, routing, parameters, invocation, timing, failure, evidence) either reduces to the existing declarative contract or forces a minimal typed extension — decided by evidence, not preference.
Pre-registered design
Executable specimens across genuinely independent execution classes (verified non-wrapper relationships; distinct failure ownership); frozen current-contract baselines; a staged protocol where candidate extraction is frozen before a sealed, independently-authored substitution case is revealed; admission requires the same property forced by at least two independent classes.
Status
A preparation draft has pinned reproducible candidate baselines; the protocol and result phase remain unratified and blocked on the kernel-realization milestone. No adequacy verdict exists.
Falsified by
(For the extension hypothesis) every candidate property reducing to profile data on the existing contract — a RETAIN verdict, published as such. (For the method) non-reproducible classification, reveal-order leakage, inability to distinguish profile data from semantic extension, or a candidate admitted from a single execution class.

← Back to the overview