Research

Everything we have published, in the order we learned it.

Papers, normative specifications and evaluation protocols on agent harnesses, memory, and how to measure what an agent actually costs. Method and instrumentation are published with every result.

7 items

2026

Kaleidoscope: Benchmarking Continual Agent Memory under Bounded Context and Computation A free-energy contextual controller with attribution-gated replay. Formulates continual agent memory as bounded context compilation over a nonstationary stream, and defines a chronological cross-domain benchmark whose matched-total-cost analysis tests whether accuracy rankings reverse once all resource use is counted. paperin progress
07 Statistical Analysis and Win Conditions The preregistered analysis plan: what counts as a result, what counts as noninferior, and the novel-task and negative-transfer protections a memory claim has to clear. protocolin progress
07 Evaluation Instrumentation The measurement contract. What every run must emit, how tokens and local compute are attributed, and what to record when a harness does not expose a field. protocolin progress
07 Benchmark Plan Task construction, chronological workstreams, prefix/suffix boundaries, and the freezing rules that keep a memory benchmark from leaking its own answers. protocolin progress
07 Matched-Cost Ablation Program The ablation ladder, run at equal total cost rather than equal task count. protocolin progress
06 Kaleidoscope normative specifications Filesystem layout and durability, process ABI and lifecycle, request/response and event contracts, trace and fixture formats. The contracts an independent implementation would need. specin progress
06 LargeRepo-Memory: construction and evaluation plan Building a memory-bearing evaluation over a large real repository, and what breaks at that scale. protocolin progress