Skip to content
Research library
Findings update 001 July 15, 2026 Locally source-verified

What survived the controls

The credible result is not one sweeping breakthrough. It is two research systems that preserve evidence well enough to accept a narrow win, reject a failed model, and keep both conclusions useful.

Waggle + Kea

8/8 exact parity

Fresh native-state Executors matched fresh full-text controls semantically and byte for byte.

Measured utility

Break-even at task 4

The single-host cold path was 30.7124% faster by task eight in one controlled local run.

HEATWAKE

0 models promoted

Both synthetic candidates lost to uniform in all five folds; the frozen action was ABSTAIN.

Jump to a findings section
Executive finding

A narrow positive result and a decisive negative result.

Waggle + Kea established exact, source-grounded native-state continuation on one local Qwen3-14B runtime and measured a same-host fan-out latency break-even at the fourth task. The same experiment also showed that the retained state was roughly 279.3 MB for an 8.5 KB shared prefix, so no storage, transport, or overall-efficiency advantage is supported. HEATWAKE trained and evaluated two architectures on a deterministic authored synthetic corpus. Both lost to a uniform baseline across all five held-out folds, failed the frozen promotion gates, and were rejected. A later one-command fixture journey then replayed custody through operator display and remained fail-closed, but no empirical market data has been admitted. Across both programs, the strongest accomplishment is a working discipline for preregistration, exact replay, independent qualification, refusal, recovery disclosure, and bounded claims.

How this was verified

Claims, controls, receipts, reruns.

1

Primary evidence

Claims were traced to frozen JSON reports, preregistrations, receipts, content hashes, source commits, and the exact source trees that produced them.

2

Fresh local rerun

The decisive Waggle/Kea gates and HEATWAKE promotion, walk-forward, empirical-machinery, service, and operator suites were rerun from the named research worktrees on July 15, 2026.

3

Claim grammar

A result is labeled supported only inside the tested fixture and runtime. Negative results are retained as results. Anything outside the tested envelope is listed as not demonstrated.

4

Public audit surface

This report is paired with a downloadable evidence ledger and SHA-256 checksum. Private repositories are not presented as publicly inspectable source.

Program 01

Waggle + Kea

Machine coordination with an independent decipher plane

Full program

Verdict

Supported at fixture scale: exact continuation, deterministic replay, same-host reuse utility, and pre-inference abstention. Not supported: a universal language, compression, or overall efficiency.

Waggle is the machine coordination path; Kea is the separate qualification, lineage, replay, interpretation, and refusal path. The research ladder deliberately moved from symbolic fixtures to a real local model state, then from parity to reuse utility, and finally to held-out continuation and abstention.

Finding 1.1 Boundary finding

Serialized channels were consumable, but not uniquely latent

8/8 numeric · 8/8 one-byte symbolic · 0/8 no-channel

In a frozen 24-call blinded study, both provided channels produced every target artifact and the no-channel arm safely abstained eight times. The numeric arm used 13,631 total tokens versus 11,662 for the symbolic control—1,969 more, or 16.9% by that descriptive count.

Boundary: A follow-up proved that every tested 32-byte vector was losslessly recodable and that the fixture semantics fit in one byte. The next paid study was therefore scientifically ineligible and made zero calls.

Finding 1.2 Supported in scope

Native model state preserved the tested capability

8/8 semantic parity · 8/8 byte parity

A Planner serialized Qwen3-14B Q6_K state and exited. Fresh Executors restored the state without receiving the original human text, and their outputs exactly matched the full-text controls.

Boundary: The eight states totaled 333,229,432 bytes versus 7,270 original prompt bytes—about 45,836× larger. Kea qualified integrity, lineage, interface, and output; it did not semantically decode the opaque state itself.

Finding 1.3 Supported in scope

Same-host reuse produced a narrow latency result

Task-4 break-even · 30.7124% lower task-8 cold wall

One serialized 1,701-token shared context fed eight fresh native Executors and eight paired fresh full-text controls. All eight pairs had exact output and token parity. At task eight, cumulative native cold wall was 36,909.212250 ms versus 53,269.589208 ms.

Boundary: This was one local counterbalanced run. OS cache, thermal state, energy, and cross-host behavior were not controlled. The retained 279,322,905-byte object was about 3,922.2× all eight logical full-text handoffs combined.

Finding 1.4 Supported in scope

Held-out work could be refused before inference

3 supported continuations · 2 pre-Executor abstentions

Across three new contexts and five tasks, the system completed the three source-supported tasks exactly. One ambiguous task and one missing-source task stopped before an Executor launched, with zero model processes for those two requests.

Boundary: The full gate used six local model processes and 76 decode calls, with zero paid-provider, external, or authority effects. This is bounded declared-source retrieval, not broad reasoning or generalization.

Not demonstrated

  • No learned Waggle protocol, universal neuralese, cross-model language, or independent semantic decoding of opaque native state.
  • No token, dollar, storage, network, energy, or overall-efficiency win.
  • No production traffic, customer result, revenue, deployment, or authority grant.
  • The reusable library surface is stable but fixture-only; the live authority gate remains unsigned.

Next legitimate gates

  1. 1. Repeat the timing study across seeds, context sizes, hardware states, and multiple hosts while measuring energy and I/O.
  2. 2. Test whether any compact, inspectable representation preserves utility after strict symbolic and recoding controls.
  3. 3. Keep Kea independent and require a signed capability envelope before any live Agent path exists.
Program 02

HEATWAKE

A liquidity research system that retained its negative result

Full program

Verdict

Supported at synthetic fixture scale: deterministic training and evaluation, exact lineage, frozen promotion gates, and safe refusal. Not supported: synthetic generalization, an empirical forecast, market edge, or profitability.

HEATWAKE began as a dual-view liquidity-model proposal and became a research operating system: content-addressed data, causal features, leakage-aware partitions, frozen controls, two model families, exact recovery, promotion policy, and a local operator that cannot turn failed evidence into a signal.

Finding 2.1 Supported in scope

The synthetic corpus and model runs were reproducible

17 bundles · 51 branches · 2 architectures · 5 folds

A deterministic authored corpus produced 17 superposition bundles, 51 branch examples, 17 tensors, and four partitions. A causal visual TCN and causal attention-fusion model were trained locally and evaluated on ten later bundles each through five expanding-origin folds.

Boundary: The corpus and target were fabricated for mechanism development. They are not exchange observations and cannot establish market behavior.

Finding 2.2 Negative result

Both learned models lost to the simplest valid baseline

Candidate 1.098683870252 · challenger 1.098700995688

The walk-forward table's uniform comparator was 1.098612321409; the promotion replay separately records the theoretical ln(3) value 1.098612288668. Under either comparator both candidates were worse, and neither beat uniform in any of five folds. Their reported Brier sums were 0.000047706305 and 0.000058968539.

Boundary: The authored target had constant one-third support, making uniform the theoretical optimum. This diagnoses a non-identifying synthetic target; it says nothing positive or negative about real-market predictability.

Finding 2.3 Negative result

The promotion system chose no model

0/2 eligible · final test sealed · ABSTAIN

Both architectures failed predictive-identifiability, every-baseline, uniform-fold, shuffled-label-discrimination, calibration-transfer, and mandatory-policy-coverage gates. The preregistered policy forbade selecting the less-bad model.

Boundary: No capital, order, trade, deployment, paid request, external send, or final-test open occurred. A safe local operator can display the research state, but its forecast provenance is mock and its action rail is disabled.

Finding 2.4 Supported in scope

One command replayed the full fixture journey and stayed fail-closed

60 fits · 30 untouched evaluations · 620 controls · 5,376 receipts · 48 envelopes

Two cold deterministic fixture builds traversed custody, environment, training, evaluation, comparison, signal, and operator replay. Every path ended at ABSTAIN. A self-consistent checkpoint-lineage rewrite was refused, exact accepted bytes were restored, and current and stale projections remained fail-closed.

Boundary: This is a real trained fixture-mechanics and recovery result, not an empirical market model. The accepted run used one worker, peaked at 996,638,720 bytes RSS, made zero network requests, and did not establish calibrated direction, edge, profitability, or trading readiness.

Finding 2.5 Boundary finding

The empirical pipeline exists, but contains no empirical result

0 admitted third-party bytes · 0 empirical samples

Source admission, an offline Hyperliquid adapter, causal feature volumes, a 900-second leakage-safe target split, baselines, walk-forward evaluation, negative controls, paper economics, and signal-envelope machinery have implementation and test evidence.

Boundary: No third-party archive has been admitted, imported, trained, evaluated, or opened as a final test. Commercial/public data rights are not cleared. The selected archives exceeded internal free space by 7,156,787,136 bytes and the custody policy requires a dedicated encrypted external SSD.

Not demonstrated

  • No live or admitted historical market corpus, empirical fit, empirical generalization, validated forecast, market edge, or profit result.
  • No trading, capital, orders, external execution, deployment, or production readiness.
  • Synthetic run counts are engineering evidence, not evidence of market-scale volume or model quality.
  • Research-use risk was accepted internally; commercial and public redistribution rights remain unresolved.

Next legitimate gates

  1. 1. Attach and qualify the encrypted external research volume before requesting any archive bytes.
  2. 2. Complete rights and jurisdiction review, admit one provider-neutral historical source, and publish the exact provenance record.
  3. 3. Run the frozen empirical split, baseline, negative-control, calibration, and paper-economics gates before opening any final test.
Cross-program finding

Cross-program conclusions.

Controls that can kill the preferred story

The one-byte Waggle control eliminated a latent-language interpretation; HEATWAKE's uniform baseline eliminated a synthetic signal interpretation. Neither control was weakened after the result.

Independent qualification before action

Kea validates message and state lineage without granting authority. HEATWAKE's promotion policy similarly separates model output from permission to produce a signal or action.

Evidence-preserving recovery

Both programs record amendments and recovery boundaries instead of silently rewriting formal runs. Reproduction can validate frozen findings without re-spending inference or opening sealed data.

Research translated into operable systems

The work includes stable library entrypoints, lifecycle cleanup, local operators, typed contracts, and refusal paths. Product engineering is part of the evidence, but is not substituted for scientific validity.

Limitations and disclosures

Limits on the reported findings.

  1. 1Waggle C6v's packaging step removed three required private ledgers after the formal result. They were reconstructed only where the receipts explicitly declared them unhash-bound and receipt-equivalent. No additional inference occurred; artifact, receipt, evaluation, and report identities were unchanged. The recovery is part of the evidence record.
  2. 2HEATWAKE's top-level legacy verifier initially stopped because a newly discovered pytest module was run by a Python environment that did not declare pytest. The affected assertions were then run with system Python plus a temporary pytest installation; 16/16 passed. The repository was not modified to hide the harness mismatch.
  3. 3The Waggle timing result is descriptive, single-host evidence. The HEATWAKE result is synthetic-only evidence. Neither has been peer reviewed.
  4. 4The public source repositories remain private. This release therefore links to public papers and a public evidence ledger, not to inaccessible GitHub URLs presented as inspectable source.
Fresh local rerun ledger

Rechecked from the source trees.

Waggle + Kea

PASS

Held-out context gate, native fan-out replay, latent-equivalence gate, library API build/import test, TypeScript check, and production build all passed from source commit 180bd10fe17142fce6e6f4a1a79ae2d85b777ed7.

HEATWAKE research

PASS

Promotion and walk-forward unit tests passed 9/9; empirical acquisition and paper-economics tests passed 16/16 from source commit 57ca6bb4209e9679e02bb7aa06cc3dfed27f95d4.

HEATWAKE operator

PASS

The operator service suite passed, and the web operator passed lint, typecheck, 78 tests across 25 files, and a production build.

External effects

ZERO

The reruns made no paid-provider calls, sent no email, placed no order or trade, opened no HEATWAKE final test, granted no authority, and admitted no new empirical bytes.

The work is ready to be judged as engineering research, not as a claim of solved agent communication or solved markets. Waggle + Kea has a reproducible mechanism, a narrow utility result, and a clearly measured bill. HEATWAKE has a complete negative promotion result and an empirical pipeline that has not yet touched empirical data. The common standard—exact lineage, independent checks, refusal, and public limits—is the part designed to scale.

Author and research contact

William Keenan · K&E Studios Research

william@kestudios.dev
Author profile