Jump to a findings section
A narrow positive result and a decisive negative result.
Waggle + Kea established exact, source-grounded native-state continuation on one local Qwen3-14B runtime and measured a same-host fan-out latency break-even at the fourth task. The same experiment also showed that the retained state was roughly 279.3 MB for an 8.5 KB shared prefix, so no storage, transport, or overall-efficiency advantage is supported. HEATWAKE trained and evaluated two architectures on a deterministic authored synthetic corpus. Both lost to a uniform baseline across all five held-out folds, failed the frozen promotion gates, and were rejected. A later one-command fixture journey then replayed custody through operator display and remained fail-closed, but no empirical market data has been admitted. Across both programs, the strongest accomplishment is a working discipline for preregistration, exact replay, independent qualification, refusal, recovery disclosure, and bounded claims.
Program record
Waggle + Kea
Supported at fixture scale: exact continuation, deterministic replay, same-host reuse utility, and pre-inference abstention. Not supported: a universal language, compression, or overall efficiency.
Open programProgram record
HEATWAKE
Supported at synthetic fixture scale: deterministic training and evaluation, exact lineage, frozen promotion gates, and safe refusal. Not supported: synthetic generalization, an empirical forecast, market edge, or profitability.
Open programClaims, controls, receipts, reruns.
Primary evidence
Claims were traced to frozen JSON reports, preregistrations, receipts, content hashes, source commits, and the exact source trees that produced them.
Fresh local rerun
The decisive Waggle/Kea gates and HEATWAKE promotion, walk-forward, empirical-machinery, service, and operator suites were rerun from the named research worktrees on July 15, 2026.
Claim grammar
A result is labeled supported only inside the tested fixture and runtime. Negative results are retained as results. Anything outside the tested envelope is listed as not demonstrated.
Public audit surface
This report is paired with a downloadable evidence ledger and SHA-256 checksum. Private repositories are not presented as publicly inspectable source.
Waggle + Kea
Machine coordination with an independent decipher plane
Verdict
Supported at fixture scale: exact continuation, deterministic replay, same-host reuse utility, and pre-inference abstention. Not supported: a universal language, compression, or overall efficiency.
Waggle is the machine coordination path; Kea is the separate qualification, lineage, replay, interpretation, and refusal path. The research ladder deliberately moved from symbolic fixtures to a real local model state, then from parity to reuse utility, and finally to held-out continuation and abstention.
Serialized channels were consumable, but not uniquely latent
8/8 numeric · 8/8 one-byte symbolic · 0/8 no-channel
In a frozen 24-call blinded study, both provided channels produced every target artifact and the no-channel arm safely abstained eight times. The numeric arm used 13,631 total tokens versus 11,662 for the symbolic control—1,969 more, or 16.9% by that descriptive count.
Boundary: A follow-up proved that every tested 32-byte vector was losslessly recodable and that the fixture semantics fit in one byte. The next paid study was therefore scientifically ineligible and made zero calls.
Native model state preserved the tested capability
8/8 semantic parity · 8/8 byte parity
A Planner serialized Qwen3-14B Q6_K state and exited. Fresh Executors restored the state without receiving the original human text, and their outputs exactly matched the full-text controls.
Boundary: The eight states totaled 333,229,432 bytes versus 7,270 original prompt bytes—about 45,836× larger. Kea qualified integrity, lineage, interface, and output; it did not semantically decode the opaque state itself.
Same-host reuse produced a narrow latency result
Task-4 break-even · 30.7124% lower task-8 cold wall
One serialized 1,701-token shared context fed eight fresh native Executors and eight paired fresh full-text controls. All eight pairs had exact output and token parity. At task eight, cumulative native cold wall was 36,909.212250 ms versus 53,269.589208 ms.
Boundary: This was one local counterbalanced run. OS cache, thermal state, energy, and cross-host behavior were not controlled. The retained 279,322,905-byte object was about 3,922.2× all eight logical full-text handoffs combined.
Held-out work could be refused before inference
3 supported continuations · 2 pre-Executor abstentions
Across three new contexts and five tasks, the system completed the three source-supported tasks exactly. One ambiguous task and one missing-source task stopped before an Executor launched, with zero model processes for those two requests.
Boundary: The full gate used six local model processes and 76 decode calls, with zero paid-provider, external, or authority effects. This is bounded declared-source retrieval, not broad reasoning or generalization.
Not demonstrated
- No learned Waggle protocol, universal neuralese, cross-model language, or independent semantic decoding of opaque native state.
- No token, dollar, storage, network, energy, or overall-efficiency win.
- No production traffic, customer result, revenue, deployment, or authority grant.
- The reusable library surface is stable but fixture-only; the live authority gate remains unsigned.
Next legitimate gates
- 1. Repeat the timing study across seeds, context sizes, hardware states, and multiple hosts while measuring energy and I/O.
- 2. Test whether any compact, inspectable representation preserves utility after strict symbolic and recoding controls.
- 3. Keep Kea independent and require a signed capability envelope before any live Agent path exists.
HEATWAKE
A liquidity research system that retained its negative result
Verdict
Supported at synthetic fixture scale: deterministic training and evaluation, exact lineage, frozen promotion gates, and safe refusal. Not supported: synthetic generalization, an empirical forecast, market edge, or profitability.
HEATWAKE began as a dual-view liquidity-model proposal and became a research operating system: content-addressed data, causal features, leakage-aware partitions, frozen controls, two model families, exact recovery, promotion policy, and a local operator that cannot turn failed evidence into a signal.
The synthetic corpus and model runs were reproducible
17 bundles · 51 branches · 2 architectures · 5 folds
A deterministic authored corpus produced 17 superposition bundles, 51 branch examples, 17 tensors, and four partitions. A causal visual TCN and causal attention-fusion model were trained locally and evaluated on ten later bundles each through five expanding-origin folds.
Boundary: The corpus and target were fabricated for mechanism development. They are not exchange observations and cannot establish market behavior.
Both learned models lost to the simplest valid baseline
Candidate 1.098683870252 · challenger 1.098700995688
The walk-forward table's uniform comparator was 1.098612321409; the promotion replay separately records the theoretical ln(3) value 1.098612288668. Under either comparator both candidates were worse, and neither beat uniform in any of five folds. Their reported Brier sums were 0.000047706305 and 0.000058968539.
Boundary: The authored target had constant one-third support, making uniform the theoretical optimum. This diagnoses a non-identifying synthetic target; it says nothing positive or negative about real-market predictability.
The promotion system chose no model
0/2 eligible · final test sealed · ABSTAIN
Both architectures failed predictive-identifiability, every-baseline, uniform-fold, shuffled-label-discrimination, calibration-transfer, and mandatory-policy-coverage gates. The preregistered policy forbade selecting the less-bad model.
Boundary: No capital, order, trade, deployment, paid request, external send, or final-test open occurred. A safe local operator can display the research state, but its forecast provenance is mock and its action rail is disabled.
One command replayed the full fixture journey and stayed fail-closed
60 fits · 30 untouched evaluations · 620 controls · 5,376 receipts · 48 envelopes
Two cold deterministic fixture builds traversed custody, environment, training, evaluation, comparison, signal, and operator replay. Every path ended at ABSTAIN. A self-consistent checkpoint-lineage rewrite was refused, exact accepted bytes were restored, and current and stale projections remained fail-closed.
Boundary: This is a real trained fixture-mechanics and recovery result, not an empirical market model. The accepted run used one worker, peaked at 996,638,720 bytes RSS, made zero network requests, and did not establish calibrated direction, edge, profitability, or trading readiness.
The empirical pipeline exists, but contains no empirical result
0 admitted third-party bytes · 0 empirical samples
Source admission, an offline Hyperliquid adapter, causal feature volumes, a 900-second leakage-safe target split, baselines, walk-forward evaluation, negative controls, paper economics, and signal-envelope machinery have implementation and test evidence.
Boundary: No third-party archive has been admitted, imported, trained, evaluated, or opened as a final test. Commercial/public data rights are not cleared. The selected archives exceeded internal free space by 7,156,787,136 bytes and the custody policy requires a dedicated encrypted external SSD.
Not demonstrated
- No live or admitted historical market corpus, empirical fit, empirical generalization, validated forecast, market edge, or profit result.
- No trading, capital, orders, external execution, deployment, or production readiness.
- Synthetic run counts are engineering evidence, not evidence of market-scale volume or model quality.
- Research-use risk was accepted internally; commercial and public redistribution rights remain unresolved.
Next legitimate gates
- 1. Attach and qualify the encrypted external research volume before requesting any archive bytes.
- 2. Complete rights and jurisdiction review, admit one provider-neutral historical source, and publish the exact provenance record.
- 3. Run the frozen empirical split, baseline, negative-control, calibration, and paper-economics gates before opening any final test.
Limits on the reported findings.
- 1Waggle C6v's packaging step removed three required private ledgers after the formal result. They were reconstructed only where the receipts explicitly declared them unhash-bound and receipt-equivalent. No additional inference occurred; artifact, receipt, evaluation, and report identities were unchanged. The recovery is part of the evidence record.
- 2HEATWAKE's top-level legacy verifier initially stopped because a newly discovered pytest module was run by a Python environment that did not declare pytest. The affected assertions were then run with system Python plus a temporary pytest installation; 16/16 passed. The repository was not modified to hide the harness mismatch.
- 3The Waggle timing result is descriptive, single-host evidence. The HEATWAKE result is synthetic-only evidence. Neither has been peer reviewed.
- 4The public source repositories remain private. This release therefore links to public papers and a public evidence ledger, not to inaccessible GitHub URLs presented as inspectable source.
Rechecked from the source trees.
Waggle + Kea
PASSHeld-out context gate, native fan-out replay, latent-equivalence gate, library API build/import test, TypeScript check, and production build all passed from source commit 180bd10fe17142fce6e6f4a1a79ae2d85b777ed7.
HEATWAKE research
PASSPromotion and walk-forward unit tests passed 9/9; empirical acquisition and paper-economics tests passed 16/16 from source commit 57ca6bb4209e9679e02bb7aa06cc3dfed27f95d4.
HEATWAKE operator
PASSThe operator service suite passed, and the web operator passed lint, typecheck, 78 tests across 25 files, and a production build.
External effects
ZEROThe reruns made no paid-provider calls, sent no email, placed no order or trade, opened no HEATWAKE final test, granted no authority, and admitted no new empirical bytes.