Skip to content
Research library
Findings update 008 August 1, 2026 Locally source-verified

Timing crossed; the quality gate failed

The apparent selector win does not survive its own quality requirement. The useful result is the contradiction: readable and native representations remained byte-identical to each other, yet repeated the same eight task errors. HEATWAKE added replay and adapter mechanics, not empirical model evidence.

Task quality

88/96 expected

Both informed representations repeated the same eight errors, so byte parity did not satisfy correctness.

Selector timing

5.538% descriptive crossover

The selector measured below both static policies, but the failed 96/96 quality gate prevents an adaptive-utility claim.

Historical mechanics

2,555,132 accepted rows

Four event-time replay modes completed with zero future-event releases; 883,154 rows remained quarantined.

Review status
Source-verified local research update; not peer reviewed or independently validated
Author
William Keenan
Publisher
K&E Studios Research
Information cutoff
August 1, 2026 at 09:05:25 EDT
Evidence scope
Frozen committed snapshots, preregistered heterogeneous-regime controls, exact zero-model replay, bounded historical event-time mechanics, quarantined adapter checks, and a frozen launch record
Jump to a findings section
Executive finding

The timing rule crossed; the exact task-quality rule did not.

Across 12 frozen regime cells, the selector matched the lower measured representation each time and recorded 484,457.695 milliseconds, 5.538% below the better static policy. That descriptive crossover did not qualify: native-restored and cached-prefix arms were byte-identical on all 96 paired decisions but each matched only 88 of 96 expected decisions, below the preregistered 96/96 quality gate. Causal-replacement and no-state controls each returned 0 of 24 target passes, and exact local replay preserved the failure. The adaptive selector line is therefore terminated with the null hypothesis retained. HEATWAKE separately passed 10 focused historical-replay checks over one closed event-time hour: 3,438,286 rows parsed, 2,555,132 accepted, and 883,154 quarantined across four deterministic mechanics modes with zero future-event releases. A 7,293-row adapter sample passed 6 focused checks but remained quarantined for missing identifiers, ordering reversals, schema drift, and an upstream revision mismatch. A frozen source record states that a fresh seven-day raw collector launched at 0 of 7 complete days; this review did not inspect its growing state. No source admission, model fit, final-test opening, signal, edge, profitability, order, trade, or production result followed.

What changed

Atomic claim deltas.

Swipe horizontally to compare all four columns.

ClaimPrior statusNew evidenceCurrent status
A heterogeneous-regime selector can beat both static representations without reducing task quality.Open. The prior uniform-regime pilot found no adaptive choice and required a heterogeneous crossover study.The selector matched the lower representation in all 12 regime cells and measured 5.538% below the better static policy, but both informed arms returned only 88/96 expected decisions instead of the required 96/96.Rejected for this study; the quality gate failed and the adaptive selector line is terminated.
Matching native and readable outputs establish correct task behavior.Earlier bounded studies had exact paired parity and exact expected decisions.The two informed arms were byte-identical on 96/96 paired decisions while repeating the same eight deterministic task errors; replacement and no-state controls each passed 0/24 targets.Contradicted as a general inference: representation parity is not the same as task correctness.
HEATWAKE historical data can run through the existing book mechanics without future-event release.No bounded historical event-time mechanics result had passed the public review.One closed train-only hour parsed 3,438,286 rows, accepted 2,555,132, quarantined 883,154, and completed four deterministic mechanics modes with zero future-event releases; 10/10 focused checks passed.Verified for event-time mechanics only; unavailable receive time, incomplete sequence context, training admission, live latency, fills, and scientific claims remain excluded.
A fresh HEATWAKE collection changes source-admission or scientific status.A fresh contract existed without committed launch evidence.The frozen snapshot records a fresh raw collector launch at 0/7 complete days. This review excluded active and later collector state, and state-dependent prelaunch checks correctly refused to reinterpret a post-launch environment.Recorded launch only; source admission and every model, final-test, signal, edge, trading, and production gate remain closed.
How this was verified

Claims, controls, receipts, reruns.

1

Frozen information cutoff

The review fixed one committed Waggle/Kea snapshot and one committed HEATWAKE snapshot at August 1, 2026 at 09:05:25 EDT. Uncommitted, growing, lock-held, active, and later source-writer bytes were excluded.

2

Preregistered heterogeneous comparison

Short, middle, and long regimes used the same pinned model, runtime, deterministic task family, equal source information, fixed expected decisions, static baselines, and a selector frozen before outcome access. Support required both timing crossover and 96/96 expected decisions in each informed arm.

3

Quality and information controls

Byte-identical paired output was checked separately from correctness against frozen expected decisions. Causal replacement and absent-state controls tested whether unsupported information could pass the target.

4

Exact local replay

Fourteen focused checks and a formal replay rehashed the accepted package, recomputed qualification and consumer output, and preserved the negative result without another model or provider call.

5

Historical event-time mechanics

One exact closed train-only hour ran through four deterministic replay modes. Rows with unsupported structure or timing were quarantined; future-event release was forbidden. The test does not supply receive-time latency, queue position, fill realism, or model eligibility.

6

Launch-custody separation

A frozen launch artifact is Recorded evidence of a start, not Verified evidence of current collection health or completion. Growing collector state was excluded, and seven full UTC days plus a terminal receipt remain prerequisites for a separate source-admission audit.

Program 01

Waggle + Kea

Descriptive crossover rejected by exact task quality

Full program

Verdict

Verified in scope: 96/96 byte-identical paired decisions, a descriptive 12-cell timing crossover, information-control refusal, Kea qualification, and exact replay. Rejected: adaptive selector utility because both informed arms reached only 88/96 expected decisions. Not established: general timing, overall efficiency, cross-model transfer, unique semantics, production value, or authority.

The selector found the intended heterogeneous timing regimes, but the experiment's stronger safeguard worked: a faster policy cannot qualify when its inputs reproduce the same task errors. This closes the adaptive selector line rather than moving its target.

Finding 1.1 Supported in scope

The timing crossover appeared

12/12 cells selected · 5.538% below best static

The frozen selector matched the lower measured representation in every short, middle, and long regime cell.

Boundary: The interval is a descriptive block bootstrap on one deterministic local fixture. It is not confirmatory, independent, or evidence of general efficiency.

Finding 1.2 Negative result

The task-quality gate stopped the result

88/96 expected per arm · 96/96 byte parity

Native-restored and cached-prefix outputs matched each other exactly but repeated the same eight deterministic errors.

Boundary: Representation parity cannot substitute for correctness against a frozen task target. The preregistered gate required 96/96 in both arms.

Finding 1.3 Supported in scope

The information controls still refused

0/24 replacement · 0/24 no-state target passes

Neither causally replaced information nor an absent handoff produced a target pass, while the qualified projection remained auditable.

Boundary: These controls support information necessity only for this fixture; they do not establish a language, semantic compression, or cross-task transfer.

Finding 1.4 Boundary finding

The negative result replayed exactly

14 focused checks · formal replay PASS

The frozen package, source pins, qualification, and consumer outputs reproduced without another model or external call.

Boundary: Local replay verifies internal consistency, not peer review or independent reproduction. Three failed preliminary trials remain preserved.

Not demonstrated

  • No adaptive selector utility, and no further selector rerun is admitted by this completed line.
  • No general speed, token, Credit, provider-cost, memory, storage, network, energy, economic, or overall-efficiency claim.
  • No learned language, hidden-state semantics, compression, unique capability, higher quality, or causal explanation for timing.
  • No cross-model, cross-runtime, cross-hardware, production, customer, deployment, or authority result.
  • No peer review, independent replication, or confirmatory inference.

Next legitimate gates

  1. 1. Retain the null hypothesis and stop the adaptive selector line; do not retarget or rerun the failed study.
  2. 2. If the separately paused source program resumes, audit only a sanitized zero-model replication package before any external publication decision.
  3. 3. Any broader representation claim requires a new, independently preregistered task family and exact task-quality controls.
Program 02

HEATWAKE

Historical mechanics verified; admission remains closed

Full program

Verdict

Verified in scope: bounded historical event-time replay mechanics, quarantine behavior, and adapter controls. Recorded only: a fresh collection launched at 0/7 complete days. Not supported: source admission, training eligibility, live latency, queue fills, a forecast, signal, edge, profitability, trading, or production operation.

Historical rows can exercise the existing book mechanics without future-event release, but missing receive-time and sequence context prevent the replay from becoming a live-latency, fill, training, or scientific result. The fresh raw collector remains a separate incomplete custody process.

Finding 2.1 Supported in scope

Historical event-time mechanics completed

3,438,286 parsed · 2,555,132 accepted

Four deterministic replay modes completed over one closed train-only hour, with zero future-event releases and 10/10 focused checks passing.

Boundary: The source lacks receive or availability time, an initial snapshot, and complete sequence context. This is mechanics evidence only.

Finding 2.2 Boundary finding

Quarantine remained material

883,154 rows · 25.686% quarantined

Unsupported or incomplete rows were withheld rather than silently promoted; missing joins, unused eligible statuses, and timestamp regressions remain disclosed.

Boundary: A passing mechanics replay does not repair source gaps or establish representativeness, causal availability, queue position, or fill realism.

Finding 2.3 Boundary finding

The adapter sample remains non-scientific

7,293 rows · 6 focused checks

Parser and custody checks passed while missing identifiers, physical ordering reversals, schema drift, and an upstream revision mismatch remained explicit.

Boundary: The sample is quarantined and cannot enter training, evaluation, labels, calibration, final testing, or source admission.

Finding 2.4 Boundary finding

The new collection is only a launch record

0/7 complete days at frozen launch

The committed artifact records a fresh forward-only collection start under the unchanged seven-day contract.

Boundary: Active and later collector bytes were excluded. No current-health, terminal-custody, source-admission, model, signal, edge, trading, or production claim follows.

Not demonstrated

  • No complete seven-day collection, terminal receipt, passing coverage audit, admitted source, or model-input eligibility.
  • No receive-time or availability-time latency, initial snapshot, complete sequence context, queue position, or fill realism.
  • No feature or label construction, calibration, model fit, candidate selection, or sealed final-test result.
  • No forecast, signal, market edge, profitability, execution result, trading readiness, or production reliability.
  • No raw market payload, identity-bearing table, private receipt, reconstruction detail, credential, order, trade, capital, or spend data is published.

Next legitimate gates

  1. 1. Let the sole source-owned collector continue without steering, restart, or review-lane access to growing state.
  2. 2. Require seven qualifying full UTC days and an exact terminal receipt before the separately frozen source-admission audit can run.
  3. 3. Keep historical and adapter artifacts quarantined unless a future preregistered role passes rights, causal-clock, coverage, and source-admission gates.
Cross-program finding

Cross-program conclusions.

Correctness outranks an attractive timing result

A selector that beats static timing baselines still fails when both informed arms miss frozen expected decisions.

Mechanics and scientific evidence remain separate

Historical rows can test event-time replay and quarantine without becoming admitted training data or a market claim.

Frozen records bound active work

A committed launch record can be reviewed without reading, steering, or promoting a growing collection.

What would change our mind

Falsifiers and revision conditions.

  1. 1Revise the C17b conclusion only if an independent reproduction of the exact frozen method reaches the preregistered 96/96 expected-decision gate; do not rewrite or rerun the completed source result.
  2. 2Revise broader representation conclusions only through a separately preregistered task family, model, runtime, and host with exact task-quality and information controls.
  3. 3Revise HEATWAKE historical scope only after receive-time or availability-time custody, initial-book and sequence completeness, and separately frozen queue and fill controls pass.
  4. 4Revise HEATWAKE source admission only after seven qualifying full UTC days, exact terminal custody, and the already-frozen admission audit pass without opening the sealed final test.
Limitations and disclosures

Limits on the reported findings.

  1. 1All verification described here is local source verification by K&E Studios Research, not peer review or independent validation.
  2. 2The selector study uses one local model, runtime, workstation, deterministic numerical task family, and descriptive timing measurements. Direct energy was unavailable.
  3. 3The selector's block-bootstrap estimate is descriptive. The quality gate failed before any adaptive-utility claim could be admitted.
  4. 4Native retained state was 705,898,303 bytes versus 5,914 readable-prefix bytes, so lower measured wall time is not an overall-efficiency result.
  5. 5Three failed preliminary C17b trials are retained. The accepted run's errors were not corrected, dropped, or relabeled after outcome access.
  6. 6The historical replay lacks receive or availability time, an initial book snapshot, and complete source sequence context. It cannot support live-latency, queue, fill, or market-validity claims.
  7. 7The frozen collection artifact records a launch at zero completed days. This review excluded the active collector and makes no current-runtime claim.
  8. 8The citation record is first-party attribution metadata. No DOI, archive deposit, external repository publication, peer review, or independent endorsement exists.
  9. 9This verification made no paid request, provider call, new model run, source-program write, final-test open, order, trade, capital change, credential change, or authority grant.
Fresh local rerun ledger

Rechecked from the source trees.

Waggle + Kea heterogeneous controls

PASS

Fourteen focused checks preserved 96/96 paired parity, the 88/96 expected-decision failure in each arm, selector accounting, replacement refusal, and no-state abstention.

Waggle + Kea formal replay

PASS

The frozen package, implementation pins, artifact pins, prior checkpoint, Kea qualification, and consumer outputs replayed with zero model, provider, network, or authority effects.

HEATWAKE historical mechanics

PASS

Ten focused checks covered deterministic replay modes, quarantine accounting, event-time release, source bindings, and zero downstream science or execution effects.

HEATWAKE adapter controls

PASS

Six focused checks preserved the sample's hash, schema, custody, refusal, and quarantine boundaries without source admission or model use.

HEATWAKE launch-state controls

REFUSED AS DESIGNED

Two state-coupled prelaunch suites refused in the post-launch environment. The frozen launch therefore remains Recorded only and active collector state remains outside this review.

The August 1 review closes the adaptive selector line with a useful negative result: heterogeneous timing crossed, but exact task quality did not. HEATWAKE advanced bounded historical replay and adapter mechanics while refusing every leap to admitted data, model evidence, signals, trading, or production. The next legitimate work is an optional sanitized zero-model replication audit only if the separately paused source program resumes, and a future HEATWAKE source-admission audit only after complete terminal custody.

Author and research contact

William Keenan · K&E Studios Research

william@kestudios.dev
Author profile