Alphature
Running research record · Updated September 17, 2026

What we tested. What we learned.

A dated record of experiments, corrections, negative results and decisions to stop.

Our current question is how to explain an uncertain cross-system operation from fragmented evidence. The entries below describe local work. Production diagnostic accuracy, measured customer savings and paying demand remain unestablished.

We publish our research findings, not private correspondence. Names, company details, quotations and identifying incident descriptions require permission. These summaries contain no customer incident narratives or endorsements.

Read the structured journal · Follow updates

Five constructed-record scenarios; no live runtime capture

Use existing execution records before building another collector

Observed. A modeled deny, revised call, timeout and later tool return could be grouped by explicit invocation identifiers. Removing a separately retained policy decision left that decision unknown; a missing identifier left an event unjoined; a reused identifier in another conversation remained separate.

Limits. This exercise used constructed records informed by published schemas. It did not verify runtime event delivery or administrative logs. Tool success did not establish a downstream commit, and a hook policy decision was not treated as authenticated human approval.

Decision. Investigate configuration of existing tools before building a replacement collector. Full validation requires an authorized operator capture and independent effect evidence.

Fifteen local before-and-after cases

Identity can be known while the effect remains unconfirmed

Observed. Adversarial inputs exposed three defects in our earlier evaluator: duplicate record IDs redirected a write claim, duplicate native IDs erased an attempt claim, and a shared resource name hid an unassigned evidence record. A separate revision rejects duplicate canonical node and relationship IDs, keys claims by evidence record, separates identity from write authority, and retains accepted partial mappings. The original eight cases and seven added variations passed the revised checks.

Limits. These are local fixture results, not production validation. Mapping truth and source authority remain assumptions. A confirmed attempt-to-record link does not establish membership in the intended logical operation when its upstream link is missing. Effect equivalence is not inferred from a shared resource.

Decision. Preserve the earlier result and this correction. Explain unknown outcomes by cause, retain partial evidence, and test identity and effect claims separately before adding inference.

Eight local synthetic cases

When a mapping disappears, attribution must stop

Observed. Three writes committed in a temporary database. Explicit mapping records attributed two distinct resources to two attempts of one intended operation; a third write with the same timestamp and payload fingerprint stayed unassigned. Removing the second dispatch mapping left that attempt UNKNOWN. Cross-tenant and non-authoritative evidence stayed unknown, competing mappings stayed CONFLICTING, and reversing input order preserved results.

Limits. Mapping truth and database authority are fixture assumptions. This does not authenticate production evidence, establish causality from similarity, or measure diagnostic accuracy on real incidents.

Decision. Keep the small example reviewable. Test the representation against independently supplied evidence before expanding it into a general graph engine.

Six local evidence variants

A confirmed write is only one claim

Observed. The revised model distinguishes minimum positive write evidence from a later corroborating lookup. Both observations share a source failure domain. Removing the lookup preserves the narrow write claim; removing the commit leaves it unknown. Payload and lookup disagreements remain conflicting. Attempt attribution is withheld when the downstream record lacks an attempt identifier.

Limits. Write confirmation does not establish business completion, caller reconciliation or the response-loss cause. Multiple records from one producer are not independent confirmation.

Decision. Carry support record identifiers through each claim; preserve unresolved questions separately.

Synthetic diagnostic

A timeout can follow a committed write

Observed. A temporary database committed a write before the harness deliberately raised a timeout. A read-only lookup subsequently found the record. The modeled retry was held; the database contained one row and no retry was executed. Partial evidence left the outcome unknown.

Limits. This is a local simulation, not a network-failure reproduction or production incident. The supplied records do not by themselves establish the injected response-loss cause.

Decision. Investigate how to distinguish recorded facts, supported conclusions and unresolved outcomes.

Six source-body checks with stub clients

Run status can omit the pending decision

Observed. The inspected check-tool code fetched output for successful runs and exposed errors for failed runs. Interrupted runs returned status and thread identity without fetching the pending interrupt context, even when the stub could supply it.

Limits. Only selected release function bodies were exercised. This was not a full SDK, hosted server or resume-operation test. No lost checkpoint was established.

Decision. Require decision context separately during diagnostic intake. Pause implementation pending an affected integration.

Sixteen scripted local cases

The reported approval failure did not reproduce

Observed. Every modeled flow paused before its write. After one resume, approval produced one synthetic write and rejection produced none. The child resumed at its pending gate rather than restarting its plan; no case interrupted a second time.

Limits. The test did not recreate the reported redispatch path, process crashes or concurrent clients. Failure to reproduce neither disproves a report nor establishes a fix.

Decision. Pause until a minimal configuration or trace explains what triggers redispatch.

Sixteen comparisons and four configuration checks

Workflow shape changed preparation replay

Observed. Nested workflow shapes repeated a completed preparation task after resume in both tested versions. A top-level shape did not. Two configuration alternatives on the newer tested version ran preparation once and preserved approval and rejection behavior.

Limits. This was not an approval bypass. Local counters stood in for preparation work; no provider expense or real duplicate downstream effect was measured. In-memory persistence does not test restart durability.

Decision. Prefer the demonstrated configuration alternatives for these fixtures. Do not start another replay product without a concrete residual need.

Local session-lifetime control; follows September 10 observation

A deadline bounded reconnection by ending the session

Observed. In the initial six-second observation, a normally closed stream triggered six GET requests; a refusal response triggered two. A later 2.5-second caller deadline ended the session. Three synthetic notifications arrived before cancellation in the useful-event case, and no new GET occurred during the 1.3-second post-exit window.

Limits. An observation window does not prove infinite reconnection. Ending the session also ends useful notifications; this is not a repair for a persistent stream.

Decision. Pause standalone product work unless a real workflow needs more than caller-owned session containment.

Five synthetic cases

A request key needs a scope

Observed. An unscoped key replayed a response across synthetic tenants when payloads matched. A structured tenant, operation and request tuple separated those cases while retaining matching retries. Simple delimiter concatenation collided for the chosen delimiter-bearing inputs.

Limits. The fixture does not supply tenant authentication or prove production isolation. Identity scoping remains a caller responsibility.

Decision. Retain the negative control and document structured scope; no product patch was inferred from caller configuration.

Five acceptance cases and two limitation controls

A missing budget file must not reset authority

Observed. The baseline met two of five acceptance cases; the candidate met all five by refusing normal startup without the existing store and its matching identity. It still accepted an older valid snapshot with the same identity. Explicit reinitialization could also create a new budget after deleting both files.

Limits. These were temporary synthetic ledgers, not provider charges. Identity matching does not establish rollback-resistant accounting or account-wide spend enforcement.

Decision. Keep initialization outside automatic recovery. Do not describe the local accounting control as a provider billing cap.

Nine synthetic scenarios

Measurement controls need evidence of their own

Observed. A clear synthetic signal reached a candidate awaiting human acceptance. Weak, competing and invalid signals stopped for review. Missing authorization performed no measurement. Unsupported actions, out-of-range parameters and exhausted resources were refused. Altered evidence failed local verification.

Limits. No physical equipment or live inference was used. A Python audit hook is not an operating-system sandbox; local hashes do not establish independent custody or authenticated human acceptance.

Decision. Keep the distinction between a simulated control exercise and a validated laboratory.

Earlier studies remain available

The Double-Trigger Test, experimental guard release and SDK compatibility study remain part of the record. Their original scope and limitations still apply. The July archive documents an earlier direction.

Corrections and publication notes

: the early selection note and feed described an experiment that had not yet run. That wording described the selection stage; the local run is now complete. The updated note links forward to the result while retaining the original selection history.

The narrow synthetic write report has also been clarified: its later lookup is additional corroboration, not independent proof. The minimum write evidence and additional support now have distinct roles in the local model.

This page is an editorial summary, not the complete experimental evidence bundle. Results are not externally replicated unless explicitly stated. Future updates should add dated entries or correction notes rather than silently replacing earlier outcomes.