Problem

Agent skills can sound persuasive while leaving behavior unchanged. The problem is to separate an apparent improvement from a reproducible, inspectable change under controlled conditions.

What I built

A persistent lab connecting canonical artifacts, controlled experiments, preserved run evidence, histories, reports, and verification tooling. The record keeps the materials behind a conclusion available for review.

Architecture / decisions

The sequence is deliberately linear: control the inputs before interpreting the result.

  1. ControlHold prompt, baseline, model, and harness conditions constant.
  2. IsolateKeep immutable Mother artifacts separate from disposable Active run state.
  3. CompareRun original, no-skill, and candidate conditions with one global run identity.
  4. ReplicateRepeat across models instead of treating one favorable run as a broad effect.
  5. VerifyHash result state and transcripts, then retain command output for verification claims.

Evidence

Preserved failures remain in the record before cleanup. Exact result state supports outcome claims; transcripts support process claims; command output supports verification claims. SHA-256 manifests make the saved materials checkable.

A failed condition has the same evidentiary weight as a favorable one: it can narrow or reject a claim.

Evidence map showing control, isolation, comparison, replication, and hash verification, with a captured saved-record status block
The 0031 block is a captured saved record from the supplied evidence image, not live or current repository status for the public repository.

Result

Evidence can support, narrow, or reject a behavioral claim. Effects that do not reproduce do not become general claims.

Technologies

Current boundary / unfinished work

Experimental candidates are not promoted as general improvements. The image’s 0031 block is a captured saved record, not live or current repository status for the public repository. Its confirmed public evidence release records runs 0001–0015.