The next AI does not know what tests passed or failed

A handoff without verification state is dangerous. The next AI may assume the build passed, retry the wrong test, or ship a half-fix.

Short version: The next agent needs the evidence, not just the edit.

Why this hurts.

When test state is lost, the next session wastes time rediscovering whether it is implementing, debugging, or verifying.

How ShardStitch solves it.

01 / ShardStitch

Record the exact check and scope

Include verification state in the continuation packet.

02 / ShardStitch

Separate passed, failed, and unrun

Separate passed checks, failed checks, and checks not yet run.

03 / ShardStitch

Tie results to the current change

Tie failures back to changed files and next smallest action.

04 / ShardStitch

Rerun the smallest relevant check

Rerun the narrowest safe check against the current files; keep any broader unrun checks explicitly pending.

What the next AI receives.

  • Commands run
  • Pass/fail state
  • Manual checks
  • Next verification step

What stays out.

  • Assumed success
  • Unclear test claims
  • Old failures already fixed
  • Logs unrelated to the current task

Re-establish what has actually been verified

For each check, preserve the command, target, result, and commit or diff it covered. "Tests passed" without scope or output can become misleading after another edit.
Label each check passed, failed, interrupted, or not run. Keep failures and pending checks visible in the handoff.
Compare the recorded result with changes made since it ran. Any relevant later edit may invalidate the old verification.
Rerun the smallest check that covers the current change, then broaden testing based on risk. Report the actual output rather than inferred success.

FAQ.

Does a handoff prove that a test passed?

No. A handoff can carry a recorded result, but the claim is only as reliable as its command output, scope, and freshness. Re-run the test when changes may have invalidated it.

What if I cannot find the old test output?

Mark the check unverified and rerun it if safe. If it cannot be rerun, state that limitation rather than claiming it passed.

Related pages.