Blog
Case studies from building hiway-kit: what broke, how we found it, and what the kit does about it now. Each post is checked against the current kit and says what changed since.
The Last Task Never Closes: Completion Marks That Evaporate at the Turn Boundary, and Three Layers of Defense
We tracked down a bug where the final task stayed in_progress after the work was done. The culprit was the order of a single instruction; this is how we stacked discipline, mechanical detection, and detector hardening on top of the fix, and what has changed since.
Green Lights Lie: A Six-Way Adversarial Audit of the Completion Gates I Built
My completion gate was green. Then I pointed six fresh-context adversarial reviewers at the whole project and found those very gates reporting failures as success. A record of verifying the verification machinery, and why 'the author is contaminated' hits hardest when you wrote the gate.
Done Is a Command's Output, Not a Claim: Designing a Durable Completion Gate for Long-Running Loop Agents
Let a loop agent mark its own work as done and unverified code sails through. How the discovery that native Tasks are session-scoped killed our first design, and how we moved completion state into a single file judged by the raw exit code of a verify command.
Harness Engineering in Practice: Hardening Git Isolation in a Parallel-Agent Toolkit with Two Adversarial Reviews
A case study in applying 'Agent = Model + Harness' to a real Claude Code toolkit: fixing race conditions in parallel work, and how a second adversarial review caught a verification bypass the first one missed.