Harness Inspector: making agent delivery observable, inspectable, and traceable
A structured summary of the Harness Inspector article: how to connect intent, agent sessions, file activity, and Git commits into an evidence-backed delivery chain.
Source and reading boundaries
The article discusses Better Harness and its Harness Inspector, a local read-only workspace for examining how an AI coding task moved from an initial intent to a code contribution. The central idea is broader than a session viewer: a session only shows part of the work, while delivery review needs evidence about why the task started, what the agent did, and what finally entered the repository.
The current project overview is available in the Better Harness repository. Use the repository as the source of truth for current adapters, installation commands, and supported output formats.
From a session to a delivery chain
The article models an agent delivery as three connected but distinct layers:
- Intent: the semantic starting point, such as a user story, Issue, specification, or architecture constraint.
- Process: how the change actually happened, including the session, searches, file reads, tool calls, edits, and validation.
- Output: what the engineering system retained as the result, with a Git commit serving as the clearest current anchor.
Story, Session, and Commit are therefore not three interchangeable views of the same object. They are observable proxies for Intent, Process, and Output. A single story may span multiple sessions, and one session may produce multiple commits, so the relationship is better understood as an evidence graph than as a perfectly linear timeline.
Workbench, Trace, and Replay
Harness Inspector uses three complementary views. Each answers a different question and preserves the boundary between observed evidence and inferred relationships.
Workbench: intent, process, and output
Workbench is the overall delivery view. It places the triggering request, the agent's session activity, file changes, and observed commits in one workspace so a reviewer can inspect the relationship between them.
The important design choice is that weak relationships stay visible as candidates or unmapped evidence. Inspector does not silently invent a complete delivery path merely because several records happen to exist near one another.
Trace: reading a session as a work trajectory
Trace expands one session into turns, user inputs, intermediate responses, tool calls, and file activity. A timeline connects events in time, repeated operations can be folded, and a reviewer can jump from a summary segment to the corresponding call.
Trace does not reconstruct private model reasoning. It reorganizes recorded behavior into a readable trajectory so teams can inspect how the agent searched for context, changed files, and performed validation.
Replay: reviewing events in order
Replay follows the retained events in sequence: user input, agent response, tool call, file activity, and commit. It helps reviewers see when a direction was formed and when modifications or checks happened.
Replay is read-only. It does not rerun tools, restore a workspace, or continue the original session. If an event has no precise timestamp, the ordering can still be preserved without manufacturing timing information.
From delivery evidence to reusable Skills
The article's original question is how to identify reusable work patterns from agent sessions. Its answer is deliberately stricter than counting frequent tool calls. Repeated file reads may indicate missing context, and repeated failed commands may be noise rather than a useful engineering habit.
A stronger Skill candidate is a work path that recurs on similar tasks and is supported by the final output and validation evidence: defining the change boundary, collecting the required context, editing, validating, and checking the result. The same path should then be tested in later deliveries to see whether it actually improves the workflow.
- Start with multiple comparable deliveries instead of one impressive session.
- Keep the task intent, evidence links, validation result, and final output together.
- Separate stable reusable steps from accidental retries, environment noise, and host-specific details.
- Validate the proposed Skill on a later task before treating it as a default workflow.
Using the idea with GPT88 workflows
Harness Inspector is a separate open-source workflow tool, not a GPT88 product or a replacement for API logs. The same evidence model is useful when reviewing coding agents that call GPT88 through an OpenAI-compatible route.
- Use a stable task identifier in the issue, agent session, commit message, and evaluation record.
- Keep the exact model ID, route, prompt version, tool permissions, and validation commands in the delivery record.
- Review the final code and test output together with the session instead of judging the agent only by its final answer.
- Use the GPT88 integration guide and model list API to confirm the active route and account access.
/better-harness analyze this project's AI coding workflow and generate an evidence-backed reportThe command above is the current Qoder-style entry described by the project README. For host-specific installation and invocation, follow the current instructions in the Better Harness repository.
Limitations and safe adoption
- A linked Story, Session, and Commit do not prove that the agent made the correct decision; they make the evidence easier to inspect.
- Observed events are not the same as hidden reasoning. Reviewers should avoid reading intent into unrecorded model state.
- A read-only inspector does not replace code review, tests, CI, permissions, secrets management, or rollback procedures.
- Different coding hosts expose different session and tool evidence. Treat missing evidence as an explicit limitation, not as proof that an event did not happen.
- Do not turn one successful delivery into a universal Skill. Compare repeated tasks and measure quality, rework, time, and operational risk.