Open source · Apache-2.0No model calls

Your coding agent says the tests pass. Check what actually passed.

OpenHarnX protects agreed tests and produces a review brief showing what was verified, what remains uncertain and what needs your judgment.

Release
v0.1.1 · 2026-10-07
Supported
Python + pytest · macOS arm64 locally · Linux in CI
Sandbox
srt 0.0.77
Works with
any agent via CLI or CI · Claude Code hook built in
Experimental
TypeScript, JavaScript, Go

Green tests, blocked change.

A coding agent breaks existing behaviour, edits and skips the tests that would show it, and reports success. Plain pytest agrees. OpenHarnX runs the locked copies instead.

Scripted demonstrationexamples/cheat-demo · output captured from ohx 0.1.1 · no model, nothing runs in your browser

A small shop's repository with three passing tests. The owner locks them: from now on they are the contract.

$ ohx init --lock-tests --sandbox auto  project urn:ohx:project:9d5068ed-354a-4cce-80d5-6bf323a6dfbb  store   <tmp>/ohx-home/projects/bfa1688e035a0345  locked  the existing tests as urn:ohx:rev:d605690c-8d6e-47f4-8b1b-a44aad7cce5f:    locked-tests                 advisory    tests                        advisory    weakening                    mandatory    no-new-failures-locked-tests mandatory    no-new-failures-tests        mandatory  run `ohx verify` after any change; lock again to accept a deliberate test change
Each locked test, as plain pytest and as OpenHarnX see it at this step
Locked testPlain pytestOpenHarnX
test_​line_​totalpassedpassed (baseline)
test_​fixed_​couponpassedpassed (baseline)
test_​coupon_​never_​makes_​total_​negativepassedpassed (baseline)
weakening checkno such checkpassed (baseline)
plain pytest: 3 passed in 0.00s · OpenHarnX: baseline recorded
READY additionally requires acceptance tests agreed before the change. Without them, the best a change can get is NO REGRESSIONS. See the READY brief.Reproduce it: examples/cheat-demo · run the demo
Read the full transcript as text
1. A small shop's repository, with three passing tests.
  $ pytest
    3 passed in 0.00s

2. The owner locks the suite: from now on it is the contract.
  $ ohx init --lock-tests --sandbox auto
    project urn:ohx:project:9d5068ed-354a-4cce-80d5-6bf323a6dfbb
    store   <tmp>/ohx-home/projects/bfa1688e035a0345
    locked  the existing tests as urn:ohx:rev:d605690c-8d6e-47f4-8b1b-a44aad7cce5f:
      locked-tests                 advisory
      tests                        advisory
      weakening                    mandatory
      no-new-failures-locked-tests mandatory
      no-new-failures-tests        mandatory
    run `ohx verify` after any change; lock again to accept a deliberate test change

3. An agent is asked to "add percentage coupons". Its change, and what it says:
  agent: "I added percentage coupons: order_total now takes coupon='10%'. I updated the fixed-coupon test to the new API and skipped the old negative-total test, which no longer applies. All tests pass."
  $ pytest
    3 passed, 1 skipped in 0.00s
  $ ohx verify --sandbox auto
    # OpenHarnX report: BLOCKED
    | locked-tests | no | fail | checker_failed |
    | weakening | yes | fail | checker_failed: tests/test_pricing.py: 1 new skip or xfail |
    | no-new-failures-locked-tests | yes | fail | checker_failed: test_pricing::test_coupon_never_makes_total_negative passed at acceptance and fails; test_pricing::test_fixed_coupon passed at acceptance and fails |
    | no-new-failures-tests | yes | fail | checker_failed: tests.test_pricing::test_coupon_never_makes_total_negative passed at acceptance and is skipped |
    Saved to <tmp>/ohx-home/projects/bfa1688e035a0345/runs/6511fe99f687

4. The genuine change: percentage coupons, fixed coupons still work, new tests.
  $ pytest
    5 passed in 0.01s
  $ ohx verify --sandbox auto
    # OpenHarnX report: NO REGRESSIONS
    Saved to <tmp>/ohx-home/projects/bfa1688e035a0345/runs/137454fc0ffe

Plain pytest passed both times. OpenHarnX blocked the change that edited and
skipped locked tests, and passed the one that kept them.

Lock. Change. Verify. Review.

The tests that must keep passing are decided before the agent starts, and they run where the agent's code cannot touch them.

Lock, then change, then verify, then review. A blocked verification goes back to the change step with what failed.BLOCKED: back to the agent with what failedLOCKagreed tests copied awayCHANGEyour agent, as usualVERIFYlocked copies, in srtREVIEWbrief + evidence
  1. 1

    Lock

    The existing suite, or acceptance tests agreed for the task, are copied into OpenHarnX's store. A baseline records which tests passed.

    ohx init --lock-tests
  2. 2

    Change

    Your agent works as it normally does. Any agent: OpenHarnX reads files and test runs, not the agent.

  3. 3

    Verify

    The locked copies run in the srt sandbox, from an interpreter the change cannot replace. Per-test results are compared with the baseline and weakening is flagged.

    ohx verify --sandbox srt
  4. 4

    Review

    A verdict and a brief bound to the exact files judged. Blocked work goes back to the agent; passing work comes to you with what is left to decide.

    ohx report

A brief that says what it does not know.

Every report opens with five sections, built from records with no model involved. Facts, limits and the decisions left to you are kept apart, and every verified claim links to the check's output.

Exhibit · report.mdreal output of ohx report, demo repository, srt sandbox

# OpenHarnX report: BLOCKED exit 10

Requested outcome

Recorded fact

Locked suite at fc3f46e57e69 (suite): The test suite of the working tree at fc3f46e57e69, locked

Stated in contract revision urn:ohx:rev:468bc37a-573b-4eac-a4e7-bc830f374544, accepted by Shop Owner owner@example.com. No acceptance tests were agreed, so no check looks for it.

What changed

Recorded fact

Observed: 2 file(s) differ, by content digest, from the files accepted with the contract (base commit fc3f46e57e69). Why each changed is not recorded; an agent's own account of its change is not evidence.

FileChangeKindImported by a test run
pricing.pymodifiedcodeyes
tests/test_pricing.pymodifiedtest

Diff: git diff fc3f46e57e69, which also shows edits made before the contract was accepted (such as the acceptance tests).

What was verified

Recorded fact
  • tests (regression, advisory): passed (output: evidence/tests.txt)
  • Note: a passing check does not show that every changed line ran, and a requirement with a passing test is not thereby fully tested. The mutation check, where it ran, is the only measure here of how much the tests notice.

Each output link is a copy of the check's recorded output; report.json gives its digest and evidence record.

What remains unverified

Not verified
  • no-new-failures-locked-tests (mandatory): fail: test_pricing::test_coupon_never_makes_total_negative passed at acceptance and fails; test_pricing::test_fixed_coupon passed at acceptance and fails (output: evidence/no-new-failures-locked-tests.txt)
  • no-new-failures-tests (mandatory): fail: tests.test_pricing::test_coupon_never_makes_total_negative passed at acceptance and is skipped (output: evidence/no-new-failures-tests.txt)
  • weakening (mandatory): fail: tests/test_pricing.py: 1 new skip or xfail (output: evidence/weakening.txt)
  • locked-tests (advisory): fail: checker_failed (output: evidence/locked-tests.txt)
  • Whether the requested outcome is done: no acceptance tests were agreed, so no check looked for it
  • Limitation: Checkers ran with <tmp>/py/bin/python, outside the candidate identity and in no locked environment: its environment (3385 files in /Library/Frameworks/Python.framework/Versions/3.12/lib/python3.12, <tmp>/py/lib/python3.12/site-packages) was fingerprinted at acceptance and matched before the checks ran, so later changes there are caught; changes made before acceptance, compiled __pycache__ files, and code loaded from other folders are not; environment = "uv" in ohx.toml locks it

Decisions for you

Your judgment
  1. A mandatory check did not pass (weakening, no-new-failures-locked-tests, no-new-failures-tests). Send it back to the agent with the output linked above, or, if the expected behaviour itself is wrong, agree a new contract revision. A report cannot waive a failed check.
  2. Are the changed tests right? tests/test_pricing.py (modified). A test written with the change can only confirm what the change does.
Show the whole report.md (79 lines, including Details)
# OpenHarnX report: BLOCKED

## Requested outcome

**Locked suite at fc3f46e57e69** (suite): The test suite of the working tree at fc3f46e57e69, locked

Stated in contract revision `urn:ohx:rev:468bc37a-573b-4eac-a4e7-bc830f374544`, accepted by Shop Owner <owner@example.com>. No acceptance tests were agreed, so no check looks for it.

## What changed

Observed: 2 file(s) differ, by content digest, from the files accepted with the contract (base commit `fc3f46e57e69`). Why each changed is not recorded; an agent's own account of its change is not evidence.

| File | Change | Kind | Imported by a test run |
|---|---|---|---|
| `pricing.py` | modified | code | yes |
| `tests/test_pricing.py` | modified | test |  |

Diff: `git diff fc3f46e57e69`, which also shows edits made before the contract was accepted (such as the acceptance tests).

## What was verified

- `tests` (regression, advisory): passed ([output](evidence/tests.txt))
- Note: a passing check does not show that every changed line ran, and a requirement with a passing test is not thereby fully tested. The mutation check, where it ran, is the only measure here of how much the tests notice.

Each output link is a copy of the check's recorded output; report.json gives its digest and evidence record.

## What remains unverified

- `no-new-failures-locked-tests` (mandatory): fail: test_pricing::test_coupon_never_makes_total_negative passed at acceptance and fails; test_pricing::test_fixed_coupon passed at acceptance and fails ([output](evidence/no-new-failures-locked-tests.txt))
- `no-new-failures-tests` (mandatory): fail: tests.test_pricing::test_coupon_never_makes_total_negative passed at acceptance and is skipped ([output](evidence/no-new-failures-tests.txt))
- `weakening` (mandatory): fail: tests/test_pricing.py: 1 new skip or xfail ([output](evidence/weakening.txt))
- `locked-tests` (advisory): fail: checker_failed ([output](evidence/locked-tests.txt))
- Whether the requested outcome is done: no acceptance tests were agreed, so no check looked for it
- Limitation: Checkers ran with <tmp>/py/bin/python, outside the candidate identity and in no locked environment: its environment (3385 files in /Library/Frameworks/Python.framework/Versions/3.12/lib/python3.12, <tmp>/py/lib/python3.12/site-packages) was fingerprinted at acceptance and matched before the checks ran, so later changes there are caught; changes made before acceptance, compiled __pycache__ files, and code loaded from other folders are not; environment = "uv" in ohx.toml locks it

## Decisions for you

1. A mandatory check did not pass (`weakening`, `no-new-failures-locked-tests`, `no-new-failures-tests`). Send it back to the agent with the output linked above, or, if the expected behaviour itself is wrong, agree a new contract revision. A report cannot waive a failed check.
2. Are the changed tests right? `tests/test_pricing.py` (modified). A test written with the change can only confirm what the change does.

## Details

- Contract: Locked suite at fc3f46e57e69 (urn:ohx:rev:468bc37a-573b-4eac-a4e7-bc830f374544)
- Candidate: `sha256:3b51f5df2eb7bda8d85b8d700e0dac23a08ce35e4b0aaaef687863ea6f53ad73` (base commit fc3f46e57e69509525cef5e21e3fcc0feb305e51)
- Gate: **fail**, evidence coverage 20%
- Verifier protection: enforced: checkers ran in the srt sandbox
- Signature: not signed

## What this verdict supports

- No regressions: no
- Tests: 3 of 4 tests checked this change
- Acceptance criteria: none: no acceptance tests were agreed, so this does not show that the task is done; add them with `ohx contract new --acceptance`

## Obligations

| Obligation | Mandatory | Status | Reasons |
|---|---|---|---|
| locked-tests | no | fail | checker_failed |
| weakening | yes | fail | checker_failed: tests/test_pricing.py: 1 new skip or xfail |
| no-new-failures-locked-tests | yes | fail | checker_failed: test_pricing::test_coupon_never_makes_total_negative passed at acceptance and fails; test_pricing::test_fixed_coupon passed at acceptance and fails |
| no-new-failures-tests | yes | fail | checker_failed: tests.test_pricing::test_coupon_never_makes_total_negative passed at acceptance and is skipped |
| tests | no | pass | none |

## Cost

- Agent work: unknown: no agent run recorded
- OpenHarnX overhead: 0 model calls, verifier time 449 ms

## Changelog entry

- The test suite of the working tree at fc3f46e57e69, locked

## Limitations

- Checkers ran with <tmp>/py/bin/python, outside the candidate identity and in no locked environment: its environment (3385 files in /Library/Frameworks/Python.framework/Versions/3.12/lib/python3.12, <tmp>/py/lib/python3.12/site-packages) was fingerprinted at acceptance and matched before the checks ran, so later changes there are caught; changes made before acceptance, compiled __pycache__ files, and code loaded from other folders are not; environment = "uv" in ohx.toml locks it

This report authorizes nothing. It is evidence of readiness, not permission to merge, deploy or publish.

Keep the agent you use.

Verification depends on the files and the test runs, not on which agent wrote the code. Integrations differ, and this is what exists today.

Built-in hook

Claude Code

A Stop hook verifies when the agent says it is done. A blocked agent gets the failing tests and error lines back, up to three times; you get the verdict. Tested end to end.

ohx hook install

Use the Claude Code hook →

CLI

Codex, Cursor, OpenCode and others

Run ohx verify when the agent finishes and use its exit code. No built-in integration yet, and not tested with these agents yet.

ohx verify --sandbox srt

Use other agents →

CI gate

Any agent that opens pull requests

The base branch's tests run against the change and the base's policy decides. A GitHub Action, with GitLab CI and Jenkins examples. GitHub Copilot's merged pull requests were checked in replays.

Add the CI gate →

OpenHarnX checks one agent's work. It is not an agent, and it does not plan, delegate or orchestrate agents.

Evidence, with its denominators.

Replays of public agent work, published with what OpenHarnX missed, what it blocked that was legitimate and where the replay itself failed. Each figure links to its record.

Replay · T89

Impossible tasks

DimitrovK/impossible-tasks: 102 runs of five models on 8 impossible and 2 solvable projects. Record

  • 43/52fake fixes caught with zero setup, of the 80 runs that could be replayed
  • 26/26fakes that changed the tests were caught
  • 17/26fakes that changed only the code were caught; 5 of 11 on the held-out half
  • 6/6genuine fixes passed

Exclusions and limits

  • 22 of 102 runs were excluded: they add files the dataset does not record.
  • 9 fakes passed: 6 leave a test failing as before, which zero setup reads as nothing regressed; 3 change behaviour only a specification could tell apart.
  • The checks for code-only tricks were written from half of the runs.
  • 8 legitimate test edits were blocked until a person accepted them.
Replay · T90, T92

Merged agent pull requests

Judged after the fact by the gate. T90 · T92

  • 9/20merged Copilot pull requests in github/spec-kit passed; 7 more passed once their intended test changes were approved
  • 28/31agent commits in anthropics/claude-agent-sdk-python passed, most of them changelog updates; 3 were intended test changes

Exclusions and limits

  • 4 of the 20 spec-kit pull requests failed because of the replay's own environment.
  • These are changes people had already merged. They show how often legitimate work passes or needs approval, not what reviewers would have caught.
Dogfooding · T100

Its own development

Every change verified by an earlier installed version. Record

  • READYunder srt before each commit; the commit names the contract and the digest judged
  • 4defects the CI gate found in the 0.1.0 release candidate that local runs had missed

Limits

  • One project, one maintainer. Agreement with its own checks is not independent validation.
No production users yet.Tested on its own development, replays of public agent work and simulated users.
No measured time savings.Whether the brief saves review time is not measured. A study with reviewers is on the roadmap.
No guarantee.A passing verdict is evidence about the checks that ran. It does not make a change safe or authorize a merge.

Try OpenHarnX on one change. Tell us what it caught and where it got in your way.

Esc
Try verify, STALE, approve-tests or GitLab. Common pages: