Skip to content

Verification engine

Done means checked by someone else.

How a different vendor checks every agent task before it counts as done.

One agent works. A different vendor checks.

An agent that marks its own work tends to agree with itself. In OmniConflux, every task gets a verifier from a different vendor, plus static checks and a sandbox run that do not depend on any model at all. Here is the loop, step by step.

  1. Step 1

    Write the check first

    Before work starts, the task and its acceptance check are written down: what done means, which tests must pass, which files may change.

  2. Step 2

    Work in isolation

    The specialist agent works in its own git worktree on its own branch. Nothing touches your main checkout while the work is under review.

  3. Step 3

    Static checks

    The diff goes through type checks, lint, syntax tree analysis and static security analysis. Each result is recorded as evidence, pass or fail.

  4. Step 4

    Sandbox run

    The build and the tests run in an isolated sandbox, not on your working copy. The verifier reads the real output, not the agent's summary of it.

  5. Step 5

    A different vendor reviews

    A verifier from a different model vendor reads the task, the diff and the evidence, then returns approve or rework with reasons. A second vendor family brings different blind spots, which is the point.

  6. Step 6

    Consensus on high-impact changes

    For changes you mark as high impact, more than one independent reviewer must agree. Disagreement goes to a person.

  7. Step 7

    Rework, then a decision

    A rework verdict sends the work back with the reasons attached. After two rework rounds, a deciding review settles it, and if that is not enough, the decision comes to you.

  8. Step 8

    Keep the record

    Every verdict is stored with its evidence: the checks that ran, the sandbox output and the reviewer's reasons. You can read why any change counted as done.

The same loop, for defensive security

For authorized, defensive engagements, the verification engine is what keeps findings honest:

  • Deterministic Vulnerability Verification: a finding is recorded only after it is reproduced in a sandbox, so reports carry evidence instead of guesses.
  • Continuous Patch and Remediation Testing: each fix is re-tested by the same check that found the issue before it is marked resolved.
  • Security Control Efficacy Testing: agreed checks confirm whether your existing controls behave as configured.
  • Continuous Adversarial Exposure Validation: scheduled, in-scope checks re-run as your systems change.

Every engagement needs written authorization and a defined scope before any work begins.

Read further

See how agents get their own worktrees in the fleet architecture guide, and what data leaves your machines on the trust page.

Questions

Why use a different vendor as the verifier?
Models from the same vendor tend to share training and blind spots. A reviewer from a different vendor family is more likely to question what the first one took for granted.
Does verification mean every result is correct?
No. Verification is an independent check with recorded evidence, not a promise. It makes weak work visible early and sends it back with reasons.
Where does the review run?
Static checks and the sandbox run on your fleet. If the verifier is a local model, the whole review stays on your hardware.
How does this relate to security work?
The same loop backs authorized, defensive security work: a finding counts only after it is reproduced in a sandbox, and a fix counts only after it is re-tested.