Method

What Certainty Labs builds, and how.

01

What a governed system requires.

Governed systems: software in which a model may inform a decision and never owns it. Four properties we build toward; each is shown where it is implemented today.

Evidence before a verdict

A model can read, summarise and propose. It does not get to conclude. A verdict is issued by a deterministic runtime from evidence and counter-evidence that a person can inspect afterwards.

Where it is implemented today
Built

A deterministic runtime issues the verdict

Each account whose research completes receives PASS, RESEARCH_MORE or KILL, written as an immutable decision record with its evidence. AI gathers evidence; it never decides. When research fails or a claim cannot be verified, the outcome is never a PASS or a KILL.

Source: A deterministic runtime issues the verdict

Uncertainty is an outcome

When the evidence does not carry the decision, the system says so and names what is missing. Holding is a result, not an error to be smoothed over with a guess.

Where it is implemented today
Built

An MCP server with four procurement decision tools

A stateless MCP server exposes four tools: analysing a public tender notice, evaluating a commercial policy, calculating landed cost and comparing bill-of-material options. Each returns a bounded outcome: a tender analysis ends in PASS, RESEARCH_MORE or REJECT, a policy evaluation in ALLOW, DENY or ESCALATE, and a landed cost is labelled verified, estimated or unknown.

Source: An MCP server with four procurement decision tools

Authority is granted, bounded and separate

What may act is decided apart from what is clever enough to propose acting. Authority is granted by a person, scoped to named effects, and some effects cannot be authorized at all.

Where it is implemented today
Built

Consequence limits what can be authorized at all

The internal authority contract defines nine classes of effect. Two can be authorized: read-only external access and internal durable writes. The other seven, including external communication, financial transactions and production deployment, are denied whoever asks.

Source: Consequence limits what can be authorized at all

Action is gated, recorded and observed

A consequential action needs a recorded approval before it can be carried out. Observing the outcome and feeding it into the next evaluation completes the chain; that part is not yet implemented in the example shown.

Where it is implemented today
Built

Twelve classes of action sit behind a recorded human approval

TerminalCor3's system of record for property acquisition and development separates preparing an action, approving it and executing it. Twelve classes of consequential action need a recorded human approval: paid registry research, paid technical due diligence, an outside adviser engagement, owner outreach, a broker engagement, lender outreach, investor or joint-venture outreach, an indicative offer, a letter of intent, exclusivity, an acquisition commitment and a capital commitment. None is executed automatically: this version has no executor.

Source: Twelve classes of action sit behind a recorded human approval

02

Operating method.

We do not begin with a platform and look for a use. We begin with a problem we operate ourselves.

  1. BUILD INTERNAL

    We build the system for a problem we have ourselves, and we are its first operator.

  2. PROVE IN REAL USE

    We run it against real problems until its limits are known: what it decides well, where it must hold, what it must never do.

  3. EXTRACT A CLEAN BOUNDARY

    Reusable infrastructure is extracted only when a boundary has been justified by use, never before.

  1. build systems internally
  2. operate them against real problems
  3. prove their boundaries
  4. extract reusable infrastructure only when the boundary is justified

The authority, review and deployment machinery described on this page is internal infrastructure. It is shown as evidence of method. It is not offered as a product.

03

Engineering discipline.

Seven steps, in this order: the method for a change that can reach production. A step says what the method requires; where a step has a recorded instance, the record follows it, with its limits.

  1. mission

    Work starts from a stated mission and its limits: what must be true when it is done, and what it must not touch.

  2. architecture reconciliation

    Before any code, the repository, the deployed state and the ownership of each part are read again and reconciled. What the repository shows wins over what anyone remembers.

  3. agent implementation

    AI agents write the implementation in an isolated branch, under the repository's own rules. Being able to write the change gives an agent no authority to ship it.

  4. independent adversarial review

    A reviewer with no part in the change tries to break it. Findings are classified by consequence and closed in code, not argued away.

    Built

    A recorded independent review, tied to one commit, and where its record ends

    One recorded example: a reviewer separate from the author reviewed one exact commit and returned changes required; the findings were closed in new commits, and the new head was reviewed again. That re-review left one wording finding, and the commit that closed it was merged after the re-review. The repository's record ends at the re-review: it holds no review of that last commit.

    Source: A recorded independent review, tied to one commit, and where its record ends

  5. mutation proof

    Each guard is disabled in turn. If the tests still pass, the guard was never proven, and the missing test is the finding.

    Built

    Guards are proven by mutation

    In one recorded check, 97 guards were each disabled or weakened in turn and the test suite was run against every mutant: 96 were killed and 1 survived, and the survivor is recorded with the result.

    Source: Guards are proven by mutation

  6. exact-head verification

    A verdict belongs to one commit. A new commit, however small, is a new head and is verified again.

  7. governed production gate

    Reaching production is its own recorded decision with its own evidence. A build that passes is not a release.

    Built

    Consequence limits what can be authorized at all

    The internal authority contract defines nine classes of effect. Two can be authorized: read-only external access and internal durable writes. The other seven, including external communication, financial transactions and production deployment, are denied whoever asks.

    Source: Consequence limits what can be authorized at all

    Proving

    The internal authority store runs deployed, in shadow

    The internal authority store has run on a deployed database over verified TLS, in shadow mode. Restart, forced-kill recovery, writer takeover, halt and revocation drills passed there. Shadow means no live external effect is authorized, and the gate closed as a conditional pass.

    Source: The internal authority store runs deployed, in shadow