Decision infrastructure

A direct answer, first

When should an AI system abstain?

An AI system should abstain — decline to act or recommend, and instead report that more evidence, review, or authority is required — when the evidence, authority, policy fit, or uncertainty behind a proposed action is insufficient to support it. There is no single, universal, scientifically settled threshold for when that line is crossed; it depends on the consequence of the specific decision and the policy governing it.

01 / When abstention is appropriate

Four conditions that can each justify abstaining

Insufficient evidence, unresolved conflict, high uncertainty, or missing authority.

  1. 01

    Evidence sufficiency

    The available evidence does not meet the bar required to justify the specific action being proposed, even if it supports a weaker or more limited action.

  2. 02

    Unresolved contradictions

    Two or more pieces of evidence conflict and the conflict has not been resolved — proceeding would mean silently picking a side.

  3. 03

    Uncertainty

    The system's own estimate of confidence is too low, or too poorly calibrated, to support the consequence of acting.

  4. 04

    Authority boundaries

    The action falls outside what the system — or the human operating it in that moment — is authorized to approve.

02 / Abstention is an outcome, not a failure

"More evidence required" is a valid result

A dependable system must be able to say it doesn't yet know.

Treating "more evidence required" as a legitimate, first-class outcome — rather than an error to be papered over with a confident-sounding answer — is one of the design principles behind Certainty Labs' architecture. In Certainty Labs' applied systems, a deterministic policy can return an explicit abstain outcome alongside build, iterate, or kill, rather than forcing a decision when the evidence does not support one.

03 / Preserving the abstention

The record should show why the system stopped

An abstention belongs in the decision record, not just the logs.

If a system abstains but nothing durable records why — which evidence was missing, which contradiction was unresolved, which authority was absent — the abstention is invisible to later review. Preserving that reasoning alongside the evidence, claims, and eventual outcome is what makes an abstention auditable rather than just silent.

04 / Open questions

Not yet settled

What remains an open research question.

  • What uncertainty threshold is appropriate for a given decision, and does that threshold generalize across domains, or does it have to be set per decision class?
  • How should an abstention be distinguished from a system simply failing or timing out?
  • How does a system remain useful in practice if it abstains too often, without loosening the thresholds that make abstention meaningful?

These questions are part of Certainty Labs' research and are tested in bounded applied systems.

05 / Collaboration

Working on abstention and uncertainty yourself?

We're interested in hearing from teams and researchers working on uncertainty, calibration, and abstention in AI-supported systems.

Collaborate with Certainty Labs