NOT LIVE Nothing on this page is a running service. Every capability below is planned.

A place where a claim
arrives with its evidence.

ClaimGarden is the planned hub for the Verifier Standard: somewhere a verifier, the claim it checks, the environment it ran in, and the benchmark it was measured against can be found together instead of scattered across four systems that disagree.

Four things, kept together

The split below is the whole idea. Each of these is normally somebody else's problem, which is why a claim so rarely survives contact with the environment that produced it.

01 PLANNED

Verifiers

A registry of checking mechanisms — SAT and SMT solvers, type checkers, provenance validators — each declaring what it can decide, what it cannot, and what it costs to run.

A verifier that will not state its own limits is not a verifier.

02 PLANNED

Claims

The propositions themselves, bound to specific artifact bytes rather than to descriptions of them, carrying their outcome and the bounds under which that outcome holds.

Including the ones that came back UNKNOWN. Especially those.

03 PLANNED

Environments

The declared conditions a result was produced under — interpreter, dependencies, hardware, and the exclusions somebody decided not to mention.

Most irreproducibility is an environment nobody wrote down.

04 PLANNED

Benchmarks

Task families with held-out sets and independently constructed counterexamples, so that a score means something more specific than a number that went up.

A benchmark you can overfit is a benchmark you will overfit.

What exists today, and what does not

This distinction is the product. It would be strange to blur it on the homepage.

OBSERVED

The Verifier Standard and its reference implementation

Public, open, and installable today as verifier-standard on the Python Package Index. Version 1.3.0 was published on 8 September 2026.

PLANNED

The ClaimGarden hub itself

No registry, no query service, no evaluation environment, and no accounts. This page serves static text and nothing else.

UNKNOWN

Coverage, adoption, and independent evaluation

How wide a task family this approach usefully covers has not been measured. No third-party evaluation, customer, or certification is claimed. Those are open questions, and a placeholder page is not the place to quietly answer them.

Following along

There is no mailing list, because there is nothing yet to mail about. The work happens in the open and the repository is the honest place to watch it.