Skip to main content
← nxtg.ai
/platformPLATFORM

Observability and Experimentation for AI Agent Workflows

NXTG.AI logs every agent action to a durable execution trail, correlates it through a metabolism engine, and evaluates changes through a pre-registered eval-rail — so performance claims are measured, not asserted.

WHAT IT DOES

A durable execution trail has logged 7,700+ real agent events.

A metabolism-correlator runs instrument → correlate → converge → eval.

Changes are evaluated via a pre-registered eval-rail.

A documented A/B result showed +0.314 improvement (4.67×), N=35, 95% CI [+0.143, +0.500] excluding zero.

NXTG publishes negative and null results — a routing tie, and a pulse-outcome measurement that found 93.3% waste (532/570), which led to disabling that component (dataset Zenodo doi:10.5281/zenodo.21229464).

Test-suite integrity is separately audited by CRUCIBLE.

HONEST COMPARISON

Opper observability reports an uptime percentage; NXTG measures change through pre-registered experiments. Those are different things — we report the experiment, positive or null.

RECEIPTS

Every claim on this page traces to a publicly probeable surface. These are the links.

the origin reportmeasured, pre-registered results
10.5281/zenodo.21229464570-row OnePulse corpus (the 93.3%-waste dataset)
OSF osf.io/8mh2xpre-registered predictions
Faultlineclaim forensics
HOW IT WORKS

How to measure AI agent performance honestly

  1. Log every action

    Write every agent action to a durable execution trail — 7,700+ real events to date — so nothing is measured from memory.

  2. Correlate the signal

    Run a metabolism engine that goes instrument → correlate → converge → eval to turn raw events into a measurable change.

  3. Pre-register the eval

    Register the prediction and the measurement method before the change is measured, so the result cannot be reverse-fit.

  4. Audit test integrity

    Have CRUCIBLE separately check that the tests guarding the change are real rather than hollow or gamed.

  5. Report positive and null

    Publish the result whatever it is — a documented A/B showed +0.314 (4.67×, N=35), and a pulse-outcome measurement found 93.3% waste (532/570) and led to disabling that component.

FREQUENTLY ASKED

How do you measure AI agent performance?

Every agent action is logged to a durable execution trail — 7,700+ real events to date — and correlated through a metabolism engine that runs instrument → correlate → converge → eval. Changes are then evaluated on a pre-registered eval-rail, so a performance claim comes from a measured experiment rather than an assertion. One documented A/B result showed a +0.314 improvement (4.67×) at N=35 with a 95% CI of [+0.143, +0.500].

Does NXTG publish negative results?

Yes, prominently. A pre-registered routing test returned a statistical tie, and a pulse-outcome measurement found 93.3% waste (532 of 570 events) — a null-value finding that led NXTG to disable that component. The underlying data is deposited openly on Zenodo (doi:10.5281/zenodo.21229464). Publishing the null and negative results is a deliberate trust signal.

What is an eval-rail?

An eval-rail is a pre-registered evaluation path: the prediction and the measurement method are registered before the change is measured, so the result cannot be reverse-fit to a hoped-for outcome. Test-suite integrity is audited separately by CRUCIBLE, which checks that the tests guarding a change are real rather than hollow or gamed.