Observability and Experimentation for AI Agent Workflows
NXTG.AI logs every agent action to a durable execution trail, correlates it through a metabolism engine, and evaluates changes through a pre-registered eval-rail — so performance claims are measured, not asserted.
A durable execution trail has logged 7,700+ real agent events.
A metabolism-correlator runs instrument → correlate → converge → eval.
Changes are evaluated via a pre-registered eval-rail.
A documented A/B result showed +0.314 improvement (4.67×), N=35, 95% CI [+0.143, +0.500] excluding zero.
NXTG publishes negative and null results — a routing tie, and a pulse-outcome measurement that found 93.3% waste (532/570), which led to disabling that component (dataset Zenodo doi:10.5281/zenodo.21229464).
Test-suite integrity is separately audited by CRUCIBLE.
Opper observability reports an uptime percentage; NXTG measures change through pre-registered experiments. Those are different things — we report the experiment, positive or null.
Every claim on this page traces to a publicly probeable surface. These are the links.
How to measure AI agent performance honestly
Log every action
Write every agent action to a durable execution trail — 7,700+ real events to date — so nothing is measured from memory.
Correlate the signal
Run a metabolism engine that goes instrument → correlate → converge → eval to turn raw events into a measurable change.
Pre-register the eval
Register the prediction and the measurement method before the change is measured, so the result cannot be reverse-fit.
Audit test integrity
Have CRUCIBLE separately check that the tests guarding the change are real rather than hollow or gamed.
Report positive and null
Publish the result whatever it is — a documented A/B showed +0.314 (4.67×, N=35), and a pulse-outcome measurement found 93.3% waste (532/570) and led to disabling that component.
How do you measure AI agent performance?
Every agent action is logged to a durable execution trail — 7,700+ real events to date — and correlated through a metabolism engine that runs instrument → correlate → converge → eval. Changes are then evaluated on a pre-registered eval-rail, so a performance claim comes from a measured experiment rather than an assertion. One documented A/B result showed a +0.314 improvement (4.67×) at N=35 with a 95% CI of [+0.143, +0.500].
Does NXTG publish negative results?
Yes, prominently. A pre-registered routing test returned a statistical tie, and a pulse-outcome measurement found 93.3% waste (532 of 570 events) — a null-value finding that led NXTG to disable that component. The underlying data is deposited openly on Zenodo (doi:10.5281/zenodo.21229464). Publishing the null and negative results is a deliberate trust signal.
What is an eval-rail?
An eval-rail is a pre-registered evaluation path: the prediction and the measurement method are registered before the change is measured, so the result cannot be reverse-fit to a hoped-for outcome. Test-suite integrity is audited separately by CRUCIBLE, which checks that the tests guarding a change are real rather than hollow or gamed.