Skip to main content
Tag

AI reliability

2 articles tagged with “AI reliability

2
Articles
ArticleResearch10 min

Quality Is a Harness Property, Not a Model Property

One night our agent fleet produced at least 9 wrong or almost-wrong claims the harness caught before they shipped — and at least one reached the running server before a second layer of defense masked the impact. An honest account of the harness — probes at the authority source, verifiers separate from generators, hooks that block, typed state over prose — making individually fallible agents collectively trustworthy, with the line drawn exactly where the evidence stops.

agent harnessmulti-agent systemsAI reliability+3
Asif Waliuddin
Aug 9, 2026
Read more
ArticleResearch8 min

A Governance-Catch Census: What Our Agent Harness Caught in One Day — and What Got Through First

A single-day census of what our production multi-agent harness caught: at least 9 wrong or almost-wrong claims caught before they shipped, and at least one that reached a running server before it was caught. Disclosed counting rule, per-item record including the miss, and an honest statement of what the numbers do and do not support.

agent harnessmulti-agent systemsAI reliability+2
NXTG.AI
Jul 27, 2026
Read more