Skip to main content
Insights•14 min read

Faultline: Four Days on Kaggle, Nine Months to Make a Verifier Tell the Truth

September 30, 2026 by Asif Waliuddin

AIAI verificationdeveloper tools
Faultline: Four Days on Kaggle, Nine Months to Make a Verifier Tell the Truth

On December 12, 2025, I posted a Kaggle hackathon entry for an AI claim checker that had zero tests. Nine months later Faultline Pro has 4,887 counted tests, and several of the ones that matter most exist because the product failed me first.

The first commit on Faultline landed on December 9, 2025. On December 12 I posted the writeup, "Faultline: Seismic Stress-Testing for AI Hallucinations", to the Overall Track of Google DeepMind - Vibe Code with Gemini 3 Pro in AI Studio, a Kaggle hackathon whose page invited builders to compete for $500,000 in credits. Empty repo to submission took four days, December 9 through December 12.

The idea fit in one line of the writeup: "Faultline is a seismic stress-test for information integrity." The tagline was "Structural engineering for the generative age." The risk I cared about was AI text that reads as true, the kind nobody thinks to question because nothing about it sounds off.

What I built in those four days was a React and Vite web app on Google's @google/genai SDK, calling gemini-3-pro-preview. It worked in three passes. The first asked Gemini to break a piece of text into atomic claims, each typed as fact, opinion, or interpretation, with an importance score from 1 to 5. The second took the load-bearing facts (importance 3 or higher) and verified each one against live Google Search grounding, keeping up to three sources per claim. The third wrote a critique of the text plus a "reinforcement prompt" you could use to regenerate it on firmer ground. Verdicts rolled up into a stability score.

One wrinkle I still remember: the Gemini API would not combine structured JSON output with the Google Search tool, so the verification pass asked for JSON in the prompt and recovered it from free text with a hand-written cleaner that stripped markdown fences and hunted for braces. When recovery failed, the claim fell back to "mixed." It was brittle, and I knew it.

It also had zero tests.

I count Faultline among the most technically challenging builds in my portfolio. That is my own judgment, and I have no metric behind it. I kept going because I believe the community needs a tool that takes what an AI told you and shows, claim by claim, what holds up. The original repo is still public. Its tag, kaggle-demo-v1, is dated December 16, four days after I posted the writeup, and marks a badge update. The submission was the December 12 writeup; the tag is a later snapshot.

Two quiet months, then a test suite

After that tag the repo went silent. There are no commits between December 16 and February 16. On February 20 I froze the Kaggle entry and decided to split a product out of it.

The rebuild moved fast once it started. On February 22 the first 73 unit tests and a CI pipeline landed, along with a provider abstraction so the engine could run on Gemini, Claude, or OpenAI. A CLI followed within days. By February 24 a commit subject recorded 868 tests.

On March 3 I promoted that branch into its own repository, Faultline Pro. The two share their first 48 commits. There was no rewrite at the split: the provider rewrite had already happened inside the original repo in February. On March 8 I published @nxtg/faultline 0.1.0 to npm under the Apache-2.0 license.

Two out of ten

Two days before that publish, on March 6, I sat down and used the CLI the way a new user would. I followed the Quick Start.

The Quick Start ran --provider mock. Every claim I fed it came back "Mock verification: supported" at 0.30 confidence.

I scored it 2/10 and called it a no-go. I spent about 45 minutes working out what the tool was supposed to do. All 868 tests were green, and CI was green. None of them asked the question I was asking, which was what a stranger sees first.

On a verification product, a mock verdict in the first-run path is a fabricated verdict. The tool exists to tell you which claims are supported, and its own front door told me everything was.

The engine had an early defect of its own. Claim extraction sometimes merged two separate sentences into a single claim, so one statement could ride through verification attached to another. The March 20 fix is a deterministic floor: after the model extracts claims, code checks every sentence against them using a normalized 40-character fingerprint and adds a claim for any sentence left uncovered. The model proposes the claims. Code guarantees that every sentence gets one.

The web app scored 42

Also on March 6, I started the web app by copying the Pro engine into it read-only. It became a Next.js app with Clerk for sign-in, Stripe for billing, and Vercel KV for storage, and it later moved to calling the deployed Pro API directly. It was live at faultline.nxtg.ai by March 31. That day the first design evaluation scored it 42/100 and blocked it.

The site's early commit log reads like a confession. On March 20 a commit removed a product mockup and fake data from the homepage. On March 29 another removed a fake testimonials section. Both had been live on a product whose entire job is telling you what is real.

April brought a launch gate. Before a planned Show HN on April 20, I required p95 scan latency under 20 seconds. A synthetic user walk-through measured 72.8 seconds end to end. Verifying claims concurrently, eight at a time with a 200-token cap on each verdict, brought it to 7.9 seconds measured on production. The same day I released faultline-action v1.0.0, a GitHub Action that installs the published CLI from npm and scans files in CI. That design choice paid off in July.

May 31: the engine was real

On May 31 I had two independent reviews run against the whole thing, both adversarial on purpose. Their verdict split cleanly. The verification engine was real: it extracted claims, searched for evidence, and judged. The product wrapped around it was not yet real.

The specifics were hard to read. The homepage showed "0 Providers" while the changelog quoted a test count that belonged to my whole portfolio. The GitHub link returned a 404. A risk tier was computed by regex. A paid tier that had been drafted did not exist anywhere in the billing code. The same day, one commit pulled 8 public falsehoods off the site.

The positioning had wandered. In its first seven months Faultline had been "Seismic Stress-Testing for AI Hallucinations," an "AI Trust & Safety Platform," "AI Claim Forensics," "the only AI safety CLI not owned by an AI lab," and "Agent Governance for AI Outputs." On June 2 I approved a smaller line that came out of the reviews, and it is still the hero on the site: "We check the receipts on AI output." Later in June a SOC2 badge that implied a certification the product had not earned came off the site.

Two days after the reviews I was in the engine's error handling, where I found a quieter lie. When a provider returned a 429 or a 503, the engine swallowed the error and reported the claim as "unverified." So "unverified" sometimes meant "never checked." The June 2 fix gave each result an explicit API-error field and each scan a degraded flag. On June 13 I closed the next hole in the chain: a fully degraded scan used to exit with code 0, which in CI means pass, which a pipeline reads as safe to publish. The --fail-on gate now fails closed when a scan is degraded.

June 21: zero sources

On June 21 I found that the default path returned no evidence at all. The live streaming scan endpoint defaulted to the OpenAI provider, and that provider, like the Claude one, had sources: [] hard-coded. The demo cards on the site rendered real sources, so anyone who clicked a demo saw grounding. Anyone who pasted their own text got verdicts with nothing behind them, on a product that said it checked live sources.

On a three-claim scan, OpenAI took about 4.8 seconds and returned 0 grounded sources. Gemini took about 14.0 seconds and returned 3. The fast path was the empty one.

That same day I wrote a code-grounded audit of the engine against what I had been saying about it. It found a two-stage claim-verification pipeline using a single provider per scan by default. The adversarial five-provider engine I had been describing did not exist yet.

I had two honest options: shrink the claim to match the code, or build the code up to the claim. I decided to build the full vision, a grounded, multi-provider, cross-model consensus engine. That same day the default path switched to grounded Gemini and the first consensus engine landed. Several providers judge the same retrieved sources. The plurality verdict wins. A real tie between statuses becomes "mixed." A provider that errors out casts an "unavailable" vote and never aborts the scan. It shipped as an opt-in mode in v0.9.0 on June 26.

Six days of a dead package

Version 0.9.1 went out to npm in late July. Its files list left out two directories that its own code imports, consensus/ and governance/. Every command died with ERR_MODULE_NOT_FOUND. The release was broken for six days.

Every test passed for all six of them. The tests ran against the working tree. The version-parity gate compared manifests. Nothing installed the published tarball and ran it.

The GitHub Action caught it, because the Action installs the published package from npm the way a user would. The July 28 fix added a packaging test that asserts against the published file list, and the Action now proves the install works before it scans anything. The same day surfaced a sibling bug: the version check regex /^0\.[4-9]\.\d+$/ rejected 0.10.0, because [4-9] matches one digit.

The Action got its own reckoning that day. Its v1.0.0 CI was a mock scan of the README that could not fail, and hardening it showed that Action inputs reached the shell unescaped. They now pass through environment variables, and CI runs an injection canary. The fix shipped in v1.1.0.

Chocolate cures cancer

On July 27 I made a call on the agent-governance claim, which had been quietly narrowed a week earlier. My answer was "Built to earn it!" The claim stays, and the product has to grow into it. The next day faultline guard shipped: pipe an agent's output into it, and it gates on claims that come back refuted or unsupported.

July 28 was a long day. The hosted verifier behind the web app had died without a sound: the Gemini billing account it ran on had closed, and nothing raised an alarm. I fixed the billing. A live probe of guard afterward checked 3 claims, verified 2, refuted 1, and attached evidence URLs to all of them.

Then came the worst one. In the CLI as it stood that day, with no API key set, it silently fell back to the mock, and the mock returns "supported" for every claim. The fix commit records the reproduction:

echo "Chocolate cures cancer and the Earth is flat." | faultline guard

The output said [OK] VERIFIED and Overall Risk: LOW.

The commit message describes it as a confident, fabricated verdict from the tool that exists to catch confident, fabricated verdicts. I read it as the March 6 failure again, four months later, through a different door. In March the docs told new users to run the mock. In July the code reached for the mock on its own, quietly, when a key was missing. The fix separates an implicit mock from an explicit one, and a dedicated test locks the behavior in. It shipped as 0.10.1 on July 29.

The silence is gone. With no key, faultline guard now reports that no provider key was found, prints "NOTHING WAS CHECKED", and says it "will not invent verdicts without one." The mock runs only when you ask for it with --provider mock, and even then it announces that it "returns SYNTHETIC results and verifies nothing" and is meant for CI wiring only.

The web app had its own July. Claims are paraphrases, so exact-string highlighting matched 0 of 8 claims on a recorded real scan; a three-layer anchor brought that to 6 of 6. And on July 29 I found that the free-tier meter counted scan history, which a clear-history button empties. The free tier had been unlimited.

What Faultline does now

You install it with npm install -g @nxtg/faultline. The latest version is 0.10.1, one of 20 published, under Apache-2.0. It runs in two modes. In local mode you bring your own provider key and the scan runs on your machine. In hosted mode you use a Faultline API key and the scan runs on the Faultline API; that is the paid path. With neither key, it checks nothing and tells you so. Guard is advisory by default and exits 0 in that case; add --fail-on and it exits 1. The live verification providers are Gemini, OpenAI, Claude, and Perplexity, and an explicit mock exists for wiring up CI.

The engine takes a block of text, extracts atomic claims, and guarantees at least one claim per sentence. It verifies up to eight load-bearing fact claims (importance 2 or higher now, down from 3 in the Kaggle version), eight at a time. Each claim comes back supported, contradicted, mixed, or unverified, and a scan that could not reach its providers says it is degraded. The risk score is computed by deterministic rules with no model involved. Output goes to JSON, Markdown, HTML, or SARIF, and --fail-on turns the exit code into a CI gate. Consensus verification is an opt-in mode.

The GitHub Action wraps the same CLI in two modes: scan a file into SARIF for code scanning, or guard a block of text such as a pull request body. The web app at faultline.nxtg.ai takes pasted AI output and shows you, claim by claim, what it could and could not verify.

Across the four codebases, the original Faultline repo has 210 commits, Faultline Pro has 1,220 (48 of them shared with the original), the web app has 323 commits on its main branch, and the Action has 5. I counted tests as of September 30 by having the test runner collect them, without running the suites: 893 in the original, 4,887 in Pro, and 1,039 unit tests on the web app's current development branch, plus 15 browser end-to-end tests. Collection proves the tests exist. It does not prove they pass today, and after this year I am not going to blur that line in my own writeup.

What I took from it

The first-run path is the product. Twice, in March and again in July, a new user could have met a mock verdict before anything real. In March, 868 passing tests did not catch it. My own first session did.

Test the thing people install. For six days, 0.9.1 was dead on arrival while every check I owned stayed green. The only check that saw it was the one that installed from npm.

Give "never checked" its own name. Until June, a rate-limit error and a genuinely inconclusive claim produced the same word, and a fully degraded scan passed the CI gate.

Test the default path separately from the demo path. The demo cards showed sources while the default endpoint returned none.

Watch dependencies that can die quietly. A closed billing account took the hosted verifier down, and nothing told me.

When the claim outruns the code, either retract it in the open or build up to it. On May 31 the right move was pulling 8 falsehoods off the site. On June 21 and July 27 it was building. The one move I refuse now is narrowing a claim quietly and hoping nobody compares.

Pipe that chocolate sentence into 0.10.1 with no key set and it refuses. It prints NOTHING WAS CHECKED and says it will not invent verdicts without a provider key. For a verification tool, that refusal is the correct answer.

Ready to build?

Ship AI you can trust

Forge gives you agents, governance, and verification — so your AI ships with confidence, not hope.

Newsletter

Enjoyed this article?

Get more insights like this delivered straight to your inbox.

Email subscription coming soon. Follow along on LinkedIn in the meantime.

Follow on LinkedIn