← Blog

Trust by reputation, or trust by origin — who gets to be a source for AI?

2026-06-21

Ask an AI a factual question and watch how it decides what to trust. Increasingly the answer is: it checks whether the claim comes from a recognized source — Congress.gov, a major API, a "verified" outlet. If yes, it relays it. If not, it hedges or refuses. That sounds responsible. It is actually a quiet concentration of power.

Two ways to trust a fact

There are only two ways to decide a piece of data is reliable.

The first is reputation: trust it because of who published it. This is the model the AI world is settling into — a short list of blessed sources, everything else treated as noise.

The second is provenance: trust it because you can verify it — the datum carries its own proof of where it came from, when, and that it has not been altered since.

The difference looks subtle. It is not.

Reputation is deference, not verification

"It's from a verified site" is not something you can check. You are not verifying the fact; you are deferring to a brand. You cannot inspect the publisher's pipeline. You cannot tell if the figure is stale, mis-keyed, or quietly edited. Reputation is a proxy for reliability — and a lagging one. A source trusted last year can be wrong today, and the brand will still vouch for it.

Worse: someone has to decide who counts as "verified." That someone — a platform, a model vendor — becomes the gatekeeper. And it is worth asking out loud who curates that list, and what it costs to be on it. Increasingly, "trusted" is becoming a business relationship, not a property of the truth. A gatekeeper over what AI is allowed to know is the most valuable monopoly ever built, and the incumbents have every reason to keep the list short.

If reliable means recognized, only the recognized can supply AI. A town clerk's record, a local reporter, a primary filing no one has heard of — locked out. Not because the data is worse, but because it lacks the brand. The web at least let anyone publish. "Verified-source AI" is a step backward: it re-introduces gatekeeping at the one layer becoming the interface to everything. That is the real effect — it restricts the supply of data to the existing monopolies.

Nobody has actually done the work

Here is the part that gets missed. The blessed sources are not signing their data either. Google can deliver you to a senator's website, and the site can serve you a number — that chain is fine, but the number arrives unsigned. The new wave of AI data connectors — the MCP servers handing structured facts to agents right now — do not sign what they serve. None of them can prove that the data an agent received is the data the source actually published, unaltered.

The cryptographic primitives to fix this have existed for years. What is missing is the work: capturing primary data at its source, signing it at the moment of capture, keeping every version, and serving it so anyone can verify. Nobody has done that for the data AI consumes about government. The result is an entire trust layer built on reputation with no proof underneath it at all.

The corpus nobody inventories

It gets worse where the recognized sources run out. The data these models fall back on is often Wikipedia and Reddit — corpora anyone can edit and nobody fully inventories. A patient, years-long manipulation campaign there would be effectively invisible: there is no signature to check, no capture history, no way to tell what changed or when. You are trusting that millions of anonymous edits net out to truth.

The model vendors know this. You can watch them quietly steering their systems away from Reddit and toward "the sources a newsroom uses." That is treating the symptom. The cure is the same first principle — data that carries its own provenance, so manipulation has to defeat math, not just win an edit war on a page nobody is watching.

And it cuts the other way too: a signed, primary-source archive becomes a yardstick. You can measure the open corpora against it and detect where they have drifted from the record. Origin does not just make our own data trustworthy — it gives you a way to catch when everyone else's has been moved.

Two levels of trust — in the right order

Trust should be a stack, and the order matters.

Level 1 — origin. The primary question is not "do I recognize the brand," it is "can I validate the signature that this came from the US government, or the state, or whoever produced it — unaltered?" That is math. It does not care who is on anyone's list.

Level 2 — reputation and curation. Useful, but secondary — a layer that rides on top of a verifiable foundation, not a substitute for one.

Today the stack is inverted: it is reputation-only, with no Level 1 underneath. We are building the missing primary layer.

Origin moves trust from the brand to the math

Origin-anchored data carries its proof with it. Each fact is bound to its exact source URL and the moment of capture, hashed so any change is evident, and signed so anyone can confirm it without trusting the messenger. Don't trust, verify, made literal.

We do both — and we fight with math

The honest standard — does it trace to a source of truth? — is exactly right. The argument is only about how you establish the trace. So we pull from the primary records themselves — Congress.gov, the House Clerk's roll calls, the FEC, the state legislatures — and we bind every datum to that source with a signature and a timestamp. You get the recognized source and the proof.

Reputation asks, "should you trust the publisher?" — a question only a handful of incumbents can answer yes to.

Origin asks, "can you verify this fact?" — a question any honest producer can answer yes to, and anyone can check.

You do not beat an approved-list oligarchy by lobbying your way onto the list. You make trust mathematical, so the list stops mattering. One of these is a permission system. The other keeps the supply of truth from collapsing into a few hands. We are building the second one — and we would rather fight with math.

Build a watchlist Why origin matters