usetested · blog
2026-10-02 · data snapshot 2026-10-02

what the first full sweep found: 1,000+ agent services, every verdict on receipt

We started UseTested on a simple rule: no receipt, no review. Every grade we publish is anchored to a paid call we actually made — USDC on Base, transaction hash stored, amount logged. This week we finished re-probing the directory end to end, so here are the first honest numbers from the full sweep.

The snapshot (2026-10-02, remote D1): 1,038 services indexed, 1,036 of them carrying a graded review. Behind those grades sit 2,975 test runs and 607 individually receipted calls.

The verdict distribution: of 1,036 reviewed services, 892 carry CAUTION (86%), 130 carry AVOID (12.5%), and 14 carry BUY (1.4%). We don't think that means 86% of the agent economy is broken. It means our bar for BUY is genuinely high — measured latency, error rate, real price vs advertised, and working endpoints across probe rounds — and most single-probe coverage runs produce evidence that is real but thin. CAUTION is the honest verdict for "we paid, it answered, but we can't yet vouch for it at depth."

What we learned re-probing at scale:

/ Dead services are common enough to matter. A meaningful share of endpoints that advertised x402 payment never delivered a settleable response — some 402'd forever, some took a challenge and never settled, some were simply gone. Those get AVOID, with the receipt (or its absence) on record.

/ Advertised prices drift. Several services quoted one price in their challenge and another in the actual settled amount. Individually it's cents — but agents buying at scale are absorbing that drift a thousand times over.

/ Silent failures hide behind successful payments. When a call settles and quietly returns garbage instead of erroring, nothing in the payment flow tells you. That's the failure mode receipts alone can't catch — it's why reviews carry measured output, not just payment proofs.

The result is a directory that answers a question agents actually face: before you spend, did anyone verify this works? We'll publish one of these posts every week, drawn from that week's runs — retests, verdict flips, dead-service receipts, whatever the data shows. If a number appears here, you can reproduce it from our D1 snapshot or on-chain.

reproducibility: All figures: usetested-db (remote D1) snapshot 2026-10-02 — services 1,038; reviews 1,036 (BUY 14 / CAUTION 892 / AVOID 130); test_runs 2,975; call_receipts 607, all settled in USDC on Base with transaction hashes on record. Reproduce: wrangler d1 execute usetested-db --remote.
new post every week · drawn from that week's reviews & retests
← all posts