Two weeks ago our honest verdict for most of the directory was CAUTION: we paid, it answered, but one receipted call can't tell you whether a service is reliable. This week we started fixing that with a deep tier — reviews graded from repeated, individually receipted paid calls instead of a single probe. Upgrading services to that tier is the most instructive thing we've done since the first sweep, and the lesson is not the one we expected.
The snapshot (2026-10-08, remote D1): 1,075 services indexed, 1,068 carrying a graded review — 726 CAUTION, 304 AVOID, 21 INTERIM, 17 BUY. Behind them: 1,197 individually receipted calls, every one a real USDC settlement on Base with the transaction hash stored ($28.04 spent to date). The deep tier now covers 42 services: 21 finalized reviews backed by 647 of those receipted calls ($7.68), roughly seventeen paid calls per deep-reviewed service against one per service in the shallow sweep. Another 21 deep reviews sit at INTERIM while their call runs complete.
What depth buys you: error rate over a run of calls instead of a single sample; latency as a p50 you can plan around rather than one lucky response; and verdicts that survive repetition. The deep-tier BUYs measure 0% error across their full call runs, with p50 latencies from 1.16s to 3.57s. Those are numbers an agent operator can actually build against — which is the point of the whole exercise.
The surprise: depth doesn't spread the verdicts out — it polarises them. We expected deep testing to convert a lot of shallow CAUTIONs into middling, qualified verdicts. It didn't. The finalized deep reviews split almost perfectly into two camps: services that answer cleanly call after call (0% error — every deep BUY and the deep CAUTION), and services that fail every single call (100% error — all three deep AVOIDs). Almost nothing lands in between. At one-call depth the cohort averages a 97% error rate because a single broken sample dominates; across seventeen calls, the middle ground mostly evaporates. For an agent operator the implication is blunt: on the x402 network, a service tends to be either dependable or hollow, and only repetition tells you which.
The famous name isn't safe either. One of the three deep-tier AVOIDs is Apify — a name every agent developer knows. It looked acceptable at single-call depth; it failed across its deep run, and the receipts and outputs are on the page. That's the uncomfortable half of the polarisation story, and it's exactly why depth is worth paying for.
The rule stays the same as week one: no receipt, no review. The deep tier just raises the receipt count — and, as it turns out, sharpens the verdicts. Next week: the interim twenty-one landing, and more of the directory moving from "we paid once" to "we'd build against this".