August 28, 20263 min read

The floor that isn’t a floor

Recon renders its partial subdomain count as a floor. Checking the claim, the floor leaks in three places, and none of them show up in the number.

Last week I wrote about Recon reporting subdomain coverage in three states instead of two, and about the state describing what the evidence supports rather than whether the request succeeded. When coverage is partial the count renders as ≥ 41 instead of 41. I was pleased with that. A number that admits it’s a lower bound.

This week I went to write about how well it works and checked the claim first. The floor leaks in three places.

The cache is exempt from the marking. Coverage counts a response as successful if it came back live or came from cache, and both yield comprehensive. So a crt.sh answer served from a six-hour-old cached copy renders as an exact number with no on it. The cache is disclosed, in body text, next to the state label. The number itself doesn’t carry it. If you read the figure and not the sentence beside it, you get an exact count for a domain that may have changed twice since anyone looked.

That cache is load-bearing and I’d keep it. crt.sh is the source that establishes a total and it’s also the source that hangs; the cache is what turns that from fatal into stale. But the marking convention says the number tells you what it can support, and here it doesn’t. Staleness is a limit on the claim exactly like partial coverage is, and only one of them made it into the digit.

The floor is a first page. Certspotter’s API is paginated by a cursor. Recon never sends one. It issues a single request and consumes whatever comes back, so ≥ N doesn’t mean at least what Certspotter knows about, it means at least what Certspotter’s first response contained. Whether those differ, and by how much, isn’t recorded anywhere in the codebase. I built a floor on top of an unexamined floor and shipped the top one with a symbol on it.

The evidence for the floor is one domain. The reason Certspotter can never be comprehensive is that its free tier under-counts large domains. That’s true and it’s measured, and the measurement is a single observation: twenty names for one domain where crt.sh returned what a commit message calls hundreds. The comparison figure was never written down. That’s enough to justify never trusting Certspotter as a total. It isn’t enough to characterise anything as a rate, and if I’d published a percentage this week it would have come from memory rather than from that.

None of these makes the wrong. It’s still a lower bound and it’s still more honest than a bare number. The problem is narrower and more annoying: the mark says the count is limited by source coverage, and there are two other things limiting it that the mark doesn’t distinguish. Someone reading ≥ 41 learns that Certspotter answered and crt.sh didn’t. They don’t learn whether they’re looking at the first page of Certspotter, and someone reading 41 doesn’t learn whether it’s live.

The thing I keep landing on is that honesty markers accumulate the same way any other feature does. I added one, put it in the right place, and then let two more limits appear underneath it without extending the convention. A mark that covers one of three limits reads, to whoever’s using the thing, like it covers all of them. It’s more useful than nothing, and it’s most misleading precisely where it looks most careful.

A qualified number tells you about the limit you thought of. It says nothing about the limits underneath it, and it looks the same either way.

Interested in how this maps to your work?

If this kind of thinking fits a problem you're facing, let's talk it through.

Start a conversation EAT (UTC+3) · Typically responds within 24 hours