Use case — Long-tail publisher discovery

Surface the long-tail publishers no directory has ever profiled

Below the CrUX top-1M sits a vast band of real, engaged, on-topic audiences that every directory and outreach tool ignores. We read all 120M+ classified domains and screen them against your ICP — so you can recruit the long-tail publishers where your competitors have exactly zero coverage.

~180,000
long-tail publishers below the top-1M a full run typically surfaces (estimate)
120M+
classified domains read — not just the ranked head
500
real long-tail domains delivered free in your pilot
The job to be done

The head is crowded; the long tail is uncontested

Every advertiser recruits the same famous, high-ranked publishers, so commissions climb and slots fill. The long tail — sites below the CrUX top-1M with small but genuinely engaged, on-topic audiences — is where uncontested capacity lives. It's invisible to directory tools precisely because it isn't ranked, which is exactly why no competitor is there.

Tools only see the head

Rank-based databases and directories index the visible top of the market. Below the top-1M, coverage collapses — and that's most of the real web.

Small but engaged

A niche site with a devoted audience often out-converts a giant generalist. Low rank is not low value when the topic fit is total.

Zero competitor coverage

Because no directory profiles these sites, no competitor is recruiting them. You reach them first, on your terms, without a bidding war.

How the pipeline solves it

Reading the whole web finds what ranking hides

Because we start from a classified-domain database rather than a rank index, the long tail is in scope from the first step — not an afterthought.

1

Start from classification, not rank

The 120M+ database is built by reading domains, not by ranking traffic. Long-tail sites are present from the start, on equal footing with the head.

2

Pull the whole vertical

We pull every classified domain in your category — the ranked head and the unranked long tail together — instead of whatever a directory happened to index.

3

Read for real audience signals

The LLM reads each site for engagement evidence — comments, community, newsletter, publishing cadence — separating a real small audience from an empty shell.

4

Screen against your ICP

Every long-tail candidate is scored on your required, preferred, and disqualifying signals, so small doesn't mean unqualified — it means on-profile and uncontested.

5

Attach honest traffic proxies

We report CrUX presence and rank group — reproducible proxies, never invented visitor counts — so you can filter by scale without pretending to precise numbers.

6

Deduplicate & deliver

One canonical row per publisher, with ICP fit, engagement signals, and rank group — a long-tail recruitment list ready for your outreach stack.

Low rank is a filter, not a verdict. A site outside the top-1M can still hold a devoted, tightly-focused audience that converts above a giant generalist — the engagement signals prove it, and the rank group keeps you honest about scale. Because no directory indexes these publishers, reaching them is a durable advantage rather than another bidding war.
A worked example

~180,000 long-tail publishers below the top-1M (estimate)

Most of a vertical's qualifying publishers sit outside the ranked head. Screening the whole category and keeping the on-profile sites below the CrUX top-1M surfaces tens of thousands to hundreds of thousands of uncontested publishers. The figures below are estimates for illustration.

Illustrative full run

Long-tail slice of one vertical

Rank group comes from CrUX presence; the long-tail count is an estimate. Your pilot returns real below-the-head domains and a projection tuned to your ICP.

0
classified domains as the source
0
ICP-fit publishers in the vertical (estimate)
0
of them below the top-1M (estimate)
0
delivered free in your pilot
120M+
classified domains
~210K
ICP-fit (est.)
~180K
below top-1M (est.)
500
free pilot slice

Reference precision on a vertical screen: ~96% on a reviewed sample. Long-tail figures above are estimates extrapolated from calibration slices, shown for illustration.

The tail is bigger than the head. Across a vertical, the below-the-top-1M cohort dwarfs the ranked publishers everyone already fights over. That's the uncontested capacity full-web reading unlocks.
Long-tail signals

The specific signals that prove a small site is real

The risk in the long tail is noise — abandoned or thin sites. These scored signals separate a genuine small audience from an empty domain, each carrying its own confidence.

CrUX presence & rank group Comments / community activity Newsletter present Active social presence Publishing cadence Fresh content (recency) Tight niche focus Relevant to your vertical Individual or small team run Affiliate links present Abandoned / stale (disqualifier) Thin AI-spun content (disqualifier)
Deliverable snapshot

The long-tail CSV, one row per publisher

An illustrative slice. Rank group sits beside real engagement signals, so you can recruit small-but-engaged sites with eyes open — never guessing at traffic.

A note on confidentiality

The rows below are illustrative and generalized. On public pages we withhold real domains and generalize descriptions, so any profile shown here is intentionally not traceable to a specific site — including through a web search — which protects independent publishers without changing the underlying data. Your pilot and client deliverables contain the actual domains and every signal, so everything can be verified directly.

DomainRank groupEngagementNicheICP fitConf.Freshness
publisher—t1Below top-1MNewsletter + commentsTight nicheFit0.90< 14 days
publisher—t2Below top-1MActive communitySub-verticalFit0.87< 30 days
publisher—t3Not in CrUXSocial + cadenceMicro-nicheFit0.82< 45 days
publisher—t4Not in CrUXNone detectedThin/staleExcluded0.91> 12 months
publisher—t5Below top-1MNewsletterTight nicheFit0.84< 60 days
Honest proxies, not invented numbers. The full CSV adds CrUX rank group, community and newsletter signals, publishing cadence, individual-vs-company run, and a per-signal confidence score — so you filter the long tail on evidence, not guesswork.
Questions

Long-tail publisher discovery — FAQ

How can you find sites that aren't in any directory?
Our starting point is a 120M+ classified-domain database built by reading domains, not by ranking traffic. Long-tail sites are in scope from the first step, so discovery never depends on a site being popular enough to appear in a rank index.
How do you avoid recruiting dead or thin long-tail sites?
Real-audience signals — comments, community, newsletter, publishing cadence, freshness — separate an engaged small site from an empty shell, and abandoned or thin AI-spun content is a disqualifier. Each signal carries a confidence score you can filter on.
How do you report traffic for sites this small?
Only through reproducible proxies: CrUX presence and rank group. We never invent visitor numbers. For long-tail sites that means an honest "below top-1M" or "not in CrUX" rather than a fabricated figure.
Where does the ~180,000 long-tail figure come from?
It is an estimate: the ICP-fit share of a vertical extrapolated from calibration slices, restricted to domains below the CrUX top-1M. It is labeled as an estimate. Your pilot replaces it with a projection built from your own ICP and a real 20-publisher sample.
What does the free pilot include?
The full pipeline configured for your ICP with a long-tail focus, stopped at the first 20 qualified below-the-head publisher domains — free — plus a projected full-run yield and quotes. It is the fastest way to see uncontested names in your vertical.

Recruit the long tail before anyone else finds it

The free pilot runs the full pipeline on your vertical with a long-tail focus, tuned to your ICP, and delivers your first 20 below-the-head publisher domains plus a full-run projection. No cost, no obligation.

Request a Free Pilot