Network directories only list the publishers who happened to sign up. The rest of your vertical — tens of thousands to hundreds of thousands of qualifying sites — never appears there. This guide walks the full-web discovery method: read the whole category, screen it against your ICP, and export one clean row per canonical domain.
An affiliate network can only show you sites that opted in, filled out a profile, and self-selected a category. That is a useful roster — but it is a recruitment queue, not a map of your vertical. The publishers your competitors have not already signed are, by definition, the ones not sitting in the same directory.
Marketplaces index the publishers who joined a network. High-intent niche sites often monetize directly and never enroll, so they are structurally invisible to that lens.
If every advertiser pulls from the same enrolled roster, you are all bidding for the same few hundred names. Overlap is high; uncontested capacity is elsewhere.
Directory categories are what a publisher typed about themselves, not what an LLM reads from the site today. Stale tags hide real fits and surface poor ones.
Starting from a classified index of the entire web — not a network roster — pulling every domain in a vertical, then reading each site against a written ideal partner profile. The output is a screened publisher universe, deduplicated to one canonical domain per site, with a confidence score on every signal.
The method is deliberately boring and inspectable — no black box, no scraped mailing list of unknown provenance.
1 — Universe: pull the whole category 2 — ICP screen: read each site with an LLM 3 — Dedup: one row per canonical domainEach step is explicit: you can see the universe size, the exact ICP definition, and the confidence on every field. This is the same pipeline that ran end to end for a large-scale production run on our own classification platform.
Your industry maps to categories in the 120M+ domain classification base. We pull every active domain in scope — including the long tail no directory has ever profiled.
"Our ideal publisher" becomes required, preferred, and disqualifying signals. The definition is written down and agreed before anything runs at scale.
A sample is screened, your team reviews it domain by domain, and the signal wording is tuned until measured precision clears your bar.
The LLM reads every domain in scope and extracts each signal with its own confidence score. ICP fit is computed from the signal set, not guessed from a keyword.
Mirrors, redirects, and country variants collapse to one canonical domain per publisher, so outreach never contacts the same site twice under different names.
A structured CSV, one row per canonical domain, ready for your partner CRM. Optional refresh runs catch new publishers, dead domains, and pivots.
In a large-scale production run on our own classification platform, the pipeline processed the travel vertical from the whole web down to a screened list. The client reviewed a sample of the output against their own judgment.
Marketplace rosters in the same vertical list a tiny fraction of these publishers — the sign-up bias, quantified.
Both have a place — a network handles payouts and tracking. But for building a recruitment pipeline that reaches beyond the enrolled roster, the two approaches start from different universes.
| Full-web discovery | Network marketplace / directory | |
|---|---|---|
| Starting set | Every classified domain in the category | Only sites already enrolled |
| Long-tail sites | Surfaced, including single-operator hubs | Mostly absent |
| Categorization | LLM reads the live site against your signals | Self-reported category tags |
| Competitor overlap | Reach names rivals never see | Everyone recruits the same list |
| False positives | Removed by disqualifying signals | Not distinguished |
| Output | Deduplicated CSV, one row per domain | In-platform list, limited export |
The uncontested capacity is rarely on page one of a search. It lives in formats and languages that directories under-index. These are the pockets a whole-web read reaches.
One-person "best X" hubs that monetize directly and never joined a network. High intent, zero directory footprint.
Country-specific review sites and blogs that English-first rosters barely index. Detected language and geo focus surface them.
Audiences that live in an inbox with a thin public site. The newsletter-presence signal flags them where a crawl-only view would skip past.
Deep single-topic sites too specific for a broad directory bucket, but a near-total audience match for a focused advertiser.
The free pilot runs full-web discovery on your vertical, tuned to your ICP, and delivers your first 20 qualified publishers plus a full-run projection. No cost, no obligation.
Request a Free Pilot