Guide — publisher discovery

How to find affiliate publishers beyond the network marketplaces

Network directories only list the publishers who happened to sign up. The rest of your vertical — tens of thousands to hundreds of thousands of qualifying sites — never appears there. This guide walks the full-web discovery method: read the whole category, screen it against your ICP, and export one clean row per canonical domain.

120M+
classified domains as the discovery source
Tens of thousands+
qualifying publishers a full run can surface per vertical (estimate)
1 row
per canonical domain in the deduplicated CSV
The blind spot

Why marketplaces only surface a sliver of the field

An affiliate network can only show you sites that opted in, filled out a profile, and self-selected a category. That is a useful roster — but it is a recruitment queue, not a map of your vertical. The publishers your competitors have not already signed are, by definition, the ones not sitting in the same directory.

Sign-up bias

Marketplaces index the publishers who joined a network. High-intent niche sites often monetize directly and never enroll, so they are structurally invisible to that lens.

Everyone recruits the same list

If every advertiser pulls from the same enrolled roster, you are all bidding for the same few hundred names. Overlap is high; uncontested capacity is elsewhere.

Self-reported tags

Directory categories are what a publisher typed about themselves, not what an LLM reads from the site today. Stale tags hide real fits and surface poor ones.

The reframe: a marketplace answers "who has signed up in this category?" Full-web discovery answers a different, larger question — "who on the entire web fits my partner profile, whether or not they have ever heard of me?"
Full-web discovery

Starting from a classified index of the entire web — not a network roster — pulling every domain in a vertical, then reading each site against a written ideal partner profile. The output is a screened publisher universe, deduplicated to one canonical domain per site, with a confidence score on every signal.

Universe → screen → CSV

Three moves, in plain terms

The method is deliberately boring and inspectable — no black box, no scraped mailing list of unknown provenance.

1 — Universe: pull the whole category 2 — ICP screen: read each site with an LLM 3 — Dedup: one row per canonical domain
The method

The full-web discovery step chain

Each step is explicit: you can see the universe size, the exact ICP definition, and the confidence on every field. This is the same pipeline that ran end to end for a large-scale production run on our own classification platform.

1

Map the vertical to categories

Your industry maps to categories in the 120M+ domain classification base. We pull every active domain in scope — including the long tail no directory has ever profiled.

2

Write the ICP as signals

"Our ideal publisher" becomes required, preferred, and disqualifying signals. The definition is written down and agreed before anything runs at scale.

3

Calibrate on a reviewed slice

A sample is screened, your team reviews it domain by domain, and the signal wording is tuned until measured precision clears your bar.

4

Run the full universe

The LLM reads every domain in scope and extracts each signal with its own confidence score. ICP fit is computed from the signal set, not guessed from a keyword.

5

Dedup and canonicalize

Mirrors, redirects, and country variants collapse to one canonical domain per publisher, so outreach never contacts the same site twice under different names.

6

Export and refresh

A structured CSV, one row per canonical domain, ready for your partner CRM. Optional refresh runs catch new publishers, dead domains, and pivots.

Proven at scale

One vertical, read end to end

In a large-scale production run on our own classification platform, the pipeline processed the travel vertical from the whole web down to a screened list. The client reviewed a sample of the output against their own judgment.

0
domains in the source dataset
0
travel-related domains identified
0
domains ICP-screened in one run
0
precision on a reviewed sample
30M
source domains
1.73M
travel-related
1.2M
ICP-screened
~96%
precision confirmed

Marketplace rosters in the same vertical list a tiny fraction of these publishers — the sign-up bias, quantified.

Side by side

Marketplace vs. full-web discovery

Both have a place — a network handles payouts and tracking. But for building a recruitment pipeline that reaches beyond the enrolled roster, the two approaches start from different universes.

 Full-web discoveryNetwork marketplace / directory
Starting setEvery classified domain in the categoryOnly sites already enrolled
Long-tail sitesSurfaced, including single-operator hubsMostly absent
CategorizationLLM reads the live site against your signalsSelf-reported category tags
Competitor overlapReach names rivals never seeEveryone recruits the same list
False positivesRemoved by disqualifying signalsNot distinguished
OutputDeduplicated CSV, one row per domainIn-platform list, limited export
Field notes

Where the long tail actually hides

The uncontested capacity is rarely on page one of a search. It lives in formats and languages that directories under-index. These are the pockets a whole-web read reaches.

Solo operators

One-person "best X" hubs that monetize directly and never joined a network. High intent, zero directory footprint.

Non-English publishers

Country-specific review sites and blogs that English-first rosters barely index. Detected language and geo focus surface them.

Newsletter-led sites

Audiences that live in an inbox with a thin public site. The newsletter-presence signal flags them where a crawl-only view would skip past.

Sub-niche authorities

Deep single-topic sites too specific for a broad directory bucket, but a near-total audience match for a focused advertiser.

Questions

Full-web publisher discovery — FAQ

Does this replace my affiliate network?
No. Your network still handles tracking, payouts, and the sign-up flow. Full-web discovery is upstream of that: it builds the recruitment list you then invite into whatever program you run, reaching publishers the roster never contained.
How many publishers can a full run realistically surface?
It varies by vertical, from tens of thousands to hundreds of thousands of qualifying publishers, and into the millions of candidate domains once you span multiple countries and languages. Your pilot returns a projection tuned to your exact ICP rather than a generic number.
Isn't this just scraping search results?
No. It starts from a 120M+ classified-domain database and reads each candidate site against a written signal set. Search scraping returns whatever ranks today; a category read returns the whole vertical, including sites that rank for nothing yet still fit your profile.
How do you avoid contacting the same site twice?
Dedup and canonicalization collapse mirrors, redirects, and country variants into a single canonical domain per publisher. The delivered CSV holds one row per site, so your outreach team never double-contacts the same operator under different URLs.
What does the free pilot include?
The full pipeline configured for your ICP, stopped at the first 20 qualified publishers — free — plus a projected full-run yield and quotes. It is the fastest way to see real names outside the marketplace roster before committing to anything.

See the publishers the marketplace never listed

The free pilot runs full-web discovery on your vertical, tuned to your ICP, and delivers your first 20 qualified publishers plus a full-run projection. No cost, no obligation.

Request a Free Pilot