Use case — website-type classification

Recruit content publishers. Not the OTAs, aggregators, and operators pretending to be them.

In a travel program — or any category with a booking layer — a keyword filter can't tell a genuine content publisher from an online travel agency, a metasearch aggregator, or the operator itself. They all mention flights and hotels. The website_type field reads what a site is, not what it talks about, so your recruitment list holds partners — never competitors.

1.2M
domains ICP-screened in one anonymized travel run
1.73M
travel-related domains isolated from a 30M dataset
~96%
precision on the client's own sample review
30M
domains in the source dataset
The job to be done

The word "hotel" appears on every one of these sites

The problem isn't finding travel domains — it's that most of them aren't publishers. Four very different business types crowd the same search results, and only one of them is a partner you can recruit. Confuse them and you email your own competitors.

Recruit

Content publisher

Writes reviews, guides, and comparisons and earns via affiliate links — it sends readers to a booking platform.

Tell: editorial voice, no native checkout, outbound affiliate links
Exclude

OTA / booking site

Sells the inventory itself with a native cart and checkout. It is the advertiser's competitor, not a channel.

Tell: own checkout, price + book buttons, inventory database
Exclude

Metasearch / aggregator

Compares prices across sellers and monetizes on referral, but competes for the same intent your advertiser wants.

Tell: live price tables, "compare deals", supplier feeds
Exclude

Operator / supplier

The airline, hotel chain, or tour operator's own site. Travel-relevant, but a brand, not a publisher.

Tell: single-brand inventory, corporate/press content, own booking
How the separation works

Five steps from a noisy category to clean publishers

The website_type classification runs alongside the ICP screen, so every domain lands with an explicit type label before a fit verdict is ever computed.

1

Pull the whole category

Every travel-related domain is isolated from the 120M+ classified universe — publishers and non-publishers alike, because you can't exclude a type you never surfaced.

2

Read structural signals, not keywords

The LLM inspects what the site does: is there a native checkout? A price-comparison table? A single-brand inventory? Outbound affiliate links? These reveal the business model.

3

Assign a website type

Each domain gets one website_type — content publisher, OTA, aggregator, operator, directory, forum, and so on — each with its own confidence score.

4

Apply type as a hard gate

Only the publisher types pass into ICP screening; OTAs, aggregators, and operators are removed structurally — not left for a human to catch one row at a time.

5

Calibrate to your precision bar

You review a sample of the type calls, we tune the boundary cases — e.g. a review site with a booking widget — and the full run only proceeds once measured precision clears your threshold.

6

Deliver with the type field kept

The website_type and its confidence stay in every row, so you can audit any verdict and re-slice the boundary later without a re-run.

A note on confidentiality

The classification examples below are drawn from a real screening run. Because the publishers involved are independent sites, we follow standard confidentiality practice on public pages: domains are withheld and descriptions generalized, so the profiles shown here are intentionally not traceable to specific sites — including through a web search. That protects the publishers without changing the underlying data. Your pilot and client deliverables carry the actual domains with every signal and type label, so every verdict can be verified directly.

A worked example

The same query, four verdicts

Four generalized domains that all rank for the same "best beach hotels" intent. A keyword filter keeps all four; the website-type gate keeps one. Confidence scores are illustrative.

Generalized profileStructural signals readwebsite_typeVerdict
Solo-run guide to a coastal region with hotel round-upsEditorial posts, own photos, outbound affiliate links, no checkoutcontent_publisherKeep (0.94)
Large "compare hotel prices" portalLive price tables from multiple suppliers, referral paramsaggregatorExclude (0.91)
Site that lists rooms and takes the booking on-pageNative cart, checkout flow, own inventory databaseotaExclude (0.96)
A single resort chain's official siteOne brand's properties, corporate/press pages, own bookingoperatorExclude (0.93)
Why this is the hard part: all four are unambiguously "travel." The difference between a partner and a competitor is the business model — and that only shows up when you read the site's structure, which is exactly what the type classifier is built to do.
Proven at scale

This is the exact problem one travel run solved

In a large-scale production run on our own classification platform, the pipeline separated content publishers from the booking layer across the whole travel vertical — then the client reviewed a sample of the output against their own judgment.

0
domains in the source dataset
0
travel-related domains isolated
0
domains ICP-screened in one run
0
precision on a reviewed sample
30M
source domains
1.73M
travel-related
1.2M
ICP-screened
~96%
precision confirmed

These are the case-study figures for this anonymized engagement. The website-type gate is what let the run isolate content publishers from OTAs, aggregators, and operators at that precision.

The structural signals

What the type classifier actually reads

Website type is decided on business-model evidence, not vocabulary. These are the concrete signals that separate a publisher from a booking or operator site — each extracted with its own confidence score.

Native checkout / cart present Live price-comparison tables Own inventory database Outbound affiliate / referral links Original editorial voice Single-brand vs multi-supplier Booking / reservation widget Corporate / press / investor pages Supplier price feeds Forum / UGC vs authored content First-hand media vs stock catalog Directory-listing structure

The same logic transfers to any vertical with a transaction layer — insurance carriers vs comparison blogs, brand stores vs review sites, marketplaces vs niche guides. The types change; the read-the-structure principle doesn't.

The deliverable

The type field ships in every row

A deduplicated CSV, one row per canonical domain, with the website type kept alongside the ICP verdict so you can audit and re-slice the publisher/non-publisher boundary at will.

ColumnExample valueWhat it does for you
website_typecontent_publisherThe hard gate — publishers pass, OTA/aggregator/operator don't
type_confidence0.94Slice review to boundary cases; auto-accept the certain ones
is_icp + confidencefit · 0.88Fit verdict, computed only after the type gate passes
monetizationaffiliate_linksConfirms a real affiliate channel, not a seller
language / geo_focusen / US-UKRoute by market without a separate enrichment step
alive_statusaliveReachability at run time, so outreach skips dead domains
Questions

Content vs OTA classification — FAQ

How is this different from filtering on keywords or categories?
Keywords and category tags describe what a site talks about; an OTA and a travel blog talk about identical things. The type classifier reads business-model structure instead — native checkout, price feeds, own inventory, outbound affiliate links — so it separates seller from publisher regardless of shared vocabulary.
What about a hybrid — a review site with a booking widget?
Those boundary cases are exactly what calibration is for. During the reviewed sample you decide which side of the line a widget-plus-editorial site falls on, and we tune the type definition to match your call. The lower type_confidence on such rows also lets you send just the ambiguous ones to a human.
Where do the 30M, 1.73M, 1.2M, and ~96% figures come from?
They are the case-study numbers from one real, anonymized engagement — a large-scale production run on our own classification platform. A 30M dataset yielded 1.73M travel-related domains, 1.2M were ICP-screened in a single run, and the client measured ~96% precision on a reviewed sample. We never name the client.
Does this only work for travel?
No. Any vertical with a transaction layer has the same publisher-vs-seller problem: insurance carriers vs comparison blogs, retail brands vs review sites, marketplaces vs niche guides. The website types are re-defined per vertical, but the method — classify the business model, then gate the ICP screen on it — is identical.
Can I keep the type field to audit verdicts later?
Yes — website_type and type_confidence stay in every delivered row. You can filter to any type, spot-check the classifier's calls against the real domains, and re-slice the publisher boundary without commissioning a new run.
Can I see the format before committing?
Yes — the free travel samples show the exact output format, and the free pilot runs the full pipeline on your vertical, gated on website type, and returns your first 20 qualified publishers plus a full-run projection.

Get a list of publishers — with the competitors filtered out

The free pilot runs the full pipeline on your vertical, gates it on website type, and delivers your first 20 qualified publishers plus a full-run projection. No cost, no obligation.

Request a Free Pilot