In a travel program — or any category with a booking layer — a keyword filter can't tell a genuine content publisher from an online travel agency, a metasearch aggregator, or the operator itself. They all mention flights and hotels. The website_type field reads what a site is, not what it talks about, so your recruitment list holds partners — never competitors.
The problem isn't finding travel domains — it's that most of them aren't publishers. Four very different business types crowd the same search results, and only one of them is a partner you can recruit. Confuse them and you email your own competitors.
Writes reviews, guides, and comparisons and earns via affiliate links — it sends readers to a booking platform.
Sells the inventory itself with a native cart and checkout. It is the advertiser's competitor, not a channel.
Compares prices across sellers and monetizes on referral, but competes for the same intent your advertiser wants.
The airline, hotel chain, or tour operator's own site. Travel-relevant, but a brand, not a publisher.
The website_type classification runs alongside the ICP screen, so every domain lands with an explicit type label before a fit verdict is ever computed.
Every travel-related domain is isolated from the 120M+ classified universe — publishers and non-publishers alike, because you can't exclude a type you never surfaced.
The LLM inspects what the site does: is there a native checkout? A price-comparison table? A single-brand inventory? Outbound affiliate links? These reveal the business model.
Each domain gets one website_type — content publisher, OTA, aggregator, operator, directory, forum, and so on — each with its own confidence score.
Only the publisher types pass into ICP screening; OTAs, aggregators, and operators are removed structurally — not left for a human to catch one row at a time.
You review a sample of the type calls, we tune the boundary cases — e.g. a review site with a booking widget — and the full run only proceeds once measured precision clears your threshold.
The website_type and its confidence stay in every row, so you can audit any verdict and re-slice the boundary later without a re-run.
The classification examples below are drawn from a real screening run. Because the publishers involved are independent sites, we follow standard confidentiality practice on public pages: domains are withheld and descriptions generalized, so the profiles shown here are intentionally not traceable to specific sites — including through a web search. That protects the publishers without changing the underlying data. Your pilot and client deliverables carry the actual domains with every signal and type label, so every verdict can be verified directly.
Four generalized domains that all rank for the same "best beach hotels" intent. A keyword filter keeps all four; the website-type gate keeps one. Confidence scores are illustrative.
| Generalized profile | Structural signals read | website_type | Verdict |
|---|---|---|---|
| Solo-run guide to a coastal region with hotel round-ups | Editorial posts, own photos, outbound affiliate links, no checkout | content_publisher | Keep (0.94) |
| Large "compare hotel prices" portal | Live price tables from multiple suppliers, referral params | aggregator | Exclude (0.91) |
| Site that lists rooms and takes the booking on-page | Native cart, checkout flow, own inventory database | ota | Exclude (0.96) |
| A single resort chain's official site | One brand's properties, corporate/press pages, own booking | operator | Exclude (0.93) |
In a large-scale production run on our own classification platform, the pipeline separated content publishers from the booking layer across the whole travel vertical — then the client reviewed a sample of the output against their own judgment.
These are the case-study figures for this anonymized engagement. The website-type gate is what let the run isolate content publishers from OTAs, aggregators, and operators at that precision.
Website type is decided on business-model evidence, not vocabulary. These are the concrete signals that separate a publisher from a booking or operator site — each extracted with its own confidence score.
The same logic transfers to any vertical with a transaction layer — insurance carriers vs comparison blogs, brand stores vs review sites, marketplaces vs niche guides. The types change; the read-the-structure principle doesn't.
A deduplicated CSV, one row per canonical domain, with the website type kept alongside the ICP verdict so you can audit and re-slice the publisher/non-publisher boundary at will.
| Column | Example value | What it does for you |
|---|---|---|
website_type | content_publisher | The hard gate — publishers pass, OTA/aggregator/operator don't |
type_confidence | 0.94 | Slice review to boundary cases; auto-accept the certain ones |
is_icp + confidence | fit · 0.88 | Fit verdict, computed only after the type gate passes |
monetization | affiliate_links | Confirms a real affiliate channel, not a seller |
language / geo_focus | en / US-UK | Route by market without a separate enrichment step |
alive_status | alive | Reachability at run time, so outreach skips dead domains |
The free pilot runs the full pipeline on your vertical, gates it on website type, and delivers your first 20 qualified publishers plus a full-run projection. No cost, no obligation.
Request a Free Pilot