Use case — custom taxonomy classification

Classify the whole publisher universe against your taxonomy — not just our standard fields.

Your program already has a category tree, a partner-tiering scheme, and fields your CRM expects. Instead of forcing your universe into our standard signal set, we engineer bespoke fields to your exact taxonomy, calibrate them on a reviewed sample, and classify 120M+ classified domains straight into the schema your team already works in.

120M+
classified domains available to your custom schema
millions of domains
classifiable in a large multi-vertical taxonomy
bespoke fields
engineered and calibrated per engagement
The job to be done

Your data model didn't come from us — the classification shouldn't either

A generic "is_icp + confidence" row is a fine start, but mature programs run on their own vocabulary. When the deliverable already speaks your schema, it drops into your stack with no re-mapping and no lossy translation.

Your category tree

Multi-level taxonomies, sub-verticals, and content-type buckets specific to how you organize partners — not a flat industry label. Each publisher is placed in your tree, with confidence.

Your definitions

Your "premium publisher," your "editorial vs. deals" line, your tier thresholds — encoded as extraction rules and calibrated to match your team's judgment, not a generic default.

Your field names

Output columns named and typed to match your CRM or PIM. The delivered CSV imports without a mapping spreadsheet — the fields already line up.

How bespoke fields get built

Six steps from your taxonomy to a classified universe

Custom classification is field engineering, not a checkbox. Each of your fields becomes an extraction task with a written definition, a confidence model, and a calibration pass before anything runs at scale.

1

Intake your taxonomy

We take your category tree, tier rules, and field list — a spreadsheet, a data dictionary, or a schema export — and map every field you want populated.

2

Translate to extraction rules

Each field becomes a precise, written definition of what evidence on a site makes it true — so "premium" or "enterprise-focused" means something specific and repeatable, not a vibe.

3

Calibrate on a reviewed sample

We classify a sample against your fields, your team reviews it row by row, and we tune the definitions until the calls match your judgment at the precision you set.

4

Run against the universe

The calibrated field set runs over the whole vertical from 120M+ classified domains — each domain read once and populated across every one of your custom fields with confidence scores.

5

Dedup to your key

Mirrors and country variants collapse to one canonical domain, keyed the way your CRM expects, so records join cleanly on import.

6

Deliver in your schema

The CSV arrives with your field names, your value sets, and a confidence column per field — ready to load without a translation layer.

Standard plus bespoke

Start from our fields, add whatever you actually need

Every engagement keeps the standard signal set as a foundation; custom fields sit on top. The teal chips are always available; the indigo chips are examples of bespoke fields defined per engagement.

Relevant to industry Website type Monetization signals Language / geo focus Traffic proxies Content freshness Your 3-level category tree Editorial vs. deals vs. tools Covers enterprise segment Staffed editorial team size band Publishes in DE / FR / ES Already on your named networks Runs a podcast / video channel Your "premium publisher" flag

The indigo fields are illustrative. Whatever your team tracks — a compliance flag, a brand-safety tier, a regional bucket — can be defined as a field and calibrated the same way as any standard signal.

A note on confidentiality

The mapped example below is drawn from real classification work. Because the publisher is an independent site, we follow standard confidentiality practice on public pages: the domain is withheld and details generalized, so the profile is intentionally not traceable to a specific site — including through a web search. That protects the publisher without changing the underlying data. Your pilot and client deliverables carry the actual domains populated across every custom field, so each classification can be verified directly.

A worked example

One publisher, classified into a client's own schema

A generalized publisher mapped to an illustrative custom taxonomy — the standard fields plus the client's bespoke ones, each with confidence. Domain withheld; values illustrative.

Classified to the client's taxonomy
Standard fields
ICP fit · 0.90 Type: content_publisher EN · UK
Client's custom fields
Category: Home › Smart-home › Security · 0.87 Content type: Editorial reviews · 0.92 Team band: 2–5 editors · 0.71 Premium flag: Yes · 0.80 On named network: Not detected · 0.83

Placed in the tree, not a flat label

Instead of "home & garden," the publisher lands three levels deep in the client's own category tree — the granularity their routing and reporting actually run on.

Confidence on every bespoke field

The lower-confidence fields — team-size band, premium flag — are exactly the ones to spot-check first. Confidence makes a custom field auditable, just like a standard one.

Same engine, your schema

Bespoke fields, engineered to the same standard

A custom field isn't a looser field. Every one you define gets a written extraction rule, a confidence model, and a reviewed-sample calibration — the same rigor as the standard signal set that reached ~96% precision in the reference engagement.

0
domains available to your schema
0
taxonomy levels supported
0
confidence score per custom field
0
domains free in your pilot
Your taxonomy
intake
Field rules
+ calibration
Full run
120M+ domains
Your schema
delivered CSV

Reference engagement (a large-scale production run on our own classification platform): a 30M dataset → 1.73M travel-related domains → 1.2M ICP-screened → ~96% precision on a reviewed sample. Custom-taxonomy figures are illustrative.

The deliverable

A CSV that already speaks your data model

One row per canonical domain, columns named and typed to your schema, with a confidence score beside every custom field so your team can trust and audit each classification.

ColumnExample valueOrigin
canonical_domainyour CRM keyStandard — dedup join key
category_l1 / l2 / l3Home / Smart-home / SecurityCustom — your category tree
content_typeeditorial_reviewsCustom — your value set
premium_flag + confidencetrue · 0.80Custom — your definition
team_size_band + confidence2–5 · 0.71Custom — calibrated bands
is_icp + confidencefit · 0.90Standard — kept alongside custom
Scope note: bespoke field engineering is part of a custom-scope engagement — the most popular way clients run this. See pricing for how custom scope is quoted.
Questions

Custom taxonomy classification — FAQ

How do I hand over my taxonomy?
However you already keep it — a spreadsheet of categories, a data dictionary, a schema export, or a walkthrough call. We map every field you want populated, confirm the value sets, and write an extraction definition for each before any calibration runs.
Can you handle a multi-level category tree, not just flat tags?
Yes. Each publisher can be placed several levels deep in your hierarchy — category, sub-category, and content type, for example — with a confidence score at the level you care about. The tree is part of the extraction spec, not a post-hoc rollup.
How do you keep a subjective field like "premium" consistent?
By turning it into a written rule and calibrating it. During the reviewed sample your team marks which sites are "premium," disagreements pinpoint where the definition is fuzzy, and we sharpen it until the calls match your judgment at your precision bar — the same loop we use for any signal.
Do custom fields come with confidence scores too?
Every custom field ships with its own confidence, exactly like a standard signal. That lets you auto-accept high-confidence classifications and route the uncertain ones to a human — so a bespoke field is just as auditable as anything in the standard set.
Will the file import into my CRM without re-mapping?
That's the goal. Columns are named and typed to your schema and keyed on the canonical domain your CRM already uses, so records join on import without a mapping spreadsheet or a lossy translation step.
Is this a custom-scope engagement?
Yes — bespoke field engineering runs as a custom-scope project, the most popular way clients work with us, because the field set and calibration are unique to each program. The free pilot is the best first step: it configures the pipeline for your ICP and returns your first 20 qualified domains before you commit to a full custom build.

Get the universe classified in your own language

The free pilot configures the full pipeline for your ICP and returns your first 20 qualified publishers plus a full-run projection — the fastest way to scope a bespoke custom-taxonomy build. No cost, no obligation.

Request a Free Pilot