Guide — LLM screening, explained

How LLM classification reads a website against your ICP — and how to trust the result

Partner teams don't need to know how a model is built; they need to know why its verdict is safe to act on. This guide explains how LLM screening reads a site against a plain-language ideal-partner profile, why every field ships with its own confidence score, and how precision is measured on a reviewed sample before a single full-run credit is spent. It also states the limits plainly.

0
domains screened in one anonymized run
0
precision on a reviewed sample
1
confidence score on every extracted field
Start here

Your ICP, in plain language — then in fields

You describe your ideal partner the way you'd brief a new hire. That description becomes a structured set of required, preferred, and disqualifying fields — each one something the model can look for on an actual page, not a vibe.

What you say:
"We want independent blogs in our niche that publish real reviews, have an engaged audience, already use affiliate links, and post regularly. We don't want retailers who sell competing products, and we don't want thin AI-spun sites."
ICP as fieldswhat the model actually checks per site
Required
Niche relevance Genuine review content Editorial independence
Preferred
Affiliate links present Fresh content Newsletter / audience
Disqualifying
Sells competing products Thin AI-spun content
The method

How the model reads a single site

Each domain goes through the same five steps. The verdict is derived from field evidence, never from a keyword match on the domain name.

1

Fetch the live site

The current pages are read — home, key articles, about, and monetization markers — not a cached title tag or a directory blurb.

2

Extract each field

For every ICP field, the model finds the on-page evidence: an actual review, a byline, affiliate link patterns, a publish date, a shopping cart.

3

Score confidence

Each field gets its own confidence score reflecting how clear the evidence was — strong, ambiguous, or thin.

4

Apply your logic

Required, preferred, and disqualifying rules combine the fields into an ICP verdict. A disqualifier overrides, however relevant the rest.

5

Emit a row

One structured row: verdict, every field value, and its confidence — auditable, sliceable, and ready for your CRM.

The part that builds trust

Why a per-field confidence score changes everything

A single yes/no verdict hides its own uncertainty. Scoring each field separately lets your team sort, threshold, and audit — and it turns "trust me" into something you can check.

One screened publisherillustrative field readout
Niche relevance
0.97
Genuine review content
0.92
Affiliate links present
0.88
Fresh content
0.71
Editorial independence
0.64

Illustrative values — not a real domain. The point is the shape: strong evidence and weak evidence are visible, not hidden inside a single verdict.

Threshold to your risk appetite

Work the high-confidence fits first for a fast win, or lower the bar to widen reach when you have outreach capacity to spare. The dial is yours.

Audit the disagreements

When your team doubts a verdict, the low-confidence field usually explains why — and tells us exactly which definition to sharpen during calibration.

Precision, measured not promised

How precision is measured before the full run

No full-scale credit is spent on faith. A reviewed sample proves the screen works on your vertical first, and only a passing sample unlocks the full run.

1. Screen a sample

A representative slice of your vertical is screened with the draft ICP. First-pass errors surface fast — sites that match words but not intent.

2. You review it

Your team marks each sampled domain fit or not fit. The disagreements are the signal — they name the definitions to sharpen.

3. Tune, then run

Field definitions and thresholds are adjusted and the sample is re-screened until measured precision clears your bar. Then the full run starts.

Precision, defined plainly: of the publishers the screen labels a fit, the share your own review confirms as genuine fits. It is counted on a sample you check — not asserted by us.
Proven at scale

One run, measured end to end

In a large-scale production run on our own classification platform, the screen read the travel vertical against the ICP. The client reviewed a sample of the output and confirmed the precision below.

0
domains in the source dataset
0
travel-related domains identified
0
domains ICP-screened in one run
0
precision on the client's sample
30M
source domains
1.73M
travel-related
1.2M
ICP-screened
~96%
precision confirmed

Figures from one anonymized engagement. A vertical of this scale typically yields tens of thousands to hundreds of thousands of qualifying publishers per run (estimate); your own numbers come from your pilot.

Said plainly

The honest limits

A screen that hides its failure modes can't be trusted. Here is where LLM classification is weak, and how the workflow contains each weakness rather than pretending it away.

It is not 100%

No screen is. Precision is measured, reported, and tunable — never claimed as perfect. The confidence scores exist precisely so the uncertain cases are visible.

It reads a moment in time

A site can pivot or lapse after screening. Freshness is captured at read time, and scheduled refresh runs catch drift — but a run is a snapshot, not a live feed.

Some evidence is hidden

Login walls, heavy JavaScript, and blocked pages limit what any reader sees. Those cases get lower confidence and are flagged rather than guessed at.

Judgment stays yours

The screen surfaces and scores; brand-fit and final approval are human calls. We deliver evidence, not a mandate to email everyone on the list.

Where it fits

What LLM screening replaces — and what it doesn't

The screen is not a magic verdict machine; it is a way to do the reading no human team could do at scale, with the uncertainty made visible. Seen against the alternatives, its role is specific.

Replaces keyword matching

A keyword filter can't distinguish intent from vocabulary. The screen reads what a page does, applies your disqualifiers, and removes the false positives a text match keeps.

Scales manual review

Hand-checking a million sites is impossible; the screen reads them all consistently, then focuses your team's judgment on the sample and the low-confidence rows that need it.

Doesn't replace your call

Brand fit, commercial terms, and final approval stay human. The screen hands your team evidence and confidence, not a decision it never had the context to make.

The division of labor: the model does the reading at a scale no team can match; your team does the judging on the cases that carry real risk. Confidence scores are what route each case to the right one.
Questions

LLM classification — FAQ

Do I need to understand prompts or models to use this?
No. You describe your ideal partner in plain language and review a sample of results. Translating that description into fields, running the screen, and calibrating it is our job. You interact with verdicts and confidence scores, never with model internals.
What does a confidence score actually mean?
It reflects how clear the on-page evidence for that specific field was. A site with an obvious, dated review and visible affiliate links scores high on those fields; a site where the signal is ambiguous scores lower. It lets you prioritize certainty and audit the borderline cases.
How is this different from a keyword or category filter?
A keyword filter matches text; it can't tell a genuine reviewer from a retailer that mentions the same words. The screen reads what the site does — tests products, sells inventory, spins content — and applies your disqualifiers structurally, which is why it removes false positives a keyword pass keeps.
What if the screen gets one wrong?
Some errors are expected — precision is measured, not perfect. That's what the reviewed sample is for: your team catches the misses, we sharpen the field definitions, and the run only proceeds once measured precision clears your bar. The confidence scores also let you review the least-certain rows first.
Can I see it run on my vertical before paying?
Yes — that's the free pilot. We configure the screen for your ICP, calibrate on a reviewed sample, and deliver the first 20 qualified publishers with their field values and confidence scores, plus a projected full-run yield. You judge the output before committing.

See the screen run on your ICP

The free pilot configures LLM classification for your ideal partner, calibrates it on a reviewed sample, and delivers your first 20 qualified publishers with per-field confidence — plus a full-run projection. No cost, no obligation.

Request a Free Pilot