Partner teams don't need to know how a model is built; they need to know why its verdict is safe to act on. This guide explains how LLM screening reads a site against a plain-language ideal-partner profile, why every field ships with its own confidence score, and how precision is measured on a reviewed sample before a single full-run credit is spent. It also states the limits plainly.
You describe your ideal partner the way you'd brief a new hire. That description becomes a structured set of required, preferred, and disqualifying fields — each one something the model can look for on an actual page, not a vibe.
Each domain goes through the same five steps. The verdict is derived from field evidence, never from a keyword match on the domain name.
The current pages are read — home, key articles, about, and monetization markers — not a cached title tag or a directory blurb.
For every ICP field, the model finds the on-page evidence: an actual review, a byline, affiliate link patterns, a publish date, a shopping cart.
Each field gets its own confidence score reflecting how clear the evidence was — strong, ambiguous, or thin.
Required, preferred, and disqualifying rules combine the fields into an ICP verdict. A disqualifier overrides, however relevant the rest.
One structured row: verdict, every field value, and its confidence — auditable, sliceable, and ready for your CRM.
A single yes/no verdict hides its own uncertainty. Scoring each field separately lets your team sort, threshold, and audit — and it turns "trust me" into something you can check.
Illustrative values — not a real domain. The point is the shape: strong evidence and weak evidence are visible, not hidden inside a single verdict.
Work the high-confidence fits first for a fast win, or lower the bar to widen reach when you have outreach capacity to spare. The dial is yours.
When your team doubts a verdict, the low-confidence field usually explains why — and tells us exactly which definition to sharpen during calibration.
No full-scale credit is spent on faith. A reviewed sample proves the screen works on your vertical first, and only a passing sample unlocks the full run.
A representative slice of your vertical is screened with the draft ICP. First-pass errors surface fast — sites that match words but not intent.
Your team marks each sampled domain fit or not fit. The disagreements are the signal — they name the definitions to sharpen.
Field definitions and thresholds are adjusted and the sample is re-screened until measured precision clears your bar. Then the full run starts.
In a large-scale production run on our own classification platform, the screen read the travel vertical against the ICP. The client reviewed a sample of the output and confirmed the precision below.
Figures from one anonymized engagement. A vertical of this scale typically yields tens of thousands to hundreds of thousands of qualifying publishers per run (estimate); your own numbers come from your pilot.
A screen that hides its failure modes can't be trusted. Here is where LLM classification is weak, and how the workflow contains each weakness rather than pretending it away.
No screen is. Precision is measured, reported, and tunable — never claimed as perfect. The confidence scores exist precisely so the uncertain cases are visible.
A site can pivot or lapse after screening. Freshness is captured at read time, and scheduled refresh runs catch drift — but a run is a snapshot, not a live feed.
Login walls, heavy JavaScript, and blocked pages limit what any reader sees. Those cases get lower confidence and are flagged rather than guessed at.
The screen surfaces and scores; brand-fit and final approval are human calls. We deliver evidence, not a mandate to email everyone on the list.
The screen is not a magic verdict machine; it is a way to do the reading no human team could do at scale, with the uncertainty made visible. Seen against the alternatives, its role is specific.
A keyword filter can't distinguish intent from vocabulary. The screen reads what a page does, applies your disqualifiers, and removes the false positives a text match keeps.
Hand-checking a million sites is impossible; the screen reads them all consistently, then focuses your team's judgment on the sample and the low-confidence rows that need it.
Brand fit, commercial terms, and final approval stay human. The screen hands your team evidence and confidence, not a decision it never had the context to make.
The free pilot configures LLM classification for your ideal partner, calibrates it on a reviewed sample, and delivers your first 20 qualified publishers with per-field confidence — plus a full-run projection. No cost, no obligation.
Request a Free Pilot