SupplyFinder

    Guide ·

    How to Build a Sitelist for Programmatic Deals

    Before you open a spreadsheet, settle one question, and settle it with whoever owns the brief: is this list here to guarantee quality, or to secure coverage?

    They are two different objects sharing a name. A quality list is narrow, defensible, and exists so you can say exactly where the budget went. A coverage list is broad, built around where the audience actually is, and exists so the campaign delivers. Build one when the brief needed the other and it fails on both counts, too restrictive to hit the numbers and too loose to promise anything.

    Once that's answered, the build is straightforward. Keeping the list alive afterwards is what decides whether the work was worth doing at all.

    The trade-off isn't yours to make alone

    Here's the part that makes this genuinely hard, and it isn't a construction problem.

    The brief asks for both. Premium, defensible environments the client would be happy to see in a report, and volumes that a short list of premium domains cannot deliver. Those two requirements are partially incompatible, and the incompatibility arrives from the advertiser and the media plan, not from you.

    So what happens next is familiar. A quality list gets judged on delivery: someone asks why pacing is behind, the answer is that the list holds 200 domains, and it gets loosened until it no longer guarantees what it existed to guarantee. Or a coverage list gets judged on quality: a client asks why their ad appeared next to something unfortunate, the list gets tightened, delivery drops, and nobody revisits the original purpose.

    Either way the buying team absorbs a trade-off that was never explicitly made, and takes the blame for whichever side gives first.

    The list is the place to surface it. A sitelist is the only artefact in this chain a non-technical stakeholder can actually read. Not a deal ID, not a supply path, not a bid strategy. A list of site names is legible to anyone, which makes it the right object to put in front of a client when the trade-off needs a decision.

    In practice that means showing two: a short list you can fully defend, and a wider one with a stated minimum standard. Here is what we can guarantee at this scale, here is what we can reach if we accept this floor, which do you want. Most advertisers have never been shown the choice, because nobody framed it in terms they could evaluate.

    That conversation moves the decision to where it belongs, and it changes what the buying team is doing: not absorbing an impossible brief, but explaining a real constraint. Write down whichever answer comes back, because it defines the list.

    Building a quality list

    Start from publishers you can name and defend, not from a data export. A quality list is usually short, and its length is a feature.

    Before the checks, a word on how far manual review gets you. Most teams already open a few domains by hand, occasionally, when something looks off. It feels like control and it isn't: a handful of pages inspected out of several hundred gives you an impression, not a standard. Manual review is for calibration. Look closely at fifteen or twenty domains, decide where your thresholds sit, then apply those thresholds to the whole list with data.

    Content substance. Read something. Original reporting, real editorial, an identifiable reason for an audience to visit. This is the check automation handles worst, which is exactly why it belongs in the calibration pass rather than the coverage pass.

    Ad experience. Ad density and refresh rate are the two metrics that most reliably separate a publisher from an arbitrage operation. Both are published at domain level by open data sources, so once your manual pass has told you where the line sits, you can apply it to every domain on the list at once.

    Selling relationship. Read the domain's ads.txt file, public at the root of the site. It lists who is authorised to sell the inventory and marks which relationships are direct. A domain you reach only through resellers is not disqualified, but you should know that's the situation before you commit.

    What a real visitor sees. Open the page the way an actual user would arrive, through a referral rather than by typing the URL. Some properties render differently for direct visits.

    Building a coverage list

    Different starting point: begin from where the audience is, then apply a floor.

    Take the domains your audience actually visits, from research, from past campaign data, from your own analytics. Then remove what fails a minimum standard rather than selecting what passes a high one. The distinction matters, because a coverage list built by selection will always be too small.

    Set the floor explicitly and write it down: a density threshold, a refresh limit, an authorisation requirement. When someone later asks why a particular domain is in the list, "it cleared the floor we agreed" is an answer. "It seemed fine" is not.

    And accept the trade openly. A coverage list will contain domains you wouldn't put in a client deck. That's the deal you made when you chose coverage, and it's worth saying so at the start rather than discovering it in a review.

    What people get wrong

    Treating the list as a deliverable. This is the expensive one. A list is built for a brief, sent, applied, and then abandoned when the campaign ends. Three months later a similar brief arrives and the work starts from zero, minus everything the last campaign taught you about which domains actually performed.

    The list should be the thing that survives, not the campaign. Its value compounds: a domain you've run three times, with three sets of results attached, is worth more to you than a domain you found in an export last week. That accumulated knowledge is the asset, and discarding the list discards it.

    Never removing anything. The mirror image. Lists that only grow eventually contain sites that were sold, redesigned, loaded with ad slots or quietly stopped attracting anyone. Additions get scrutiny, removals get none, so the list decays invisibly.

    Confusing the list with the outcome. A sitelist constrains where ads can serve. It does not guarantee they serve well. Good domains have bad placements, and a clean list applied to poor formats produces poor results.

    Building it once per campaign instead of once per audience. Most agencies run recurring briefs for similar audiences. Organising lists by audience rather than by campaign means each one gets used and improved repeatedly, instead of being rebuilt from scratch each time.

    What to do next

    Put the trade-off in front of whoever owns the brief before you build anything. Two lists, one defensible and one wider with a stated floor, and a straight question about which outcome they want. Write down the answer at the top of the list.

    Then build. Calibrate on fifteen or twenty domains by hand, set your thresholds, and apply them to the rest with open data rather than opening pages one at a time.

    Then give it a home and a review date. After each significant flight, add what performed, remove what didn't, and note why. That habit is what turns a list from a task you redo every quarter into something that gets better each time you use it.

    Questions people actually ask

    What is a sitelist in programmatic advertising?
    A sitelist is a defined set of domains or apps that a campaign is allowed to buy, used to constrain where ads can appear. It can be applied as an inclusion list in a DSP, or used to assemble a curated deal with a supply partner. It differs from a blocklist in that anything not named is excluded by default.
    How do I choose which sites to include?
    Start by deciding whether the list exists to guarantee quality or to secure coverage, because the criteria differ. For quality, look at ad density, refresh behaviour, content substance and whether the publisher sells directly. For coverage, start from where the audience actually is and apply a quality floor rather than a quality ceiling.
    How can I verify a domain before adding it to a sitelist?
    Read the domain's ads.txt file, which is public and sits at the root of the site, to see who is authorised to sell its inventory and which relationships are direct. Check ad density and refresh rate, which open data sources publish at domain level. And open the page through a referral link to see what a real visitor experiences.
    How do I explain to a client why a strict sitelist limits delivery?
    Show two lists rather than explaining a constraint in the abstract. A short list you can fully defend, and a wider one with a stated minimum standard, alongside what each can realistically deliver. A list of site names is readable by anyone, which makes it the one artefact in programmatic a non-technical stakeholder can evaluate, and it turns an argument about pacing into a decision they can make.
    How often should a sitelist be updated?
    Often enough that campaign learnings feed back into it, which in practice means after each significant flight rather than on a fixed calendar. Sites change ownership, redesign, add ad slots or decline in traffic quality, and a list assembled six months ago describes a web that has moved on.
    What is the difference between a sitelist and a blocklist?
    A sitelist names what is allowed and excludes everything else by default. A blocklist names what is forbidden and allows everything else. The first protects quality at the cost of reach, the second protects reach and is always one step behind new domains.

    Related reading