Skip to content
Programmatic SEO

Programmatic SEO agency work. Without the page farm.

Generating ten thousand pages has been trivial for two years. The difficult, valuable, unglamorous half is deciding which two hundred deserve to exist and deleting the rest before anyone sees them.

This site runs 71 pages built exactly that way. You can read them and judge for yourself, which is not something most programmatic SEO services will offer you.

The gate97% deleted
Keyword combinations the data allows12,400
Someone actually searches this3,180
We hold data nobody else has910
Survives the duplicate check470
Worth a human editing it380

Illustrative ratios, not a forecast. The number that matters is the last one, and every gate above it exists to make that number smaller.

What the pages get built from, in order of defensibility

  • Your own data

    The defensible one

  • Public datasets

    Needs an angle

  • Research you run

    Slow, unrepeatable

  • Scraped listings

    Everyone has it

The gate

Six tests, and every one of them deletes pages.

A programmatic SEO strategy is mostly a deletion policy. These run before a template is written, because the cheapest page to remove is the one that was never generated.

  • 01

    Does anybody actually search this?

    The combinatorial explosion is seductive: five hundred cities times forty services is twenty thousand pages. Most of those combinations have never been typed by a human being. Volume data at the long tail is unreliable, so this test is about plausibility as much as it is about a number.

    74% cut
  • 02

    Do we hold something nobody else has?

    If the page can be assembled from the same public source your competitors used, it will read like theirs because it is theirs. The pages that survive are the ones sitting on proprietary data, original research, or a judgement only your team can make.

    71% cut
  • 03

    Is it different from its own siblings?

    This is the one that fails quietly. Two pages built from one template with a noun swapped are near-duplicates no matter how the template is written. Catching that needs a measurement across every pair, not a spot check on three of them.

    48% cut
  • 04

    Is a human going to edit it?

    Not proofread. Edit, with the authority to delete. If nobody on the project has the time or the standing to throw a page away, the gate does not exist and the page count is the only thing anyone will manage.

    19% cut
  • 05

    Does it reach the reader from somewhere?

    A page whose only inbound link is the sitemap is a page nobody visits and a page crawlers deprioritise. Internal linking has to be designed with the template, not bolted on once the pages exist and the graph has already gone flat.

    structural
  • 06

    Would you publish this if it were your only page?

    The blunt version of every test above, and the one worth asking last. It is also the question a quality rater is effectively asking, which makes it less of a moral stance than it sounds.

    the point

The percentages are illustrative of the shape, not a promise about your dataset. What is not illustrative is the direction: almost everything the data permits should not be published.

Read this before you sign anything

What a programmatic SEO agency cannot promise you.

Nowhere else in search is the gap between what is easy to sell and what is good for the client this wide. A programmatic SEO agency is paid to publish pages, and you are trying to acquire customers, and those two goals stop agreeing at about page three hundred.

  1. 01

    Google named the failure mode, and it is this one.

    Scaled content abuse entered Google's spam policies in March 2024, defined as generating many pages primarily to manipulate rankings rather than to help people. That policy was written about the standard version of this work. Anyone selling page volume as the deliverable is selling you the thing the policy describes.

  2. 02

    Page count is not a result and I will not report it as one.

    Twelve thousand pages published is an activity metric that feels like an outcome. The numbers worth reading are how many pages get impressions at all, how many earn a click, and how many the crawler bothered to return to. Most programmes never look at the third one.

  3. 03

    Some of this will be deleted, including pages you paid for.

    A template that works at three hundred pages can stop working at three thousand, and the correct response is pruning rather than pushing. If the plan has no deletion step in it, it is not a plan, it is a publishing schedule.

  4. 04

    It is a poor fit for most businesses, and I will say so early.

    This only works when there is a real dataset and real long-tail demand underneath it. Plenty of companies have neither, and for them a smaller number of genuinely good pages beats anything programmatic. That answer costs me the retainer and it is still the right one.

The version of this that works is slower, smaller and considerably less impressive on a slide. It also does not evaporate on the next spam update.

The work

Six workstreams, and the second one deletes most of your pages.

Notice the order. Data before keywords, one hand-written page before any template, and the checks built before the volume rather than after the first problem.

  1. 01

    Weeks 1–2

    Find the data, then the keywords

    Backwards from how it is usually pitched. What structured, defensible data do you actually hold or can you generate, and does anybody search along its axes? Most programmes start with keyword modifiers and reverse-engineer a dataset to fit, which is how you end up with pages that have nothing to say.

    Data sources on hand1 defensible
    Product usage aggregatesunique
    Pricing across the categorypartial
    Public registry datacommodity
    Competitor listingscommodity

    If every row reads commodity, the honest recommendation is not to do this.

  2. 02

    Weeks 2–4

    Run the gate

    Take every candidate combination through the six tests and expect to delete the overwhelming majority. This is the step that makes the difference between a directory people use and thin content that quietly drags the rest of the domain down with it.

    Gate run97% deleted
    Candidates in12,400
    Real demand3,180
    Defensible data910
    Shipped380

    Illustrative ratios. The last number is the only one worth managing.

  3. 03

    Weeks 3–5

    Design one page properly first

    Write a single page by hand, at full length, as the reference the rest are held against. If a section on it could be pasted onto its neighbour with a noun swapped, that section is wrong and gets rebuilt before anything scales. One good page is the specification.

    Reference page reviewthe bar
    Sections that survive a noun swaprebuild
    Fields only this entity can answerkeep
    A claim the data supportskeep
    Filler introductioncut

    The reference page is written before the template, not extracted from it.

  4. 04

    Weeks 4–8

    Build the template and the graph together

    The template, the link model and the schema get designed as one thing. Related-entity links are chosen on editorial fit rather than by filling a slot, because a link grid that exists to be a link grid is visible to a reader and to a crawler.

    Link graph checkno orphans
    Pages with zero inbound links0
    Pages with only oneflagged
    Reciprocal pairstracked
    Links resolving to a 4040

    Checked against what the page RENDERS, not against the array behind it.

  5. 05

    Ongoing

    Automate the checks, not the writing

    Scripts that measure repetition between every pair of pages, validate every cross-link, verify keyword placement against the built HTML and enforce a paragraph-length ceiling. The generation is the cheap half. The regression net is what lets you keep publishing without the set degrading.

    pre-publish checks4 scripts
    repetition · 13 measurespass
    links · every cross-linkpass
    keywords · built HTMLpass
    crossover · sibling pagespass

    These are the real checks running on this site's own directories.

  6. 06

    Month 3+

    Prune on a schedule

    Pages that never earned an impression after two quarters get improved or removed. This is the least popular part of the work and the one that keeps the set healthy, because a large body of pages nobody reads is a liability rather than an asset.

    Monthly readleading · business

    Leading

    Pages with impressions, crawl rate

    Business

    Assisted signups, qualified sessions

    Shape of a report, not a forecast. Published count is deliberately not on it.

Proof you can check

This site is the case study, and it is one click away.

Most programmatic work is invisible behind an NDA. Two directories on this domain were built with exactly the process above, and you can open any page and judge whether it should exist.

  • 71

    Pages published, across two directories

  • ~8–10k

    Rendered words on each one

  • 13

    Repetition measures in the checking script

  • 0

    Pages with no inbound internal link

Nine review passes ran over those pages before this sentence was written, and each one found defects the automated checks could not see. That is the honest state of the art, not a finished system.

The alternatives

Four ways to build pages at scale. One of them prunes.

You can hire a programmatic SEO company, buy a generator and run it yourself, ask a programmatic SEO consultant for the plan, or write everything by hand. The difference is not the generation step, which is now trivial everywhere.

busyless
Programmatic SEO agency
A pSEO tool
Writing by hand
What gets delivered
A gate, a reference page, a template, the link graph, and the scripts that keep it honest
A page count, usually quoted up front
A generator you still have to feed and edit
Individually written pages, slowly
Attitude to deletion
Built in. Most candidates never ship, and pages get pruned
Against the incentive, since volume is the invoice
Not a feature
Not needed at this scale
Duplicate detection
Measured across every pair of pages, and again per section
Spot-checked, if at all
Left to you
Not an issue
Who edits the output
A human with the authority to throw a page away
Sometimes a junior, sometimes nobody
You
The writer
Proof it works
71 pages of it published on this domain, which you can read
Case studies
Demo
Portfolio
Cost shape
$2,500 diagnosis, then from $5,000/mo to own it
Retainer, often priced per page
Tool subscription plus your time
Per article

Writing by hand is genuinely the right answer more often than this page's existence suggests. If your market is a few hundred buyers searching a dozen terms, seo landing pages at scale will not help you and I will say so in the first fortnight.

Engagement

Find out whether your data supports this. It might not.

The first question is not how many pages we can build, it is whether you hold anything worth building them from. That answer takes two weeks and is occasionally no, which is a genuinely useful outcome for the price.

Start here

Discovery

$2,500one-off · two weeks

A read of your data, the demand underneath it, what your competitors have already built, and an honest verdict on whether this is the right instrument at all. You keep the gate criteria and the plan either way.

  • Which of your datasets are defensible and which are commodity
  • Long-tail demand along the axes the data actually has
  • What competitors have already published, and where the gaps are
  • A verdict, including the possible verdict that you should not do this
Book the diagnosis

Most common

Ownership

From $5,000per month

The gate, the reference page, the template and link graph, the checking scripts, and the pruning schedule. Everything lands in your repository, so the programme keeps working whether or not I am still on it.

  • One hand-written reference page before any template exists
  • Template, link graph and schema designed as one thing
  • Repetition, link and keyword checks running before every publish
  • Quarterly pruning, with a stated deletion policy
Talk it through

Programmatic pages are also unusually good at being cited by AI answer engines, because they answer one specific question in one place. That side has its own page.

Let's build

One call. Real plan, not a pitch.

30 minutes. We talk about what's already working, who owns content today, and whether a fractional Head of Content is actually the right move. If it isn't, I'll say so.

Direct calendar

Book a 30-min intro call

Loading scheduler…

Calendar busy?

Send a note instead.

One sentence on the bottleneck. I'll reply within 24h with a sharper next step.

Or send a note

FAQ

Questions, ahead of time.

  • Is programmatic SEO against Google's guidelines?+
    Not inherently. Google's own guidance has long accepted database-driven pages that genuinely help someone, and directories, listings and comparison sites have worked this way for decades. What entered the spam policies in March 2024 is generating pages primarily to manipulate rankings without regard for whether they help anyone. The distinction is real and it is the whole design problem.
  • How many pages should a programmatic project actually publish?+
    Far fewer than the data allows, and I would be sceptical of anyone naming a number before looking at your dataset. The useful question is how many combinations survive a demand test and a defensibility test, which is usually a small fraction of what is technically generatable. On this site's own directories the answer has been dozens per batch, not thousands.
  • Will these pages hurt the rest of our site?+
    They can, and that risk is the main reason to run a gate. A large set of near-duplicate pages competes with itself, dilutes the crawl budget on a big site, and gives a quality assessment plenty to catch. Pages that each answer a distinct question with distinct information do not carry that problem, which is why the deletion step is load-bearing rather than fastidious.
  • Can AI just write all of these pages?+
    It can produce the draft, and on this site it does. What it cannot do is decide which pages deserve to exist, notice that two of them argue the same thing in different words, or accept the reputational risk of publishing something wrong. Automating the generation and leaving the judgement unautomated is the arrangement that works. Doing the reverse is how page farms happen.
  • What kind of business is this a bad fit for?+
    One with no proprietary dataset and no long-tail demand, which is more companies than the pitch decks suggest. If your market is a few hundred high-value buyers searching a dozen terms, this is the wrong instrument entirely and a small number of excellent pages will beat it. I would rather tell you that in the diagnosis than eight months in.
  • How long until programmatic pages rank?+
    Indexing is the first hurdle and it is not guaranteed at all: a large new set gets crawled selectively, so expect partial indexing for weeks and design the link graph to help it. Meaningful traffic on the long tail usually starts in month three or four, and the shape is a slow accumulation across many pages rather than a jump on any one of them.
  • What does programmatic SEO work cost here?+
    A two-week Discovery is $2,500 and tells you whether your data supports this at all, which is a genuinely possible answer. Ongoing ownership starts at $5,000 a month and covers the design, the gate, the checks and the pruning, with production capacity priced separately if the volume needs it.
  • Do we own the templates and scripts afterwards?+
    Yes. The template, the gate criteria, the link model and the checking scripts live in your repository and stay there. That is deliberate: a programme that only runs while somebody keeps invoicing you is a dependency rather than an asset, and this work is supposed to leave you with something that keeps working.