How To Tell If Your AI SEO Agency Is Working
The Signals That Mean Something Four things are hard to fake and worth watching closely. Your own pages beginning to appear in cited sources, which is directly observable in any assistant that shows citations.
Each individual inconsistency looks trivial. Collectively they prevent a set of mentions from resolving to one confident record, and the symptom is a brand that gets described vaguely or hedged around rather than recommended.
The Types That Rarely Earn Their Keep Elaborate breadcrumb hierarchies, speakable markup, deeply nested item lists and most of the specialised types outside their intended vertical produce little observable difference in how a brand is understood or recommended.
Distinguish between a supplier who is failing and one who is reporting badly, because the remedies differ entirely. Ask for the raw answers and read them yourself before deciding. It is not unusual to find that sound work has been buried under a dashboard nobody understands, and fixing the reporting is far cheaper and less disruptive than replacing a team that is actually doing the job.
The distinction to draw is between flat results with the inputs completed, and flat results with the inputs missing. The first is a category or timing problem and may be worth persisting with. The second is a delivery problem.
Ask what was done, not what happened. If listings were corrected, pages rewritten and outreach attempted, and the numbers are still flat, that is information about the market. If none of it happened, the numbers were never going to move.
What to Do First Run five prompts describing a purchase your best customer would be making, from a signed out session, and see what gets named and cited. Then check whether your product data survives with scripts disabled, and whether your name and identifiers are consistent across every listing you can find.
One check is worth running independently once a quarter, without telling anyone. Take ten prompts from the agreed set, run them yourself in a signed out session, and compare what you find against the most recent report. Broad agreement is reassuring. A consistent gap in the agency's favour is the single most informative finding available to you, and it is not something a report will ever surface.
The defensible position is to spend an hour on it if you like, and to spend the rest of the week on the things every system already reads: accessible pages, accurate Organization markup, consistent identity and content a machine can quote.
Be wary of proposals where the largest line is content production. It is the easiest work to scale, the easiest to bill and the least likely to be the constraint, particularly before a baseline exists. A proposal weighted toward diagnosis, technical fixes and third party corrections is usually cheaper and almost always sequenced better.
Reviews Do Disproportionate Work For products more than for services, review content is the evidence base. Volume matters, recency matters more, and llm seo detail matters most, because a review that describes a specific use gives a model something to match against a specific question.
Where a Real Tension Exists Two places, and they are worth naming honestly rather than pretending everything aligns. The first is the hero section. A large image with six words over it is a legitimate design choice and it gives a machine nothing to work with.
Pricing in this field is unusually opaque, partly because the work is new and partly because the absence of an independent scoreboard makes it hard for a buyer to tell whether they are getting value. That combination invites vague scoping.
And do not let anyone rewrite your entire site in the flat, listicle heavy register that is currently fashionable in this discipline. It reads as machine assembled to human beings, and content that reads that way tends to be treated as low quality by both audiences.
None of them are harmful. They just consume implementation and maintenance time that would achieve more if spent making the Organization markup accurate everywhere, or correcting the directory listing that has your old address on it.
Every usability study for thirty years has said readers scan, look for the relevant section, and want the conclusion before the reasoning. Extraction wants the same thing for different reasons. When somebody claims that writing for machines requires sacrificing readability, they are usually describing keyword stuffing, which is a separate and obsolete practice.
Where Marketplaces Fit Marketplace listings are frequently cited, and they are a mixed blessing. They provide corroboration and structured data you did not have to build, and they put a description of your product in circulation that you only partly control.
What Not to Do in the Name of Legibility Hidden text intended only for machines fails on every axis. It is detectable, it violates most guidelines, and it produces exactly the uniform low quality signal you were trying to avoid.