Commercial Evaluation Guide
GEO Agency Services Explained: What a Serious Provider Actually Owns
Evaluate a GEO agency by the decisions it owns, the evidence it produces, what actually ships and how it separates visibility from commercial results.
Brenden, Founder and search operator
7 min read
A GEO agency should own the hard part: turning a messy market, inconsistent evidence and changing search behaviour into a controlled set of decisions your team can implement and measure.
If all it owns is prompt tracking and content production, it is not running the system. It is reporting on one surface and feeding another.
The direct answer
A serious GEO agency should:
- identify the search and AI decisions that matter commercially;
- prove which source classes shape those decisions;
- find the access, fact, page, proof and corroboration gaps;
- choose the right intervention;
- implement it or give an exact, testable instruction to the owner;
- verify what changed;
- retest the same decision set;
- connect visibility to qualified visits, enquiries, pipeline or revenue where the data exists.
The agency cannot own your legal approvals, invent customer proof, force a publisher to include you or guarantee a platform recommendation. It should expose those dependencies instead of hiding them inside a vague retainer.
What the provider should own
Market and decision modelling
The agency should translate your offer into real decision families:
- category and service discovery;
- provider shortlists;
- product or vendor comparisons;
- cost, suitability and risk questions;
- location and serviceability;
- alternatives and switching;
- technical or procurement evaluation;
- branded verification.
Each family needs a defined audience, market and commercial action. Tracking “What is our brand?” alongside “Who should I hire for an enterprise migration?” as though they are equivalent makes the reporting meaningless.
Search and source evidence
The provider should inspect conventional search results, AI answers and the source sets behind them. It should record the system, date, prompt or query, geography limitations and source class.
The output must distinguish:
- a page being available to crawl;
- a page being fetched;
- a brand being mentioned;
- a source being cited;
- a link being shown;
- a person visiting;
- a person converting.
No composite score should erase those states.
Intervention design
For each priority gap, the provider should decide whether the answer is:
- repair technical access;
- clarify an official fact;
- improve an existing page;
- build a missing money page;
- publish a guide, comparison, tool, methodology or proof asset;
- strengthen internal links;
- correct a profile or directory;
- earn independent coverage or corroboration;
- improve the conversion destination;
- do nothing yet because the evidence or authority is insufficient.
“Publish four GEO articles” is not intervention design.
Implementation control
The agency should name:
- the exact URL or external surface;
- the current problem;
- the proposed change;
- the owner;
- the proof or approval required;
- the acceptance test;
- the retest date;
- the expected commercial role.
If your team or another vendor must implement the work, the instruction still needs to be specific enough to verify. A 70-slide deck that leaves your developer guessing is not strategy. It is outsourced ambiguity.
Proof and reporting
The agency should keep a change ledger showing what was recommended, approved, implemented and verified.
Reporting should answer:
- What did the market evidence show?
- What changed because of it?
- What is now crawlable, indexed or represented differently?
- What was fetched, mentioned, cited or linked on retest?
- What qualified visits or commercial actions followed?
- What remains blocked, by whom and why?
That is enough to make decisions. A graph moving from 42 to 48 “AI visibility” is not.
What stays with your team
The provider cannot operate honestly without client ownership.
| Responsibility | Provider role | Your role |
|---|---|---|
| Commercial truth | Structure and test it | Approve services, markets, prices and fit |
| Proof | Identify the missing evidence | Supply and authorise real results, reviews and case studies |
| Claims | Draft within known boundaries | Legal, compliance or executive approval where required |
| Product knowledge | Extract and organise expertise | Provide operators and source material |
| Development | Implement or specify and QA | Provide access, capacity and release approval |
| PR and third parties | Identify source opportunities | Approve outreach, relationships and public statements |
| Measurement | Define and instrument | Approve analytics, CRM and data access |
| Publication | Prepare and verify | Authorise public release |
A mature provider makes these boundaries clearer. A weak one pretends it can replace missing business evidence with better copy.
What should exist after the first operating cycle
Avoid arbitrary promises about how many days everything takes. Access, site complexity, data quality and approval speed change the sequence.
Before the first cycle is considered complete, you should have:
- a defined set of commercially important prompts and queries;
- a dated baseline with source evidence and limitations;
- a source-class map;
- an official-facts and entity contradiction list;
- a prioritised intervention queue;
- exact page and external-surface owners;
- at least one completed intervention;
- proof that the change exists in the final surface;
- a matched retest;
- separate visibility and commercial reporting states;
- an unresolved-dependency list.
The point is not the document count. The point is proving one complete loop.
The evidence packet to demand
For every material recommendation, ask for:
| Evidence field | What good looks like |
|---|---|
| Decision | The exact customer or procurement question |
| Observation | Dated search, answer and source evidence |
| Limitation | Geography, personalisation, retrieval or access caveat |
| Diagnosis | Access, official fact, content, proof, authority or conversion gap |
| Intervention | Exact page, source or workflow change |
| Claim status | What is verified, opinion, estimate or prohibited |
| Owner | One accountable person or team |
| Verification | URL, screenshot, source diff, crawl or platform evidence |
| Retest | Same method, with changed and unchanged states separated |
| Commercial trace | Visit, enquiry, opportunity, sale or honest attribution gap |
This packet prevents a common failure: the provider changes the test, the prompt, the market and the scoring method after the work, then declares improvement.
How Google and OpenAI change the brief
Google’s current generative AI search guide is blunt: normal SEO remains foundational; unique, non-commodity content matters; there is no special GEO schema; and creating pages for every query variation is not the answer.
OpenAI’s publisher guidance says not to block OAI-SearchBot if you want site content eligible for ChatGPT search summaries and snippets. It also says ChatGPT search referral links include a trackable UTM source.
Those official boundaries kill a lot of agency theatre:
- crawler access is not a ranking guarantee;
- schema is not a citation button;
- “AI-ready formatting” cannot rescue commodity content;
- a tracked referral is evidence of a visit, not every appearance;
- third-party monitoring does not reveal platform internals.
The provider should know the difference.
Provider red flags
Do not sign until you can explain away each of these:
- guaranteed citations, rankings or recommendations;
- a proprietary score with no underlying observations;
- no distinction between fetched, mentioned, cited, linked and visited;
- content volume fixed before market research;
- special AI schema presented as the main lever;
- no access to the evidence behind the report;
- anonymous “expert” content with no source lineage;
- fake community, review or directory activity;
- no client claim-approval process;
- recommendations without named implementation owners;
- no matched retest;
- no connection to a commercial destination;
- every solution somehow becoming another blog post.
Match the engagement to the stage
Not every company needs a full managed programme.
| Your current state | Sensible engagement | Bad purchase |
|---|---|---|
| You do not know whether AI-assisted discovery matters | Bounded market and source-set diagnostic | Twelve-month content retainer |
| Your site has crawl and indexation failures | Technical repair with matched search retest | Prompt monitoring alone |
| Your official facts are inconsistent | Entity and commercial-page reconciliation | Mass third-party mentions |
| Your pages are strong but corroboration is weak | Source-class and authority programme | Another onsite glossary |
| You are visible but visits do not convert | Destination, proof and conversion work | More visibility reports |
| You have a mature search team | Specialist research, QA and selected implementation | A provider duplicating internal capability |
You may not need a GEO agency at all if your important decisions do not trigger retrieval, your customers do not use these discovery paths, or the first commercial constraint is product, offer or sales execution. The provider should be willing to say that.
Compare proposals by decision coverage
Two proposals with the same monthly fee can contain radically different work. Compare them against the operating loop:
- How many distinct commercial decisions are in scope?
- Which markets and platforms will be tested?
- Does the fee include implementation or only recommendations?
- Who writes, approves, publishes and verifies each change?
- Is third-party outreach included, or only a list of targets?
- Are developer, PR, media or placement costs separate?
- Does measurement use first-party data where available?
- Is the evidence exportable if the engagement ends?
- What event causes the provider to stop, change course or expand?
Do not compare article counts. Compare how much of the decision-to-evidence-to-change-to-retest loop the provider truly owns.
Three provider questions that expose the operating model
Ask:
- Show me one decision from prompt to source set to shipped change to retest.
- Which outcomes do you control, which do you influence and which can you only observe?
- What will you refuse to claim if the evidence is not there?
The answers tell you more than the service label.
For the full scope checklist, read what generative engine optimisation services should include. If your immediate problem is agency recommendation visibility, use the agency shortlist and comparison guide.
Make the agency prove one complete loop
We will show you the decisions worth winning, the sources shaping them and the work required to move them—without pretending a dashboard is the result.
See Searchmaxxed's AI search optimisation service or show us the market.
Primary sources
- Optimizing for generative AI features on Google Search — Google Search Central.
- Google Search Essentials — Google Search Central.
- Publishers and developers FAQ — OpenAI.
Keep solving the problem
Agency Selection
Compare SEO and AI-search partners by diagnosis, senior ownership, proof, implementation, measurement and commercial fit.
Related resources
Your next move
Fix the layer blocking the result.
We find the first broken layer, fix it and measure what changes.