AI CommerceJul 31, 2026

How to Choose a GEO Agency: Four Types and Twelve Questions (2026)

How to select a GEO or LLMO agency. This guide sorts firms into four types by the layers they work on, sets out what engagements cost, gives twelve questions to ask before signing with notes on what good answers sound like, and flags the pitches that conflict with published evidence, plus three extra checks for retail and e-commerce.

Key takeaways

  1. Choose a GEO agency by how far down the stack they work, not by what they claim to specialize in. A proposal that never touches crawling or rendering is building on ground nobody has checked
  2. Lock the metric before you sign. In our own testing, one retailer was mentioned 72 times yet finished first zero times. Report the first number and the engagement looks successful; report the second and it does not
  3. Three common pitches conflict with the published evidence: guaranteed rankings, llms.txt as a centerpiece, and schema markup sold as an AI citation lever

Where this article comes from

We sell a GEO service (Stella LLMO). Writing a selection guide as a vendor means saying so first.

With that on the table, this article does not rank agencies. Lists of seventeen, eighteen and twenty-three firms already exist, and adding to the pile does not help anyone decide. What follows is the decision framework and the question list instead. We place ourselves in one of the four categories toward the end.

The choice comes down to how far down the stack they work

GEO breaks into four layers: being reachable by AI systems, being understood correctly, being chosen on the strength of your evidence, and being measured. Our overview of the discipline covers the model.

Agencies differ in which layers they take on, and proposals rarely make this legible. Everyone writes that they optimize for AI search, so a content shop and a firm that will rewrite your rendering pipeline describe themselves in the same words.

A proposal that never touches layers one and two is building on unchecked ground. If your product pages render price in JavaScript only, another hundred articles will not put that price in front of an AI. Establish this before anything else.

Agencies fall into four types

TypeLayers coveredTypical costBest fitWhat to verify
ContentLayers 2 to 3: being understood and being chosen200,000 to 500,000 yen per monthCompanies with little published material to work fromWho handles the technical layer underneath
TechnicalLayers 1 to 2: being reachable and being understoodAn implementation fee plus a retainerSingle-page apps, dynamically rendered sites, large catalogsWhether they implement or only specify
Measurement toolingLayer 4: measurementTens of thousands of yen per month upwardCompanies that want their current position in numbers firstWhich AI systems are covered, and accuracy in your language
Earned coverageLayer 3: being chosenPer engagement, sometimes performance basedCompanies pursuing branded search and reputationThat the target placements exist and stay maintained

Plenty of firms span more than one type, and some present themselves as full service. The broader the claim, the more each layer deserves separate verification. Given how young this market is, few organizations genuinely hold all four at the same depth.

The type most often overlooked is the technical one. GEO reads like a content discipline, but AI crawlers do not execute JavaScript. GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot each fetch the raw HTML, read whatever text is there, and move on without waiting for rendering or coming back. Google's Gemini is the exception, since it uses Googlebot's rendering infrastructure. Which means you can rank in Google while ChatGPT sees neither your price nor your stock. Whether a firm can fix that varies sharply.

What it costs

The following ranges aggregate the pricing that agencies publish on their own sites in the Japanese market. This is not an industry survey, so treat it as a spread rather than a benchmark.

One-off audits run from roughly 200,000 to 1,000,000 yen, with other tallies putting the floor at zero and the ceiling at 500,000. Monthly retainers run from 200,000 to 1,000,000 yen, with the 200,000 to 300,000 band the most common.

The range is wide because scope varies so much. Content production alone stays cheap; rendering work, product feed cleanup and earned third-party coverage push it up. Align the scope across quotes before comparing the numbers. Two proposals at the same monthly rate can cover entirely different layers.

!
These figures are not a statistical survey

They aggregate what firms choose to publish. Many do not publish at all, so the real distribution is wider. Do not disqualify a proposal on the grounds that it falls outside this range.

Twelve questions to ask before you sign

Written to be usable in the meeting itself. Each carries a note on what a solid answer sounds like.

How far down the stack do you work?

A good answer starts at crawling and rendering. A concerning one keeps returning to content production whenever you ask about technical work

Which metric will you report?

A good answer separates mention rate from recommendation position. A concerning one offers AI-referred traffic alone, which is the last number to move and stays flat early on

When and how do you establish the baseline?

A good answer measures before any work starts. Measuring after the fact removes the comparison entirely

How many prompts, across which AI systems, run how many times?

A good answer gives all three numbers immediately. AI answers vary between runs, so a single pass proves nothing

What is the evidence for this tactic?

A good answer points to a controlled experiment, platform documentation, or their own measurement. "It is standard practice" is not evidence

How do you explain llms.txt and structured data?

A good answer distinguishes tactics by evidence quality. Treating either as the centerpiece is a warning sign, for reasons set out below

How much of the third-party work do you take on?

Placement in comparison media and review sites is out of reach for on-site work alone. If they do not cover it, decide now who will

How does your own site appear in AI search?

This reveals whether the firm practices what it sells. If they cannot produce numbers for themselves, the case for handing them your visibility is thin

What is the term, and what ends it early?

Technical fixes surface within weeks; earned reputation takes months. Check that the term matches the layers being worked

Who owns the deliverables, and what remains afterward?

Measurement sheets, implemented code, placements earned. If nothing stays with you, every switch restarts from zero

Who actually does the work?

Whether the person pitching is the person delivering. If not, ask about the delivery team's experience

If results do not come, what do you try next?

A good answer has a hypothesis and a testing order ready. "We produce more content" is not a plan

You do not need all twelve. If time is short, ask the second and the eighth. Those two alone reveal most of what you need to know about whether a firm works from measurement.

Pitches that conflict with the evidence

Only claims with published counter-evidence appear here.

Guaranteed rankings or guaranteed first place. AI answers depend on each provider's model and search implementation, and no outside party controls them. What can be promised is a higher probability of making the shortlist. We do not offer guarantees either.

llms.txt as the centerpiece. Ahrefs examined 137,000 domains and found that 97% of published llms.txt files had never been requested. Google's guidance for AI features states plainly that creating one is unnecessary. Publishing it costs little, but nothing supports building a program around it.

Schema markup sold as an AI citation lever. In a controlled experiment published in June 2026, Ahrefs added JSON-LD to 1,885 pages and compared them against a 4,000-page control group, finding no significant increase in AI citations. Google's guidance likewise says no special schema.org markup is required for generative AI search. This does not make structured data pointless. It remains necessary for rich results and for consistency with Merchant Center feeds. The purpose is simply different from what the pitch claims.

An SEO proposal with the vocabulary swapped. If the table of contents matches a standard SEO deck, the substance probably has not changed. Look for the AI-specific items: rendering, per-platform sourcing, prompt measurement.

Silence about platform change. This field moves quickly. In March 2026 OpenAI discontinued in-chat checkout and moved to surfacing recommendations that route shoppers to merchant storefronts. Ask what has changed recently and see whether the proposal was built on assumptions from six months ago.

Three extra checks for retail and e-commerce

Selling products adds three questions.

Do they make product data machine readable? Whether price, stock and specifications appear as text in the HTML. This sits outside content production, so it will not appear in a proposal unless you ask.

Can they work with product feeds? Google Merchant Center on one side, OpenAI's merchant feed on the other. Google's own guidance recommends Merchant Center feeds for product visibility in AI answers.

Do they cover reviews and third-party surfaces? What an AI cites is rarely limited to your own site.

What happens when these three go unaddressed is visible in the field. In our audit of nine major Japanese electronics retail brands, Product and Offer structured data was present on product pages at one brand out of nine. That is the state of the largest players, which also means the gap is still open.

The metric you choose changes what the result looks like

Here is why question two matters, measured rather than argued.

Our study of air conditioner purchase consultations collected 75 responses comparing electronics retailers. One major chain was mentioned 72 times, among the highest in the set. In the same study, that chain finished first zero times.

Report that engagement on mention rate and the exposure looks healthy. Report it on first-place recommendation rate and the chain never reaches the front of the shortlist. Same study, same retailer, opposite conclusions.

The remedy differs too. Low mention rate calls for exposure work; low first-place rate calls for stronger reasons to be preferred. Agree on which number you will be shown before the contract starts.

Where in-house ends and outsourcing begins

Nothing requires outsourcing all of it. The layers differ.

Layer one usually fits in-house. Checking robots.txt, verifying indexation and removing JavaScript dependencies are bounded tasks for a team with developers.

Layers three and four benefit most from outside help. Earning placement in comparison media and review sites depends on relationships and stalls when handled part-time. Weekly prompt measurement also survives better outside the day job.

A middle path works well: outsource measurement, keep implementation in-house. Once the current state is visible in numbers, deciding where to intervene is something your own team can do. On a constrained budget this shape tends to return the most.

Where we sit

Across the four types, we span the technical and measurement categories. The focus on retail and e-commerce reflects where we see the largest blockages: machine-readable product data and product feeds.

As practice, we publish recurring studies that put purchase questions to five generative AI services, covering retailer recommendations, hotel recommendations, the accuracy of returns and warranty answers, and machine-readability audits of official sites. Our research library holds them, and every number in this article came from that work.

We do not guarantee outcomes. What we commit to is identifying and removing the factors working against you at each stage of retrieval, comprehension and recommendation.

Frequently asked questions

Should I distrust a firm that quotes below the range?

Not on price alone. A narrower scope is cheaper for a reason. If the number looks low, check whether the technical layers are included, since their absence is usually what makes a quote look cheap.

How long should the initial term be?

Long enough to see the layers being worked actually surface. Judging after a few weeks means cutting before the work reaches the answer.

Can our existing SEO agency handle this?

Often partly, since SEO fundamentals are a precondition for GEO. The question is whether they can take on the AI-specific work: rendering, per-platform sourcing and prompt measurement.

Does location matter?

Almost all of the work is remote. Select on measurement discipline and scope rather than geography.

How long until results appear?

Clearing technical faults in layer one can change answers within weeks once pages are re-fetched. Layer three, the reputational work, runs in months. An initial audit that identifies the blocked layer gives you the timeline.

Is there a trial?

Standalone audits are widely available, and ours is free. Turning the current state into numbers first, then deciding whether ongoing support is warranted, is the order we recommend.

Summary

Walk into the meeting with criteria rather than a shortlist of names. Three of them carry most of the weight.

Establish how far down the stack the engagement goes, because proposals that skip layers one and two are building on unchecked ground. Agree the metric before signing, because mention rate and first-place rate turn the same result into opposite conclusions. Ask for the evidence, because guaranteed rankings, llms.txt centerpieces and schema-as-citation-lever all conflict with what controlled experiments and platform documentation show.

If you need your current position before you can judge any of it, we offer a free AI visibility assessment for retail and e-commerce businesses. Which AI, which questions, and how you compare against competitors. Numbers first, then quotes.