What Is LLMO (GEO)? A Practical Guide to Getting Recommended by AI [2026]
A complete guide to LLMO (GEO): what it means, how the labels differ, why it matters now, how AI decides what to recommend, five self-checks you can run today, the four layers of implementation, which tactics hold up in controlled experiments, and how to measure results.
Key takeaways
- LLMO means getting your products and your company named when someone asks an AI like ChatGPT for a recommendation. A search results page lists ten links; an AI answer usually names about three. Miss that shortlist and you may as well not exist
- AI answers are far more fixed than people expect. When we measured it, ChatGPT and Gemini each named the same electronics retailer first in 15 out of 15 refrigerator consultations. Waiting for the order to shuffle is not a plan
- The work comes down to three things: get read, get understood correctly, and give the model a reason to pick you. This guide separates the tactics that hold up in controlled experiments from the ones that circulate without evidence
LLMO means preparing to be named as a candidate
Preparing your information so that when an AI like ChatGPT or Gemini builds an answer, your products and your company appear among the options it puts forward. It stands for Large Language Model Optimization, the term that took hold in Japan from around mid-2025. English-speaking markets usually call the same discipline GEO.
The work comes down to three things: getting found by AI, being understood correctly, and giving the model a reason to recommend you. The goal resembles what search engine optimization has always aimed at, but the change of audience swapped out which tactics actually work.
Humans open a page and take in the design and the copy. An LLM reads text and data, summarizes it, and reassembles an answer in its own words. When the reader never sees the page, polishing the visual presentation stops paying off.
There is one more structural difference worth internalizing. A search results page lists ten links, so ranking fifth still earns clicks. But an AI answer usually names about three options. If you are not among them, you are functionally invisible. That all-or-nothing quality is why the topic has become urgent.
GEO, LLMO, AIO, AEO: how do they differ?
The short answer: the four labels describe nearly the same thing. They differ in where they originated and how they spread, while the practical work overlaps substantially.
LLMO (Large Language Model Optimization) is the term that settled in Japan, and Japanese vendors have converged on it. GEO (Generative Engine Optimization) dominates in English and originates from a 2024 research paper. AIO refers to optimizing for AI answer blocks such as AI Overviews, though its definition varies by author. AEO (Answer Engine Optimization) means being cited as the direct answer to a question, and has since receded into a secondary label.
You do not need to be precise about the vocabulary, but it matters when you are searching for information. Look for "GEO" in English sources and "LLMO対策" in Japanese ones to reach the material you want faster. Our piece on how the four terms relate sets out where the labels genuinely diverge. If you arrived via AIO, start with what AIO means; if via AEO, start with AEO (AI Engine Optimization).
Why this matters now
Demand is rising, referral traffic is thinning, and fewer than one company in ten has started. That is where things stand as of July 2026.
Demand-side behavior has already shifted. Generative AI usage in Japan reached 51% as of February 2026, roughly double the 27% recorded a year earlier. Google launched AI Mode in Japanese in September 2025, making AI-summarized answers the default experience rather than an experiment.
Websites are feeling the other side of that shift. Browser log analysis found that the share of Google searches sending a visitor to a site fell to 41.1% in 2025. Close to six in ten searches now end without a single site visit. The contest has moved from competing for clicks to competing for a mention inside the answer. On our own site, AI referrals now account for roughly 5% of all visits.
Corporate response lags well behind. In a December 2025 survey of 1,036 marketers, just 7.82% said they had implemented LLMO measures, while 49.61% reported no interest at all. Among those who had started, 36.5% saw improved search rankings and 34.1% saw more inquiries arriving through AI search. Demand is running ahead of supply, and that gap favors whoever moves first.
How this differs from SEO: from ranking to being in the answer
| Dimension | Traditional SEO | LLMO / GEO |
|---|---|---|
| What comes back | A list of links | One answer, plus product candidates |
| Goal | Rank high in search results | Get referenced inside the AI's answer |
| Who compares | A person reads and compares | An AI reads and compares |
| What gets chosen | Sites and pages | A product, and where to buy it |
| Assets that work | Backlinks, content volume, E-E-A-T | First-party data, third-party mentions, consistency of information |
| Stability | Rankings shift daily | The same question tends to get the same answer |
| Measurement | Search Console, rank trackers | Repeated prompt testing (automated tooling is still immature) |
| Relationship | The foundation. Not something to drop | A layer built on top of SEO |
Gemini, Google's AI Mode, and AI Overviews all generate answers grounded in the Google search index, which means a page Google has not indexed cannot appear in an AI answer either. Google's own guide states that its generative AI features are rooted in core search ranking and quality systems, and that established SEO practice remains relevant.
Everything you built through SEO keeps working as a precondition, and LLMO-specific work stacks on top of it.
Inside the AI, something like a shopping proxy is at work
Being told that "AI searches and then answers" leaves the interesting part invisible. Swap in a human proxy and the picture sharpens immediately. Picture someone asked to "pick me a quiet cordless vacuum, I live alone."
An apartment building means it has to be usable at night; living alone makes weight matter. The unstated constraints get inferred
Noise level, weight, nighttime use, price band. The angle changes with each pass
Manufacturer claims, comparison-site rankings, complaints in reviews. Everything gets cross-checked
Whether you are among those three is what decides the outcome in the AI era
What the AI does internally is close to exactly that. Framed this way, the work to do becomes obvious. The material that proxy needs in order to judge has to exist where search will surface it, in a form that makes judging easy. Strip away the jargon and that is all LLMO is.
The architecture behind an answer

How an AI answer gets built. What you optimize is not the model, but the material it reads
At the center sits the LLM itself, holding knowledge acquired during pre-training. When that stored knowledge is not enough, the model calls the search subsystem. Search does not always run. It is one tool among several, invoked only when the model judges it necessary.
The public web shown along the bottom of the diagram is where the material originates. Two distinct crawlers draw from it: training crawlers (GPTBot and others) feed the next model's training data, and search crawlers (OAI-SearchBot and others) feed the search index. Those two paths are the only ways your information reaches an AI.
There are only two routes in
| Route 1: Training (long term) | Route 2: Search (near term) | |
|---|---|---|
| How it arrives | Training crawlers (GPTBot and others) collect it into the next model's training data | Search crawlers (OAI-SearchBot and others) collect continuously into the search index |
| When it lands | Only when a model is updated. Parameters never change mid-conversation | Days to weeks at best. You can measure a fix and iterate |
| What moves it | Volume of brand mentions, third-party writing, accumulated first-party data | Content answering the decomposed angles, fetchable pages, structured data, product feeds |
| Timescale | Months to years | Days to weeks |
| Where it sits in practice | A long-term accumulation | The main battleground. Start here |
Separating the routes matters because they pay off on completely different timescales. The training route accumulates over months to years and only surfaces when a model is updated. The search route shows the effect of a fix within days or weeks, and lets you measure it again. That is why the search route is where the practical work concentrates.
Blanket-blocking AI crawlers in robots.txt to prevent training will also seal off the search route, effectively withdrawing you from consideration in AI search.
One question becomes many searches
The model does not paste your question into a search box. It decomposes the question into subtopics and issues several queries in parallel.
Google calls this query fan-out and describes AI Mode as running roughly a dozen searches in the time a single search would take. Independent analyses have observed around nine sub-queries per question in AI Mode, and two or three for ChatGPT's standard search.
Treat those as snapshots, not constants. The count depends heavily on the model and its reasoning depth. Reasoning models read results and then search again, so a single question can balloon into dozens of searches, and deep-research modes run into the hundreds. The structure to internalize is decomposition plus iteration, not any fixed number.
For the vacuum example, one sentence splits into noise level in decibels, weight comparisons, nighttime use in apartments, price bands, and review sentiment, all searched in parallel. Three consequences follow.
The AI picks the search terms. What the user typed and what actually gets searched do not match, so optimizing for a single keyword no longer works
Every decomposed angle needs an answer. A page that addresses the headline topic but misses the sub-query angles drops out of consideration
Conversation context changes the queries. The same product draws different sub-queries depending on what was said just before
Does AI read the whole page?
It depends on the platform. Treating them as one behavior leads to the wrong optimization plan.
Google works from passages pulled out of its search index. Its guide to generative AI features describes retrieving relevant pages from the index and then confirming specific information from those pages.
Uses search ranking systems to retrieve relevant pages from the search index, then verifies specific information from the retrieved pages to generate a more reliable response.
ChatGPT behaves more like an agent. OpenAI's web search tool specification defines three distinct actions: searching, opening a page, and searching within an opened page. The sequence closely resembles how a person works through source material.
Claude is the most explicit of the three. Anthropic's web search tool documentation states that with basic web search, every search result is loaded into the context window. The separate web fetch tool retrieves the full text of a page as extracted plain text, and the documentation states plainly that JavaScript-rendered sites are not supported.
Those differences translate directly into different tactics. For Google, a question-shaped heading followed by a short direct answer matches the unit that gets extracted. For ChatGPT and Claude, the whole page is read, so padding and inconsistencies with your other pages both become part of the judgment.
One more thing worth knowing: being read does not mean being cited. OpenAI returns consulted URLs separately from citations and notes that the former usually outnumbers the latter. Measuring by citations alone undercounts your results.
Every platform reads from a different source
This is why optimization cannot be centralized in one place. Each system relies on its own index and its own crawler.
| AI service | Where answers come from | What matters for optimization |
|---|---|---|
| ChatGPT (with search) | Its own crawler OAI-SearchBot plus partner indexes | Make sure OAI-SearchBot is not blocked. Product feeds are application-based |
| Google AI Mode / AI Overviews | The Google search index | Standard Google SEO is a prerequisite. Unindexed pages are never candidates |
| Gemini | Grounding via Google Search | Same as above. Google's assessment carries straight through |
| Perplexity | Its own crawler PerplexityBot | Cites more sources per answer, so comparison and review content gets picked up |
| Microsoft Copilot | The Bing index | Bing Webmaster Tools and IndexNow are the practical levers |
Google surfaces (AI Mode, AI Overviews, Gemini) inherit your existing SEO position directly. ChatGPT and Perplexity operate their own crawlers, so robots.txt and WAF rules need to be verified for each. Copilot depends on the Bing index, which makes Bing Webmaster Tools and IndexNow the practical shortcut. Our comparison of AI search engines covers citation counts and answer accuracy across services, and product recommendations in ChatGPT covers that platform's routes specifically.
Where candidates quietly drop out
The retrieval process builds a query from the question, hits an index, fetches pages, then reads and compares what came back. Each step has conditions that remove you from consideration without any signal.
Pages that are not indexed never appear in the lookup at all. This is upstream of every other problem
Blocked crawlers cause fetch failures. WAF and bot-mitigation rules that reject AI crawlers are common in practice
JavaScript-only content may not be read. Prices and stock status are frequent casualties
Contradictory information gets excluded as unreliable. Specs or prices that disagree across your site, marketplaces, and comparison articles fall into this trap
When a fetch fails, the model does not stop. It simply reads a different site. If your information never arrives, the answer gets assembled from what competitors and third-party media wrote instead. How these conditions play out on product pages is covered, with measurements, in GEO for e-commerce.
Polishing your own site does not supply enough material
Traditional SEO was a contest to lift your own pages. With AI, your pages are not the only thing being read. Comparison sites, reviews, personal blogs and social posts, video, and news coverage all get laid side by side, and what multiple sources agree on becomes the basis for the recommendation.
Return to the proxy analogy and it lands. Nobody handed that request would reach a conclusion by reading the manufacturer's site alone. Cross-checking the manufacturer's claims against what reviewers say is precisely the job.
This is the largest practical departure from SEO. Placement in comparison articles, accumulated reviews, and media mentions become primary components of the strategy rather than optional extras. Where to draw the line between in-house work and outside help is covered in how to choose a GEO agency.
Measurement shows recommendations are already locked in
We run an ongoing study that puts identical shopping questions to five generative AI services and repeats them. In our refrigerator shopping study, when the comparison was limited to electronics retailers, ChatGPT and Gemini each placed K's Denki first in 15 out of 15 runs. Claude put Yamada Denki first in 11 of 15. Answers diverge between models, but within a single model they barely move.
Recommendations also concentrate at the top. In our Okinawa hotel study, 84 properties appeared across 625 recommendation slots, yet the top five alone took 44.2% of them.
Our air conditioner study surfaced a different lesson. One major retailer was mentioned 72 times yet recommended first zero times. The name comes up; the brand is not chosen. Measuring AI visibility purely by whether you were mentioned hides that gap entirely.
Buyer behavior is moving in parallel. In a survey of 400 people who purchased home appliances within the previous six months, 29.5% had used AI while researching, rising to 54.0% once you include those who considered it or wished they had.
Start by finding your position: five self-checks you can run today
Before any optimization work, you can establish how AI currently treats you. No tools required.
Ask the major AI services about your category. In a temporary chat, ask ChatGPT (search on) and Gemini "what do you recommend for [your category]?" three times each. Record whether you appear, in what position, and which sites are cited
Confirm you are indexed. Search "site:yourdomain.com" on Google and check that your key pages appear. Repeat on Bing, which underpins the Copilot route
Open your robots.txt. Visit yourdomain.com/robots.txt and check whether AI crawlers such as GPTBot and OAI-SearchBot are being blocked wholesale
Disable JavaScript and open a product page. Turn JavaScript off in your browser and verify that price, stock, and specifications still render
Check referrals in GA4. Is traffic from chatgpt.com or perplexity.ai being measured? Without configuration, AI referrals hide inside "direct"
Whichever check fails is your starting point. Each of the five maps to one of the layers described next.
How to do LLMO: four layers, bottom up
Individual tactics run into the dozens, but arranged by dependency they collapse into four layers. Investing in an upper layer while a lower one is broken produces nothing. Respecting that order is what determines return on spend.
AI crawlers can arrive at the page and read it
Product and company information resolves consistently to the same thing
You appear in the results of each query the question decomposed into
You measure continuously with prompts and feed it back into the work
The same ground is available in two other shapes: as parallel checks in our 22-item GEO checklist, and as an ordered process in the ten steps for AI search.
Layer 1: Reachable — open up crawling and indexing
If no AI names you at all, the cause is usually here. Audit robots.txt, WAF, and CDN bot-mitigation so AI crawlers can reach product and company pages. Then confirm indexing status on both Google and Bing, registering with Search Console and Bing Webmaster Tools where needed. Bing supports IndexNow for immediate update notifications.
The key is treating crawlers by role rather than as one category.
- Training: GPTBot (OpenAI), Google-Extended (Google), ClaudeBot (Anthropic). Blocking these does not affect visibility in search-based answers
- Search: OAI-SearchBot (ChatGPT search), PerplexityBot, Bingbot (which underpins Copilot). Blocking these removes you from AI search
- User-initiated: ChatGPT-User and similar, which visit on the spot to answer a user's question
Google's AI features (AI Overviews, AI Mode, Gemini) build on the normal search crawl through Googlebot. If you want to opt out of training only, name the training crawlers individually rather than blocking everything. Finally, identify pages where price or stock render only through JavaScript, and move that content server-side.
Layer 2: Understandable — make the information consistent
If your name appears but the details are wrong or outdated, the problem sits here. Align product names, model numbers, prices, availability, and specifications across your own site, the marketplaces you sell on, and your catalogs. Because AI cross-checks multiple sources, contradictions get you dropped as an unreliable candidate.
At the page level, keep one topic per page and give common questions a question-shaped heading with a short answer directly beneath. That is the form that answers fan-out sub-queries at passage level. Structured data work belongs to this layer too. As the next section shows, it is not the main lever for AI citations, but it earns rich results in classic search and serves as useful discipline for making information machine-readable. How to redesign a whole store around this is covered in e-commerce SEO in the AI era.
Layer 3: Selectable — appear in the results of each decomposed query
When you are reachable and understood but the recommendation order still will not move, the question is whether you appear in search results at all.
As the earlier section showed, AI does not search your question as written. It rewrites it into several queries and runs them in parallel. The AI decides the search terms, and they do not match the words the user typed. The results for each of those queries hold your pages and your competitors' alike, and AI reads that whole set as context, summarizes it, and judges from there.
So this layer has two requirements: appearing in the results for each decomposed angle, and having substance on the page that appears. Winning a single keyword no longer does the job.
The numbers bear this out. In a study of 863,000 keywords and 4 million AI Overview URLs published in March 2026, Ahrefs found only 38% of cited pages ranked in the top 10 for the original query. In the same study in July 2025 it was 76%. The other 62% were picked up from the results of a different, decomposed query.
The substance half is how you write. Weave statistics, numbers, and cited sources into the copy, and anchor claims in time ("as of July 2026"). That pattern lifts visibility in generative engines by up to 40%, per the controlled experiment covered in the next section. Beyond that, first-party data that only you can produce — surveys, measurements, statistics — is the strongest asset for becoming the thing that gets cited.
And for the angles where you do not rank, appearing on the third-party pages that do is the only route in. That mechanism is why placement in comparison articles, accumulated reviews, press releases and media mentions works. How AI perceives a brand is discussed under the Share of Model framework, and preparing for the risk of incorrect perceptions hardening belongs to this layer's design as well.
Layer 4: Measurable — track your position and how it moves
Recommendations do not come back identical every time. Unlike a search position, this is not one fixed number, so you read it as a share of recommendations across repeated runs. Design prompts, run them repeatedly, and record mentions, recommendation order, cited URLs and factual accuracy. Since you cannot improve what you cannot measure, this layer is worth starting first rather than last. The method is detailed below.
Which layer is blocking you is not something you can guess. Our service, Stella LLMO, starts by diagnosing exactly that.
Beyond the recommendation — connecting the path to purchase
Separate from the four layers, retail and e-commerce have a continuation, because being recommended and being bought are different milestones.
Routes for delivering product feeds to AI platforms are emerging: Google Merchant Center, OpenAI's product feeds (application-based), and Perplexity's Merchant Program. The setup steps, and how Google's feed differs from ChatGPT's, are in our product feed guide. Beyond those sits agentic commerce, where AI agents complete the purchase. Necessity varies sharply by category and scale, so this is not something every company should start on, but retail and e-commerce should keep it in view.
Which tactics hold up under testing?
Most articles on this subject take the form of "30 tactics you should try." The problem is that the listed tactics are not verified to remotely the same standard. When measures backed by controlled experiments sit alongside measures that merely sound reasonable, allocating a limited budget becomes guesswork. Here is where the published evidence stood as of July 2026.
| Tactic | Evidence status | Source |
|---|---|---|
| Include statistics and cited sources in the copy | Works (up to 40% visibility lift) | Controlled experiment in the GEO paper (KDD 2024) |
| Answer the question directly in the heading | Works | Consistent gap across multiple citation-rate analyses |
| Allow crawling and ensure pages are fetchable | Prerequisite (fail this and you are not a candidate) | Published specifications from each AI platform |
| Mentions in third-party media and comparison articles | Works (a primary driver of recommendations) | Analyses of citation sources in AI answers |
| Adding JSON-LD schema | No incremental effect on AI citations observed | Ahrefs controlled experiment (1,885 pages, June 2026) Google: not a requirement for generative AI search |
| Publishing llms.txt | Almost never read | Ahrefs study (137,000 domains, 97% got zero requests) Google's own guide states it provides no benefit |
The clearest win is weaving statistics, numbers, and cited sources into the content. The GEO paper presented at KDD 2024 tested nine content rewrites across 10,000 queries and found that adding statistics, adding quotations, and citing sources lifted visibility in generative engines by up to 40%.
Structured data calls for more care. In a controlled experiment published by Ahrefs in June 2026, 1,885 pages that newly added JSON-LD were compared against 4,000 control pages, and no significant increase in AI citations appeared. One detail matters for interpretation: the pages studied were already being cited heavily by AI. The correct conclusion is not that schema is worthless, but that adding schema to a page AI already cites will not push it higher. It still earns rich results in classic search.
llms.txt fares worse. Ahrefs examined 137,000 domains and found that 97% of published llms.txt files were never requested once. Google goes further, stating in its official guide that llms.txt provides no benefit.
Measuring LLMO: three tiers
The top tier is AI visibility: mention rate, recommendation rank, citation adoption, and factual accuracy. No official instrument exists here the way Search Console exists for Google. The dependable method today is designing prompts and running them repeatedly. Define roughly twenty questions your target customers would realistically ask, run them across each AI service multiple times, and record mentions, recommendation order, cited URLs, and factual accuracy for both your brand and competitors.
Mix three types of question to see the full picture: recommendation ("what do you recommend for [category]?"), comparison ("which is better, [A] or [B]?"), and branded ("what is [your company] known for?").
Measurement tools such as Profound, Otterly.AI, and Semrush AI Toolkit have appeared, but coverage and methodology vary and the category is still immature. Build a baseline with your own measurements first, then evaluate tooling when scale demands it.
Below AI visibility sit traffic (AI-referred sessions, breakdown by source) and outcomes (inquiries, purchases, AI-referred conversion rate).
They frequently arrive without a referrer and get absorbed into "direct." Without a configuration that isolates AI referrals as their own channel, you will systematically undercount whatever your work achieves.
Logging cited URLs pays a second dividend. You end up with a list of the outlets AI actually reads, which tells you where coverage would influence recommendations. Layer 3 decisions stop being guesswork and start being driven by measurement.
Common questions about LLMO
The opposite. Google's AI surfaces ground their answers in the Google index, which makes SEO a precondition for LLMO. This is a budget allocation question, not a replacement question.
They mean nearly the same thing, so either works. In practice, search "LLMO対策" for Japanese material and "GEO" for English material to reach what you want faster.
Fixing technical faults in Layer 1 can change answers within weeks, once crawling and re-fetching catch up. Building the external evidence in Layer 3 runs on a multi-month timeline. Because the horizon depends on which layer you are working in, the initial diagnosis is what sets expectations.
No. Answers depend on each vendor's model and retrieval implementation, and no outside party can promise placement. What can be committed to is removing the factors that disadvantage you at each stage and raising the probability of making the shortlist. Treat any "guaranteed number one" pitch as a claim requiring evidence.
It varies widely by vendor and scope. Some agencies only produce content; others implement crawl infrastructure and pursue external placement. Rather than comparing prices first, confirm which of the four layers a proposal actually covers. A proposal that never touches Layers 1 and 2 is stacking work on an unbuilt foundation. Published price ranges and the questions to ask before signing are in how to choose a GEO agency.
The cost is low enough to justify as insurance, but it is not a core tactic. A study of 137,000 domains found 97% were never requested, and Google states in its official guide that it provides no benefit. If a vendor puts it at the center of a proposal, ask what evidence they are relying on.
Layer 1 (crawl access, indexing, removing JavaScript dependencies) is usually within reach of an internal development team. Layer 3 is the hard part, since exposure in comparison media and review sites requires relationships outside the company. Measure first, find which layer is blocked, then decide how to staff it.
Arguably more so. With implementation at 7.82%, there is still a window to establish a position before larger competitors move in earnest. Given how rigid the recommendation order proves to be in our measurements, arriving early is an advantage that compounds.
Summary
Three points carry the rest of this guide. Start from the bottom layer and work up: adding content while your pages cannot be fetched accomplishes nothing. Choose tactics by evidence quality: statistics with sources have controlled-experiment support, while schema and llms.txt show no measured lift in AI citations. Measure recommendation order rather than mentions, because a name that appears without being chosen does not produce revenue.
The first practical step is turning your current state into numbers. Which AI, which questions, and how you compare against competitors in each. We offer a free AI visibility assessment for retail and e-commerce businesses. Start by finding out how AI currently sees you.



