
Generative Engine Optimization: A Practitioner’s Study of AI Search Visibility
INSTITUTE EDITION: PAIR-CRT-001
SCHOOL: School of Creation · LEVEL: Practitioner · EDITION TYPE: Core
READING TIME: 22 minutes · EVIDENCE LABEL: Evidence-Based
LAST REVIEWED: September 2026
COMPANION POST: Getting Recommended by AI: How Businesses Earn Visibility in ChatGPT, Claude, and AI Search
1. Learning Objectives
By the end of this study, you will be able to:
1. Explain the two ways AI assistants acquire information: model training and live retrieval.
2. Classify AI crawlers by function and configure robots.txt to allow search access while controlling training access.
3. Apply the five layers of the PrimalMogul Recommendation Stack to a business website.
4. Distinguish between non-commodity and commodity content.
5. Design a monthly prompt panel to measure AI share of voice.
6. Evaluate common GEO claims against published platform guidance.
2. Prerequisites and Context
No formal prerequisites. Access to your website’s robots.txt file and Google Search Console is recommended for the exercises.
Customers increasingly ask AI assistants for recommendations instead of scanning search results. Search platforms have responded with new measurement tools.
Google launched generative AI performance reports in Search Console in June 2026, and Bing introduced an AI Performance report earlier in 2026, focused on citation frequency. The discipline of earning visibility in these systems is called generative engine optimization.
3. Key Terms
| Term | Definition |
|---|---|
| Generative Engine Optimization (GEO) | The practice of improving how often and how accurately AI systems mention and cite a source. |
| Answer Engine Optimization (AEO) | A related term for optimizing content to appear as direct answers in AI and search features. |
| Retrieval-Augmented Generation (RAG) | A method in which an AI system searches for current information and uses it to compose an answer. |
| Query Fan-Out | A technique in which an AI system breaks one question into several related searches and combines the results. |
| Crawler | Software that reads web pages automatically on behalf of a search engine or AI company. |
| robots.txt | A text file at the root of a website that tells crawlers which pages they may access. |
| Entity | A distinct, identifiable thing, such as a business, person, or product, that search systems recognize and connect to facts. |
| Citation | A link or attribution to a source within an AI-generated answer. |
| Share of Voice | The percentage of relevant AI answers that mention a given business. |
| Commodity Content | Information widely available across many sources, offering no unique value. |
| Non-Commodity Content | Original information, analysis, data, or experience unavailable elsewhere. |
| AI Overviews / AI Mode | Google’s generative AI features that summarize answers within search results. |
| llms.txt | A proposed file format summarizing a site’s content for AI systems; not used by Google Search. |
4. Core Instruction
4.1 How AI Answers Are Assembled
AI systems draw on two sources:
| Source | How It Works | Owner’s Control |
|---|---|---|
| Training data | Text collected before the model was trained | Low; reflects past web presence |
| Live retrieval | Real-time searches during the conversation (RAG) | High; depends on current accessibility and quality |
Google describes its generative AI features as using retrieval-augmented generation and query fan-out to highlight content from its search index. A single customer question may trigger multiple searches. A page that answers several related questions clearly is more likely to be retrieved across those searches.
Rule: Live retrieval is where most business recommendations originate. Prioritize retrievability over training presence.
4.2 Layer One: Access
Crawler Classification
| Company | Training | Search Indexing | User-Initiated Fetch |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | — | PerplexityBot | Perplexity-User |
| Google-Extended (Gemini training controls) | Googlebot | — |
Each crawler requires its own robots.txt directive. Blocking ClaudeBot does not block Claude-SearchBot or Claude-User.
Anthropic states that all three of its crawlers honor robots.txt. OpenAI and Perplexity note that robots.txt rules may not apply to user-initiated fetchers.
Policy: Allow AI search, block AI training
# Allow AI search and user-requested access
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
# Block AI model training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
# Traditional search
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /Warning: Blocking training crawlers is a legitimate business decision, but it may reduce how well future models know your business. Allowing them increases long-term familiarity. Choose deliberately.
Google’s AI opt-out. Search Console now includes a control that removes a site’s content from Google’s AI features. Sites that opt out receive no traffic or impressions from those features.
4.3 Layer Two: Clarity
AI systems must confidently identify a business before recommending it.
1. Consistent core details: identical name, address, phone, hours, and service area across every listing.
2. A definitive About page: plain statements of what the business does, whom it serves, where, and since when.
3. Google Business Profile and Bing Places: complete, verified, and current.
4. Organization structured data: helpful for entity clarity, though Google states no special schema is required for its AI features.
4.4 Layer Three: Credibility
AI systems weigh third-party evidence.
| Signal | Example |
|---|---|
| Reviews | Detailed customer reviews naming specific services |
| Earned media | Local news features, industry publication quotes |
| Directories | Industry associations, chambers of commerce, licensing boards |
| Partnerships | Mentions on supplier, lender, or partner websites |
Rule: Earn mentions; never manufacture them. Google states that seeking inauthentic mentions to influence its AI features is unlikely to help, because those features rely on the same spam safeguards as core search.
4.5 Layer Four: Content
Non-commodity content is the central content principle. Google’s 2026 guidance emphasizes providing a unique point of view and helpful, people-first content that goes beyond commodity information.
| Commodity (Low Citation Value) | Non-Commodity (High Citation Value) |
|---|---|
| “5 Tips for Choosing an HVAC Company” | “What Furnace Replacement Actually Costs in the Inland Empire: 2026 Price Data From 140 Installations” |
| Generic service descriptions | Documented process, pricing ranges, and warranties |
| Rewritten industry articles | Original data, case studies, and firsthand experience |
Answer-first structure:
- State the direct answer in the first one or two sentences.
- Support it with specific evidence: numbers, sources, examples.
- Address related follow-up questions on the same page.
- Display a visible last-updated date.
Research supports evidence density. In a 2024 academic study, GEO: Generative Engine Optimization (Aggarwal et al., presented at KDD 2024), adding citations, quotations, and statistics to content improved source visibility in generative engine responses by up to roughly 40% in the researchers’ test environment. Results in live systems vary.
4.6 Layer Five: Measurement
| Tool | What It Measures | Limitation |
|---|---|---|
| Search Console generative AI report | Impressions in AI Overviews, AI Mode, and Discover AI features, by page, country, and device | Shows no click or query data |
| Bing Webmaster Tools AI Performance | Citation frequency in Bing AI experiences | Bing ecosystem only |
| Prompt panel | Mentions across ChatGPT, Claude, Perplexity, and Google AI Mode | Manual; results vary by session |
| Analytics referrals | Visits arriving from AI platforms | Undercounts when AI answers without a click |
The Prompt Panel Method:
1. Write 10 questions a real customer would ask, mixing broad and specific.
2. Ask each question in four assistants monthly, in a fresh session.
3. Record: mentioned (yes/no), position in the answer, accuracy of details, and competitors named.
4. Calculate share of voice: mentions Ă· total checks.
4.7 Evaluating GEO Claims
| Claim | Evidence-Based Verdict |
|---|---|
| “You need an llms.txt file.” | Google states its AI features do not require one. Other systems may read it; optional. |
| “Chunk content into small pieces for AI.” | Google states chunking is unnecessary for its AI features. Organize for readers. |
| “Special schema guarantees AI inclusion.” | Google states no special schema is required. Structured data aids entity clarity only. |
| “We guarantee ChatGPT rankings.” | No party controls AI outputs. Treat guarantees as a warning sign. |
| “Block all AI crawlers to protect content.” | Blocking search crawlers can remove the site from AI answers. Separate training and search decisions. |
5. Worked Example
Inland Comfort Services is a fictional business used for illustration. Figures are illustrative.
Situation. An HVAC company in Southern California wants to appear when customers ask AI assistants for local recommendations.
Step 1: Access audit.
The robots.txt file, installed by a previous web developer, blocks all bots except Googlebot. OAI-SearchBot, Claude-SearchBot, and PerplexityBot cannot read the site. Fix: apply the Section 4.2 policy.
Step 2: Clarity audit.
The business appears under three name variations across 14 listings, with two outdated phone numbers. Fix: standardize all listings and publish a definitive About page.
Step 3: Baseline prompt panel.
| Metric | Result |
|---|---|
| Questions | 10 |
| Assistants | 4 |
| Total checks | 40 |
| Mentions | 3 |
| Share of voice | 7.5% |
| Competitor most mentioned | Named in 22 of 40 checks |
Step 4: Competitor analysis.
The leading competitor publishes a financing page with real payment examples, has 600 detailed reviews, and was featured in a regional news story about heat-wave preparedness.
Step 5: 90-day plan.
Publish a furnace and AC replacement cost guide using the company’s own installation data; publish a financing page with real payment examples; request detailed reviews from 60 recent customers; pitch a seasonal maintenance story to local media.
Step 6: Day-90 prompt panel.
| Metric | Baseline | Day 90 |
|---|---|---|
| Mentions | 3 of 40 | 11 of 40 |
| Share of voice | 7.5% | 27.5% |
Lesson: the largest single gain came from Step 1. Content cannot be cited if crawlers cannot read it.
6. Case Study
PrimalMogul’s Search Performance, June–September (Founder Protocol, written in third person)
The founder of PrimalMogul reviewed three months of Google Search Console data for primalmogul.com. Original mogul profiles and rankings dominated performance.
The ranking of the richest Black women entrepreneurs earned 755 clicks from 22,331 impressions, a click-through rate of 3.4%. Original profiles of lesser-known executives such as Samiel “Blacc Sam” Asghedom earned tens of thousands of impressions.
By contrast, generic AI how-to posts on widely covered topics recorded few or no clicks. One post on AI agents for small business recorded 44 impressions and no clicks in the period.
Principles demonstrated:
- Content with original analysis on underserved subjects earned visibility; commodity topics did not.
- Search performance data reveals which content systems treat as unique.
- Measurement must precede strategy.
7. Common Errors
| Error | Why It Happens | Correction |
|---|---|---|
| Blocking AI search crawlers accidentally | Blanket bot-blocking rules or security plugins | Audit robots.txt by crawler function |
| Treating GEO as separate from SEO | Marketing hype around new terms | Build on SEO fundamentals first |
| Publishing commodity content at volume | Volume feels productive | Publish fewer pages with original information |
| Inconsistent business details | Listings created over years by different people | Standardize every listing |
| Buying reviews or mentions | Pressure for fast results | Earn authentic evidence; violations risk penalties |
| Measuring only traffic | Traditional analytics habits | Add impressions, citations, and share of voice |
8. Implementation Protocol
Phase 1: Foundation (Weeks 1–2)
- Audit robots.txt by crawler function and apply a deliberate policy. (1 hour)
- Verify Google Search Console and Bing Webmaster Tools. (1 hour)
- Run the baseline prompt panel. (2 hours)
Phase 2: Clarity and Credibility (Weeks 3–6)
- Standardize business details across all listings. (4 hours)
- Rewrite the About page as a definitive entity statement. (2 hours)
- Launch a review request process for recent customers. (2 hours setup)
Phase 3: Content (Weeks 7–12)
- Identify the five questions customers ask most before buying.
- Publish one non-commodity page per question, answer-first, with original evidence.
- Add visible last-updated dates and review each page quarterly.
Phase 4: Measurement (Monthly)
- Rerun the prompt panel and record share of voice.
- Review the Search Console generative AI report.
- Adjust content priorities toward questions where competitors are mentioned and you are not.
9. Practice Exercise
Deliverable: A completed AI visibility baseline report.
Estimated time: 90 minutes
1. Open yourdomain.com/robots.txt and list every AI crawler rule.
2. Classify each rule as training, search, or user-initiated.
3. Write 10 customer questions.
4. Test each in ChatGPT, Claude, Perplexity, and Google AI Mode.
5. Record mentions, accuracy, and competitors named.
6. Calculate share of voice and identify the top competitor.
Self-check:
1. Every search crawler’s status is confirmed.
2. Questions reflect real customer language.
3. Share of voice is calculated as mentions Ă· total checks.
4. At least one specific content opportunity is identified.
10. Knowledge Check
Q1. Which crawler indexes content for Claude’s search results?
A. ClaudeBot
B. Claude-SearchBot
C. GPTBot
D. Google-Extended
Answer: B. ClaudeBot handles training; Claude-SearchBot handles search indexing.
Q2. A site blocks ClaudeBot only. What happens to Claude-SearchBot?
A. It is also blocked
B. It is unaffected; each crawler needs its own directive
C. It is blocked for 30 days
D. It switches to training
Answer: B.
Q3. According to Google’s 2026 guidance, optimizing for its generative AI features is primarily:
A. A separate discipline requiring llms.txt
B. Still SEO
C. Achieved through special schema
D. Achieved through content chunking
Answer: B.
Q4. A business is mentioned in 6 of 40 prompt panel checks. What is its share of voice?
A. 6%
B. 12%
C. 15%
D. 40%
Answer: C. 6 Ă· 40 = 15%.
Q5. What does the Search Console generative AI report currently not provide?
A. Impressions
B. Page data
C. Click data
D. Country data
Answer: C.
Q6. Which page offers the highest citation value?
A. “Top 10 Tips for Hiring a Plumber”
B. A rewritten industry article
C. Original local pricing data from the company’s own jobs
D. A keyword-rich services list
Answer: C. Original, non-commodity information gives AI something unique to cite.
Q7. An agency guarantees first-position ChatGPT recommendations. The correct response:
A. Sign immediately
B. Treat the guarantee as a warning sign; no party controls AI outputs
C. Request a larger package
D. Block all AI crawlers
Answer: B.
11. Summary
AI assistants answer from training data and live retrieval, and most business recommendations come from retrieval.
The PrimalMogul Recommendation Stack organizes visibility into five layers: access for crawlers, clarity of business identity, third-party credibility, non-commodity content, and ongoing measurement.
Google’s 2026 guidance confirms that optimizing for its AI features remains SEO and dismisses llms.txt, chunking, and special schema as requirements. Monthly prompt panels and Search Console’s generative AI reports provide the evidence needed to direct effort.
12. References and Further Study:
- Google Search Central. Optimizing Your Website for Generative AI Features on Google Search. May 2026, updated July 2026.
- Google Search Central Blog. Introducing Search Generative AI Performance Reports in Search Console. June 3, 2026.
- Anthropic. Crawler documentation: ClaudeBot, Claude-User, Claude-SearchBot.
- OpenAI. Overview of OpenAI crawlers.
- Microsoft Bing. Introducing AI Performance in Bing Webmaster Tools. February 2026.
- Aggarwal, P., et al. GEO: Generative Engine Optimization. Proceedings of KDD 2024.
- PrimalMogul. Internal Google Search Console data, June–September 2026.
Further Study in the Vault:
Mogul Content Traffic Protocol principles
· Choosing a Business Structure
· Starting a Business in the AI Era
This study is educational. AI platforms change their systems frequently; verify current crawler documentation with each provider before changing robots.txt.











