# A Guide to Search Discovery - Content Cucumber

> The lenses, weights, and checks in the Content Cucumber Search Discovery Audit, with what earns a Pass and the order of work that returns points fastest.

*Source: [https://contentcucumber.com/search-discovery-guide/](https://contentcucumber.com/search-discovery-guide/)*

*By Content Cucumber. Updated October 2026*

---

**A guide**

The Content Cucumber audit grades seven lenses and applies the same checks to each site it audits. This guide names each check, says what earns a Pass, and orders the work by the points it returns, so you can score your own site before anyone audits it.

[Download the PDF](https://contentcucumber.com/wp-content/uploads/2026/09/aeo-seo-audit-checklist.pdf)

## How the Audit Scores You

The Content Cucumber Search Discovery Audit answers one question. When a buyer looks for what a business does, by searching Google or by asking an AI assistant, does that buyer find the business and choose it? The audit grades seven lenses, publishes the checks inside each lens, and adds the lens scores into one composite out of 100.
This guide covers the whole audit. Each section below names the checks in a lens, says what earns a Pass, and lists the places where sites lose points, so you can score your own site before anyone audits it.
LensWeightThe question it answersCited by AI25Do six AI answer engines name you when a buyer asks?Findable15Can search engines and AI crawlers reach and read your site?Understood15Can a machine tell who you are and what you offer?Substantial15Does your content cover what buyers ask, and is the content current?On-voice10Does your copy read as written by a person, and does it point to a next step?Trusted10Do named people, reviews, and outside sources back your claims?Reddit10When buyers ask other buyers, are you in the thread?**Composite****100**The weighted total of the seven lens scores.How a lens becomes a score
Each check inside a lens grades Pass, Partial, or Fail. A Pass counts as 1, a Partial counts as one half, and a Fail counts as 0. The lens score is the average of its checks, multiplied by 100. A lens with 16 checks that returns 5 Pass, 10 Partial, and 1 Fail adds up to 10 out of 16, which is a score of 62.
The composite multiplies each lens score by its weight and adds the results. A lens scoring 62 at a weight of 15 contributes 9.3 points. A check the audit cannot measure is marked Manual and left out of the average, so a measurement gap on our side does not land on your score as a Fail.
Cited by AI follows its own arithmetic, covered in the next section, because the engines produce the score and a person does not grade it.
What sits beside the score
The audit reports several measurements next to the composite without adding them into it. Performance and platform, search data, local and accessibility checks, Google AI Overviews and AI Mode, and the AI mentions trend each get their own section later in this guide. Fixing those items still helps, because many of them feed a check inside a scored lens, and a buyer feels a slow page whether or not the page changes a composite.

## Cited by AI (25 Points)

Cited by AI carries the heaviest weight in the audit. The audit writes 20 buying questions that a buyer would type, splits them between premium and everyday intent, and includes a few that name the company. Six answer engines then answer the questions with live web search switched on. The engines are OpenAI, Gemini, Perplexity, Copilot, Kimi, and Claude. An engine that cannot search the live web during the run is left out and named in the report. The audit leaves an unmeasured engine out and does not score it as zero.
The five parts of an engine score
Each engine earns a score out of 100 built from five parts. A judge model reads the engine's 20 answers for the first four parts, and the fifth part is counted straight from the answers.
PartPointsWhat earns the pointsBrand Sentiment40How the engine describes you. 40 means the descriptions are positive across the answers, 20 means neutral, and 0 means negative.Brand Recognition20Whether the engine brings you up. A brand the engines do not surface sits near the bottom, and a brand that surfaces across the question set sits near the top.Presence Quality20How rich the mentions are once they happen. The judge looks at depth of the mention, the quality of outside sources behind it, and specifics such as process steps, numbers, and published terms.Market Score10Where the engine places you in the category. 10 means category leader, 5 means mid-pack, and 0 means invisible.Share of Voice10Your mentions divided by your mentions plus your competitors' mentions across the open buying questions. Each percentage point earns one point, and 10 percent of the category mentions earns the full 10.The Cited by AI lens score is the average of the engines that were measured. In Content Cucumber's own audit, the six engines scored 57, 55, 57, 62, 68, and 58, which average 59.5 and round to 60.
How to raise each part
**Sentiment.** The engines describe you from pages they retrieve, so read what the top results say about you. Resolve complaints that sit on review sites, and publish a plain description of what you do and who you serve on a page you control.
**Recognition.** Publish a page for each buying question, with your name inside a self-contained answer. In one audit, the engines retrieved the client's own page twice and still credited a competitor in the answer, because the page answered the question without naming the client.
**Presence quality.** Give the engines specifics to quote, such as your process in numbered steps, your numbers with a source, and your terms in plain words. Outside sources count, so a review platform listing or an industry roundup that names you lifts this part.
**Market score.** Engines place a brand using the category lists they find. Appearing in the roundups and directories that the engines cite moves you from invisible toward mid-pack.
**Share of voice.** This part rewards being named next to competitors on open questions, so the questions without your name in them carry the points. Write for those questions first.
Google AI Overviews and AI Mode
The audit measures Google's AI Overviews on the same buying questions. The report lists who Google cited, in order, whether you appear among the citations, and which domains own the category across the whole question set. Where no AI Overview fires, the audit reports that as a finding, because classic ranking still decides the result on those questions. The report also records your organic position for each question, since a site missing from both the citations and the top ten has a visibility problem that comes before an AI problem.

## Findable (15 Points)

Findable asks whether search engines and AI crawlers can reach your pages and read them. The checks in this lens are one-time fixes, and the fixes keep paying after the work ends.
CheckWhat earns a Passrobots.txt present and readableThe file returns a 200, parses without errors, carries a Sitemap line, and leaves the site, its CSS, and its scripts open to crawlers.AI crawler accessGPTBot, ClaudeBot, PerplexityBot, and Google-Extended each receive the full page with a 200 when the audit tests them live, and no gate or login stands in front of the pages you want cited.Explicit AI crawler policyrobots.txt names the AI crawlers and states your intent for each one.llms.txt/llms.txt returns a 200 and points an engine at the pages that matter, such as what you do, what you offer, and your key policies. /llms-full.txt returns a 200 as well. A site-builder file that lists one link grades Partial.XML sitemapA valid sitemap declared in robots.txt, with lastmod dates that match the dates of edits.Sitemap content qualityThe sitemap lists live, indexable pages and leaves out staging URLs, redirects, placeholder author pages, and query-string duplicates.Server-rendered contentA fetch with JavaScript switched off returns the headline and body copy, blog posts included.IndexabilityNo unintended noindex tags. Gated pages are the exception.Canonical tagsIndexable pages carry a self-referencing canonical that matches the preferred host.HTTPS and redirect hygieneHTTP, www, and non-www versions resolve in one 301 hop to one host, and legacy domains resolve.Security headersStrict-Transport-Security, Content-Security-Policy, X-Content-Type-Options, and Referrer-Policy appear on the homepage response.Mobile page load within crawl budgetThe median of three mobile runs finishes in a time a crawler will wait for, on a page weight that fits the crawl budget.Blog crawlableThe blog archive is reachable from server-rendered HTML and listed in the sitemap.Crawl trapsFaceted navigation and search parameters do not generate endless crawlable link loops.Staging sealedStaging and preview environments block crawlers and stay out of the sitemap.Navigation and link integrityMenus point to live pages, and one path leads to each section of the site.Language tree parityOn a multilingual site, each translated tree mirrors the main tree and declares its alternates.Where sites lose points
A blog loaded through a client-side widget, so a no-JavaScript fetch returns navigation and footer text and no article. GPTBot and ClaudeBot do not execute JavaScript, which leaves the posts invisible to them.
An auto-generated llms.txt that lists the homepage and little else.
Two live blog paths, where the navigation points at one and the posts live at the other, which splits whatever authority the posts earn.
A sitemap that lists pages with no links from the site, template parts that redirect, and uncategorized archives.
Fix these first
Open robots.txt to the AI crawlers you want, and say so by name.
Publish a complete llms.txt and llms-full.txt, then check the output by fetching both files.
Confirm the blog renders without JavaScript.
Clean the sitemap down to live, indexable pages.

## Understood (15 Points)

Understood asks whether a machine can tell who you are and what you sell. Structured data answers that question in a format built for machines.
CheckWhat earns a PassJSON-LD present and validStructured data appears on the templates the audit samples and parses without errors.Organization schemaA sitewide Organization block with name, a logo URL that resolves, a description, and your social profiles in sameAs.WebSite schemaA WebSite node on the homepage, with a SearchAction when the site has search.One coherent entity graphNodes connect through @id values, and the site has no duplicate or conflicting blocks for the same entity.Service, Product, or Offer markupOffer pages carry the type that matches what they sell, with identifiers such as sku, gtin, or mpn on products.aggregateRating substantiatedThe rating sits on a Product or Service, reviewCount is a number, and the rating on the page links to a source a buyer can check.Article or BlogPosting with a named authorEach post carries one consistent article graph that names a Person as author.Person schema for named peopleTeam members and practitioners with bios carry Person markup that includes their credentials.BreadcrumbListInterior pages carry breadcrumbs with position numbers and working URLs.FAQPage markupFAQ markup matches the questions and answers visible on the page.LocalBusiness markupBusinesses with a location publish address, geo, a phone number in international format, and opening hours.Schema hygienehttps contexts, absolute URLs, no deprecated types, and no conflicting duplicates.Titles, meta descriptions, and H1Each indexable page carries a populated title, meta description, and a single H1 with copy in it.Open Graph and Twitter cardsThe core pages carry an Open Graph image, title, and description plus a Twitter card, and the image URL resolves.Where sites lose points
A strong Organization block on the homepage and no markup for the fourteen named people with bios on the team page.
An aggregateRating that stores reviewCount as a comma-formatted string, which strict parsers drop.
Two WebSite nodes describing one site, and a phone number stored without a country code.
A rating of 4.9 across thousands of reviews claimed in the hero and carried by no markup.
Pages that ship no Open Graph image, so a link shared to Messages arrives without a photo.
Fix these first
Publish one Organization node with sameAs links and tie the other nodes to it with @id.
Add Person markup for authors and named team members.
Add Service, Product, or Offer markup to each offer page.
Mark up the FAQ content that already appears on the page.

## Substantial (15 Points)

Substantial asks whether your content covers what buyers ask and whether the content is current. The checks measure depth, coverage, and rhythm with counts taken from your own pages.
CheckWhat earns a PassPublishing cadenceDated posts appear at a steady rhythm, and the visible dates move because edits happened. A bulk lastmod refresh with unchanged visible dates grades Partial.Content freshnessDated posts with visible publication dates exist, and dated pages have been retired or updated.Post depthSampled posts clear a 1,500-word floor for full topic coverage.Page depthCore pages clear a 400-word floor, and service pages describe format, length, and what a buyer receives.Topical coverage of buyer questionsA page exists for each question buyers ask, and the head terms for your category have a page of their own.A page for each named offerEach platform, product line, or location the homepage names has its own page.FAQ-format contentQuestion headings with self-contained answers appear in the page HTML.Original insightPosts carry first-hand material, such as your own tests, named sources, and a stated point of view.Case studiesCase studies name the client, state a measured outcome, and link to the client's site.Comparison contentPages compare you with the alternatives a buyer weighs.Duplicate and thin contentPages share little copy. Shingle overlap near 45 percent grades Partial, and near 90 percent grades Fail.Fact consistency across pagesNumbers, dates, and claims agree wherever they repeat.AuthorshipPosts carry a visible byline and a date.Where sites lose points
Gated pages that serve one template with about 90 percent overlap.
Posts that follow one formula, with unnamed anecdotes and no sources, which reads as machine-scaled writing under one human byline.
A homepage that names a platform with no page for that platform.
A service page with about 100 words per offering and no mention of format, session count, or length.
A blog that published in bursts with gaps of four to seven months.
Fix these first
Write the comparison pages and the per-offer pages that are missing.
Bring core pages above the depth floors.
Add named case studies with measured outcomes.
Retire or update dated posts before writing new ones.

## On-voice (10 Points)

On-voice asks whether your copy reads as written by a person and whether a visitor knows the next step. Engines quote clear passages, and buyers act on clear pages.
CheckWhat earns a PassCopy reads as written by a personThe sentences carry a specific voice, and the blog does not share one formula across posts.Clear value proposition above the foldThe hero states what you sell and who it serves in one sentence.Primary call to actionOne main action is visible and reachable from the pages a buyer visits, and the page behind the action has enough copy to finish the job.Title tagsTitles name the category and read as a value claim, not as a page label.Meta descriptionsEach page carries a unique description, written as a plain sentence.H1 qualityThe H1 names the product and the place, and the post title and H1 match.Answer-shaped passagesParagraphs under question headings run about 134 to 167 words and stand alone, the length engines lift whole.Superlatives come with evidenceClaims such as industry leader link to an award, a citation, or a third-party source.Proof points specific and sourcedNumbers appear with the source beside them.Brand-voice consistencyThe voice holds across the homepage, service pages, and posts.Testimonials read as written by customersTestimonials carry a name, a role, and detail a customer would know.Free of typos and template leftoversPages carry no placeholder copy, broken sentences, or leftover template text.Pricing clarityWhere a buyer expects to find terms, the page states them in plain words.Where sites lose points
A homepage hero built as a brand film with no sentence that says what the company sells.
A homepage title that reads like a page label, such as Company Homepage.
Paragraphs of about 48 words under headings, which stand alone and fall well short of the 134 to 167 word range engines lift.
Industry leader and top-rated claims with no award or source behind them.
Meta descriptions that close with an exclamation point.
Fix these first
Rewrite the hero so one sentence says what you sell and who it serves.
Rewrite titles and H1s so each names the category.
Expand answer paragraphs on the key pages to the 134 to 167 word range.
Attach a source to each number and a name to each testimonial.

## Trusted (10 Points)

Trusted asks whether a buyer or an engine can verify who stands behind the site. The checks reward names, addresses, outside reviews, and claims with a source.
CheckWhat earns a PassNamed authorsPosts carry a visible byline that links to an author page, and the markup names the same person.Author depthAuthor pages carry a bio, credentials, and the author's other work.About pageThe About page tells the company story, states how long you have operated, and is the page engines quote.Address, phone, and contact routeA street address and phone number appear on the site, in PostalAddress markup, and match across the footer and contact page.Credentials as dataTeam pages list names, titles, tenure, and credentials, and the same facts sit in the markup.Customer proofCase studies and testimonials name clients with measured outcomes.Third-party reviewsReviews live on platforms a buyer can verify, such as Google, Clutch, G2, or BBB, and the site links to the source.Claims come with a visible sourceA rating or statistic on the page links to where the number came from.Independent entity recordA Wikidata entry exists for the company.Social profiles machine-declaredProfiles appear in sameAs and agree on the company name and description.Roundup inclusionRoundup lists that the engines cite include the company.Backlink profile healthReferring domains are earned and varied, and the spam score stays low. A ratio of thousands of links from dozens of domains points to sitewide placements.Brand disambiguationSearches for the company name return the company and not a namesake.Legal and policy pagesPrivacy, terms, and return or service policies exist and are linked from the footer.Where sites lose points
1,486 posts with no byline and no author markup, which leaves the content with no named author for an engine to assess.
A rating claimed on the site that no third-party platform confirms.
A street address missing from the site, with a Chicago profile on BBB that the site does not link.
Two different phone numbers across the contact page and the footers.
A leadership page with names and no bios, credentials as suffixes, and no date.
Fix these first
Put a byline, a date, and an author page on each post.
Publish the address and phone number in text and in markup, and match them across the site.
Claim the review profiles, link them from the site, and ask clients for reviews.
Add a source link beside each rating and statistic.

## Reddit (10 Points)

Reddit counts as a scored lens because Google routes buying searches into Reddit threads and the engines cite those same threads. The audit measures three things.
**Search placements.** The audit lists the reddit.com results on page one of the buying searches, records the position of each, and reads the thread title, body, and top-level comments to see whether your brand is named inside.
**Brand mentions.** The audit searches Reddit for your name and your domain across the past 12 months and records up to 25 mentions.
**Engine citations.** The audit counts how often the six engines cite reddit.com in their answers, and whether the threads they cite favor your brand or a competitor.
CheckWhat earns a PassBrand named in first-page Reddit threadsThreads that rank on page one for your buying searches name your brand in the post or the comments.Brand mentions across RedditPeople mention your brand in the past 12 months, in more than a handful of threads.Company participation in threadsSomeone from the company answers questions as a named person and discloses the affiliation.Engine citations of Reddit favor the brandThe Reddit threads that engines cite name you, and do not name competitors alone.Reddit position on your reviews queryA search for your name plus reviews returns threads that read well for you.Community threads outside RedditForums and community sites in your field carry threads that name you.Where sites lose points
Page-one threads that discuss the category and do not name the brand.
Zero brand mentions in a year of Reddit search results.
Engines that lean on Reddit for the category while citing threads that name competitors.
Fix these first
List the threads that rank for your buying searches and read each one.
Answer the question asked in the thread, as a named person from the company, and state the affiliation.
Point to a page on your site when the page answers what the thread asks.

## Performance and Platform (Reported Beside the Score)

The audit measures speed and reports the numbers next to the composite. A slow page costs you visitors whether or not the audit adds the speed into a lens, and the Findable lens includes a check on mobile load within crawl budget.
What the audit measures
Lighthouse runs on the homepage and key templates for the four categories, which are performance, accessibility, best practices, and SEO, on mobile and desktop. Mobile figures are the median of three runs, because a single lab run varied up to fivefold on the same URL.
Google publishes the thresholds for Core Web Vitals on [web.dev](https://web.dev/articles/vitals). A good Largest Contentful Paint finishes within 2.5 seconds, a good Interaction to Next Paint stays under 200 milliseconds, and a good Cumulative Layout Shift stays under 0.1.
The recon step records your platform, how the pages render, your CDN, your analytics, the schema present, and the search-relevant absences.
Where sites lose points
A 28 MB homepage that carried 22.6 MB of hero video, with a mobile load time of 28.3 seconds against 2.9 seconds on desktop. The 2.9 second desktop result hid the problem from the people who work on the site.
Render-blocking scripts in the document head, and images served at a size far above what the layout needs.
Fix these first
Compress or remove autoplay video on mobile.
Serve images in modern formats at the displayed size, and lazy-load the images below the fold.
Defer scripts that do not draw the first screen.
Content Cucumber does not sell performance engineering, so the audit names a development partner for these fixes and says so in the report.

## Search Data (Reported Beside the Score)

The search data section reports how the site performs on Google. The audit pulls the numbers from DataForSEO for the United States, and the numbers reach the report as measured values with the pull date beside them.
**Rankings.** The keywords your domain ranks for, with positions, so the report shows how many sit inside the top 10 and the top 30.
**Keyword demand.** Search volume per month and difficulty for the terms in your category, matched against the pages you own.
**Who ranks on page one.** The competitors in the positions you want.
**Backlinks and authority.** Referring domains, total links, domain rank, and spam score. These numbers also feed the backlink check in the Trusted lens.
**Your brand query.** Your homepage holds position 1 for your own name, and the rest of page one belongs to you or to sources that describe you well.
Where sites lose points
A brand name that ranks at position 2 because a third-party page took position 1.
A domain with hundreds of ranked keywords and one inside the top 30, which is the brand name.
Fix these first
Own the results you can on your brand query, with the homepage, key pages, and social profiles.
Match each high-demand term in your category to a page, and write the page where the page is missing.

## Local and Accessibility Sections

The Local section and the Accessibility section appear when the business or site calls for each. The Local section runs for businesses with a physical location or a service area, and the Accessibility section runs from the Lighthouse accessibility results.
Local
**Google Business Profile.** The profile is verified and carries hours, phone, website, services, a description, and photos uploaded within the past 90 days.
**Name, address, and phone consistency.** The top citation sources show the same details, apart from trivial differences such as Suite and Ste.
**Local schema.** LocalBusiness markup, or the matching subtype, includes address, geo, telephone, and opening hours.
**Reviews.** New reviews arrive about once a month, and the owner answers at least 80 percent of reviews within a week.
Accessibility
The Lighthouse accessibility score and the issues behind it.
Touch targets at least 48 by 48 CSS pixels with room between them.
A viewport tag on each page, with no horizontal scroll on a phone.

## AI Mentions Over Time and Progress

Two sections show movement. The AI mentions section charts a month-by-month trend of how often the answer layer mentions your brand, starting in August 2025, along with the brands the answer layer names for your category and the source domains it cites. The section appears on a first audit.
The progress section appears on a re-audit and compares the new run with a stored baseline. The audit freezes its question set on the first run, so the second run asks the same 20 questions and the comparison holds. Re-run the audit once a month and the trend lines show which fixes moved the engines.

## The 90-Day Order of Work

The order below puts fast, one-time fixes first and the slow-building work last. Inside each window, fix the checks marked Serious before the checks marked Low.
WindowWorkLenses it movesDays 1 to 30Open robots.txt to the AI crawlers by name. Publish llms.txt and llms-full.txt. Clean the sitemap. Confirm server-rendered content. Publish the Organization node with sameAs, add Person and Article markup, and mark up the FAQ content already on your pages. Populate titles, meta descriptions, and H1s.Findable, UnderstoodDays 31 to 60Add bylines, dates, and author pages. Publish the address and phone number in text and markup. Claim review profiles and link them. Rewrite the hero, titles, and calls to action. Attach sources to ratings and statistics.Trusted, On-voice, UnderstoodDays 61 to 90Write the per-offer pages, comparison pages, and named case studies. Bring core pages and posts above the depth floors. Rewrite answer paragraphs to 134 to 167 words. Earn placements in the roundups and directories the engines cite. Start answering threads on Reddit as a named person.Substantial, Cited by AI, RedditCited by AI carries 25 points and responds to the whole stack, because the engines read your pages and the outside sources that describe you. Work in the first two windows lays the ground, and the third window supplies the pages and mentions the engines retrieve.

## Score Yourself Before We Do

Print the worksheet, mark each check Pass, Partial, or Fail, and run the arithmetic from the first section. Each Pass has a stated condition, which keeps two people who score the same site close to the same answers.
For the AI half of the score, ask the engines yourself. Open ChatGPT, Gemini, Perplexity, and Claude, ask the engines the questions buyers type, without naming your company, and record who each engine names. Our [free AI Visibility Check](https://contentcucumber.com/ai-search/llm-discovery/) runs that test across the engines for you, and the [56-point checklist](https://contentcucumber.com/free-56-point-aeo-seo-audit-checklist/) covers the site-side checks in a printable format.

[Get your AI Visibility Check](https://contentcucumber.com/ai-search/llm-discovery/)

## Questions About the Audit

### What is the Search Discovery Audit?

The Search Discovery Audit is the report Content Cucumber runs to measure whether buyers can find and choose a business, through Google search and through six AI answer engines. The report grades seven lenses and adds them into a composite score out of 100.

### How is the composite score calculated?

Each check inside a lens grades Pass (1), Partial (0.5), or Fail (0), and the lens score is the average of its checks times 100. The composite multiplies each lens score by its weight and adds the results. The weights are Cited by AI 25, Findable 15, Understood 15, Substantial 15, On-voice 10, Trusted 10, and Reddit 10.

### Can a site score 100?

A 100 requires a Pass on each check and full marks from the six engines. Across the 20 scored audits we have run, composites range from 40 to 72, so the top of the scale leaves room to grow for each site.

### Which AI engines does the audit query?

The audit queries OpenAI, Gemini, Perplexity, Copilot, Kimi, and Claude, each with live web search. An engine that cannot search during the run is named as unmeasured and left out of the average.

### Does the audit score Core Web Vitals?

The audit reports Core Web Vitals and the four Lighthouse categories next to the composite, and the Findable lens includes a check on mobile load within crawl budget. The numbers matter to buyers whether or not they change a lens.

### Which fixes return points first?

Start with the one-time fixes in Findable and Understood, such as AI crawler access, llms.txt, server-rendered content, and Organization markup. Then add the trust and voice work, and finish with the content depth and outside mentions that move Cited by AI.

### Is llms.txt a ranking factor?

llms.txt is an emerging standard and has no proven ranking effect. The audit grades llms.txt because publishing llms.txt costs little and gives an engine a clear list of the pages that describe your business.

### How do I get the audit run on my site?

Request the free AI Visibility Check for a first look at the engine results, or ask Content Cucumber to run the full Search Discovery Audit. Both start from the form on the AI Visibility Check page.
