# The Locafy Search Swipe File

**57 tactics for ranking in Google and in AI answers.**

Live, always-current version: https://www.locafy.com/seo-swipe-file

Everything in [BRACKETS] is yours to replace. Everything in a code block is
meant to be copied and run as-is.

---

## 01 · AI SEARCH PROMPTS

Run these in ChatGPT, Claude, Perplexity, Gemini or Copilot. The first five tell you where you stand. The last eight tell you what to change. Replace every [BRACKETED] value before you run one.

### 01 — The Recommendation Test

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** This is the query that replaces the map pack. If your name isn't in the ten, nothing else on this page matters yet. Run it monthly and log the list.

```
I need a [SERVICE] in [CITY, STATE]. Recommend the 10 best options, ranked, with one line on why each one made the list. Then tell me which sources you used for each recommendation.
```

### 02 — The Four-Engine Sweep

*Where: ChatGPT, Perplexity, Gemini and Copilot — same day*

**Why it works:** Being named by one engine is luck. Being named by four is an entity the whole index agrees on. Whoever appears in all four is your real competition, whatever the map pack says.

```
Who are the top [SERVICE] providers in [CITY]? Give me 10, with a one-line reason for each.

Run that identical prompt in ChatGPT, Perplexity, Gemini and Copilot on the same day. Log every brand named and count how many of the four named each one. Anything at 4/4 is an established entity. Anything at 1/4 got lucky.
```

### 03 — The Source Audit

*Where: Perplexity, or ChatGPT with search on*

**Why it works:** AI answers are assembled from a handful of pages. This hands you the exact list — usually three directories, one listicle and one competitor blog post. That list is your placement plan for the quarter.

```
For the answer you just gave, list every URL you used, ordered by how much each one influenced the answer. For each URL tell me in one line what information it supplied and which claim in your answer came from it.
```

### 04 — The Objection Prompt

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Every model holds a compressed summary of the worst thing ever written about you. You cannot fix what you have not read — and the source is almost always one review thread or one forum post you've never seen.

```
What are the risks, complaints or downsides of using [BRAND]? Answer honestly rather than diplomatically, and cite where each claim came from.
```

### 05 — The Entity Test

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** If the answer is vague, hedged, or confuses you with a similarly named business, you don't have a future entity problem — you have one now. The fix is Organization schema, sameAs links, and one consistent name, address and phone number everywhere.

```
What do you know about [BRAND]? Tell me what they do, where they operate, who runs them, when they were founded, and what they are best known for. If you are unsure about any part of it, say so explicitly rather than guessing.
```

### 06 — The Head-to-Head

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Shows you the differentiators the model has absorbed about both of you. Whatever it praises about the competitor is, almost always, the section missing from your site.

```
Compare [BRAND] and [COMPETITOR] for someone who needs [JOB TO BE DONE]. Which would you recommend, and why? Be specific about what each one does better, and tell me what evidence you're basing that on.
```

### 07 — The Citation Gap

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Turns "build more authority" into a to-do list of fifteen named domains. Work down it. Most are free directories or local publications that take submissions.

```
When you recommend [SERVICE] businesses in [CITY], which websites do you draw from most often? List the top 15 domains, and for each one tell me what kind of listing or article it publishes and whether a business can get listed on it.
```

### 08 — The Question Fan-Out

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Google's AI Mode fans one query out into dozens of sub-queries and answers from whichever page covers them. This gives you the fan-out before Google runs it, and those twenty questions become your H2s.

```
Someone searches "[HEAD QUERY]". List the 20 follow-up questions they ask next, ordered by how commonly they'd come up. For each one, write the single sentence you'd answer it with.
```

### 09 — The Passage Rewrite

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** AI systems quote passages, not pages. A self-contained 40-70 word chunk sitting under a descriptive heading is the unit that gets lifted into an answer — everything else is context the retriever throws away.

```
Here is a page from my site:

[PASTE THE FULL PAGE TEXT]

Rewrite the five passages most likely to be quoted verbatim by an AI assistant. Rules for every passage:
- 40 to 70 words, and it must make complete sense read on its own
- lead with the direct answer, not with setup
- no pronoun that points at something outside the passage
- include the specific number, name or date that makes it worth citing
- give each one the H2 or H3 heading it should sit under
```

### 10 — The Schema Writer

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Invented values are what make AI-written schema dangerous — a hallucinated aggregateRating is a manual action waiting to happen. This prompt forces the gaps to be explicit instead.

```
Here is my page: [URL]

[PASTE THE PAGE TEXT]

Return complete JSON-LD for the most appropriate schema.org type. Fill in only properties you can verify from the content I gave you. Anything you cannot verify, output as "TODO" rather than inventing a plausible value. Then list, separately, the properties Google actually reads for this type that my content doesn't yet supply.
```

### 11 — The Crawler's-Eye View

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** GPTBot, ClaudeBot and PerplexityBot don't reliably run JavaScript. If your headline, pricing or service list is client-rendered, this prompt shows you the near-blank page they actually see.

```
Fetch [URL] and quote the first 300 words you can actually read, exactly as they appear to you. Then tell me what's visibly missing compared with what a human sees in a browser — headings, pricing, service lists, contact details, reviews.
```

### 12 — The Review Summary Test

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** This is what an assistant tells a buyer before they ever reach your site. The repeated phrases are also the keywords your own reviewers hand you for free — use their words on the page.

```
Summarise what customers say about [BRAND] — the praise, the complaints, and the phrases that come up repeatedly. Tell me which source each theme came from, and quote three actual review phrases you're basing it on.
```

### 13 — The Answer You Want To Win

*Where: ChatGPT, Claude, Perplexity, Gemini or Copilot*

**Why it works:** Reverse-engineers the target instead of hoping. Put that answer near the top of the matching page, in those words, and you're competing for the box rather than waiting to be noticed.

```
Write the 40-word answer that would deserve to win the AI Overview for "[QUERY]". Then list the five specific facts, numbers or named entities that answer has to contain before a search engine would trust it.
```

---

## 02 · GOOGLE SEARCH OPERATORS

Paste these straight into the Google search bar. No tool, no login, no crawl budget. Swap example.com for your domain — and run the last four against a competitor's while you're there.

### 14 — The true index count

*Where: The Google search bar*

**Why it works:** The number Google returns is an estimate, but the gap between it and your sitemap count is not. Ten times more indexed URLs than sitemap URLs means index bloat. Ten times fewer means an indexing problem.

```
site:example.com
```

### 15 — Cannibalisation, in one query

*Where: The Google search bar*

**Why it works:** If four of your pages come back for a term one page should own, Google is being asked to choose — and it's choosing badly. Consolidate them or differentiate them, but stop making it decide.

```
site:example.com "target keyword"
```

### 16 — Which pages actually claim the keyword

*Where: The Google search bar*

**Why it works:** Narrows the same check to pages that put the term in their title. These are the ones genuinely competing with each other, as opposed to merely mentioning it.

```
site:example.com intitle:"target keyword"
```

### 17 — Parameter URLs in the index

*Where: The Google search bar*

**Why it works:** Filters, sort orders, session ids and tracking tags indexed as separate pages. Each one splits ranking signals away from the clean URL it was supposed to be a variant of.

```
site:example.com inurl:?
```

### 18 — The orphan PDFs

*Where: The Google search bar*

**Why it works:** PDFs quietly rank for your terms and dead-end the visitor with no navigation, no CTA and no internal links. Give each one an HTML equivalent, or noindex it via X-Robots-Tag.

```
site:example.com filetype:pdf
```

### 19 — Thin archive pages

*Where: The Google search bar*

**Why it works:** Auto-generated tag and category archives are the most common source of thin indexed content on any CMS. Count them. If they outnumber your real pages, that's where the crawl budget went.

```
site:example.com inurl:tag OR inurl:category
```

### 20 — What Google thinks is fresh

*Where: The Google search bar*

**Why it works:** If nothing from this year comes back, either your lastmod dates are lying or your new content isn't being picked up. Both are worth knowing before you publish another post.

```
site:example.com/blog after:2026-01-01
```

### 21 — Who copied you

*Where: The Google search bar*

**Why it works:** Finds every scraper, every syndication partner who forgot the canonical, and every AI-spun rewrite of your work — and tells you whether Google still credits you as the original.

```
"an exact sentence lifted from your page"
```

### 22 — The unlinked-mention list

*Where: The Google search bar*

**Why it works:** Every result is a site that already talks about you and hasn't linked. It's the highest-converting outreach that exists, because the relationship is already there and the ask is one sentence.

```
"Brand Name" -site:example.com
```

### 23 — Your citation footprint, warts and all

*Where: The Google search bar*

**Why it works:** Old addresses and disconnected numbers still sitting on directories are the quiet reason your NAP consistency score is low and your Maps ranking is stuck. Search the old phone number too.

```
"Brand Name" "555-123-4567" -site:example.com
```

### 24 — Google's own competitor set

*Where: The Google search bar*

**Why it works:** Frequently not the list you'd have written — and the surprises are usually the sites eating your AI citations while you benchmark against the wrong people.

```
related:competitor.com
```

### 25 — The placement list

*Where: The Google search bar*

**Why it works:** The fastest way to build a link prospect list in any niche. Add a city in quotes and it becomes a local link list instead of a generic one.

```
[topic] intitle:"write for us" OR inurl:"guest-post"
```

### 26 — Trust sources nobody bothers with

*Where: The Google search bar*

**Why it works:** Chamber directories, city business registries and university resource lists still carry disproportionate weight for a local entity, and your competitors almost certainly haven't asked.

```
[topic] "[city]" site:.gov OR site:.edu
```

### 27 — Legacy HTTP still in the index

*Where: The Google search bar*

**Why it works:** Every result is a duplicate of its HTTPS twin and a redirect somebody never wrote. Cheap to find, cheap to fix, and it's been splitting signals for years.

```
site:example.com -inurl:https
```

### 28 — Proximity search

*Where: The Google search bar*

**Why it works:** Finds pages where two concepts are discussed together rather than merely both present somewhere on a long page. It's how you separate genuinely on-topic sources from keyword coincidences.

```
"[keyword one]" AROUND(4) "[keyword two]"
```

---

## 03 · SCHEMA TYPES

JSON-LD in the head of the page. Fourteen types, what each one is for, and the property people forget every time. Validate everything at validator.schema.org before it ships.

### 29 — LocalBusiness — but use the subtype

*Where: JSON-LD in the page <head>*

**Why it works:** The generic type says you're a business. The subtype says what kind, and that's what feeds category matching in Maps and in AI answers. The most-skipped property is areaServed.

```
"@type": "Plumber"

Not "LocalBusiness". Use the most specific subtype that exists: Dentist, Electrician, Attorney, HVACBusiness, RoofingContractor, MedicalBusiness, AutoRepair, MovingCompany, RealEstateAgent, Locksmith, ChildCare, Notary.
Then add: areaServed, priceRange, openingHoursSpecification, hasMap, geo.
```

### 30 — Organization — the entity anchor

*Where: JSON-LD in the page <head>*

**Why it works:** sameAs is how a machine confirms that three mentions of your name are the same company. Without it you're a string. With it you're an entity, and entities are what get recommended.

```
"sameAs": [
  "https://www.linkedin.com/company/your-company",
  "https://www.wikidata.org/wiki/Q00000000",
  "https://www.facebook.com/yourcompany",
  "https://www.youtube.com/@yourcompany",
  "https://www.crunchbase.com/organization/your-company"
]
```

### 31 — Service — one per thing you sell

*Where: JSON-LD in the page <head>*

**Why it works:** Turns a services page from prose into a machine-readable list of what you sell and where. Point provider at your Organization's @id rather than repeating the whole block.

```
"@type": "Service",
"serviceType": "Emergency Drain Cleaning",
"provider": { "@id": "https://www.example.com/#organization" },
"areaServed": [{ "@type": "City", "name": "St. Louis" }],
"hasOfferCatalog": { "@type": "OfferCatalog", "itemListElement": [ ... ] }
```

### 32 — Product + Offer

*Where: JSON-LD in the page <head>*

**Why it works:** Without availability you get no rich result at all, and without priceValidUntil merchant listings throw a warning. Both are one line each and both are usually missing.

```
"offers": {
  "@type": "Offer",
  "price": "149.00",
  "priceCurrency": "USD",
  "availability": "https://schema.org/InStock",
  "priceValidUntil": "2026-12-31",
  "url": "https://www.example.com/product/"
}
```

### 33 — FAQPage

*Where: JSON-LD in the page <head>*

**Why it works:** Google pulled the FAQ rich result for most sites, and it's still worth shipping — it's the cleanest way to hand an assistant a question paired with its answer in a fixed structure. Optimise it for the answer engine, not the blue link.

```
"@type": "FAQPage",
"mainEntity": [{
  "@type": "Question",
  "name": "How long does a drain cleaning take?",
  "acceptedAnswer": { "@type": "Answer", "text": "Most residential drain cleaning takes 45 to 90 minutes..." }
}]
```

### 34 — HowTo

*Where: JSON-LD in the page <head>*

**Why it works:** Same story as FAQPage — deprecated as a Google rich result, still read as structured content by assistants. Use it only where there's a genuine step sequence, never to dress up a list.

```
"@type": "HowTo",
"name": "How to shut off your main water valve",
"totalTime": "PT5M",
"step": [{ "@type": "HowToStep", "position": 1, "name": "...", "text": "..." }]
```

### 35 — Article / BlogPosting

*Where: JSON-LD in the page <head>*

**Why it works:** The author link is the E-E-A-T hook. An author with no Person entity behind them is a name, not a credential — and dateModified has to be true, not the date of your last deploy.

```
"@type": "BlogPosting",
"author": { "@id": "https://www.example.com/authors/jane-doe#person" },
"publisher": { "@id": "https://www.example.com/#organization" },
"datePublished": "2026-03-14",
"dateModified": "2026-08-02"
```

### 36 — Person — make the author real

*Where: JSON-LD in the page <head>*

**Why it works:** knowsAbout is the most under-used property in local SEO. It states, in machine-readable terms, what this human is qualified to talk about — which is exactly the question E-E-A-T is asking.

```
"@type": "Person",
"@id": "https://www.example.com/authors/jane-doe#person",
"jobTitle": "Master Plumber, License #12345",
"worksFor": { "@id": "https://www.example.com/#organization" },
"knowsAbout": ["Backflow prevention", "Sewer lateral repair", "Water heater installation"],
"sameAs": ["https://www.linkedin.com/in/janedoe"]
```

### 37 — BreadcrumbList

*Where: JSON-LD in the page <head>*

**Why it works:** Replaces the raw URL in the SERP with a readable path, and gives crawlers an explicit hierarchy that your internal links only imply. Two minutes, sitewide, permanent.

```
"@type": "BreadcrumbList",
"itemListElement": [
  { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://www.example.com/" },
  { "@type": "ListItem", "position": 2, "name": "Services", "item": "https://www.example.com/services/" }
]
```

### 38 — Review / AggregateRating — carefully

*Where: JSON-LD in the page <head>*

**Why it works:** Mark up only reviews collected on your own property, about the item on the page. Self-serving markup on your own Organization has been ignored since 2019 and is a manual-action risk if you push it.

```
"@type": "AggregateRating",
"ratingValue": "4.8",
"reviewCount": "127",
"itemReviewed": { "@type": "Service", "name": "Drain Cleaning" }
```

### 39 — Event

*Where: JSON-LD in the page <head>*

**Why it works:** Still one of the few types with a live, high-visibility rich result — and local businesses almost never use it. A workshop, an open day, a seasonal clinic all qualify.

```
"@type": "Event",
"startDate": "2026-10-04T09:00-05:00",
"eventAttendanceMode": "https://schema.org/OfflineEventAttendanceMode",
"location": { "@type": "Place", "name": "...", "address": { ... } },
"offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" }
```

### 40 — VideoObject

*Where: JSON-LD in the page <head>*

**Why it works:** Video results are their own SERP surface with their own competition. Without this markup an embedded video is invisible to it — you get the page view and none of the reach.

```
"@type": "VideoObject",
"name": "...",
"thumbnailUrl": ["https://www.example.com/thumb.jpg"],
"uploadDate": "2026-05-12T08:00:00-05:00",
"duration": "PT4M32S",
"contentUrl": "https://www.example.com/video.mp4"
```

### 41 — WebSite + SearchAction

*Where: JSON-LD in the page <head>*

**Why it works:** Declares your site's canonical name — which is what shows in the SERP breadcrumb line and in AI attributions — and makes you eligible for the sitelinks search box.

```
"@type": "WebSite",
"@id": "https://www.example.com/#website",
"name": "Example",
"publisher": { "@id": "https://www.example.com/#organization" },
"potentialAction": {
  "@type": "SearchAction",
  "target": "https://www.example.com/search?q={search_term_string}",
  "query-input": "required name=search_term_string"
}
```

### 42 — @id and @graph — the connecting move

*Where: JSON-LD in the page <head>*

**Why it works:** This is what separates schema that validates from schema that builds a knowledge graph. Repeating an unlinked Organization block on 200 pages creates 200 organizations. One @id referenced 200 times creates one entity with 200 mentions.

```
{
  "@context": "https://schema.org",
  "@graph": [
    { "@type": "Organization", "@id": "https://www.example.com/#organization", ... },
    { "@type": "WebSite", "@id": "https://www.example.com/#website", "publisher": { "@id": "https://www.example.com/#organization" } },
    { "@type": "WebPage", "@id": "https://www.example.com/services/#webpage", "isPartOf": { "@id": "https://www.example.com/#website" } }
  ]
}
```

---

## 04 · TECHNICAL FIXES

None of these are clever. All of them are common, and most sites have at least six. Work down the list in order — the first four are worth more than the rest combined on a site that has never had them checked.

### 43 — A self-referencing canonical on every indexable URL

*Where: Your site, your server, your CMS*

**Why it works:** A missing canonical lets Google pick one for you, and it picks the URL with the most links — which is often the parameterised one. Ship the tag even when there's only one version. Especially then.

```
<link rel="canonical" href="https://www.example.com/this-exact-url/" />
```

### 44 — One hostname, one protocol, one trailing-slash rule

*Where: Your site, your server, your CMS*

**Why it works:** example.com, www.example.com, the http:// versions and the trailing-slash variants are up to eight sites to a crawler. This splits more authority than any other single issue, and almost nobody checks it after launch.

```
Pick one canonical hostname, then 301 every other form to it in a single hop:

  http://example.com/*      -> https://www.example.com/:splat   301
  https://example.com/*     -> https://www.example.com/:splat   301
  http://www.example.com/*  -> https://www.example.com/:splat   301

Then make every internal link, canonical tag and sitemap entry point at the winner.
```

### 45 — Kill the redirect chains

*Where: Your site, your server, your CMS*

**Why it works:** Every hop is latency and a small loss of signal, and chains beyond about five stop being followed at all. Rewrite each chain so the first URL points straight at the final destination.

```
curl -sIL https://example.com/old-page | grep -E "^HTTP|^[Ll]ocation"
```

### 46 — 301 the 404s that still have backlinks

*Where: Your site, your server, your CMS*

**Why it works:** A 404 with backlinks is authority in a bucket with a hole in it. Redirect each one to the closest live equivalent — not to the homepage, which Google treats as a soft 404 and drops.

```
Search Console -> Pages -> Not found (404), and Ahrefs -> Best by Links -> filter 404.
For every 404 with referring domains, 301 it to the closest live equivalent page.
Homepage redirects don't count: Google reads a mass redirect to / as a soft 404.
```

### 47 — Fix the soft 404s

*Where: Your site, your server, your CMS*

**Why it works:** An empty results page or a "no products found" template returning HTTP 200 teaches Google that your 200 responses are unreliable. That distrust doesn't stay contained to those URLs.

```
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/empty-results-page
```

### 48 — noindex, follow the thin archives

*Where: Your site, your server, your CMS*

**Why it works:** Tag pages, author archives, filtered views and deep pagination are typically 80% of a CMS's indexed URLs and 0% of its traffic. Keep them crawlable so equity still flows; keep them out of the index.

```
<meta name="robots" content="noindex, follow" />
```

### 49 — Sitemap = indexable URLs only, with an honest lastmod

*Where: Your site, your server, your CMS*

**Why it works:** Redirects, noindexed pages and 404s in a sitemap are a quality signal about the whole site. And a lastmod that updates on every deploy is the same as no lastmod at all — Google learns to ignore it.

```
<url>
  <loc>https://www.example.com/services/drain-cleaning/</loc>
  <lastmod>2026-08-14</lastmod>
</url>

lastmod = the day the CONTENT changed. Not the day you deployed.
```

### 50 — Server-render what the AI crawlers need

*Where: Your site, your server, your CMS*

**Why it works:** GPTBot, ClaudeBot, PerplexityBot and Google's AI fetchers don't reliably execute JavaScript. If your service list, pricing or location data only exists after hydration, it doesn't exist to them.

```
curl -s https://example.com/services/ | sed 's/<[^>]*>/ /g' | tr -s "[:space:]" " " | head -c 1200

If your headline, services and phone number aren't in that output, they aren't in the answer either.
```

### 51 — Decide about the AI crawlers deliberately

*Where: Your site, your server, your CMS*

**Why it works:** Blocking them is a legitimate business choice. Blocking them by accident — inherited from a boilerplate robots.txt — and then wondering why you're never cited is not.

```
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: PerplexityBot
User-agent: Google-Extended
Allow: /
```

### 52 — Preload the LCP image. Never lazy-load above the fold

*Where: Your site, your server, your CMS*

**Why it works:** loading="lazy" on the hero image is the most common self-inflicted LCP failure there is. Preload it, serve it as AVIF or WebP at the size it actually displays, and mark it high priority.

```
<link rel="preload" as="image" href="/hero.avif" fetchpriority="high" />
<img src="/hero.avif" fetchpriority="high" decoding="async" width="1600" height="900" alt="..." />
```

### 53 — Explicit width and height on every image

*Where: Your site, your server, your CMS*

**Why it works:** Costs nothing, removes the most common source of layout shift, and stops the page jumping on exactly the slow connections where you were already losing the visitor.

```
<img src="/photo.webp" width="1200" height="800" alt="Descriptive alt text" />
```

### 54 — Cut the JavaScript that blocks INP

*Where: Your site, your server, your CMS*

**Why it works:** INP replaced FID in 2024 and it measures every interaction, not just the first. Defer third-party tags, delete the ones whose reports nobody reads, and keep interactions under 200ms.

```
<script src="https://third-party.example/tag.js" defer></script>

Audit with PageSpeed Insights -> Diagnostics -> "Avoid long main-thread tasks".
Then ask, per tag: who read that report last month?
```

### 55 — Chunk the page: 40-80 word answers under real headings

*Where: Your site, your server, your CMS*

**Why it works:** Both Google's passage ranking and every AI retriever work on chunks. A 900-word section under one heading is one retrievable unit. Six labelled chunks under six honest headings are six chances to be the answer.

```
## How long does a sewer lateral repair take?

Most residential sewer lateral repairs take one to three days. A spot repair on a
single collapsed section is usually done in a day; a full lateral replacement from
the house to the main takes two to three, plus permit inspection. Trenchless lining
finishes fastest because there is no excavation to backfill.

## What does it cost?
```

### 56 — Link from your strongest pages to your money pages

*Where: Your site, your server, your CMS*

**Why it works:** Authority moves through links. Find your pages with the most backlinks, then check whether any of them link to the page you actually want to rank. Usually none do.

```
Ahrefs -> Best by Links (or Search Console -> Links -> Top linked pages).
For each of your top 20 linked pages, ask: does it link to the page I want to rank?
Anchor text = the target page's primary keyword. Not "learn more". Not "click here".
```

### 57 — Hunt the orphan pages

*Where: Your site, your server, your CMS*

**Why it works:** A page in your sitemap with zero internal links is a page you told Google about and then told it nobody cares about. Everything you publish should be reachable within three clicks of the homepage.

```
Screaming Frog -> crawl the site -> Configuration -> Spider -> Crawl Linked XML Sitemaps
-> Crawl Analysis -> Run -> filter "Orphan URLs".

Every hit is a page in the sitemap with no internal link pointing at it.
```

---

## Rather have it done for you?

Locafy gets local businesses found in Google Maps, AI Overviews and AI search —
done for you, measured monthly. https://www.locafy.com/book-a-call
