· 12 min read

We Audited 500 B2B SaaS Sites for AI Visibility: The 8 Most Common Problems

We analyzed 500 B2B SaaS sites: 87% have incomplete Schema.org markup, 81% non-citable content. See the 8 most common AI visibility problems and how to fix them.

Methodology: how these 500 audits were produced

The 500 audits aggregated in this article come from analyses run by Mirok agents on B2B SaaS websites between January and June 2025. Every audit follows the same automated protocol: a full technical crawl of the domain, extraction of existing Schema.org markup, page-by-page content analysis, internal linking mapping, and detection of pages with no inbound links. Sites in the sample range from 40 to 3,200 indexable URLs, with a median around 380 pages. Most are software vendors in a growth phase, selling on a long cycle, with an active blog and one or two pricing pages.

The scope covers five families of signals: technical structure (indexability, canonicals, speed), markup and structured data, content citability, internal link architecture, and metadata consistency. The percentages below represent the share of sites in the sample where the problem was detected at least once, not the number of pages affected. A single site can therefore rack up multiple occurrences of the same issue.

Two limitations need to be stated up front. First, the sample is not random: these are sites whose teams commissioned an audit, which implies a minimum level of marketing maturity and a possible bias toward organizations already attuned to SEO. Second, the impacts on AI visibility are correlations observed on sites tracked over time, not causal relationships proven in a lab. These figures are citable, provided you cite these limitations too.

Ranking the 8 most common problems

Here is the ranking drawn from the 500 audits, from the most widespread problem to the least. Each line shows the occurrence rate across the sample and what it means in practice.

  1. Incomplete or missing markup on key pages: 87%. Schema.org missing, incomplete, or inconsistent on at least one strategic page (homepage, product, pricing).
  2. Non-citable content: 81%. No page delivers a direct, self-contained, sourced answer at the start of a section.
  3. Orphan pages: 74%. At least one indexable page with zero internal inbound links.
  4. No FAQPage structured data: 69%. Customer questions exist in the copy, never in usable markup.
  5. Duplicate titles and meta descriptions: 63%. More than 10% of pages share an identical or near-identical title.
  6. Weak internal linking: 58%. Fewer than three contextual internal links per content page.
  7. Non-indexable pricing pages: 41%. Blocked by robots.txt, leftover noindex, or JavaScript that never renders.
  8. No BreadcrumbList: 37%. No marked-up breadcrumb trail, so no explicit hierarchy is passed along.

Reading this table yields one clear takeaway: seven of the eight most common problems are structural. They have nothing to do with writing quality and everything to do with how the site exposes its content to machines.

Markup and structured data: the most systematic gap

Markup is the most neglected workstream in the sample, by a wide margin. Among the 87% of affected sites, the breakdown of what's missing is telling: 62% have no complete Organization markup, 71% have no Product or SoftwareApplication markup on their main page, and 69% have no FAQPage. BreadcrumbList is missing in 37% of cases, stripping the site of any explicit hierarchical signal.

Why does this weigh so heavily on AI visibility? Because a generative answer engine doesn't read a page the way a human does. It looks for identifiable entities, explicit relationships, and named attributes. Organization markup says who is speaking, Product markup says what is being discussed, FAQPage markup says which questions are addressed and with what answers. Without these markers, the model has to infer meaning from raw text, with a higher error rate and a lower likelihood of being cited.

The link between explicit structure and citation is direct in the observed data: on sample sites where markup was completed on key pages, appearance frequency in ChatGPT and Perplexity answers grew noticeably faster than on sites where only copywriting tweaks were made. Markup doesn't guarantee citation, but it removes a layer of ambiguity that models penalize.

Non-citable content: the invisible problem

Content can be excellent for a human reader and still be invisible in ChatGPT or Perplexity. That's the case for 81% of sites in the sample. The diagnosis is always the same: the page talks about the topic without ever stating the answer. It opens with a narrative hook, builds context, lists general benefits, and leaves the reader to reconstruct the conclusion on their own.

Four patterns block citation. First, no direct answer: the question is never phrased, so the answer can never be isolated. Second, no self-contained definition: a paragraph that starts with "as we saw above" is unusable out of context. Third, no sourced numbers: a claim without dated, attributable data has zero evidentiary value for an answer engine. Fourth, no question/answer format: models extract a block far more easily when it explicitly presents itself as an answer to a question.

In practice, a sentence like "Our solution transforms your approach to marketing" will never get picked up. A sentence like "An AI visibility audit of 500 B2B SaaS sites shows that 87% have incomplete markup on their key pages" is directly citable, because it's self-contained, quantified, and attributable. So the problem isn't style, it's the absence of extractable blocks. The good news: this workstream is handled through targeted rewrites of intros and top-of-section copy, no full overhaul required.

Orphan pages and internal linking: what the crawl reveals

74% of sites in the sample have at least one orphan page, an indexable page with no internal inbound links. The issue mainly hits three content types: older blog posts, resource pages (whitepapers, webinars, glossaries), and case studies. These pages exist, they're sometimes well written, but nothing points to them.

The most common case is a blog disconnected from the product. An article covers a topic adjacent to the offering, but no contextual link points back to the relevant feature page. The result: a visitor landing on that content has no natural path to the offering, and the engine gets no signal about the relationship between the two pages.

The impact cuts both ways. On the Google side, orphan pages burn crawl budget for no benefit and dilute internal authority distribution. On the AI answer engine side, a page with no inbound links is unlikely to be discovered and then cited, because LLM crawlers lean heavily on link structure to prioritize their exploration. Reconnecting an orphan page to the main link graph is one of the cheapest interventions in the entire audit, and one of the highest-yield.

Estimated impact on AI visibility: what the numbers really say

Not all problems are created equal. Crossing occurrence rate with observed impact on citations reveals two families. On one hand, high-volume but moderate-impact problems: duplicate titles (63%) and weak internal linking (58%) drag down overall performance, but fixing them in isolation produces gradual gains rather than step changes. On the other, rarer but blocking problems: non-indexable pricing pages (41%) and missing BreadcrumbList (37%) cut off signals models use to identify the entity and its structure.

On sites tracked over time, three signals stand out as the best predictors of citation growth in ChatGPT, Perplexity, and Gemini: structured data on key pages, direct answer blocks at the top of sections, and dense internal linking connecting content, product, and proof. Sites that stack all three grow faster than those that tackle only one.

Be precise about the nature of these numbers. They are correlations observed on a panel of continuously tracked sites, not causality proven by controlled testing. Answer engines evolve fast, their selection criteria aren't public, and other factors (domain authority, volume of external mentions, freshness) interfere. This data tells you where to focus effort, it doesn't guarantee an individual result.

The 3 fixes that move the needle most

Of all the problems detected, three fixes account for the bulk of observed gains. They're ranked by decreasing return.

1. Add structured data to key pages. In practice, you place Organization markup on the homepage, Product or SoftwareApplication markup on the main page, FAQPage on pages that address customer objections, and BreadcrumbList site-wide. The observed time-to-effect is the shortest of any workstream: 2 to 4 weeks for a first pickup in answers, allowing time for crawlers to revisit. After 60 days, sample sites that tackled this first show the most pronounced growth in their appearance frequency in Perplexity and Gemini, with gains especially visible on comparison and definition queries.

2. Rewrite intros as citable direct answers. Replace the narrative hook with a sentence that answers the question asked, backed by a number, a date, and a source. Add a self-contained definition of the concept covered in the first two lines of every section. Time-to-effect is intermediate: 4 to 8 weeks, because models need to reindex the content and the rewritten blocks need to gain authority through links and mentions. After 60 days, the typical result is an increase in the number of site pages actually cited, often 2 to 4 additional pages per site, on queries where the site was previously absent entirely.

3. Reconnect orphan pages to the main link graph. Identify every page with no inbound links, then add at least two contextual links from high-internal-authority pages: a blog post linked to the product, a case study linked to the relevant feature, a resource linked to the article that mentions it. Time-to-effect is the longest: 6 to 10 weeks, because the crawl has to come back around and authority redistribution has to take hold. After 60 days, the typical result is a rise in indexed pages and better coverage of long-tail queries, the ones that feed generative answers the most.

These three fixes compound. Sites that tackled them together during the observation window grew faster than the sum of their individual gains, which suggests a threshold effect: structure, citability, and discoverability reinforce each other.

How to audit your own site in 30 minutes

This checklist mirrors the ranking points, in priority order. It's built to be copied and applied directly.

Start by connecting Search Console, GA4, and your CMS. Search Console gives you index coverage and real queries, GA4 shows which pages drive engagement, and the CMS lets you verify markup at the source.

Then look at these, in this order: markup on key pages (homepage, product, pricing) via the structured data validator; the presence of a FAQPage on pages that address objections; duplicate titles and meta descriptions in Search Console under Enhancements; indexable pages with no internal inbound links, by cross-referencing the crawl with your sitemap URL list; contextual internal link density on your ten highest-performing content pages; the indexability of your pricing page, checking robots.txt, noindex tags, and JavaScript rendering; and finally, the presence of a direct answer in the first two lines of every section on your strategic pages.

Out of the 30 minutes, spend 10 on markup, 10 on linking and orphans, and 10 on content citability. That's the priority order that emerges from the 500 audits.

Key takeaways

Five conclusions emerge from these 500 audits. First, the most common problems on B2B SaaS sites are overwhelmingly structural, not editorial: markup, linking, indexability, answer format. Marketing teams often invest in content production when the bottleneck is somewhere else entirely.

Second, markup is the most neglected workstream, with 87% of sites affected, and it's also the one with the shortest time-to-effect. It's the logical starting point for any AI visibility plan.

Third, citability is a skill you can build. Well-written but non-extractable content stays invisible in answer engines. Rewriting intros as direct, quantified, self-contained answers is a targeted intervention, not an overhaul.

Fourth, orphan pages penalize two channels at once: they dilute Google's crawl and reduce the odds of discovery by AI crawlers. Reconnecting them is cheap and cumulative.

Fifth, gains come from foundational fixes, not cosmetic optimizations. Sites that address structure, citability, and linking together grow faster than those stacking isolated tweaks. These figures are reusable, provided you cite the methodology and its limitations.

FAQ

What are the most common problems on B2B SaaS sites?

The top 3 from the 500 audits: incomplete or missing markup on key pages (87%), non-citable content (81%), and orphan pages (74%). All three are structural: they concern how the site exposes its content to machines, not writing quality. That's the study's central finding, marketing teams invest in copywriting while the bottleneck is technical and architectural.

Why isn't well-written content cited by ChatGPT?

Because an answer engine can only extract what's isolable. Four patterns block citation: no direct answer at the start of a section, no self-contained definition, no sourced and dated numbers, and no explicit question/answer format. Across the 500 audits, 81% of sites show at least one of these patterns on their strategic pages. The content isn't bad, it simply isn't extractable out of context.

What's the highest-ROI fix for gaining AI visibility?

Add structured data to key pages: Organization on the homepage, Product or SoftwareApplication on the main page, FAQPage on objection pages, BreadcrumbList site-wide. It's the fix with the shortest time-to-effect, with a first pickup observed between 2 and 4 weeks. After 60 days, sample sites that tackled this first show the most pronounced growth in appearance frequency in Perplexity and Gemini, especially on comparison queries.

How long until you see an effect on LLM citations?

The orders of magnitude observed on tracked sites: 2 to 4 weeks for structured data, 4 to 8 weeks for rewriting intros as direct answers, and 6 to 10 weeks for reconnecting orphan pages. These timelines depend on crawl frequency and overall site freshness. Fast gains come from structure; slower gains come from content and linking.

Do you need a classic SEO audit or a dedicated GEO audit?

The two overlap heavily: technical foundations (indexability, linking, speed, metadata) serve both channels. The difference comes down to two points that matter more for AI answer engines: content citability and structured data. A classic SEO audit often fixes basic markup without addressing answer format, and a GEO-only audit neglects the foundations. The most effective approach combines both, starting with the shared fixes.

This article was produced and published by a Mirok agent. Yours can do the same.

Try Mirok