Semantics.rs
Semantics.rs/SEO/AI search
Architecture

How is a site prepared for AI search?

AI answers are assembled from the same index that produces ordinary results. Preparation is therefore not a special markup but a set of practices that make a page extractable: an answer in the first screen, a table instead of a paragraph, a claim with a source and a date, and structured data only for what is actually visible.

This page does not describe someone else's framework. It describes practices carried out on this site and on two taxi projects, with one measured outcome and an explicit list of what must not be concluded from it.

How is an AI answer assembled?

The system takes a query, expands it into a set of related questions, searches the index for each of them, selects passages from several documents, and assembles an answer with links to sources. A page enters that answer as a passage, not as a whole document.

The consequence is practical: the unit being selected is not the page but a paragraph, a table, or a table row. A page that gives a clear, self-contained answer in one place offers that unit. A page that spreads the same answer across an introduction, context and conclusion offers nothing that can be lifted out on its own.

From query to citation: query, expansion into sub-questions, index search, passage selection, assembled answer A five-step diagram. Step one: the user submits a query. Step two: the system expands the query into several sub-questions. Step three: for each sub-question the same index used by ordinary search is queried. Step four: individual passages are selected from several documents rather than whole documents. Step five: the answer is assembled and accompanied by links to the sources the passages came from. 0102030405 Query Expansion intosub-questions Search of thesame index Passageselection Answerwith sources one query becomes several questions a paragraph or table row is selected,not a whole document Where a page drops out If the answer to the question in the heading does not exist as a self-contained unit, there is nothing to select at step 04. The system then picks a competitor document that has that unit, even if the rest of its page is weaker.
Five steps from query to citation. Selection happens at step four, at passage level.
  1. Query — the user asks a question.
  2. Expansion — one query becomes a set of sub-questions.
  3. Search — each sub-question is run against the same index that ordinary search uses.
  4. Passage selection — individual paragraphs or table rows are chosen from several documents.
  5. Answer — the text is assembled with links to sources.

Every expanded question needs a matching component on the page. If a component is missing for one of them, that is a functional gap, not a stylistic omission. How those intents are listed before writing is set out in the guide on query templates.

Why is there no special markup for AI answers?

Google points to the same basic practices that apply to ordinary search: availability for indexing, useful content and compliance with Search guidelines. There is no structured data type that guarantees entry into an AI answer.

Structured data still does its job — it describes machine-readably what is visible on the page and grants eligibility for certain result types. Eligibility is neither a guarantee of display nor a documented ranking effect. The rules for a cohesive graph are set out on the page about structured data.

What is not claimed here: no schema type guarantees a citation in an AI answer. Any offer promising “AI Overview optimisation” through added markup is selling something that is not documented.

Which components make a page extractable?

A component is extractable when it carries the whole answer without the rest of the page. A fare table, a sentence with a figure directly under the heading, and a row holding a single attribute all meet that condition. An introduction, a transition and a conclusion do not.

Components and the intent each one answers
ComponentAnswers the intentCheck
Answer in the first screenWhat the answer to the heading isThe first few sentences carry the whole answer without context
Table with caption and scopeWhich values apply and to whatA single row is understandable without the rest of the table
Claim with source and dateWhere the figure comes from and when it heldEvery figure carries a period, a tool and a data owner
Heading as a questionWhether the page answers the query at allThe first sentence below the heading answers that heading
Figures with alt and figcaptionWhat the diagram claimsThe result can be read without the image
Cohesive @graphWhich entities are on the page and how they relateThe markup describes only what is visible
Caveat next to an estimateWhat must not be concluded from the figureFor every claim there is a conceivable proof that it is wrong

Why the layout of a page by itself changes which part gets extracted is set out on the page about the main content.

What of that is applied on this site?

The following holds across the 28 pages of the Serbian version. The figures were counted in the site's source code, not estimated.

Component coverage, Serbian version of Semantics.rs, 28 pages
PracticePagesShare
Section heading phrased as a question28 / 28100%
Figures with alt and figcaption16 / 16100%
Table with caption and scope25 / 2889%
Answer in the first screen22 / 2879%
Caveat next to a claim19 / 2868%
Date in a <time> element15 / 2854%
FAQ block where an FAQ actually stands14 / 2850%

The last two rows are not shortcomings. An FAQ block stands on 14 pages and exactly 14 pages carry FAQPage in their structured data — not one more. Marking up an FAQ where none is present on the page would be a mismatch between markup and content.

How this was counted: by searching the source code of every page in the Serbian version for the presence of the corresponding HTML elements and classes. This is the state of the site, not a performance measurement. Coverage changes with every content edit.

What does our own measurement show?

On the taxi.co.rs project, pages whose main content is a fare table enter AI answers in 29.1% of their impressions. The site-wide average is 13.5%. City hubs, whose main content is an overview rather than a concrete figure, stand at 6.8%.

Share of impressions in AI features by page type on taxi.co.rs A horizontal bar chart with three values. City fare pages, whose main content is a fare table, have a share of 29.1 per cent. The site-wide average is 13.5 per cent. City hubs, which offer an overview rather than a concrete figure, have 6.8 per cent. The source is Google Search Console, the generative features report, covering 12 June to 14 August 2026, aggregated by page, owner data. SHARE OF IMPRESSIONS IN AI FEATURES · TAXI.CO.RS · GSC · 12 JUNE – 14 AUGUST 2026 City fare pages Site-wide average City hubs main content: fare table all pages combined main content: overview 29.1% 13.5% 6.8% The difference is not attributed to any single change: content, layout, internal links and structured data were altered at the same time.
Share of impressions in AI features by page type. Fare pages carry a concrete figure, hubs offer an overview.
  • City fare pages — 29.1% of impressions; main content is a fare table.
  • Site-wide average — 13.5% of impressions.
  • City hubs — 6.8% of impressions; main content is an overview, not a concrete figure.

Measurement: Google Search Console, generative features report, , aggregated by page, owner data.

What this figure is, and what it is not: the difference between page types is measured and repeats on the English version of the same pattern. What is not measured is the cause. Content, layout, internal links and structured data were changed at the same time on that project, so the difference is attributed to the system rather than to any single component. A control case on another project produces a markedly weaker result and is described in the leskovac.taxi study.

The full distribution, with every page type and the complete list of limitations, is in the taxi.co.rs case study.

What about AI search can be measured, and what cannot?

What is measured is how many times a link to the site was shown in Google's generative features, and how many visits arrive from AI platforms. What is not measured is which queries caused them, or how many times the site was cited without a click.

Data availability by source
SourceGivesDoes not give
Search Console, generative featuresImpressions per pageQueries, clicks as a separate metric, export through the API
Analytics, referrer from AI platformsVisits that arrivedCitations without a click, visits with no referrer
Server and CDN logsVisits by AI crawlersThe link between a crawl and a citation

Because of this, visibility in AI answers cannot be expressed as a single number. Every measured figure is a lower bound: visits from mobile apps often carry no referrer and end up counted as direct traffic, and a citation without a click leaves no trace in analytics.

How the AI impressions report is read and where its limits lie is set out on the page about measuring results.

Limit of the comparison: Search Console data and AI platform visit data do not share a definition of the event or a period. Impressions and visits are neither added together nor divided by one another.

How do you check your own page?

Take one page and work through eight questions. Each has a yes or no answer, with no grade in between.

  1. Does the first sentence below the heading answer the question in that heading?
  2. Can that sentence be read on its own, without the rest of the page, and still make sense?
  3. Is the most important figure in a table rather than in a paragraph?
  4. Does every table have a caption and scope?
  5. Does every figure carry a period, a tool and a data owner?
  6. Can the result of every diagram be read without looking at the image?
  7. Does the structured data describe only what is visible on the page?
  8. For every claim — is there a conceivable proof that it is wrong?

A page that fails the first two questions has nothing to offer at the passage selection step, however good the rest of it is. What this check looks like applied to a whole site is shown in the audit example.

Common questions about AI search

Is there a special schema markup for AI answers?

No. Google points to the same basic practices: availability for indexing, useful content and compliance with Search guidelines. There is no “AI Overview schema” that guarantees a citation.

Can I see the queries behind AI impressions?

Not through an official report. Search Console reports impressions in generative features but not the queries behind them. The data is exposed neither through the API nor through the BigQuery export.

How is visibility in ChatGPT measured?

Only the visits that arrive are measured, through the referrer in analytics. A citation without a click leaves no trace, and mobile app traffic often carries no referrer, so every measured figure is a lower bound.

Should AI crawlers be blocked?

It depends on the business model. Blocking protects content from being used in training, but it also removes the possibility of being cited. The decision is made per domain, not as a general rule.

Is this different from ordinary SEO?

Not fundamentally. The same index, the same guidelines, the same quality criteria. The difference is that a passage is selected rather than a document, so layout and extractability carry more weight than before.

Does shorter text stand a better chance of being cited?

It is not length but extractability. A sentence carrying a concrete figure directly under the heading can be lifted out without the rest of the text. The same figure spread across three paragraphs cannot.

Primary sources

This guide is a method. How it is carried out on a real site, with a price and a timeline, is shown by the semantic SEO audit.

Want to know where your site stands?

Send the URL and how you currently measure success.

Request a site audit