Skip to content

The readers you are not counting

A named client fetches your page. Your analytics calls it a bot and drops it. Later a person arrives from an answer written out of that same page, and that gets filed under referrals. Two measurement errors, one cause.

Argument, not a feature listWritten against the build in this repositoryNo figure on this page is a measurement of your site

What actually arrives at your server

Three kinds of client knock on a web server, not two. Micaforge keeps each in its own table, because they are answers to different questions.

A person
A human in a browser, as far as one request can show. Written to events.
An agent
A non-human client that names itself: a product token its operator published, an operator you can look up, and a purpose. Written to agent_events.
A suspected bot
Non-human, and not nameable. Scored against header signals, with the signals that fired stored beside the row so the judgement can be audited. Written to bot_events.

Almost every analytics tool recognises the first and bins the other two together. That was reasonable while the second bin held search crawlers, because the trade was legible: a crawler took a page, a results list carried a link to it, some readers followed the link. Both ends of the exchange showed up in the same dashboard.

The new occupant of that bin does not make the same trade. It takes the page for a training corpus, or to ground an answer somebody is waiting on right now, and that answer can be complete enough that nobody needs to arrive at all.

What share of your traffic is it?

This page will not tell you, and you should distrust a page that does. Micaforge has never measured your site, and the percentages in circulation are other people's sites, measured other people's ways, quoted without the method.

What can be said is where your own number is written down. It is in your web server's access log, it has been all along, and nothing in your analytics is reading it.

8k4k50250Jun 23Jul 1Aug 1Guide publishedGPTBot: 4,102 fetchesClaudeBot re-crawlFirst answer arrivals
Demonstration record, not measured trafficFetches by named AI agents 3.7kfull scale 8kReaders arriving from an answer engine 35full scale 50Separate gains. Each pen has its own full scale, so the distance between them is not a quantity.
Demonstration record. Illustrative figures, not measured traffic. Two audiences over sixty days, written from the same centre rule: machine fetches upward, readers arriving from an answer downward. Read the shape, not the totals. The machine pen steps, because a crawl arrives in bursts around a publish. The human pen is written, because reading is continuous. Each pen has its own gain and prints its full-scale value, because the two differ by about two orders of magnitude. That difference is the finding, so it is stated as a number rather than left to the eye.

What an answer-engine referral is

The path from your page to a reader now has a step in the middle.

  1. Somebody asks a question of an answer surface rather than a search box.
  2. The surface consults an index built by an earlier crawl, or fetches your page there and then to ground the reply. Both fetches are in your log; neither is in your analytics.
  3. An answer is written. Your page may be cited under it, or may not.
  4. A fraction of those readers click the citation. They land on your site with a referrer host of chatgpt.com, perplexity.ai, claude.ai or a peer.

Step four is an acquisition channel. To a tool with no such channel it is a referral, filed beside a link somebody left in a forum, and the fact that it came from an answer built out of step two is nowhere on the screen.

Micaforge makes answer_engine a top-level channel next to search, social and direct, with the same conversion and quality reporting as any other. That is the smaller half of the idea. The larger half is that it can be read against the crawl that produced it.

The returned half is a floor

Browsers default to strict-origin-when-cross-origin, so a cross-origin referrer usually carries a host and nothing else. An answer surface that shares a host with an ordinary search page cannot be told apart from that site's own results in the common case, and Micaforge does not guess: counting every search on such a host as an answer arrival would be a fabricated number.

An answer rendered inside a search results page is invisible by construction. A reader who clicks a citation in one arrives with the search engine as the referrer, indistinguishable from an ordinary click. So the returned count is at best a floor. Your real gap is no better than the one on the screen, and may be worse.

Why filtering the bots is the wrong default

Start with the concession: filtering is right for the human report. A pageview count with crawler fetches folded into it is worthless, and Micaforge defaults every screen to audience=human for exactly that reason.

The mistake is not the filter. It is the bin. Discarding is not the same as excluding, and three things go with the rubbish.

You lose the supply chain
The crawl is upstream of the answer that sends you a reader. Delete the crawl and the arrival has no cause: you can see that answer engines send you fifty sessions a week, and never that they took eight thousand pages to do it, or which pages.
You lose cost and policy
Fetches are bandwidth, cache pressure and origin load. A rule in your robots.txt is a claim about behaviour that either held or did not. Neither question can be asked of a bin.
You lose the difference between a claim and a fact
A user-agent string is an assertion, and anyone can type one. Where an operator publishes address ranges or reverse-DNS suffixes, Micaforge checks the source against them and marks the row verified, unverified or failed. A spoofed crawler reported as the company it named would be worse than not reporting it at all.

A filter is a view. A tool that offers only the filtered view has decided the question on your behalf, and it decided it in an era when the answer did not matter much.

What the Citation Gap tells an operator

For every URL, Micaforge holds two counts side by side: how often named non-human clients fetched it, and how many sessions arrived at it from an answer engine. The ratio between them is the Citation Gap.

It is one number because a decision is one decision. Not "crawl volume is up" on one screen and "referrals are flat" on another, but a single line per page saying what was taken and what came back.

With no arrivals at all, the ratio is the fetch count itself. That is the reading a person makes anyway, and it keeps the value finite: JSON cannot spell infinity, and a serialiser would quietly turn it into nothing.

The thresholds are constants, and they are published so a verdict can be checked instead of believed. Cited means at least three answer-engine arrivals in the window. Heavily crawled means at least twenty agent fetches. A cited, heavily crawled page is extracted when the ratio passes twenty-five. A page with two arrivals and nineteen fetches is quiet, not extracted.

Extracted

Crawled hard, and either never cited or cited far less than it is taken.

Decide the trade deliberately: price it, gate it, or rewrite the page so an answer has to link out to be useful. Keeping it open is also a decision, and a reasonable one.

Compounding

Crawled and cited. The exchange is working.

This is the shape that pays. Look at what the page does differently, and write more like it.

Invisible

Cited, but barely crawled. Earning readers on a trace too thin to explain them.

Usually a sitemap, a robots rule or an internal link that is missing. Cheap to fix, and the cheapest win on the report.

Quiet

Neither crawled nor cited enough to say anything.

Nothing. Thin data gets a shrug, not a verdict.

Both halves are joined against the content register, the table of pages the site is known to have, so a page nothing has crawled and nobody has read still appears with zeroes. Joining the other way round would filter away exactly the rows the screen was opened to find.

Extracted is a measurement, not an instruction. Blocking the crawler is one response. Writing the page so an answer has to link out is another. Deciding the reach is worth the trade is a third, and often the right one. The report's job is to let you make that choice on evidence.

What it costs to see it

Agents do not run JavaScript. A script tag has never seen one and never will, which is why no tool that only ships a script tag can report this, however it is configured.

The agent half of Micaforge is fed from the access log your web server already writes. One shipper process reads it and posts signed batches to the ingest endpoint. Nothing about your pages changes.

micaforge-shipper --host https://analytics.example.com \
  --site 1 --key-file ingest.key /var/log/nginx/access.log
Know this before you install

If your site sits somewhere you cannot get an access log from, this half of the product stays dark. The human analytics still work in full, every page reads quiet on the Citation Gap, and the report will say so rather than inventing the other number.