Skip to content

What the machines took.
What came back.

See which AI crawlers read your pages, which answers send readers back, and which pages are worth writing more of. Open-source, cookieless analytics that counts the machines as an audience instead of throwing them away.

8k4k50250Jun 23Jul 1Aug 1Guide publishedGPTBot: 4,102 fetchesClaudeBot re-crawlFirst answer arrivals
Demonstration record, not measured trafficMachine fetches 3,746full scale 8kArrived from an answer engine 35full scale 50Separate gains. Each pen has its own full scale, so the distance between them is not a quantity.
Taken174,917fetches by named AI agents
Returned1,358readers arriving from an answer
The Citation Gap129 : 1taken for every one that came back

Every reader counted. Including the ones that never click.

Your human traffic, the way you already read it, beside the half of your audience that never runs a line of your JavaScript: the crawlers deciding what an answer engine says about you.

The numbers you check every morning

Without a cookie, so without a consent banner standing between a reader and your page. All of it in the self-hosted build, none of it behind a tier.

  • Visitors, sessions, pageviews, bounce rate, time on page
  • Real-time, to the second
  • Pages, entries, exits, scroll depth
  • Channels, referrers, UTM campaigns, paid click IDs
  • Country, region and city, with maps
  • Browser, OS, device and screen class
  • Custom events with JSON properties
  • Goals, funnels, journeys, retention cohorts
  • Session replay and error tracking
  • Core Web Vitals by page, device and country
  • Visitor profiles and identify()
  • Forty-plus filter dimensions, saved segments
  • Feature flags and A/B experiments
  • Organisations, teams, roles, per-site access
  • Public dashboards and private share links
  • Stats API, CSV export, scheduled reports, alerts
  • Imports from Google Analytics, Plausible and Umami

What the machines do with your work

A growing share of what reads your site is not a person. Micaforge shows you who it is, what it took, and whether any reader came back for it.

  • Agent traffic as a channel. Every named AI crawler, grouped by operator and by what it is collecting for: training corpus, answering a question now, search index.
  • Verified, or honestly unverified. Source addresses checked against the operator's published ranges. A spoofed ClaudeBot is reported as failed verification, not as Anthropic.
  • Crawl coverage per URL. First fetch, last fetch, how often, and how stale the machines' copy of a page has become.
  • Answer engines as an acquisition channel. ChatGPT, Perplexity, Claude, Gemini and Copilot sit beside Search, not buried in Referrals.
  • The Citation Gap. Per page: taken versus returned. The one number that says whether writing more of this is worth it.
  • Policy versus behaviour. Your robots.txt and llms.txt, read back to you against what actually fetched.
  • Content decay. Pages whose readers fell away while the crawlers kept coming.
  • An MCP server. Point an agent at your own analytics and let it answer the question itself.

Know what to do with every page

Micaforge holds two counts against every URL: how much the machines fetched it, and how many people arrived at it from an answer. Together they say whether a page is paying you back, and what to change when it is not.

  • Extracted

    Crawled hard. Almost never cited.

    Your work is in the corpus and none of the traffic is coming back. These are the pages to price, gate, or rewrite so the answer has to link out.

  • Compounding

    Crawled and cited.

    The machines read it and readers arrive because of it. This is the shape that works. Write more of these.

  • Invisible

    Cited, but barely crawled.

    Answer engines are sending people to a page the crawlers keep skipping. Usually a sitemap, robots or internal-linking problem, and a cheap fix.

  • Quiet

    Neither crawled nor cited enough to say anything.

    Not a failure. An absence of signal: too little of either count to draw a verdict from. Leave it alone and read it again when there is more of it.

A Citation Gap report. Demonstration record, illustrative figures, not measured traffic.
PageMachine fetchesAnswer arrivalsGapVerdict
/docs/api/authentication8,42112702 : 1Extracted
/blog/why-cookieless-works3,1104866.4 : 1Compounding
/docs/self-hosting/install6,00231194 : 1Extracted
/changelog2101741.2 : 1Invisible
/docs/tracking/spa-routing1,8902407.9 : 1Compounding
/legal/dpa440none returnedQuiet

You wrote a rule. Did anyone follow it?

Micaforge reads your own robots.txt and llms.txt, then checks every fetch against them. Where an agent publishes its address ranges, the claim is verified. Where it does not, the row says so rather than guessing, and a fetch nobody can be identified as making is charged to nobody.

robots.txt
User-agent: GPTBot
Allow: /
Disallow: /docs/api/User-agent: ClaudeBot
Allow: /User-agent: Google-Extended
Disallow: /User-agent: CCBot
Disallow: /User-agent: Bytespider
Disallow: /
Observed against that policy. Demonstration record.
AgentCollecting forFetchesVerifiedFetched anyway
GPTBotOpenAITraining corpus8,421✓Verified against published address ranges118
ClaudeBotAnthropicTraining corpus6,120✓Verified against published address ranges—
PerplexityBotPerplexitySearch index3,044✓Verified against published address ranges—
Google-ExtendedNot attributedTraining corpus2,870—No published ranges to check against—
CCBotNot attributedTraining corpus1,902—No published ranges to check against—
BytespiderNot attributedTraining corpus1,466✕Claimed this agent; the address did not match—
OAI-SearchBotOpenAISearch index1,120✓Verified against published address ranges—
meta-externalagentNot attributedTraining corpus806—No published ranges to check against—

The 118 is fetches of /docs/api/, disallowed for GPTBot in the policy beside this table, from an address that verified as OpenAI's. The other 6,238 fetches against a Disallow rule came from clients that did not verify: counted, held under the user-agent string they claimed, and charged to nobody.

Free while it is young. AGPL forever.

There is no community edition with the good parts removed. The self-hosted build is the whole build. One command puts it on a small VPS with Postgres, ClickHouse and Valkey beside it.

Micaforge is early. Some of what is listed above is newer than the rest, and the comparison pages say plainly where Plausible and Google Analytics are still ahead.

Read the install
01
git clone https://github.com/micaforge/micaforge
cd micaforge && ./install.sh
02
<script defer data-site="1"
  src="https://analytics.example.com/mf.js"></script>
03
micaforge-shipper --host https://analytics.example.com \
  --site 1 --key-file ingest.key /var/log/nginx/access.log

Step three is what makes the machines visible: crawlers never run JavaScript, so their fetches come from your server's own log.