What the machines took.
What came back.
See which AI crawlers read your pages, which answers send readers back, and which pages are worth writing more of. Open-source, cookieless analytics that counts the machines as an audience instead of throwing them away.
Every reader counted. Including the ones that never click.
Your human traffic, the way you already read it, beside the half of your audience that never runs a line of your JavaScript: the crawlers deciding what an answer engine says about you.
The numbers you check every morning
Without a cookie, so without a consent banner standing between a reader and your page. All of it in the self-hosted build, none of it behind a tier.
- Visitors, sessions, pageviews, bounce rate, time on page
- Real-time, to the second
- Pages, entries, exits, scroll depth
- Channels, referrers, UTM campaigns, paid click IDs
- Country, region and city, with maps
- Browser, OS, device and screen class
- Custom events with JSON properties
- Goals, funnels, journeys, retention cohorts
- Session replay and error tracking
- Core Web Vitals by page, device and country
- Visitor profiles and
identify() - Forty-plus filter dimensions, saved segments
- Feature flags and A/B experiments
- Organisations, teams, roles, per-site access
- Public dashboards and private share links
- Stats API, CSV export, scheduled reports, alerts
- Imports from Google Analytics, Plausible and Umami
What the machines do with your work
A growing share of what reads your site is not a person. Micaforge shows you who it is, what it took, and whether any reader came back for it.
- Agent traffic as a channel. Every named AI crawler, grouped by operator and by what it is collecting for: training corpus, answering a question now, search index.
- Verified, or honestly unverified. Source addresses checked against the operator's published ranges. A spoofed
ClaudeBotis reported as failed verification, not as Anthropic. - Crawl coverage per URL. First fetch, last fetch, how often, and how stale the machines' copy of a page has become.
- Answer engines as an acquisition channel. ChatGPT, Perplexity, Claude, Gemini and Copilot sit beside Search, not buried in Referrals.
- The Citation Gap. Per page: taken versus returned. The one number that says whether writing more of this is worth it.
- Policy versus behaviour. Your
robots.txtandllms.txt, read back to you against what actually fetched. - Content decay. Pages whose readers fell away while the crawlers kept coming.
- An MCP server. Point an agent at your own analytics and let it answer the question itself.
Know what to do with every page
Micaforge holds two counts against every URL: how much the machines fetched it, and how many people arrived at it from an answer. Together they say whether a page is paying you back, and what to change when it is not.
Extracted
Crawled hard. Almost never cited.
Your work is in the corpus and none of the traffic is coming back. These are the pages to price, gate, or rewrite so the answer has to link out.
Compounding
Crawled and cited.
The machines read it and readers arrive because of it. This is the shape that works. Write more of these.
Invisible
Cited, but barely crawled.
Answer engines are sending people to a page the crawlers keep skipping. Usually a sitemap, robots or internal-linking problem, and a cheap fix.
Quiet
Neither crawled nor cited enough to say anything.
Not a failure. An absence of signal: too little of either count to draw a verdict from. Leave it alone and read it again when there is more of it.
| Page | Machine fetches | Answer arrivals | Gap | Verdict |
|---|---|---|---|---|
| /docs/api/authentication | 8,421 | 12 | 702 : 1 | Extracted |
| /blog/why-cookieless-works | 3,110 | 486 | 6.4 : 1 | Compounding |
| /docs/self-hosting/install | 6,002 | 31 | 194 : 1 | Extracted |
| /changelog | 210 | 174 | 1.2 : 1 | Invisible |
| /docs/tracking/spa-routing | 1,890 | 240 | 7.9 : 1 | Compounding |
| /legal/dpa | 44 | 0 | none returned | Quiet |
You wrote a rule. Did anyone follow it?
Micaforge reads your own robots.txt and llms.txt, then checks every fetch against them. Where an agent publishes its address ranges, the claim is verified. Where it does not, the row says so rather than guessing, and a fetch nobody can be identified as making is charged to nobody.
User-agent: GPTBot
Allow: /
Disallow: /docs/api/User-agent: ClaudeBot
Allow: /User-agent: Google-Extended
Disallow: /User-agent: CCBot
Disallow: /User-agent: Bytespider
Disallow: /| Agent | Collecting for | Fetches | Verified | Fetched anyway |
|---|---|---|---|---|
| GPTBotOpenAI | Training corpus | 8,421 | ✓Verified against published address ranges | 118 |
| ClaudeBotAnthropic | Training corpus | 6,120 | ✓Verified against published address ranges | — |
| PerplexityBotPerplexity | Search index | 3,044 | ✓Verified against published address ranges | — |
| Google-ExtendedNot attributed | Training corpus | 2,870 | —No published ranges to check against | — |
| CCBotNot attributed | Training corpus | 1,902 | —No published ranges to check against | — |
| BytespiderNot attributed | Training corpus | 1,466 | ✕Claimed this agent; the address did not match | — |
| OAI-SearchBotOpenAI | Search index | 1,120 | ✓Verified against published address ranges | — |
| meta-externalagentNot attributed | Training corpus | 806 | —No published ranges to check against | — |
The 118 is fetches of /docs/api/, disallowed for GPTBot in the policy beside this table, from an address that verified as OpenAI's. The other 6,238 fetches against a Disallow rule came from clients that did not verify: counted, held under the user-agent string they claimed, and charged to nobody.
Free while it is young. AGPL forever.
There is no community edition with the good parts removed. The self-hosted build is the whole build. One command puts it on a small VPS with Postgres, ClickHouse and Valkey beside it.
Micaforge is early. Some of what is listed above is newer than the rest, and the comparison pages say plainly where Plausible and Google Analytics are still ahead.
Read the installgit clone https://github.com/micaforge/micaforge
cd micaforge && ./install.sh<script defer data-site="1"
src="https://analytics.example.com/mf.js"></script>micaforge-shipper --host https://analytics.example.com \
--site 1 --key-file ingest.key /var/log/nginx/access.logStep three is what makes the machines visible: crawlers never run JavaScript, so their fetches come from your server's own log.