Cookieless Analytics: How We Measure Without Tracking
Cookieless analytics is the only kind of tracking we run — here's exactly what KinetixSEO measures, what it keeps, and what it throws away.
· By Rogier Bruggeman, Founder of KinetixSEO
Why we're explaining our own measurement instead of comparing tools
Cookieless analytics is the only kind of tracking we run — here's exactly what KinetixSEO measures, what it keeps, and what it throws away. This isn't a roundup of analytics vendors or a pitch for privacy-friendly analytics as a category. It's a plain account of how our own site counts page views and how our own scan tool keeps statistics on the web problems it finds, because site owners deciding whether to trust an SEO tool's numbers deserve to know where those numbers come from. We'll walk through site analytics, scan statistics, and the publishing rules that sit on top of both. Full detail behind each of these claims — retention periods, the sampling process, and scoring versions — sits in our methodology and privacy policy, which this article summarizes rather than replaces.
Test the technical signal
Want to see this on your own site?
Use your own site as the evidence. Get a free SEO and AI-citation readiness baseline, then monitor what changes.
Part 1: how we measure our own site traffic
We run a self-hosted instance of Plausible to count page views, and it sets no cookies and writes nothing to the visitor's browser storage. That's why our site has no consent banner for analytics: there is no tracking identifier stored on the visitor's device to ask permission for. According to Plausible's own data policy, the product does not use cookies, browser cache, or local storage, and it does not store raw IP addresses or User-Agent strings. Unique visitors are instead counted with a daily rotating salt that is deleted every 24 hours — a number is derived from the visit and then the material needed to reproduce that number is thrown away before the next day starts.
The one part of our site that is consent-gated is error and performance monitoring, which stores a session id so we can reconstruct what happened when something breaks. That session id persists for the length of a browsing session, which is enough to link a stack trace to the sequence of actions that caused it — analytics and error monitoring are built to do genuinely different jobs, and only one of them needs a persistent identifier.
Why we checked the script file instead of trusting the vendor page
We didn't take Plausible's documentation at face value; we checked it against the exact script file our server loads. A vendor's marketing page can describe an idealized version of the product, and a script can be updated after the page describing it was written, so the only claim worth trusting is the one made by the bytes that actually run in the visitor's browser. We pin that script with a subresource-integrity (SRI) hash, which means the browser refuses to execute the file if its contents ever change without our knowledge. That gives us a verifiable claim rather than a reported one: if the loaded script doesn't match the hash we've reviewed, nothing runs. It's the same discipline we'd want anyone to apply before taking any vendor's privacy claims about analytics without cookies at face value — check the artifact, not the brochure.
The trade-off we accept
Cookieless counting means we cannot stitch a visitor's visits across days into a single journey, and we cannot build per-person funnels. There's no identifier that persists long enough to say "this person read the pricing page on Monday and converted on Thursday." We get aggregate counts — page views, referrers, top pages, trends over time — but not individual paths through the site. That's the deliberate cost of not storing anything on the visitor's device: journey data and cookieless tracking are close to mutually exclusive, and we've chosen the latter for the main analytics on our own site.
Part 2: what our free scan tool keeps as statistics
Every free check run through our tool keeps a small statistics record, separate from the result shown to the person running it, so that we can later describe what share of websites has a given problem. What's kept is narrow and specific:
- A keyed, one-way code for the domain — not the domain name itself, but a value computed from it that can't be reversed without the key
- The scores the check produced
- Which finding types fired (for example, a missing meta description or an oversized image) without the content that triggered them
- The top-level domain (such as .com or .co.uk)
- Whether the site served over HTTPS
- The date the check ran
What's never kept is just as specific: the URL that was checked, the path, any page content, or who ran the check. We don't store the address, we don't store what was on the page, and we don't store any detail about the person running it.
Why "pseudonymous" is the correct word, not "anonymous"
Anyone holding our secret key could re-identify a domain, which is exactly why we call this data pseudonymous and never anonymous or fully anonymised. The domain code is generated with a secret key, and anyone holding that key could take a guessed domain, run it through the same keyed function, and check whether the resulting code matches a record in the statistics table. That's what separates pseudonymous data from anonymous data: the first is re-identifiable by someone with the key, the second isn't re-identifiable by anyone. It's also why the key itself lives outside the database that holds the statistics records — splitting the two means a single compromised database doesn't hand over both the lock and the key at once. Full details of this separation, and how long records are retained, are in our methodology and privacy policy.
Part 3: the publishing rules we hold ourselves to
Before any figure reaches an article, it has to clear three rules: it must come from a defined random sample rather than free-check traffic, it must be built from enough distinct websites to be stable, and it must state which version of our scoring produced it. Each rule closes off a different way a published number could mislead a reader, so we treat all three as non-negotiable rather than best-effort.
Sampling. We only publish figures drawn from a defined random sample of public websites, never from free-check traffic. People who voluntarily run an SEO check are not a representative picture of the web — they've self-selected because they already suspect a problem, they skew toward sites actively being optimized, and the sample is shaped by whoever happened to share a scan link that week. A statistic built from that pool would describe our users, not the web, so we keep the two uses of the same scan technology separate: one feeds the tool you get a result from, the other feeds a sampling process built specifically to be published.
Minimum sample size. No figure is published unless it's built from a large enough set of distinct websites to be stable. Below some threshold, a handful of unusual sites can swing a percentage significantly, and the number stops meaning anything repeatable — so a figure built from too small a sample simply doesn't run, regardless of how interesting it looks.
Versioned scoring. Every figure states which version of our scoring engine produced it. When we change how a check is scored, a figure spanning the change isn't a trend — it's two different measurements glued together and mislabeled as one line. Versioning keeps that from happening silently, and it's also why this particular article reports no statistics of its own: it explains the method, not a result. When we do publish a figure elsewhere on the site, it will say which sample it came from, how many sites it covers, and which scoring version produced it — the same inspectability we're describing here for our own analytics setup.
What this means if you're weighing cookieless analytics yourself
If you're a developer evaluating analytics without cookies for your own site, the questions worth asking are the same ones we asked of our own setup: what does the script actually write to storage, does that match what the vendor documents, and what are you giving up by not persisting an identifier. Cross-day journeys and per-session funnels are real capabilities you lose, and whether that trade-off is acceptable depends on what decisions your analytics need to support — a marketing team building multi-touch attribution needs something a cookieless tool was never built to give it, while a team that mainly wants to know which pages get read and whether a change in copy moved traffic can run entirely without cookies. The same discipline applies beyond analytics: whether you're validating structured data before publishing it or diagnosing Core Web Vitals regressions, the number is only trustworthy if the process that produced it is checkable rather than taken on faith. How you handle scan or usage statistics on your own product is worth deciding deliberately too: whether you key identifiers, where the key lives, and what minimum sample size you'll accept before calling something a trend.
Frequently asked questions
Does cookieless analytics mean no data is collected at all?
No — cookieless analytics means no identifier is written to the visitor's browser, not that nothing is measured. Our Plausible instance still counts page views, referrers, and aggregate trends; it just does so using a daily rotating salt rather than a persistent cookie, so the visitor's device holds nothing after they leave the page.
Why doesn't KinetixSEO's site show a cookie consent banner for analytics?
We don't show one because our analytics setup doesn't store anything on the visitor's device that would need consent — no cookies, no browser cache entries, no local storage writes. The one piece of our site that is consent-gated is error and performance monitoring, which stores a session id and is handled separately from page-view analytics.
What's the difference between pseudonymous and anonymous data?
Pseudonymous data can be re-identified by someone holding the right key; anonymous data can't be re-identified by anyone. Our scan statistics use a keyed, one-way code for each domain, and because that key exists, someone holding it could test a guessed domain against the records — which is exactly why we call the data pseudonymous and keep the key stored separately from the statistics database.
Does KinetixSEO use free-check traffic to publish statistics?
No — we only publish figures drawn from a defined random sample of public websites, never from the traffic of people running free checks. Free-check users are self-selected, typically because they already suspect a problem with their site, so that pool doesn't represent the broader web and we don't treat it as if it does.
How does KinetixSEO decide a sample is large enough to publish?
A figure only gets published once it's built from a large enough set of distinct websites that a few unusual outliers can't swing the result, and the exact threshold and sampling process are documented on our methodology page. Every figure we do publish also states which version of our scoring produced it, so a change in how a check is scored never gets silently blended into an existing trend.
Sources
- Plausible Analytics, Data Policy ()
Plausible does not use cookies, browser cache or local storage, and does not store raw IP addresses or User-Agent strings; unique visitors are counted with a daily rotating salt that is deleted every 24 hours
Test the technical signal
See how your own site scores on SEO and AI-search visibility — free report, no signup.
Use your own site as the evidence. Get a free SEO and AI-citation readiness baseline, then monitor what changes.
Related articles
- AI Citation Tracking: How It Actually WorksGEO & AI Search
- Largest Contentful Paint: Definition, Thresholds & FixesTechnical SEO
- Rich Results: What They Are and How to Earn ThemTechnical SEO
- Article Schema: A Step-by-Step Setup GuideTechnical SEO