# Shruwd — full documentation > Shruwd measures how a brand appears in AI-generated answers (Google AI Overviews > and ChatGPT), sets that beside first-party AI-crawler server logs, > diagnoses why the brand is or is not cited, and recommends specific fixes that > can be re-checked to prove movement. This file is every documentation page, in order. The curated index is at https://shruwd.io/llms.txt and the REST API describes itself at https://shruwd.io/openapi.json. --- # Overview Shruwd tells you how AI assistants describe your brand, why you are left out of the answers you care about, and whether the fix you made changed anything. ## What you get **A measured picture.** Your mention rate, share of voice, prominence and citation share on Google AI Overviews and ChatGPT — each with a confidence interval, so you can tell a real number from a small sample. **A reason, not just a number.** Findings name the cause and the page it applies to: a robots.txt rule blocking an AI crawler, a page that never gets cited, a competitor who owns a question you should be winning. **A fix you can check.** Mark a fix applied and Shruwd re-measures against the numbers from before you changed anything, then tells you whether the difference was real. ## How it works You give Shruwd a brand, the questions your buyers ask, and the competitors you expect to be compared against. Every week it asks those questions on each engine, records who gets named and which pages get cited, and turns what it finds into a short list of things to fix. ## Next steps - [Quickstart](https://shruwd.io/docs/getting-started/quickstart.md) — Go from an empty account to your first measurement. - [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md) — What the four numbers mean, and when to trust them. - [Prompts](https://shruwd.io/docs/tracking/prompts.md) — Choose the questions worth measuring. - [Findings and fixes](https://shruwd.io/docs/findings/how-findings-work.md) — Act on a finding and prove it moved. --- # Quickstart Add a brand, add the questions your buyers ask, and get your first measurement. Setup takes about five minutes; results arrive over the following hours. Working from code or with an AI agent? The same steps run over the API. Install the TypeScript SDK: ```bash npm install @shruwd/sdk ``` For Claude, Cursor or another MCP client, the server is `npx -y shruwd-mcp` — see [MCP server](https://shruwd.io/docs/api/mcp-server.md). Both take an API key from **Account → API keys**. The API is described at [`/openapi.json`](https://shruwd.io/openapi.json), and the [API overview](https://shruwd.io/docs/api/api-overview.md) covers authentication and errors. ## Prerequisites - A Shruwd account. If you do not have one, [create one](https://shruwd.io/sign-up) — the free plan needs no card. - Your brand's domain. - A rough idea of the questions a buyer asks before choosing a product like yours. ## 1. Add your brand Go to **Brands → Add a brand** and fill in three fields. | Field | What to enter | |---|---| | **Brand name** | Your name as it is written in answers. It becomes the first thing Shruwd matches on, so spell it the way the world does. | | **Domain** | A URL or a host. It is reduced to the registrable domain, and citations of any page on it count as yours. | | **Timezone** | Sets the day boundary for measurements. | Leave **Suggestions** on. Shruwd reads your homepage and drafts prompts and competitors for you to edit, which is faster than starting from a blank page. Nothing is saved until you add it. > **Note** > > The domain cannot be changed later — it defines what Shruwd treats as you. To track a > different domain, add another brand. ## 2. Add competitors Go to **Competitors** and add the brands you expect to be compared against. Shruwd counts a mention only when it matches exactly, so add the spellings answers actually use. Short or common-word names need **context terms** — words that make "Arc" mean the company rather than the word. Shruwd asks for them rather than guessing. Add them before your prompts. Your first measurement starts with your first prompt and counts only the competitors that exist by then, and on the free plan it is the only one. You can skip this and come back. Once answers start arriving, Shruwd lists the names that came up but are not tracked, and you can accept them in one click. ## 3. Add prompts Prompts are the questions asked on each engine, every cycle. Go to the brand's **Prompts** page and add them one per line. Write them the way a buyer would type them, not as keywords: ``` best waitlist software for a product launch launchlist alternatives how do I collect sign-ups before launch ``` Pick an **intent** for each batch. Intent decides which diagnostics apply, and `commercial` and `comparison` prompts are where competitors get named — start there. Your first measurement begins within minutes of your first prompt, so add the whole set together. ## 4. Read the first results Results appear on the brand's **Overview** as each cycle completes. Expect the first numbers within a few hours, not immediately. Until a metric has at least ten responses behind it, Shruwd shows **insufficient data** rather than a number. That is not zero and not an error — it means the sample is too small to mean anything yet. ## Next steps - [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md) — what each number means - [Prompts](https://shruwd.io/docs/tracking/prompts.md) — choosing questions worth measuring - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — turning a result into a fix --- # Reading your results Your brand's **Overview** shows four numbers, each with a range. This page explains what they mean and when they are solid enough to act on. ## The four numbers | Metric | Answers | |---|---| | **Mention rate** | How often you are named at all | | **Share of voice** | How much of the naming goes to you rather than your competitors | | **Prominence** | How early in the answer you appear, when you do | | **Citation share** | How often your domain is the source being cited | A brand named six times in one answer counts once. Counting every occurrence would reward long answers and swing on formatting, so mentions are counted per answer. ## Every number has a range A result reads like `32% (24–41%)`. The range is a 95% confidence interval: the answers Shruwd sampled put your true mention rate somewhere in that band. Read the band, not just the middle. `32% (24–41%)` and `32% (10–60%)` are the same point estimate and completely different facts. ## Insufficient data is not zero Shruwd never shows a number computed from fewer than ten responses. You will see **insufficient data** instead. This shows up most on individual prompts. Each prompt is asked three times per cycle, so a single prompt takes several weeks to clear the threshold — while your prompt set as a whole clears it immediately. A three-run result is noise, and showing it as a number would be worse than showing nothing. ## Share of voice moves for two reasons Share of voice is your mentions as a fraction of all mentions across you and your tracked competitors. It rises when you do better **and** when a competitor does worse. Always read it next to mention rate. If share of voice is up and mention rate is flat, nothing about you changed — a competitor lost ground. When an answer names nobody in your tracked set, it counts toward neither side. Shruwd reports those separately: - **Naming nobody** — answers that mention no tracked brand. High across a group of prompts means nobody owns the category yet, which is an opportunity rather than a loss. - **No AI answer** — the engine did not produce an AI answer for that prompt. If this is high, the prompt set is the problem, not your visibility. ## When a number has actually moved Week-to-week variation is large. Shruwd only calls a change movement when both of these hold: 1. The confidence intervals before and after do not overlap. 2. The change is at least 5 percentage points. It also corrects for the fact that you are watching many prompts at once — testing dozens of prompts every week would otherwise turn up a few false alarms forever. If the provider changed its underlying models between two windows, the comparison is marked **model changed**. The difference may be the model rather than you. ## Next steps - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — turn a number into a fix - [Prompts](https://shruwd.io/docs/tracking/prompts.md) — what you are measuring - [Competitors](https://shruwd.io/docs/tracking/competitors.md) — who you are measured against --- # Brands A brand is a domain, plus the prompts, competitors and findings attached to it. Most accounts have one. ## Adding a brand **Permissions:** owner. 1. Go to **Brands → Add a brand**. 2. Enter the **brand name** as it is written in answers. This becomes the first alias Shruwd matches on, so spell it the way the world does. If the name is short or an everyday word, like "Front", Shruwd also asks for a few words that mark a mention as being about you — otherwise every "front-end" would count. 3. Enter the **domain**. A URL or a bare host is fine — both are reduced to the registrable domain, so `https://www.acme.com/pricing` and `acme.com` are the same brand. Citations of any page on it count as yours. 4. Choose a **timezone**. This sets the day boundary for measurements and rollups. 5. Leave **Suggestions** on to have Shruwd draft prompts and competitors from your homepage. Nothing is saved until you add it. The first measurement waits for your first prompt, so adding a brand before you have prompts ready costs you nothing. Setup asks for [competitors](https://shruwd.io/docs/tracking/competitors.md) next, then prompts, because the measurement counts only the competitors that exist when it starts. ## What you can change later | | | |---|---| | **Name** | Editable. Display only — what Shruwd matches on lives in the brand's aliases on the [Competitors](https://shruwd.io/docs/tracking/competitors.md) page. | | **Timezone** | Editable. Applies from the next cycle. | | **Domain** | Not editable. The domain defines what Shruwd treats as you, so a new domain is a new brand. | Renaming a brand does not change its URL in the dashboard. ## Archiving a brand **Permissions:** owner. Go to the brand's **Settings → Archive brand**. Archiving stops all measurement. Nothing is deleted — history stays readable, and the domain becomes available to add again. > **Note** > > Two active brands in the same workspace cannot share a domain. If you get a duplicate > error, the domain is already on another brand — archive that one to free it. ## Limits Your plan sets how many active brands you can have: one on Crawl and Starter, three on Pro. Archived brands do not count. On Pro, all three brands draw from **one shared pool of 100 prompts**. Adding a brand does not add prompts. See [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md). ## Next steps - [Prompts](https://shruwd.io/docs/tracking/prompts.md) — the questions asked for this brand - [Competitors](https://shruwd.io/docs/tracking/competitors.md) — who you are measured against --- # Prompts Prompts are the questions Shruwd asks on each engine, every cycle. Choosing them well is the highest-leverage decision in your setup — everything else is measured against them. ## Writing a good prompt Write what a buyer would actually type, not a keyword. | Write this | Not this | |---|---| | `best waitlist software for a product launch` | `waitlist software` | | `launchlist alternatives` | `competitor comparison` | | `how do I collect sign-ups before launch` | `sign up collection` | Keep the set stable. Prompts are re-asked every cycle, and that repetition is what makes week-over-week comparison mean anything. Changing them often resets what you can compare. Aim for a mix, weighted toward commercial and comparison prompts — those are where competitors get named and where a change is worth having. ## Adding prompts **Permissions:** editor or above. 1. Go to the brand's **Prompts** page. 2. Enter your prompts in **Add prompts**, one per line. 3. Choose an **intent** for the batch. 4. Add **tags** if you want to group prompts for filtering. Optional. 5. Select **Add**. Each prompt is 1–500 characters. A batch is all or nothing: if it would take you past your plan's prompt limit, none of it is added and Shruwd tells you the numbers. Your first measurement starts within minutes of your first prompt and counts only the prompts and [competitors](https://shruwd.io/docs/tracking/competitors.md) that exist by then. Add competitors first and the prompts together. ## Choosing an intent Intent decides which diagnostics apply to a prompt, so it is worth getting roughly right. | Intent | Use it for | |---|---| | `commercial` | "Best X for Y" — a buyer ready to choose. The prompts most worth winning. | | `comparison` | "X vs Y", "alternatives to X" — weighing options. Where competitors get named. | | `informational` | "How does X work" — learning, not yet buying. | | `navigational` | Looking for a specific brand or page by name. | | `problem` | Describes the problem, not a product: "how do I stop losing sign-ups". | ## Editing a prompt Changing a prompt's **text** or **intent** starts a new version. Past measurements stay attached to the wording that produced them, so your history stays honest — but the new wording starts from zero responses and needs time to clear the ten-response threshold. Changing **tags** does not change what is asked, and takes effect immediately. ## Removing a prompt Select the remove icon on a prompt. It stops being asked from the next cycle on, and its history is kept. Removing a prompt frees a slot against your plan limit. Re-adding it later counts as a new prompt. ## Reading prompt coverage The **Current prompts** list shows each prompt's mention rate over the last 30 days for the selected engine, plus two counts worth watching: - **with no AI answer** — the engine produced no AI answer for this prompt. A high count means the prompt is not one these engines answer, whatever you do to your site. - **naming nobody** — the answer named no tracked brand. Nobody owns the question yet. A prompt showing **Not measured yet** has not been through a completed cycle. ## Limits Prompts are counted **across all brands in your workspace**: 10 on Crawl, 25 on Starter, 100 on Pro. Only active prompts count. See [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md). ## Next steps - [Competitors](https://shruwd.io/docs/tracking/competitors.md) — who you are measured against - [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md) — what comes back --- # Competitors Competitors are the brands you expect to be compared against. They set the field your [share of voice](https://shruwd.io/docs/getting-started/reading-your-results.md) is measured against. ## How matching works Shruwd counts a mention only when an alias matches **exactly, on word boundaries**. There is no fuzzy matching and no patterns. This is deliberate. A near-match that silently counts the wrong brand corrupts every number downstream and is close to impossible to spot later. The cost is that you have to add the spellings answers actually use. | You add | Matches | Does not match | |---|---|---| | `LaunchList` | "LaunchList", "launchlist" | "Launch List", "LaunchLists" | Add each spelling you expect as a separate alias. ## Adding a competitor **Permissions:** editor or above. 1. Go to the brand's **Competitors** page. 2. Select **Add a competitor**. 3. Enter the **name**, then any other **aliases** the answers use. 4. Add the competitor's **domain** so their citations are attributed to them. 5. Select **Add**. On a new brand, add competitors before prompts. The first measurement starts with the first prompt and counts only the competitors that exist by then; a competitor added later counts from the next cycle. ### Short and common-word names A name that is six characters or shorter, or that is an ordinary English word, is refused until you add **context terms** — words that must appear nearby for a mention to count. They are what make "Arc" mean the company rather than the word. Shruwd asks rather than guessing, because the alternative is a metric that quietly counts every unrelated use of a common word. ### Exclusions Add an **exclusion** to stop a specific phrase counting as a mention — useful when your competitor's name appears inside a longer, unrelated product name. ## Accepting a suggested competitor Once answers start arriving, **Add as competitor?** lists names that came up but are not tracked, with how often each appeared and a domain hint where the answers linked one. 1. Review the name and the domain hint. Check the spelling before it starts counting. 2. Select **Add** to track it, or **Dismiss** to quiet it for 90 days. Accepting goes through the same checks as adding by hand, including context terms. ## Your brand Your own brand appears at the top of the page as **Your brand**. It is matched the same way as any competitor, so if answers spell your name differently — an old name, a common misspelling, a product name used in place of the company — add those as aliases here. Editing your brand's name in settings does not change what is matched. Your brand cannot be removed. ## Removing a competitor Select the remove icon. It stops counting from the next cycle on. Past measurements keep it, and share of voice is recomputed against the remaining set. ## Next steps - [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md) — how share of voice is built - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — including when a competitor owns a prompt --- # Connecting Search Console Connecting Google Search Console gives Shruwd per-page impressions — which of your pages people actually look for. Two diagnostics need that context and stay switched off without it. ## What it enables | Diagnostic | Needs Search Console for | |---|---| | A page with demand is never crawled | Knowing the page has real search demand, so an uncrawled page is worth flagging | | The wrong page gets cited | Knowing which page people search for, to compare against the one being cited | Without the connection, neither can tell a page nobody wants from a page being missed. ## Prerequisites - **Plan:** Pro. The connection is refused before the Google consent screen on other plans, so you will not waste a round trip. - **Permissions:** admin or owner. - A Google account with access to the Search Console property for your brand's domain. ## Connecting 1. Open the brand's **Settings**. 2. Find the **Google Search Console** card. 3. Select **Connect** and sign in with Google. 4. Grant read access to the property for your domain. You are returned to Settings and the card shows the connected **Property** and **Last sync**. ## Checking the connection The card shows its state at a glance: | State | Meaning | |---|---| | *not connected* | Never connected | | *active* | Working. **Last sync** shows when data last arrived | | *revoked* | Access was withdrawn at Google's end. Reconnect to restore it | A connection can be revoked outside Shruwd — from your Google account's security settings, or by losing access to the property. Shruwd cannot restore it on its own; reconnect from this card. ## Disconnecting Select **Disconnect** on the card. The two diagnostics above stop producing new findings. Existing findings and history are kept. ## Next steps - [Crawler logs](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md) — the other source of first-party evidence - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — what these diagnostics produce - [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md) — which plan includes this --- # Choosing a log path Sending your own server logs tells Shruwd which AI crawlers actually reach your pages and which get refused. It is the only first-party evidence in the product, and it is included on every plan, free one included. This page picks the right path for your hosting. Each one is a separate guide. ## Why bother Two of the four access diagnostics work from your domain alone. The other two need logs, including the most valuable rule in the catalogue: your server returning an error to a verified AI crawler while serving the same page fine to everyone else. Most people who have this problem do not know it — it is usually a bot-protection setting nobody remembers switching on. Without logs you still get findings. With logs you get the ones that say *we watched this happen*, which are the ones worth acting on first. ## Pick your path | Your setup | Path | Code? | |---|---|---| | Vercel (Pro or Enterprise), Netlify, Fly and similar | [Platform log drain](https://shruwd.io/docs/crawler-logs/platform-log-drains.md) | None | | Behind Cloudflare, any plan | [Forwarder Worker](https://shruwd.io/docs/crawler-logs/cloudflare-worker.md) | Deploy a small Worker | | Cloudflare **Enterprise** zone | [Logpush](https://shruwd.io/docs/crawler-logs/cloudflare-logpush.md) | None | | Dynamic origin with no CDN in front | [Server middleware](https://shruwd.io/docs/crawler-logs/server-middleware.md) | A middleware file | | Vercel Hobby, no Cloudflare in front | None yet. Hobby has no log drains; put Cloudflare in front, any plan, and use the Worker | — | When two rows fit, take the higher one. Platform drains and Logpush need no code and cannot drift out of date, and Logpush also records the requests Cloudflare itself refused, which the Worker never sees. > **Warning** > > **If your site is static, prerendered or ISR, do not use server middleware.** Middleware > only sees requests that reach your origin, and the pages worth measuring are exactly the > ones your CDN answers without ever calling you. On a typical marketing site that hides > most of your traffic. Use the Worker instead — it runs before cache lookup. ## Client IP is required Every path must send the client IP of each request. A user-agent string is a claim — spoofing `GPTBot` takes seconds. Shruwd verifies each hit against the crawler vendor's published IP ranges, and only verified hits count toward anything. A log source with no client IP is refused at setup rather than accepted and quietly ignored: ``` HTTP 422 No line carried a client IP. ``` That refusal is deliberate. Accepting the data would mean you did the setup and got a dashboard of numbers that cannot be used. ## What happens after logs arrive Your brand's **Crawlers** page starts showing hits per day and a per-bot breakdown, split three ways: | | | |---|---| | **Live retrieval** | Verified crawling that feeds AI answers. | | **Training** | Batch crawling for model training. | | **Unverified** | Claimed to be a crawler and could not be verified. Shown, never counted. | Each plan includes a monthly line allowance — 100k on Crawl, 500k on Starter, 2M on Pro. Past it, lines are discarded until the month turns and the page says so rather than showing you a flat chart. ## Next steps - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — which bots are recognised and how they are verified - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — if nothing shows up --- # Platform log drains If your host can forward request logs to an HTTP endpoint — Vercel, Netlify, Fly and most others — this is the whole setup. No code, and the logs are edge logs, so they include requests your CDN answered without touching your origin. ## Prerequisites - **Permissions:** admin or owner. - A host with an HTTP log drain feature. These are often on a paid tier; check your platform's current pricing. ## 1. Get your endpoint and token 1. Open the brand's **Settings → Crawler logs**. 2. Select **Create token**. 3. Copy the **endpoint** and the **token**. > **Warning** > > The token is shown **once**. It cannot be retrieved later, only replaced — and creating a > new one immediately revokes the old one, which will stop a drain that is already working. ## 2. Add the drain on your platform Point your platform's log drain at the endpoint with the token as a header. | Setting | Value | |---|---| | URL | the **endpoint** from step 1 | | Header | `X-Shruwd-Ingest-Token: YOUR_TOKEN` | | Content type | `application/x-ndjson` | | Format | NDJSON, one JSON object per line, or one JSON array of objects | ### On Vercel Drains need a Pro or Enterprise team. On Hobby with Cloudflare in front of the site, use the [forwarder Worker](https://shruwd.io/docs/crawler-logs/cloudflare-worker.md) instead. Hobby without Cloudflare in front has no log path yet. 1. Open **Team Settings → Drains → Add Drain** and choose **Logs**. 2. Name the drain and select the project. 3. Select the sources `static`, `lambda`, `edge`, `external`, `redirect` and `firewall`. `firewall` is the one that records the requests Vercel's firewall refused. 4. Select the `production` environment. Add no sampling rule, so every request is sent. 5. Choose **Custom Endpoint**: the **endpoint** from step 1, format NDJSON or JSON, and the custom header `X-Shruwd-Ingest-Token: YOUR_TOKEN`. 6. Select **Create Drain**. Vercel tests the endpoint, which answers `200`. > **Warning** > > Keep **Team Settings → Security & Privacy → IP Address Visibility** on. With IP > addresses hidden, Vercel sends no client IP, and no hit can be verified. ## 3. Check the fields Shruwd reads these fields per line. Common alternative spellings — `ClientIP`, `clientip`, `remote_addr` — are understood, so most platforms work as they come. Vercel's format works as it is, including the request details it nests under `proxy`. | Field | Required | Used for | |---|---|---| | `client_ip` | **Yes** | Verifying the hit is really from the crawler it claims to be | | `user_agent` | **Yes** | Identifying which bot. Lines without one are dropped | | `timestamp` | No | Defaults to arrival time. ISO 8601 | | `host` | No | Telling your domains apart | | `path` | No | Per-page diagnostics | | `status` | No | Detecting crawlers being refused | | `method`, `bytes` | No | Filtering and volume | If your host offers drain sampling, keep it at 100%. Crawler numbers are presented as exact counts, and a sampled drain would make them partial. A single line looks like this: ```json {"timestamp":"2026-09-09T10:00:00Z","host":"example.com","path":"/pricing","user_agent":"Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)","status":200,"client_ip":"23.98.142.176","method":"GET","bytes":18234} ``` ## 4. Test before relying on it Add `?validate=1` to the endpoint to parse a sample and report on it without writing anything. ```bash curl -X POST "YOUR_ENDPOINT?validate=1" \ -H "X-Shruwd-Ingest-Token: YOUR_TOKEN" \ -H "Content-Type: application/x-ndjson" \ -d '{"timestamp":"2026-01-01T00:00:00Z","host":"example.com","path":"/","user_agent":"GPTBot","status":200,"client_ip":"1.2.3.4"}' ``` `usable: true` means the endpoint and token are both good. Anything else names what is missing. ## 5. Confirm it is live Within a few minutes of real traffic, **Settings → Crawler logs** moves from *awaiting first logs* to *active*, and **Last received** fills in. ## Response codes | Code | Meaning | |---|---| | `202` | Accepted (`200` to Vercel, which asks for it). Parsing happens afterwards, so this does not mean the lines were usable — use `?validate=1` to check that | | `401` | Wrong or revoked token | | `413` | Batch over the size cap | | `422` | Only with `?validate=1`: the sample can't be used, and `reason` says why | ## Next steps - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — if the card stays on *awaiting first logs* - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — what gets recognised --- # Cloudflare Worker A small Worker on a route in front of your site forwards AI-crawler hits to Shruwd. It runs **before cache lookup**, so it sees the requests Cloudflare answers from cache without calling your origin — which on a static or prerendered site is most of the pages worth measuring. Works on any Cloudflare plan. It sends one short line per crawler hit and nothing at all for ordinary visitors. > **Warning** > > The Worker only sees requests Cloudflare lets through. A request refused by Cloudflare's > own settings (AI crawler blocks, WAF custom rules, rate limiting) stops before the Worker > runs, so Shruwd never hears about it. Cloudflare does not document where Bot Fight Mode > runs, so do not count on seeing its refusals either. To see what Cloudflare refused, open > **Security → Analytics** in the Cloudflare dashboard. On an Enterprise zone, > [Logpush](https://shruwd.io/docs/crawler-logs/cloudflare-logpush.md) records those refusals too. > > What the Worker does see is a refusal made behind Cloudflare — by your host, a security > plugin or your server — and the finding says that is where it happened. ## Prerequisites - **Permissions:** admin or owner, to create the token. - A Cloudflare zone for your domain, and permission to deploy a Worker on it. - `wrangler` available (`npx wrangler` needs no install). ## 1. Get your endpoint and token 1. Open the brand's **Settings → Crawler logs**. 2. Select **Create token**. 3. Copy the **endpoint** and the **token**. > **Warning** > > The token is shown once and cannot be retrieved later. Creating a new one immediately > revokes the old one, which will stop a Worker that is already running. ## 2. Create the Worker Make a directory in your site's repository with the two files below. ```jsonc [wrangler.jsonc] { "$schema": "node_modules/wrangler/config-schema.json", "name": "shruwd-log-forwarder", "main": "src/index.js", "compatibility_date": "2026-09-01", // The route must match your canonical host ONLY. If customer custom domains // resolve through this same zone, this pattern must not catch them. "routes": [ { "pattern": "example.com/*", "zone_name": "example.com" } ], "vars": { "CANONICAL_HOST": "example.com", // Comma-separated path prefixes. "/" is the root page only; // "/blog/" is everything underneath it. "PATH_ALLOWLIST": "/,/pricing,/features,/blog/,/docs/", "SHRUWD_INGEST_URL": "YOUR_ENDPOINT" } } ``` ```js [src/index.js] // Forwards AI-crawler hits to Shruwd. Fails open: whatever happens in here, // your site's response is returned unchanged. // A broad prefilter, not the bot list. Identification proper happens on // Shruwd's side against a catalogue that is kept current, so this only has to // be loose enough not to drop something that catalogue would have caught. const BOT_UA_PREFILTER = /bot|crawl|spider|gpt|claude|perplexity|bytespider|applebot|ccbot|meta-external|google|extended/i // Static assets are never the page a crawler judges you on. .txt and .xml stay, // because robots.txt and sitemaps are interesting. const ASSET_PATH = /\.(?:js|mjs|cjs|css|map|png|jpe?g|gif|webp|avif|svg|ico|woff2?|ttf|otf|eot|mp4|webm|mp3|wav|pdf|zip|gz|json|wasm)$/i function parseAllowlist(value) { return String(value ?? '/').split(',').map((s) => s.trim()).filter(Boolean) } // Cheapest checks first — this runs on every request to the zone. function decide({ method, host, path, userAgent }, { canonicalHost, allowlist }) { if (method !== 'GET' && method !== 'HEAD') return false if (host.toLowerCase() !== canonicalHost.toLowerCase()) return false if (ASSET_PATH.test(path)) return false if (!allowlist.some((p) => (p === '/' ? path === '/' : path === p || path.startsWith(p)))) return false return BOT_UA_PREFILTER.test(userAgent) } let warned = false function maybeReport(request, response, env, ctx) { const url = new URL(request.url) const userAgent = request.headers.get('user-agent') ?? '' const report = decide( { method: request.method, host: url.hostname, path: url.pathname, userAgent }, { canonicalHost: env.CANONICAL_HOST, allowlist: parseAllowlist(env.PATH_ALLOWLIST) }, ) if (!report) return // Set by the edge; a client cannot forge it. Inside a Worker, // x-forwarded-for can be forged, so it is not used as a fallback. const clientIp = request.headers.get('cf-connecting-ip') if (!clientIp) return // The query string is never sent: it carries ids on some routes, and a // per-page crawl count needs none of it. const line = { timestamp: new Date().toISOString(), host: url.hostname, path: url.pathname, user_agent: userAgent, status: response.status, client_ip: clientIp, method: request.method, bytes: Number(response.headers.get('content-length') ?? 0) || 0, } // One request per hit, deliberately. A Worker isolate is ephemeral and // per-colo, so a buffer would be lost on eviction and split across colos. // Volume is low because the prefilter already ran. ctx.waitUntil( fetch(env.SHRUWD_INGEST_URL, { method: 'POST', headers: { 'x-shruwd-ingest-token': env.SHRUWD_INGEST_TOKEN, 'content-type': 'application/x-ndjson', }, body: JSON.stringify(line) + '\n', }) .then((res) => { if (!res.ok && !warned) { warned = true console.warn(`shruwd forwarder: ingest answered ${res.status}; check the token and URL`) } }) .catch((error) => { if (!warned) { warned = true console.warn('shruwd forwarder: ingest unreachable', error?.message ?? error) } }), ) } export default { async fetch(request, env, ctx) { // Never inside the try. The only unconditional path is returning your // origin's response — a logging pipeline that can break the site it // observes is worse than one that loses events. const response = await fetch(request) try { maybeReport(request, response, env, ctx) } catch { // deliberately silent } return response }, } ``` ## 3. Configure it Edit the four values in `wrangler.jsonc`: | Setting | Value | |---|---| | `routes` | Your canonical host, e.g. `example.com/*` with `zone_name` `example.com` | | `CANONICAL_HOST` | Your canonical host, without a scheme | | `PATH_ALLOWLIST` | Your public, crawlable pages, comma separated | | `SHRUWD_INGEST_URL` | The **endpoint** from step 1 | > **Warning** > > `PATH_ALLOWLIST` is an allowlist on purpose. Your site probably serves customer content > and token-bearing links from the same origin, and a blocklist fails open — the next route > someone adds starts forwarding data to a third party. The worst case here is a missing > datapoint for a marketing page. ## 4. Deploy Store the token as a secret rather than putting it in the config or the source. ```bash npx wrangler secret put SHRUWD_INGEST_TOKEN npx wrangler deploy ``` Paste the token when prompted. ## 5. Verify Request a **prerendered** page with a crawler user-agent. That is the case origin middleware could not see, so it is the one worth testing. ```bash curl -s -o /dev/null "https://example.com/pricing" \ -H "User-Agent: Mozilla/5.0 (compatible; GPTBot/1.2)" ``` Within a minute, **Settings → Crawler logs** moves from *awaiting first logs* to *active*. Your test hit comes from your own machine, not from OpenAI, so it is recorded as **unverified** and excluded from headline metrics. That is correct — and the card moving off *awaiting first logs* is the signal you are looking for. ## Next steps - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — how verification works - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — if nothing shows up --- # Cloudflare Logpush A Logpush job with an HTTP destination sends your zone's request logs straight to Shruwd. No code, and it captures everything the edge serves. ## Prerequisites - **Permissions:** admin or owner. - A **Cloudflare Enterprise** zone. Logpush is not available on Free, Pro or Business — on those plans use the [forwarder Worker](https://shruwd.io/docs/crawler-logs/cloudflare-worker.md) instead. ## 1. Get the destination URL 1. Open the brand's **Settings → Crawler logs**. 2. Select **Create token**. 3. Copy the **Logpush destination URL**. It already has the token embedded as a `header_Authorization` parameter, which Logpush turns into a request header on every upload. > **Warning** > > The URL contains the token and is shown once. Creating a new token revokes the old one > and will stop a job that is already running. ## 2. Create the job In the Cloudflare dashboard, go to your zone → **Analytics & Logs → Logpush → Create a Logpush job**. | Step | Value | |---|---| | Destination | **HTTP destination** | | HTTP endpoint | the destination URL from step 1, verbatim | | Dataset | **HTTP requests** | | If logs match | `ClientRequestHost` equals your domain, **or** `ClientRequestHost` equals it with `www.` in front | | Fields | `EdgeStartTimestamp`, `ClientRequestHost`, `ClientRequestPath`, `ClientRequestUserAgent`, `EdgeResponseStatus`, `ClientIP`, `ClientRequestMethod`, `EdgeResponseBytes` | | Advanced → timestamp format | Any. RFC3339, Unix and UnixNano are all read | | Advanced → sampling | 100% | On **Submit**, Cloudflare uploads a small gzipped test file and expects a `2xx`. Shruwd answers it and does not count it toward your usage. If job creation fails with `error validating destination`, the token in the URL is wrong or has been revoked. ## 3. On a busy zone, cap the batch size Shruwd reads up to 48 MB decoded per upload and drops the remainder of anything larger. On a high-traffic zone, create the job through the API with explicit limits instead. ```bash curl -X POST "https://api.cloudflare.com/client/v4/zones/YOUR_ZONE_ID/logpush/jobs" \ -H "Authorization: Bearer YOUR_CLOUDFLARE_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "shruwd-crawler-logs", "dataset": "http_requests", "destination_conf": "YOUR_DESTINATION_URL", "enabled": true, "filter": "{\"where\":{\"or\":[{\"key\":\"ClientRequestHost\",\"operator\":\"eq\",\"value\":\"example.com\"},{\"key\":\"ClientRequestHost\",\"operator\":\"eq\",\"value\":\"www.example.com\"}]}}", "output_options": { "field_names": ["EdgeStartTimestamp","ClientRequestHost","ClientRequestPath","ClientRequestUserAgent","EdgeResponseStatus","ClientIP","ClientRequestMethod","EdgeResponseBytes"], "timestamp_format": "rfc3339" }, "max_upload_bytes": 5000000, "max_upload_records": 10000 }' ``` ## 4. Verify Within a few minutes of traffic, **Settings → Crawler logs** moves from *awaiting first logs* to *active*. A hit from an IP outside the crawler vendor's published ranges is recorded as **unverified** and excluded from headline metrics. That is correct behaviour, and it still proves the pipe is working. ## Next steps - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — how verification works - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — if nothing shows up --- # Server middleware Middleware in your own app forwards AI-crawler hits to Shruwd. Use it only when no CDN answers for your site. > **Warning** > > **Read this before you start.** Middleware only sees requests that reach your origin. If > your site is static, prerendered or ISR, a CDN answers most requests without ever calling > your server — including exactly the marketing pages worth measuring. On one real Nuxt > site, eleven of twenty tracked paths never touched the origin at all. > > If a CDN sits in front of you, use the [forwarder > Worker](https://shruwd.io/docs/crawler-logs/cloudflare-worker.md) or a > [platform drain](https://shruwd.io/docs/crawler-logs/platform-log-drains.md) instead. Both see cached > requests; this cannot. ## Prerequisites - **Permissions:** admin or owner, to create the token. - A Nuxt or Nitro app whose origin genuinely serves the traffic. ## 1. Get your endpoint and token Open the brand's **Settings → Crawler logs** and select **Create token**. Copy both values — the token is shown once, and creating a new one revokes the old one. ## 2. Add the middleware Create `server/middleware/crawler-log.ts` in your app. ```ts [server/middleware/crawler-log.ts] const ENDPOINT = process.env.SHRUWD_INGEST_URL const TOKEN = process.env.SHRUWD_INGEST_TOKEN // Flush on whichever comes first. Both are small: a serverless instance can be // frozen between requests, and anything still buffered is lost. const MAX_BATCH = 20 const MAX_WAIT_MS = 10_000 // A deliberately broad prefilter. Real identification happens on Shruwd's side // against a list that is kept current, so this only has to be loose enough not // to drop something that list would have caught. const LOOKS_LIKE_BOT = /bot|crawl|spider|gpt|claude|perplexity|bytespider|applebot|ccbot|meta-external/i // An ALLOWLIST of your public, crawlable pages — not a list of exclusions. // A blocklist fails open: the next route someone adds starts forwarding // customer data to a third party. Edit this to match your site. const LOGGED_PREFIXES = ['/docs', '/features', '/pricing', '/blog', '/faq'] function isLoggedPath(pathname: string): boolean { if (pathname === '/') return true return LOGGED_PREFIXES.some((p) => pathname === p || pathname.startsWith(p + '/')) } let buffer: string[] = [] let timer: ReturnType | null = null // Complain once, then stay quiet. Without this, a missing variable, a revoked // token and "no crawlers yet" all look identical. let warned = false function warnOnce(message: string) { if (warned) return warned = true console.warn(`[shruwd] log drain not working: ${message}`) } function flush() { if (timer) { clearTimeout(timer); timer = null } if (buffer.length === 0) return if (!ENDPOINT || !TOKEN) { warnOnce('SHRUWD_INGEST_URL or SHRUWD_INGEST_TOKEN is not set') buffer = [] return } const body = buffer.join('\n') buffer = [] // Not awaited: a network stall must never reach your request path. fetch(ENDPOINT, { method: 'POST', headers: { 'X-Shruwd-Ingest-Token': TOKEN, 'content-type': 'application/x-ndjson' }, body, }) .then((res) => { if (!res.ok) warnOnce(`endpoint returned ${res.status}`) }) .catch((error) => { warnOnce(`request failed: ${error instanceof Error ? error.message : String(error)}`) }) } export default defineEventHandler((event) => { // The query string is dropped, not just ignored — it carries ids on some // routes and a per-page crawl count needs none of it. const pathname = event.path.split('?')[0] ?? '/' if (!isLoggedPath(pathname)) return const ua = getRequestHeader(event, 'user-agent') ?? '' if (!LOOKS_LIKE_BOT.test(ua)) return // cf-connecting-ip first: behind Cloudflare it is the only client IP header // a visitor cannot forge. The first x-forwarded-for entry is the fallback. const ip = getRequestHeader(event, 'cf-connecting-ip') ?? (getRequestHeader(event, 'x-forwarded-for') ?? '').split(',')[0]?.trim() ?? event.node.req.socket.remoteAddress ?? '' if (!ip) return const started = Date.now() // Status and byte count only exist once the response is finished. event.node.res.once('finish', () => { try { buffer.push(JSON.stringify({ timestamp: new Date(started).toISOString(), host: getRequestHeader(event, 'host') ?? '', path: pathname, user_agent: ua, status: event.node.res.statusCode, client_ip: ip, method: event.method, bytes: Number(event.node.res.getHeader('content-length') ?? 0), })) if (buffer.length >= MAX_BATCH) flush() else if (!timer) timer = setTimeout(flush, MAX_WAIT_MS) } catch { // Never throw into the request path. } }) }) ``` **Edit `LOGGED_PREFIXES`** to your own public pages before deploying. It is an allowlist on purpose: your app probably serves customer content and token-bearing links from the same origin, and none of that should leave your servers. ## 3. Set the variables | Variable | Value | |---|---| | `SHRUWD_INGEST_URL` | the endpoint from step 1 | | `SHRUWD_INGEST_TOKEN` | the token from step 1 | Set them in your **deployed** environment, not only locally. That omission is the most common reason nothing arrives. > **Warning** > > If you switch to `useRuntimeConfig()` instead of `process.env`, two things must change > together. Declare the keys in `nuxt.config.ts`, **and** prefix the variables with `NUXT_` > (`NUXT_SHRUWD_INGEST_URL`). Miss either and the config reads `undefined`, the middleware > returns at its first check, and nothing is ever sent — with no error anywhere. ## 4. Verify ```bash curl -X POST "YOUR_ENDPOINT?validate=1" \ -H "X-Shruwd-Ingest-Token: YOUR_TOKEN" \ -H "Content-Type: application/x-ndjson" \ -d '{"timestamp":"2026-01-01T00:00:00Z","host":"example.com","path":"/","user_agent":"GPTBot","status":200,"client_ip":"1.2.3.4"}' ``` `usable: true` means the endpoint and token are good, so any remaining problem is in your app. Then deploy and wait for real crawler traffic — the buffer flushes at 20 events or 10 seconds, so a single test request will not appear instantly. ## Known limits - **Cached responses are invisible.** A request your CDN serves never reaches the middleware. "Zero hits on `/`" can mean "cached", not "never crawled". - **Buffered events are lost** when a serverless instance is reclaimed. Crawler hits feed rates and per-page counts, so losing a small fraction evenly does not bias them. - **A brand-new crawler can slip the prefilter** until the pattern is widened. ## Next steps - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — step-by-step diagnosis - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — what gets recognised --- # Crawlers tracked Shruwd recognises the AI crawlers below, verifies each hit against the vendor's published IP ranges, and separates fetches made to answer a question from batch crawling for training. ## Recognised crawlers These are the bots Shruwd identifies today, across OpenAI, Anthropic, Perplexity, Apple, ByteDance, Meta, Common Crawl and Amazon: `GPTBot` · `OAI-SearchBot` · `ChatGPT-User` · `ClaudeBot` · `Claude-User` · `Claude-SearchBot` · `PerplexityBot` · `Perplexity-User` · `Applebot` · `Bytespider` · `Meta-ExternalAgent` · `CCBot` · `Amazonbot` The list is maintained on Shruwd's side and updated as new crawlers appear, so you never need to change anything you deployed. Your brand's **Crawlers** page lists each bot that has reached you, with its vendor. ### Google-Extended and Applebot-Extended are not crawlers Both are `robots.txt` tokens that control how your content may be used, and neither ever appears in a log. Google says `Google-Extended` "doesn't have a separate HTTP request user agent string". Apple says `Applebot-Extended` "does not crawl webpages". The fetching is done by Googlebot and by Applebot. So neither has a row on the Crawlers page, and an empty one would tell you nothing. Shruwd checks `Google-Extended` where it does exist: the `robots.txt` diagnostic reports a rule that disallows it and says what that affects, which is Gemini grounding and AI training, not AI Overviews. See the [findings reference](https://shruwd.io/docs/findings/findings-reference.md). Google's AI Overviews and AI Mode are fed by Googlebot, the same crawler as Google Search. There is no separate Google AI crawler to look for in your logs. ## Live retrieval versus training Each bot is classified by what it is for, and the distinction matters more than the totals. | | What it is | Why you care | |---|---|---| | **Live retrieval** | Fetched to answer a specific person's question, right then | A page it cannot fetch cannot be cited in that answer. This is the number to watch | | **Training** | Batch crawling to build a model | Decides what a model learns, not what an answer cites today | `ChatGPT-User` and `Perplexity-User` are live-retrieval crawlers. `GPTBot` is a training crawler. Presenting them as one undifferentiated figure would hide the most interesting signal in the data, so Shruwd never does. ## How a hit is verified A user-agent string is a claim, not evidence — anyone can send `GPTBot` in a header. Shruwd checks the client IP of every hit against the ranges the vendor publishes. Those ranges are refreshed every 24 hours. If a refresh fails, the last known-good set is kept rather than marking everything verified or everything unverified — both would quietly corrupt the number. | Result | What happens | |---|---| | **Verified** | The IP belongs to the vendor. Counted in every metric | | **Unverified** | The user-agent claims a bot from an IP the vendor does not own. Shown on the Crawlers page, never counted | Unverified hits are displayed rather than hidden, because seeing them is useful — but they never reach a headline number. Where a vendor publishes no ranges at all and no reverse-DNS check is documented, its hits stay permanently unverified. ## Reading the Crawlers page Once logs arrive you get: - **Verified hits per day** — the timeline. It appears after two days of logs. - **By bot** — per-crawler verified hits, unverified hits, and errors. A non-zero error count means that crawler was refused, which is a problem worth a look. ## Diagnostics that need logs | Diagnostic | Needs | |---|---| | The server refuses a verified AI crawler | Logs showing the refusal on at least three separate days | | A page with demand is never crawled | **30 days** of uninterrupted logs, **and** [Search Console](https://shruwd.io/docs/tracking/connecting-search-console.md) for the demand signal | The three-day requirement rules out a transient outage being reported as a block. The 30-day requirement is a guard, not a delay for its own sake. A brand whose drain broke for a week must not be told a page was never crawled when the truth is that nobody was watching. Search Console is needed alongside it because "never crawled" is only worth raising for a page people actually search for. ## Next steps - [AI crawlers](https://shruwd.io/ai-crawlers) — the same list as a public reference, with each vendor's IP-range file - [Choosing a log path](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md) — if you have not set this up yet - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — what these diagnostics produce --- # Ingest not arriving **Settings → Crawler logs** still reads *awaiting first logs* after you deployed. Work through these in order — the first two catch most cases. ## 1. Test the endpoint and token directly This separates configuration from your code entirely, and you can run it from anywhere. `?validate=1` parses your sample and reports on it without writing anything. ```bash curl -X POST "YOUR_ENDPOINT?validate=1" \ -H "X-Shruwd-Ingest-Token: YOUR_TOKEN" \ -H "Content-Type: application/x-ndjson" \ -d '{"timestamp":"2026-01-01T00:00:00Z","host":"example.com","path":"/","user_agent":"GPTBot","status":200,"client_ip":"1.2.3.4"}' ``` | Response | Means | |---|---| | `usable: true` | Endpoint and token are both fine — the problem is in your app or platform. Go to step 2 | | `401` | Wrong or revoked token. See below | | `404` | Wrong brand in the endpoint URL | | `422` | The sample can't be used. `reason` says why | > **Warning** > > **Creating a token revokes the previous one.** If you minted a new token in Settings after > setting up the drain, the one your app is using is dead. Only the newest token works. ## 2. On Vercel, check the drain In **Team Settings → Drains**: - The drain isn't marked as errored or paused. - Its sources include `static`, `lambda`, `edge`, `external`, `redirect` and `firewall`. - Its environments include `production`, and it has no sampling rule. Then check **Team Settings → Security & Privacy → IP Address Visibility** is on. Hidden IPs look like a working drain: logs arrive and the card reads *active*, but every line is refused, so the Crawlers page stays empty. ## 3. Check the variables in the deployed environment Set locally and forgotten in the host's configuration is the single most common cause. Check the values in the environment your site actually runs in, not your local `.env`. If you are using [server middleware](https://shruwd.io/docs/crawler-logs/server-middleware.md), it logs one line the first time a send fails — look for `[shruwd] log drain not working:` in your app's logs. That line names the cause. If you switched the middleware to `useRuntimeConfig()`, confirm both halves: the keys declared in `nuxt.config.ts` **and** the `NUXT_` prefix on the variables. Missing either reads `undefined` silently. ## 4. Check whether the request was served from cache This only applies to server middleware. A request your CDN answers never reaches your origin, so the middleware never sees it. Check `cf-cache-status` on the response: | Value | Reached your origin? | |---|---| | `DYNAMIC` or `MISS` | Yes | | `HIT` | No — the middleware could not have seen it | If most of your pages are cached, middleware is the wrong path. Switch to the [forwarder Worker](https://shruwd.io/docs/crawler-logs/cloudflare-worker.md) or a [platform drain](https://shruwd.io/docs/crawler-logs/platform-log-drains.md). > **Note** > > `cf-cache-status: DYNAMIC` means Cloudflare did not serve it from Cloudflare's cache. It > does **not** guarantee the request reached your application — a second CDN behind it can > still answer from static files. ## 5. Wait a little Middleware batches: it sends at 20 events or 10 seconds, whichever comes first. One test request will not appear immediately. Crawler traffic is also genuinely intermittent. A low-traffic site may wait hours between real AI-crawler hits. ## Other things worth knowing **Logs arrive but the card shows gaps.** *Days with data* on the Settings card shows how many of the last 30 days received anything. Gaps switch off the diagnostic that needs 30 uninterrupted days. **Logs arrive but nothing is counted.** Check the **Unverified** column on the Crawlers page. Hits from IPs outside the vendor's published ranges are shown but never counted — which is correct, and still proves your pipe works. **It stopped partway through the month.** Check **Lines this month** against your plan's allowance. Past the cap, lines are discarded until the month turns. The Crawlers page says so rather than showing a flat chart. **A `413` response.** Your batches are over the size cap. On Cloudflare Logpush, set `max_upload_bytes` to `5000000`. ## Next steps - [Choosing a log path](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md) — if you picked the wrong one - [Crawlers tracked](https://shruwd.io/docs/crawler-logs/crawlers-tracked.md) — how verification works --- # How findings work A finding names something specific that is costing you visibility, the page it applies to, and what to do about it. Working through findings is how you use Shruwd. Findings are re-evaluated after every cycle, and access findings every day on every plan. New ones are capped at five per brand per week — a list of forty gets nothing done. ## What a finding tells you | | | |---|---| | **What to do** | Numbered steps naming a specific page and change | | **Where** | The URL or prompt the finding applies to | | **Should move** | The metric you should expect to change if the fix works | | **Why we think so** | The evidence behind it | | **Confidence** | How much to trust it — see below | | **Severity** | `critical`, `high`, `medium` or `low` | ## Confidence: how much to trust it Confidence is shown on every finding, and the three levels deserve different responses. | Level | What it means | |---|---| | **Observed** | A first-party fact. Shruwd watched this happen in your own logs or in a live fetch of your page. | | **Inferred** | Not directly observed, but the causal chain is strong and every link is checkable in the evidence. | | **Heuristic** | A correlation from comparing your pages with the ones being cited. A lead, not a fact. | Do an observed finding first. Treat a heuristic one as a hypothesis worth testing. ## Working a finding **Permissions:** editor or above. Findings are grouped into three tabs: **To do**, **In progress** and **Closed**. 1. Open a finding from the **To do** tab. 2. Read **What to do** and **Why we think so**. If it does not apply, dismiss it with a reason. 3. Make the change on your site. 4. Select **Mark fix applied** and note what you changed. Shruwd records your current numbers as the baseline to compare against. 5. Wait. Shruwd re-measures after 14 days, so there is time for the change to be crawled and to show up in answers. 6. Read the verdict in the **Closed** tab. > **Warning** > > Mark the fix applied *after* the change is live. The baseline is captured at that moment, > and a baseline taken before the change is what makes the comparison meaningful. **Access findings skip the wait.** Shruwd confirms the fix from your own logs, robots.txt or page as soon as it shows up, on every plan, with no recheck spent. A robots.txt fix resolves even if you never mark it applied. ## The states a finding moves through | State | Meaning | |---|---| | **open** | The rule fired and nobody has picked it up | | **in hand** | You acknowledged it | | **fix applied** | You made the change; the baseline is captured | | **rechecking** | Shruwd is re-measuring against that baseline | | **resolved** | The recheck showed real movement, or your own data confirmed an access fix | | **did not move** | The recheck ran and the number did not move enough to call | | **dismissed** | You rejected it. Quiet for 90 days, then it can return | | **stale** | The rule stopped firing, with no fix recorded from you | **Resolved** and **stale** are different on purpose. Stale means the condition went away on its own — Shruwd will not claim your fix did it. A finding marked **did not move** re-opens after 30 days if the condition still holds, with a note that the previous fix did not work. That is worth knowing. ## Why a finding shows no recheck Rechecks are limited by plan: none on Crawl, one a month on Starter, six a month on Pro. If you are out of rechecks, the finding waits until next month and says so. Access findings need none. See [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md). ## Next steps - [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md) — what counts as movement - [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md) — which findings your plan shows --- # Rechecks A recheck is how Shruwd answers the only question that matters after you make a change: did it work? Access findings do not need one. Their fix shows up in your own logs, robots.txt or page, and Shruwd resolves them when it does — see [How findings work](https://shruwd.io/docs/findings/how-findings-work.md). ## What happens after you mark a fix applied 1. **Your current numbers are captured as a baseline** — the metric values, their intervals, the sample size and the timestamp. This is the "before" the recheck compares against. 2. **Shruwd waits 14 days.** AI answer surfaces re-crawl and re-index on their own schedule. Rechecking on day two measures nothing and would report a false *did not move*. 3. **A recheck cycle runs** on the prompts the finding affects, at **nine repetitions** instead of the usual three. A small prompt subset needs the extra sampling for the result to be conclusive either way, and because the subset is small it is cheap. 4. **The verdict appears** in the finding's **Closed** tab. > **Warning** > > Mark the fix applied only once the change is live. The baseline is captured at that > moment, so marking it early compares your change against itself. ## Reading the verdict | Verdict | Meaning | |---|---| | **resolved** | The numbers moved, and the movement passed both gates | | **did not move** | The recheck ran and the change was not big enough or not distinguishable from noise | The prompts are compared together and one by one. Together is the fix's own claim — it was made for those prompts — and it is the comparison a recheck can usually decide: one prompt alone rarely has enough answers to show anything but a very large change. A fix aimed at one prompt still counts if that prompt moves on its own. *Did not move* is a real result, not a failure of the product. Most single-page changes do not shift a prompt-set-wide number by five percentage points, and saying so is the whole reason for measuring. A finding marked *did not move* **re-opens after 30 days** if the condition still holds, noting that the previous fix did not work. ## Why a fix that looks like it worked is not called movement The verdict uses the same rule as everything else in Shruwd: the intervals before and after must not overlap, **and** the change must be at least five percentage points, after correcting for the fact that many prompts are being checked at once. See [When a number has moved](https://shruwd.io/docs/metrics/when-a-number-has-moved.md) for the full rule. If the provider changed its underlying models between the baseline and the recheck, the comparison is marked **model changed** — the difference may be the new model rather than your fix. ## Recheck limits Rechecks are capped per month by plan. | Plan | Rechecks per month | |---|---| | Crawl | None | | Starter | 1 | | Pro | 6 | If you are out, the finding waits until next month and says so on the fix panel. Spend them on findings where you actually changed something substantial. ## Next steps - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — the full finding lifecycle - [When a number has moved](https://shruwd.io/docs/metrics/when-a-number-has-moved.md) — the movement rule - [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md) — what your plan includes --- # Findings reference Every diagnostic Shruwd runs. For how findings behave — states, confidence, rechecks — see [How findings work](https://shruwd.io/docs/findings/how-findings-work.md). ## Reading an entry Each rule carries a **confidence** and a **severity**, and both change how you should treat it. | Confidence | | |---|---| | `observed` | A first-party fact, from your own logs or a live fetch | | `inferred` | Not directly observed, but every link in the chain is in the evidence | | `heuristic` | A correlation against the pages being cited. A lead, not a fact | ## Which findings come first Families are not equal, and one can make another pointless. **An open access finding on a URL hides the structure and citation findings for that same URL.** Recommending schema markup on a page that returns 403 to an AI crawler wastes your afternoon — fix the access problem and the rest reappear, re-evaluated. Findings are also capped at five new per brand per week. A list of forty gets nothing done. ## Where each family applies Prompt intent decides which diagnostics run. | Family | informational | comparison | commercial | navigational | problem | |---|:--:|:--:|:--:|:--:|:--:| | Access | ✓ | ✓ | ✓ | ✓ | ✓ | | Citation | ✓ | ✓ | ✓ | — | ✓ | | Structure | ✓ | ✓ | ✓ | — | ✓ | --- ## Access Whether AI crawlers can reach your pages at all. The highest-confidence family — the inputs are your own logs and live fetches, with no inference. Checked every day, on every plan. ### robots.txt blocks an AI crawler `ACCESS_ROBOTS_BLOCK` · observed · **critical** **Fires when** your `robots.txt` disallows `OAI-SearchBot`, `ChatGPT-User`, `PerplexityBot` or `Google-Extended` for a page that matters — one in your sitemap, one with search impressions, or one already being cited. **What to do** — the finding shows the offending directive with its line number and an exact replacement block as a diff. **Resolves when** the directive is gone. Shruwd rereads robots.txt daily, and within minutes once you mark the fix applied. > **Note** > > `Google-Extended` affects Gemini grounding and AI training — **not** AI Overviews. These > get conflated constantly, and the right fix differs. ### The server refuses a verified AI crawler `ACCESS_4XX_5XX_TO_AI_UA` · observed · **critical** The most valuable rule in the catalogue, and the one that needs [crawler logs](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md). **Fires when** a verified AI crawler was refused (4xx or 5xx) more often than served on a page, on at least three separate days in the last 30, while a normal browser gets a 200 for the same URL. An occasional error does not count. **What to do** — this is almost always bot protection nobody remembers enabling: Cloudflare Bot Fight Mode or AI Crawl Control, a WAF managed rule, rate limiting, a security plugin, or user-agent blocking at your origin. The finding identifies the likely blocker from the response signature and names the setting to change. Which of these your logs can show depends on your [log path](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md). Only Logpush records refusals made by Cloudflare itself: the forwarder Worker runs after Cloudflare's security rules, and your server after that. So with the Worker, look for Cloudflare's own refusals in Cloudflare (**Security → Analytics**). A refusal Shruwd does see was made behind Cloudflare, and the finding tells you so. **Resolves when**, after you mark the fix applied, that crawler is served on that path on a day it visits. ### Content only renders with JavaScript `ACCESS_JS_ONLY_CONTENT` · inferred · **high** **Fires when** the raw HTML has under 500 characters of visible text, the page has search impressions, and the markup contains an empty root mount node. **What to do** — server-render, statically generate or prerender that route. AI crawlers largely do not execute JavaScript, so content that only exists after hydration does not exist for them. **Resolves when** the raw HTML carries more than 2,000 characters of visible text. ### A page with demand is never crawled `ACCESS_NEVER_CRAWLED` · observed · **medium** **Fires when** a URL is in your sitemap, has at least 100 search impressions in 30 days, and has had **zero** verified AI-bot hits in that time. Needs 30 days of uninterrupted [crawler logs](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md) and a [Search Console connection](https://shruwd.io/docs/tracking/connecting-search-console.md). It stays silent on brands with partial log coverage rather than blaming you for a gap in the data. **What to do** — the page is reachable but undiscovered. Check how deep it sits from your homepage, whether the sitemap includes it with an accurate `lastmod`, and whether anything the crawlers already visit links to it. --- ## Citation Where answers get their sources. The highest-value family after access. ### A third-party page gatekeeps the answer `CITE_THIRD_PARTY_GATEKEEPER` · inferred · **critical** The finding that answers "why does ChatGPT recommend them and not us?" **Fires when** across a cluster of at least three related prompts, one third-party domain appears in 40% or more of the answers, you are absent from the page being cited, and at least one competitor is on it. **What to do** — getting onto that specific source is usually the highest-leverage action available to you. The finding names the exact URLs, which competitors appear on them and how they are described, and the route in for that kind of source: a vendor listing submission, an editorial contact, community participation. Rechecked after 30 days rather than 14 — third-party inclusion propagates slowly. ### The wrong page gets cited `CITE_WRONG_PAGE` · observed · **high** **Fires when** you are cited for a prompt, but the cited URL does not match the prompt's intent — a blog post cited for a buying question when you have a pricing page. **What to do** — the right page is under-optimised relative to the one being cited. Strengthen it against that intent, and link to it prominently from the page currently getting the citation. ### A competitor owns this prompt `CITE_COMPETITOR_OWNS_PROMPT` · observed · **high** **Fires when** over 30 days one competitor's mention rate is at or above 60% while yours is at or below 15%, across at least ten responses. **What to do** — the finding derives a content gap from the competitor pages being cited: what those pages cover that your target page does not. Rechecked at nine repetitions after 30 days, so the verdict is conclusive on a small prompt subset. ### Your page reads as stale `CITE_STALE_SOURCE` · inferred · **medium** **Fires when** a cited page of yours has no date signal newer than 18 months, in a category where the competitor pages being cited average under six months. **What to do** — refresh the content and the date signals together. > **Warning** > > Changing the date without changing the content is not the fix. Models extract facts, and > stale facts get superseded whatever the date says. --- ## Structure How extractable your content is. Every rule here is `heuristic` — a lead worth testing, not a fact. ### No direct answer near the top `STRUCT_NO_DIRECT_ANSWER` · heuristic · **high** **Fires when** for a prompt where a competitor is cited and you are not, your most relevant page has no 40–120 word paragraph in the first 1,200 characters sitting under a heading that matches the question. **What to do** — add a heading in the question's own form, and a short direct answer immediately beneath it, before any elaboration. The finding shows the passage from each cited competitor page that does answer it, side by side. Needs at least two cited competitor pages to compare against. ### Structured data the cited pages have `STRUCT_MISSING_SCHEMA` · heuristic · **medium** **Fires when** your page carries no JSON-LD of a type present on at least half of the competitor pages cited for the prompts it targets. **What to do** — the finding generates a JSON-LD block populated from your page's actual content, not a template with placeholders to fill in. ### Fewer concrete facts than the cited pages `STRUCT_LOW_FACT_DENSITY` · heuristic · **medium** **Fires when** your page's fact density is below 60% of the mean of the cited competitor pages. Density counts numerals, dates, percentages, currency amounts, measurement units and proper nouns per 100 words. **What to do** — add specifics of the kinds those pages carry: concrete numbers, dates, named integrations, versions. Models cite pages containing extractable facts. Never fires on pages under 300 words — that is a different problem. --- ## Next steps - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — confidence, states, acting on one - [Rechecks](https://shruwd.io/docs/findings/rechecks.md) — proving a fix worked - [Choosing a log path](https://shruwd.io/docs/crawler-logs/choosing-a-log-path.md) — unlocks two access findings --- # Metric definitions What each number counts, precisely. For an orientation to reading them, start with [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md). Every metric is computed over a **scope**: one brand, one engine, a date range, and optionally a subset of prompts. `N` is the number of responses in that scope. ## Mention rate How often you are named. ``` mention_rate(entity) = responses mentioning the entity / N ``` **A mention is counted once per response.** An answer that names you six times counts the same as one that names you once. Counting every occurrence would reward long answers and make the number swing on formatting changes that have nothing to do with you. **A competitor is counted from when you added it.** Answers from before then were never checked for it, so its rate is over the answers since. A newly added competitor shows *insufficient data* until ten answers have been checked for it. ## Share of voice How much of the naming goes to you rather than the rest of the tracked set. ``` share_of_voice(entity) = mentions of the entity / mentions of everyone tracked ``` The tracked set is your brand plus the competitors you have added. Answers that name nobody in that set count toward neither side. Shares are over the answers checked against the whole current set, so they always add up to 100%. When you add a competitor, share of voice starts again from the next answers and shows *insufficient data* until there are ten. Removing one doesn't restart it. When no tracked brand is named at all in the scope, share of voice is **undefined**, not zero. You will see *no brand mentioned*. Zero and undefined mean different things, and treating a dead prompt set as a losing one is the wrong conclusion. > **Warning** > > Share of voice rises when a competitor falls. Always read it beside mention rate — if > share of voice is up and mention rate is flat, nothing about you changed. ## Prominence Where you appear in an answer, not just whether. ``` prominence = 1 − (characters before your first mention / total characters) ``` The result runs from just above 0 to 1, where higher means earlier. It is averaged only over responses where you were actually mentioned, so it is undefined until you have been named at least once — the dashboard shows *Not mentioned yet*. Alongside it, **mention rank** is the order of first appearance among the tracked entities: who the answer names first, second, third. Prominence is never folded into share of voice or into a single composite "visibility score". A blended score cannot be debugged or argued with, so Shruwd shows the parts. ## Citation metrics A citation is a URL in the answer's list of cited sources. ``` citation_share(domain) = responses citing that domain / N self_citation_rate = responses citing any of your domains / N ``` A domain is counted **once per response**, however many of its pages are cited. Which specific page was cited is kept, because that is what several diagnostics need — a competitor citing your pricing page and a review site citing your homepage are different problems. ## Coverage metrics Two numbers describe what happened in the answers that did not name anyone. | Metric | Counts | What it means | |---|---|---| | **Naming nobody** | Responses naming no tracked brand ÷ `N` | Nobody owns the category here. Across a group of prompts, this is an opportunity rather than a loss. | | **No AI answer** | Runs where the engine produced no AI answer ÷ `N` | The engine does not answer these prompts. A prompt-set problem, not a visibility problem. | If **no AI answer** is high, nothing you change on your site will move the other numbers for those prompts. Rewrite the prompts instead. ## Next steps - [Confidence and sample size](https://shruwd.io/docs/metrics/confidence-and-sample-size.md) — why a number may be withheld - [When a number has moved](https://shruwd.io/docs/metrics/when-a-number-has-moved.md) — the rule for calling a change real --- # Confidence and sample size AI answers vary between runs. Ask the same question twice and you can get two different sets of brands. Every number Shruwd shows is therefore an estimate from a sample, and it is shown with the range that sample supports. ## Reading the range A result reads `32% (24–41%)`. That range is a 95% confidence interval. The point estimate is the best single guess. The range is how much the sample constrains it. Two results with the same middle can carry completely different weight: | Result | What it supports | |---|---| | `32% (29–35%)` | A solid number. Act on it. | | `32% (10–60%)` | Barely more than "somewhere in the middle". Wait for more data. | Intervals narrow as responses accumulate. A prompt set measured for three months has much tighter ranges than one measured for two weeks, with no change in your visibility. ## Why a number is sometimes withheld Shruwd never displays a metric computed from fewer than **ten responses**. You get **insufficient data** and a count of how many more runs are needed. This is not an error, not a loading state, and not zero. It means the sample cannot support a number yet, and printing one anyway would invite a decision the data does not justify. Where you will meet it: - **Individual prompts** hit it often. Each prompt is asked three times per cycle, so a single prompt needs several weeks of cycles to clear ten responses. - **Your whole prompt set** clears it immediately, because every prompt contributes. - **Narrow filters** — one prompt, one engine, a short date range — can drop back below the floor. Widen the range or the selection. ## How the ranges are calculated You do not need this to use the product, but it is here so the numbers can be checked. **Rates** — mention rate, citation share, self-citation rate and the coverage rates — use a **Wilson score interval**. These are proportions that sit near 0% or 100% more often than not, exactly where the textbook normal approximation produces nonsense like a lower bound below zero. **Share of voice** uses a **percentile bootstrap** over 1,000 resamples instead. It is a ratio where the same response feeds both the numerator and the denominator, so it is not a simple proportion and a Wilson interval would misstate it. The bootstrap's randomness is seeded per brand, entity, engine and day, so recomputing a past day returns exactly the same interval. An interval that shifted every time the page loaded would be worth nothing. ## Next steps - [Methodology](https://shruwd.io/methodology) — why Shruwd works this way, on one page - [When a number has moved](https://shruwd.io/docs/metrics/when-a-number-has-moved.md) — comparing two periods - [Metric definitions](https://shruwd.io/docs/metrics/metric-definitions.md) — what each number counts --- # When a number has moved Two measurements will almost always differ. Most of that difference is sampling noise. Shruwd calls a change **movement** only when it passes both gates below. ## Both gates must pass **Gate 1 — the difference is statistically distinguishable.** The two periods are compared with a two-proportion test, and the 95% intervals must not overlap. **Gate 2 — the difference is big enough to care about.** The change must be at least **5 percentage points**. Either gate alone gives the wrong answer. Gate 1 by itself will flag a change of half a percentage point once the sample is large. Gate 2 by itself will flag pure noise when the sample is small. Together they mean a reported movement is both real and worth your attention. ## Testing many prompts at once Checking dozens of prompts every week, a few will look significant by chance alone — forever, no matter how well the product works. Shruwd applies a false-discovery-rate correction across every comparison in one evaluation, so the flagged movements are the ones that survive being one of many tests. The batch is one brand's comparisons; results are never pooled across brands. The practical effect: a change you can see in the chart may not be labelled as movement. That is the correction doing its job. ## When the engine changed underneath you Providers update the models behind their answers. When that happens between two periods, the comparison is marked **model changed**. The difference may be real, or it may be the new model behaving differently. Shruwd flags it rather than silently attributing it to something you did. > **Note** > > This same rule decides the verdict on a finding you have fixed, with the fix's prompts > compared together as well as one by one. See [Rechecks](https://shruwd.io/docs/findings/rechecks.md). ## What this means in practice - Do not read week-to-week wobble as progress or decline. Look at the range, not the point. - A fix that produced no labelled movement has not been proven to work — that is honest, and it is the point of measuring. - Real movement usually needs a real change. Fixes that alter a single page rarely shift a prompt-set-wide number by 5 points. ## Next steps - [Confidence and sample size](https://shruwd.io/docs/metrics/confidence-and-sample-size.md) — where the ranges come from - [Rechecks](https://shruwd.io/docs/findings/rechecks.md) — proving a specific fix worked --- # Plans and limits ## What each plan includes | | **Crawl** (free) | **Starter** | **Pro** | |---|---|---|---| | Brands | 1 | 1 | 3 | | Prompts, shared across brands | 10 | 25 | 100 | | Engines | both | both | both | | How often | one snapshot | weekly | weekly | | Repetitions per prompt | 1 | 3 | 3 | | Rechecks per month | — | 1 | 6 | | History | 30 days | 90 days | 400 days | | Finding families | access only | all | all | | Findings shown in full | 3 | all | all | | Crawler log lines per month | 100k | 500k | 2M | | Search Console | — | — | ✓ | | Team members, owner included | 2 | 5 | 15 | Both engines — Google AI Overviews and ChatGPT — are measured on every plan. Annual billing is ten months' price. ## About the free plan Crawl is not a trial. It does not expire and it takes no card. It measures your domain **once**, ever. That snapshot is tied to the domain rather than the account, so re-registering does not produce a second one. What continues indefinitely is crawler-log ingest and the access findings built on it, checked every day. ## What happens at a limit Every limit is a hard stop — there is no overage billing. Shruwd tells you where the limit bites rather than after a failed action. | Limit | What happens | |---|---| | **Brands** | Adding is refused. Archive a brand to free the slot. | | **Prompts** | The batch is refused whole, with the numbers. Nothing is partially added. | | **Team members** | Invites are refused. Open invites count toward the cap. | | **Findings** | The top findings are readable in full; the rest are counted but locked. | | **Rechecks** | The finding waits for next month. | | **Crawler log lines** | Further lines are discarded for the rest of the month. The crawlers page says so rather than showing a flat chart. | | **History** | Older data is hidden from view, never deleted. Upgrading reveals it immediately. | ## Changing plan **Permissions:** owner. Go to **Account → Billing**. Upgrades apply immediately. A downgrade never removes anything. If it leaves you over a limit — more brands, prompts or members than the new plan includes — everything keeps working and you cannot add more until you are back under. Hidden history returns if you upgrade again. ## Next steps - [Brands](https://shruwd.io/docs/tracking/brands.md) — what counts toward the brand limit - [Prompts](https://shruwd.io/docs/tracking/prompts.md) — the shared prompt pool --- # Team and roles Everyone in a workspace can see every brand in it. Their **role** decides what they can change. ## What each role can do | | viewer | editor | admin | owner | |---|:--:|:--:|:--:|:--:| | See everything, including the team list | ✓ | ✓ | ✓ | ✓ | | Prompts, competitors, finding transitions, brand name and timezone, run a cycle | | ✓ | ✓ | ✓ | | Invites, roles, removing people, connections, ingest tokens | | | ✓ | ✓ | | Create or archive a brand, billing | | | | ✓ | An API key never exceeds the role of the person who created it. ### The owner There is exactly one owner per workspace, and a person owns at most one workspace. The owner cannot be removed or demoted, and ownership cannot currently be transferred. The owner also receives the [weekly digest and alerts](https://shruwd.io/docs/account/email-notifications.md), and is the only person who can reach [Billing](https://shruwd.io/docs/account/billing.md). ## Inviting someone **Permissions:** admin or owner. 1. Open a brand and go to its **Team** page. 2. Enter the person's email address and choose a role. 3. Select **Invite**. They get an email with a join link that expires in **seven days** and works once. > **Warning** > > **A membership comes from opening the link while signed in — nothing else.** Signing up > with the invited address does not join anyone, and a forwarded link admits whoever opens > it, under their own account. The Members list shows the email of the account that actually > joined, so you can always see who came in. ### Seats Each plan includes a number of members, the owner included: 2 on Crawl, 5 on Starter, 15 on Pro. **Open invites count toward the limit**, so revoke stale ones to free a seat. If the team is full, the invite is refused. If someone accepts when the team has since filled up, their invite stays open until a seat frees. There is a cap of 20 invites per workspace per 24 hours. Revoking and resending counts against it. ## Managing open invites The **Invites** card lists everything outstanding. You can **resend** or **revoke** any of them. Revoking takes effect immediately — the link stops working. ## Changing someone's role **Permissions:** admin or owner. Pick a new role from the member's row. You cannot change the owner's role, and you cannot change your own. ## Removing someone **Permissions:** admin or owner — or yourself, to leave. Removing someone takes effect at once, and **their API keys for this brand stop working**. Their history is kept, so a finding can still show who acted on it. To leave a workspace yourself, remove your own row. You lose access immediately, and anyone on the team can invite you back. The owner cannot be removed. ## Next steps - [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md) — how many members your plan includes - [Billing](https://shruwd.io/docs/account/billing.md) — owner only - [Email notifications](https://shruwd.io/docs/account/email-notifications.md) — who receives what --- # Billing Everything billing-related lives under **Account → Billing**. **Permissions:** owner. Billing is the one thing an admin cannot do. ## What the page shows | Card | | |---|---| | **Plan** | Your current plan and what it includes. A pending change, if you have one | | **Usage** | Runs and rechecks used against your allowance this month | | **Payment** | Your card and invoices, through Stripe | | **Cancel plan** | Ends the subscription at the end of the period | Usage resets on the first of each month. ## Changing plan Pick a plan from **Upgrade**, reached from any upgrade link in the app. Every one of those links names the limit it came from, so you land on the plan comparison with the relevant row marked. | | When it applies | Proration | |---|---|---| | **Upgrade** | Immediately | Yes — you are invoiced the difference | | **Downgrade** | At the end of the current period | No | An upgrade takes effect on your next request: new limits apply at once, the next cycle uses them, and history that was hidden by the old plan's window becomes visible immediately. A downgrade changes nothing until the period ends. You keep everything until then. ### Downgrades that need tidying first If you have more brands, prompts or team members than the new plan includes, the downgrade is accepted but **flagged**. You will be asked to get under the limit before the period ends. Nothing is ever deleted to make a downgrade fit. Bring the counts down yourself — archive a brand, remove prompts, remove members — and the flag clears the next time the page is read. Choosing your current plan again cancels a pending change. ## Annual billing Annual is ten months' price — two months free. Tax is added at checkout where it applies. ## Cancelling Cancelling ends the subscription at the end of the period. You are not deleted: the workspace moves to the free **Crawl** plan and keeps its history, findings and rollups. If you subscribe again later, it is all still there. ## When a payment fails | Day | What happens | |---|---| | 0 | The payment fails. A banner appears on every page for the owner. Measurement continues | | 5 | The owner is emailed with the date measurement stops | | 7 | Grace ends. Scheduling stops — no new cycles | | Up to 14 | Stripe retries the payment several times | | 14 | If it never succeeds, the subscription is cancelled and the workspace returns to the free plan | Fixing the card during the grace period clears everything — the banner goes and cycles continue uninterrupted. **Your data is kept throughout.** Retention does not depend on your plan, so a workspace that lapses and comes back keeps its rollups and findings. ## Next steps - [Plans and limits](https://shruwd.io/docs/account/plans-and-limits.md) — what each plan includes - [Team and roles](https://shruwd.io/docs/account/team-and-roles.md) — who can reach billing --- # Email notifications Shruwd sends few emails. This is all of them. ## The weekly digest One email a week with new findings, finished rechecks and your latest numbers. It covers every brand in your workspace, not one per brand. **It is skipped when a week has nothing new.** No "nothing happened" emails. ### Turning it on or off 1. Open the brand's **Settings**. 2. Go to the **Notifications** tab. 3. Toggle **Weekly digest**. **Permissions:** owner. Emails about a brand go to the workspace owner, so the switch is theirs — other members see a note saying so rather than a control they cannot use. You can also unsubscribe from the link at the bottom of any digest. ## Alerts These are sent to the owner when something needs attention. They are not optional, because each one means the product is about to stop doing something you are paying for. | Email | Sent when | |---|---| | **Payment failing** | Two days before measurement stops, after a failed payment | | **Crawler lines at 80%** | You have used 80% of the month's log-line allowance | | **Crawler lines on pace** | Your current rate would use the rest before the month ends | | **Crawler lines used up** | The allowance is gone and lines are being discarded | | **Log ingest stopped** | A brand that was sending logs has sent none for 48 hours | | **AI crawler blocked** | The daily check finds robots.txt blocking, or your server refusing, a crawler that answers questions | | **Reconnect Search Console** | Google stopped accepting the connection | | **Recheck verdict** | A fix was confirmed, or a recheck finished. Also sent to whoever marked it applied | The ingest alert is worth acting on quickly — a silent drain looks identical to "no crawlers visited", and the diagnostics that depend on continuous logs switch off. ## Invitations Inviting someone sends them a join link that expires in seven days. That email deliberately contains no free text — just the domain, who invited them, the role, and the button. Nobody has agreed to hear from Shruwd before they accept, and an invite email carrying typed-in content is a phishing template waiting to happen. ## Next steps - [Team and roles](https://shruwd.io/docs/account/team-and-roles.md) — who receives what - [Ingest not arriving](https://shruwd.io/docs/crawler-logs/ingest-not-arriving.md) — if you get the ingest alert - [Billing](https://shruwd.io/docs/account/billing.md) — if you get the payment alert --- # API overview Everything the dashboard does, the API does — it is the same surface, not a mirror of one. ``` https://shruwd.io/api/v1 ``` Paths in these docs are relative to that base, as in the OpenAPI document: `POST /brands` is `POST https://shruwd.io/api/v1/brands`. JSON in, JSON out, UTF-8. The contract describes itself in OpenAPI 3.1 at [`/openapi.json`](https://shruwd.io/openapi.json), which needs no key to read. The same document is also served at `/api/v1/openapi.json`. ## Authenticating Mint a key in the dashboard under **Account → API keys**, then send it as a bearer token. ```bash curl "https://shruwd.io/api/v1/workspace" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Keys can only be created from a signed-in session. A key uses the product; it can never administer the account — it cannot mint other keys, change billing, or delete anything that is not brand configuration. ### Scopes | Scope | Grants | |---|---| | `read` | Every `GET` | | `write` | Brand, prompt and entity writes; finding transitions; manual cycles; ingest tokens | | `default` | Both | A request outside its key's scopes is `403 insufficient_scope`. ### Roles A request acts with your role in the workspace. | Role | Can | |---|---| | `viewer` | Every read | | `editor` | Prompts, competitors, finding transitions, brand name and timezone, run a cycle | | `admin` | The above, plus invites, connections and ingest tokens | | `owner` | The above, plus creating and archiving brands, and billing | A key never exceeds its holder's role. Refusals are `403 not_a_member` and `403 insufficient_role`, the latter carrying `details: { required, role }`. ## Workspaces are implicit Your key determines the workspace. No path or body ever carries a workspace id. If you own no workspace, your first `POST /brands` creates one. There is no workspace resource to set up first. A brand or finding id belonging to another workspace returns **`404`**, not `403` — indistinguishable from an id that does not exist. ## Identifiers Ids are ULIDs and opaque. Two human-readable handles sit beside them, and both are accepted wherever the id is: | Handle | Used by | Example | |---|---|---| | A brand's `slug` | `{brandId}` | `/brands/waitlister` | | A finding's `number` | `{findingId}` | `/findings/7` | Finding numbers start at 1 per workspace, in first-seen order, and are never reused. ## Every metric carries its uncertainty There is no shape in this API that returns a bare number. ```json { "state": "ok", "point": 0.142, "lo": 0.081, "hi": 0.226, "n": 42 } ``` `point`, `lo` and `hi` are proportions from 0 to 1: 0.142 is 14.2%. `state` is `ok`, `insufficient_data` (fewer than ten responses — **not zero**), or `undefined`. See [Reading your results](https://shruwd.io/docs/getting-started/reading-your-results.md). ## Errors The response body **is** the error. There is no wrapper. ```json { "code": "plan_limit", "message": "…", "retryable": false, "details": { "limit": 25, "current": 25 } } ``` | Field | | |---|---| | `code` | Stable and machine-readable. Branch on this, never on `message` | | `message` | Human-readable. May change at any time | | `retryable` | Whether retrying the identical request could succeed | | `details` | Present when there is something specific to say | ### Common codes | Code | Status | Meaning | |---|---|---| | `not_found` | 404 | No such resource, or it belongs to another workspace | | `insufficient_scope` | 403 | The key lacks `read` or `write` | | `not_a_member` · `insufficient_role` | 403 | Your role is too low | | `not_entitled` | 403 | Your plan does not include this | | `plan_limit` | 409 | A plan cap is used up. `details` carries the numbers | | `rate_limited` | 429 | See below. `retryable: true` | | `invalid_domain` · `invalid_name` · `invalid_timezone` | 422 | Bad brand input | | `duplicate_domain` | 409 | Another active brand here has that domain | | `duplicate_prompt` | 409 | That prompt is already active on the brand | | `context_terms_required` | 422 | A short or common-word competitor name needs context terms | | `domain_taken` | 409 | Another entity of this brand owns that domain | | `self_entity` | 409 | The brand's own entity cannot be removed | | `snapshot_taken` | 409 | A free-plan domain has already had its one measurement | | `workspace_on_hold` | 409 | Measurement is paused pending review | | `suggestion_decided` | 409 | That suggestion was already accepted or dismissed | | `empty_patch` | 400 | A `PATCH` with no fields | ## Rate limits Per key, in fixed one-minute windows. | | Per minute | |---|---| | `GET` | 120 | | Everything else | 30 | Over the limit is `429 rate_limited` with `Retry-After` in seconds. Honour it — the [SDK](https://shruwd.io/docs/api/typescript-sdk.md) does this for you. ## Times and ranges Timestamps are ISO 8601 UTC. Days are `YYYY-MM-DD`. > **Note** > > Date ranges default to **today in the brand's timezone**, which is the calendar the > measurements are bucketed on. UTC today can be a day ahead of it. `from` is clamped to your plan's history window. When that happens the response says so in `historyFrom` rather than failing. ## Versioning The version is in the path. Changes within `v1` are additive only; anything breaking would be `/api/v2`. ## Next steps - [Brands](https://shruwd.io/docs/api/brands.md) — create and manage what you measure - [TypeScript SDK](https://shruwd.io/docs/api/typescript-sdk.md) — a typed client over all of this - [For AI agents](https://shruwd.io/docs/api/for-ai-agents.md) — machine-readable docs --- # Brands API Paths on this page are relative to `https://shruwd.io/api/v1`. | Method | Path | Scope | Role | |---|---|---|---| | `POST` | `/brands` | `write` | owner | | `GET` | `/brands` | `read` | viewer | | `GET` | `/brands/{brandId}` | `read` | viewer | | `PATCH` | `/brands/{brandId}` | `write` | editor | | `DELETE` | `/brands/{brandId}` | `write` | owner | `{brandId}` accepts the brand's id or its `slug`. ## Create a brand ```bash curl -X POST "https://shruwd.io/api/v1/brands" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Waitlister","domain":"waitlister.me","timezone":"Europe/Helsinki"}' ``` | Field | Required | Notes | |---|---|---| | `name` | Yes | How the brand is written in answers. Becomes its first matching alias | | `domain` | Yes | A URL or a bare host. Reduced to the registrable domain | | `timezone` | No | Sets the day boundary for measurements | | `contextTerms` | For short or common-word names | Words that mark a mention as being about you. See below | `https://www.acme.com/pricing` and `acme.com` resolve to the same brand. An IP address or a string with no public suffix is `422 invalid_domain`. A short or common-word name, like "Front" or "Arc", is refused with `422 context_terms_required` until you send `contextTerms` — for Front, say `["shared inbox", "customer support"]`. A mention then counts only when one of those words is near it; without them every "front-end" would count as a mention of you. The same rule applies to competitors ([Entities](https://shruwd.io/docs/api/entities.md)). Creating a brand also: - **Creates its self entity**, with the brand name as one alias and the domain attached. Nothing else can create one, and it cannot be removed. - **Plants the first cycle.** The response carries `firstCycle` as `{ cycleId, kind, scheduledFor }`, or `null`. - **Waits for prompts.** A cycle with no active prompts stays pending, so creating a brand before its prompts costs nothing. Measurement starts within minutes of the first prompts and covers only the prompts and competitors that exist then, so add [competitors](https://shruwd.io/docs/api/entities.md) first and every prompt in one call. On the free plan the response also carries `snapshot`: | Value | Meaning | |---|---| | `planned` | Your one measurement for this domain is scheduled | | `already_taken` | This domain has had its free measurement, in some workspace. The brand exists; no cycle is planned | | `null` | You are not on a snapshot plan | ### Errors | Code | Status | | |---|---|---| | `invalid_domain` · `invalid_name` · `invalid_timezone` | 422 | Bad input | | `context_terms_required` | 422 | The name is short or a common word. Send `contextTerms`; the reason is in the body | | `duplicate_domain` | 409 | An unarchived brand here already has that domain. `details` names it — `{ domain, brandId, slug, name }` | | `plan_limit` | 409 | Your plan's brand count is used up. `details` carries `{ limit, current, resource }` | Another workspace tracking the same domain is not a conflict. ## List brands ```bash curl "https://shruwd.io/api/v1/brands" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Each brand carries a `stats` block — open findings, the last completed and next pending cycle, the 30-day self mention rate per entitled engine, and a 56-day daily trend for the primary engine. Days below the display floor are `null` rather than a number. These come from rollups, so the list is cheap to poll compared with computing anything yourself from individual answers. ## Get one brand Returns the brand, its entitled engines, prompt and entity counts, and the same health block the dashboard shows — including the status of its Search Console connection and its log drain. ## Update a brand ```bash curl -X PATCH "https://shruwd.io/api/v1/brands/waitlister" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Waitlister HQ"}' ``` Only `name` and `timezone`. Both optional, at least one required — `400 empty_patch` otherwise. > **Note** > > **The domain is not editable.** It defines the brand's self entity, so a different domain > is a different brand. The `name` here is display only; what Shruwd actually matches on > lives in the brand's aliases — see [Entities](https://shruwd.io/docs/api/entities.md). A timezone change applies from the next cycle. The `slug` is stable across renames. ## Archive a brand ```bash curl -X DELETE "https://shruwd.io/api/v1/brands/waitlister" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Archives; nothing is deleted. Cycles stop being scheduled, history stays readable, and the domain becomes available to add again. ## Next steps - [Prompts](https://shruwd.io/docs/api/prompts.md) — the questions asked for this brand - [Entities](https://shruwd.io/docs/api/entities.md) — the brand and its competitors --- # Prompts API Paths on this page are relative to `https://shruwd.io/api/v1`. | Method | Path | Scope | Role | |---|---|---|---| | `GET` | `/brands/{brandId}/prompts` | `read` | viewer | | `POST` | `/brands/{brandId}/prompts` | `write` | editor | | `PATCH` | `/prompts/{promptGroupId}` | `write` | editor | | `DELETE` | `/prompts/{promptGroupId}` | `write` | editor | A prompt is addressed by its **`promptGroupId`** — the identity that survives edits. `promptId` is the current version and is returned for reference. ## List prompts ```bash curl "https://shruwd.io/api/v1/brands/waitlister/prompts?engine=google_aio&from=2026-08-01" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Returns active prompts. Pass `engine`, `from` and `to` to include per-prompt coverage for that window — mention rate, plus counts of answers that named nobody and prompts the engine produced no AI answer for. ## Add prompts Batch, and **atomic**. ```bash curl -X POST "https://shruwd.io/api/v1/brands/waitlister/prompts" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompts": [ { "text": "best waitlist software for a product launch", "intent": "commercial" }, { "text": "launchlist alternatives", "intent": "comparison", "tags": ["competitor"] } ] }' ``` | Field | Required | Notes | |---|---|---| | `text` | Yes | 1–500 characters, trimmed | | `intent` | Yes | One of the five below. It decides which diagnostics apply | | `tags` | No | Free-form, for grouping | ### Intents | Intent | For | |---|---| | `commercial` | "Best X for Y" — a buyer ready to choose | | `comparison` | "X vs Y", "alternatives to X" — where competitors get named | | `informational` | "How does X work" — learning, not yet buying | | `navigational` | Looking for a specific brand or page by name | | `problem` | Describes the problem, not a product | ### Errors | Code | Status | | |---|---|---| | `duplicate_prompt` | 409 | That text is already active on this brand | | `plan_limit` | 409 | The batch would exceed your prompt cap | > **Warning** > > `plan_limit` refuses the **whole batch**. Submitting 30 prompts against a 25-prompt plan > gets you an error and the numbers — `details` carries `limit`, `current` and `submitted` — > not 25 silently accepted. Prompts are counted across every brand in the workspace. ## Update a prompt ```bash curl -X PATCH "https://shruwd.io/api/v1/prompts/$groupId" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{"intent":"comparison"}' ``` Accepts `text`, `intent`, `tags` and `active`. | Changing | Effect | |---|---| | `text` or `intent` | **Starts a new version.** Past measurements stay attached to the wording that produced them, and the new version starts from zero responses | | `tags` or `active` | Updated in place. They do not change what is asked | Reactivating with `active: true` counts as a creation against your plan limit. ## Remove a prompt ```bash curl -X DELETE "https://shruwd.io/api/v1/prompts/$groupId" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Deactivates it. Nothing is deleted — it stops being asked from the next cycle on and its history is kept. The slot is freed against your plan limit. ## Next steps - [Entities](https://shruwd.io/docs/api/entities.md) — who you are measured against - [Measurement](https://shruwd.io/docs/api/measurement.md) — reading what comes back --- # Entities API An entity is your brand or a competitor, together with the aliases, domains, exclusions and context terms that decide what counts as a mention of it. Paths on this page are relative to `https://shruwd.io/api/v1`. | Method | Path | Scope | Role | |---|---|---|---| | `GET` | `/brands/{brandId}/entities` | `read` | viewer | | `POST` | `/brands/{brandId}/entities` | `write` | editor | | `PUT` | `/entities/{entityId}` | `write` | editor | | `DELETE` | `/entities/{entityId}` | `write` | editor | ## Matching is exact Every alias is matched as an exact string on word boundaries. The API does not accept patterns or regular expressions, and there is no fuzzy matching anywhere. A near-match that quietly counts the wrong brand corrupts every metric downstream and is close to impossible to detect later. The cost is that you supply the spellings. ## Add a competitor ```bash curl -X POST "https://shruwd.io/api/v1/brands/waitlister/entities" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "LaunchList", "aliases": [{ "alias": "LaunchList", "kind": "name", "caseSensitive": false }], "domains": ["getlaunchlist.com"] }' ``` | Field | Required | Notes | |---|---|---| | `name` | Yes | The canonical name | | `aliases` | No | Defaults to one alias equal to `name` | | `domains` | No | Reduced to registrable domains | | `exclusions` | No | Phrases that must not count as a mention | | `contextTerms` | No | Required for short or common-word names — see below | Competitors only. `isSelf` is not accepted; the self entity arrives with the brand. ### Short and common-word names A name of six characters or fewer, or one on the common-word list, is refused with `422 context_terms_required` until you supply `contextTerms`. The entity is then stored with `requires_context` set. > **Warning** > > This is not a nicety to route around. Without context terms, "Arc" counts every ordinary > use of the word as a competitor mention, and every share-of-voice number for the brand > becomes wrong in a way nobody notices. Your brand's own entity has the same rule: [creating a brand](https://shruwd.io/docs/api/brands.md) with such a name needs `contextTerms` too. An entity whose name needs them and has none reads back with `needsContextTerms: true`; add them with `PUT`. ### Errors | Code | Status | | |---|---|---| | `context_terms_required` | 422 | The name needs context terms. The reason is in the body | | `domain_taken` | 409 | Another entity of this brand already owns that domain | | `invalid_alias` · `invalid_domain` · `invalid_name` | 422 | Bad input | ## Replace an entity `PUT` takes the **desired** state — the full set of aliases, domains, exclusions and context terms you want. ```bash curl -X PUT "https://shruwd.io/api/v1/entities/$entityId" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"LaunchList","aliases":[{"alias":"LaunchList","kind":"name"},{"alias":"Launch List","kind":"name"}],"domains":["getlaunchlist.com"]}' ``` The server diffs against what exists: rows you dropped are closed as of now, new ones are inserted, unchanged ones are left alone. Nothing is ever edited in place, because a measurement taken yesterday must still resolve against yesterday's aliases. ## Remove an entity ```bash curl -X DELETE "https://shruwd.io/api/v1/entities/$entityId" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Closes the entity as of now. It stops counting from the next cycle; past measurements keep it, and share of voice is recomputed against the remaining set. The self entity is `409 self_entity` — archive the brand instead. ## Suggestions Shruwd lists names that appeared in your answers but are not tracked. | Method | Path | Scope | |---|---|---| | `GET` | `/brands/{brandId}/suggestions?state=` | `read` | | `POST` | `/suggestions/{suggestionId}/accept` | `write` | | `POST` | `/suggestions/{suggestionId}/dismiss` | `write` | Each carries the share of a prompt cluster's answers that named it, and a domain hint where the answers linked one. **Accept** creates the competitor through the same path as `POST …/entities` — the same exact-alias rule, the same `context_terms_required` check. Pass `name`, `aliases`, `domains`, `exclusions` or `contextTerms` to override the draft; the domain hint is used unless you give `domains`. **Dismiss** quiets it for 90 days. A suggestion that has already been decided is `409 suggestion_decided`. ## Next steps - [Measurement](https://shruwd.io/docs/api/measurement.md) — reading mention rate and share of voice - [Competitors](https://shruwd.io/docs/tracking/competitors.md) — the same thing in the dashboard --- # Measurement API Paths on this page are relative to `https://shruwd.io/api/v1`. | Method | Path | Scope | Role | |---|---|---|---| | `POST` | `/brands/{brandId}/cycles` | `write` | editor | | `GET` | `/brands/{brandId}/cycles` | `read` | viewer | | `GET` | `/brands/{brandId}/visibility` | `read` | viewer | | `GET` | `/brands/{brandId}/visibility/series` | `read` | viewer | | `GET` | `/brands/{brandId}/crawlers` | `read` | viewer | | `GET` | `/brands/{brandId}/answers` | `read` | viewer | > **Warning** > > **Measurement is asynchronous.** Running a cycle schedules work; results arrive over the > following hours. Nothing you can call returns a finished measurement, so do not poll in a > tight loop and do not report a number before there is one. ## Run a cycle ```bash curl -X POST "https://shruwd.io/api/v1/brands/waitlister/cycles" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{}' ``` Creates a manual cycle due now and returns its id. It draws on the same period allowance as any cycle. On a free plan this **is** the one snapshot, and is `409 snapshot_taken` once used. A workspace paused pending review is `409 workspace_on_hold`. `GET /brands/{brandId}/cycles?limit=` returns recent cycles with their state and counts. ## Read visibility ```bash curl "https://shruwd.io/api/v1/brands/waitlister/visibility?engine=google_aio" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` | Parameter | | |---|---| | `engine` | `google_aio` or `chatgpt` | | `from` `to` | `YYYY-MM-DD`. `to` defaults to today **in the brand's timezone** | Returns `{ engine, range, asOf, nResponses, entities, prompts }`. Each entity carries `mentionRate`, `shareOfVoice` and `avgProminence`, every one of them with its interval or an explicit `insufficient_data`. If `from` reaches past your plan's history window it is clamped, and the response says so in `historyFrom` rather than failing. ## Read the trend ```bash curl "https://shruwd.io/api/v1/brands/waitlister/visibility/series?window=30" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` | Parameter | | |---|---| | `window` | `7`, `30` or `90`. Default `30` | | `from` `to` | Default is the last ninety days | Returns `points` (one per cycle close) and, per entity, a `series` **index-aligned** with it. A point below the display floor is `insufficient_data`, never a number. `modelChanges` lists the moments the provider changed its underlying models. A difference spanning one of those may be the model rather than you — see [When a number has moved](https://shruwd.io/docs/metrics/when-a-number-has-moved.md). ## Read crawler activity ```bash curl "https://shruwd.io/api/v1/brands/waitlister/crawlers" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Per-bot activity with verified and unverified hits kept separate. Only verified hits count toward anything. `daily` carries a per-day series per bot, index-aligned with `daily.days`. > **Note** > > A day with no log ingest is `null`, not `0`. It means nobody was watching, not that > nothing happened — do not chart it as zero. ## Read individual answers ```bash curl "https://shruwd.io/api/v1/brands/waitlister/answers?limit=10" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` `limit` is 1–50, default 10. Newest first, measured runs only. Each answer carries `prompt`, `engine`, `state`, `executedAt`, whether you were `mentioned` and at what `rank`, the surrounding `context`, the `competitors` named, and the `citations` with `isSelf` marked. > **Warning** > > **This is evidence, not a metric.** Never compute a rate from these rows — the sample is > whatever the last few runs happened to be. It exists so you can read what an answer > actually said, which is useful while a brand is still below the display floor. ## Next steps - [Findings](https://shruwd.io/docs/api/findings.md) — the diagnosed causes behind these numbers - [Metric definitions](https://shruwd.io/docs/metrics/metric-definitions.md) — what each number counts --- # Findings API Paths on this page are relative to `https://shruwd.io/api/v1`. | Method | Path | Scope | Role | |---|---|---|---| | `GET` | `/brands/{brandId}/findings` | `read` | viewer | | `GET` | `/findings/{findingId}` | `read` | viewer | | `POST` | `/findings/{findingId}/transition` | `write` | editor | `{findingId}` accepts the finding's id or its `number` — 1, 2, 3… per workspace in first-seen order, never reused. ## List findings ```bash curl "https://shruwd.io/api/v1/brands/waitlister/findings?states=open,acknowledged" \ -H "Authorization: Bearer $SHRUWD_API_KEY" ``` Ordered the way the dashboard and the digest order them: severity, then confidence, then estimated impact. Every row carries `confidence`, which decides how much weight a finding deserves: | Confidence | Meaning | |---|---| | `observed` | A first-party fact, from your own logs or a live fetch of your page | | `inferred` | Not directly observed, but every link in the chain is checkable in the evidence | | `heuristic` | A correlation from comparing your pages with the ones being cited. A lead, not a fact | Never present a `heuristic` finding as a fact. If your plan caps visible findings, the rest are counted in `locked` rather than hidden, and fetching one is `403 not_entitled`. Findings reads are **not** clamped to your history window — a finding older than the window is still a finding. ## Get one finding Returns the full evidence, the recommendation's steps, and: | Field | | |---|---| | `baseline` | The snapshot captured at `fix_applied`: metric values, intervals, `n`, timestamp | | `stateChangedAt` | When it entered its current state | | `events` | Newest first, each with the payload it recorded — a recheck verdict's per-prompt gates are readable here | | `suppressedBy` · `suppressedByNumber` | Set when a more fundamental finding on the same URL is masking this one | Suppression matters: there is no point restructuring a page for citation while your server is refusing the crawler that would read it. ## Transition a finding ```bash curl -X POST "https://shruwd.io/api/v1/findings/7/transition" \ -H "Authorization: Bearer $SHRUWD_API_KEY" \ -H "Content-Type: application/json" \ -d '{"to":"fix_applied","note":"Removed the Disallow for OAI-SearchBot"}' ``` `to` is one of: | Value | Meaning | |---|---| | `acknowledged` | You have picked it up | | `fix_applied` | The change is live. **Captures the baseline** and returns when the recheck will run | | `dismissed` | Not a problem. Quiet for 90 days, then it may return | > **Warning** > > `resolved` and `not_moved` are **refused**. Those verdicts belong to the recheck, which > compares the baseline against a fresh measurement — they are not something a caller, human > or agent, can assert. See [Rechecks](https://shruwd.io/docs/findings/rechecks.md). Mark `fix_applied` only once the change is live. The baseline is captured at that moment, so marking it early compares your change against itself. ## Next steps - [How findings work](https://shruwd.io/docs/findings/how-findings-work.md) — the full lifecycle - [Rechecks](https://shruwd.io/docs/findings/rechecks.md) — how the verdict is decided --- # TypeScript SDK `@shruwd/sdk` is a typed wrapper over the [API](https://shruwd.io/docs/api/api-overview.md). ESM, and no runtime dependencies beyond `fetch`. Its types are generated from the OpenAPI document, so the client and the contract cannot drift apart. ```bash npm install @shruwd/sdk ``` ## Getting started ```ts import { Shruwd } from '@shruwd/sdk'; const shruwd = new Shruwd({ apiKey: process.env.SHRUWD_API_KEY! }); const brand = await shruwd.brands.create({ name: 'Waitlister', domain: 'waitlister.me' }); // Competitors first: the first measurement starts as soon as prompts exist. await shruwd.entities.add(brand.brandId, { name: 'LaunchList', domains: ['getlaunchlist.com'], }); await shruwd.prompts.add(brand.brandId, [ { text: 'best waitlist software for a product launch', intent: 'commercial' }, { text: 'launchlist alternatives', intent: 'comparison' }, ]); ``` Constructing without a key throws immediately rather than failing on the first request. ### Options | Option | Default | | |---|---|---| | `apiKey` | — | Required | | `baseUrl` | `https://shruwd.io/api/v1` | Point at a preview deployment | | `fetch` | global `fetch` | Supply your own | | `maxRetries` | `3` | See below | | `timeoutMs` | — | Per request | | `client` | — | Your app’s name and version, e.g. `acme-reporter/2.1.0`. Sent with the SDK’s own so your usage is attributable. Never a person, an account or a secret | ## Methods | Namespace | Methods | |---|---| | `workspace` | `get()` | | `brands` | `list()` · `create()` · `get()` · `update()` · `archive()` | | `prompts` | `list()` · `add()` · `update()` · `remove()` | | `entities` | `list()` · `add()` · `set()` · `remove()` | | `cycles` | `list()` · `run()` | | `visibility` | `get()` · `series()` | | `crawlers` | `get()` | | `findings` | `list()` · `get()` · `transition()` | | `suggestions` | `list()` · `accept()` · `dismiss()` | | `connections` | `createIngestToken()` | | `setup` | `get()` · `suggest()` | | `answers` | `list()` | List calls unwrap their envelope, so `brands.list()` returns an array rather than `{ brands: [...] }`. ## Reading a metric ```ts const visibility = await shruwd.visibility.get(brand.brandId, { engine: 'google_aio' }); // point, lo and hi are proportions from 0 to 1. const pct = (x: number) => (x * 100).toFixed(1); for (const entity of visibility.entities) { const rate = entity.mentionRate; if (rate.state === 'ok') { console.log(`${entity.canonicalName}: ${pct(rate.point)}% (${pct(rate.lo)}–${pct(rate.hi)}%, n=${rate.n})`); } else { console.log(`${entity.canonicalName}: not enough data yet`); } } ``` Discriminating on `state` is not optional politeness — there is no `point` to read when the state is `insufficient_data`, and treating it as `0` is the single most common way to misreport these numbers. ## Errors Every failure throws a `ShruwdError`. ```ts import { Shruwd, ShruwdError } from '@shruwd/sdk'; try { await shruwd.entities.add(brandId, { name: 'Arc' }); } catch (error) { if (error instanceof ShruwdError && error.code === 'context_terms_required') { await shruwd.entities.add(brandId, { name: 'Arc', contextTerms: ['browser', 'The Browser Company'] }); } else { throw error; } } ``` | Property | | |---|---| | `status` | HTTP status | | `code` | The stable code. Branch on this | | `message` | Human-readable | | `retryable` | Whether retrying could succeed | | `details` | Structured context, when there is any | | `retryAfterSeconds` | Set on a `429` | A failure with no JSON body becomes `http_`, so you always get a code to branch on. ## Retries Handled for you, with backoff: | Response | Retried | |---|---| | `429` | Yes, on every method, honouring `Retry-After` — unless it asks for more than 30 seconds, which is thrown at once for you to schedule | | `5xx` and network errors | Only on `GET`, `PUT` and `DELETE` | | `4xx` other than `429` | Never | `POST` is not retried on `5xx` because it is not idempotent — a brand or a batch of prompts might have been created before the failure. Handle those yourself if you need to. Up to `maxRetries` attempts, then the error is thrown. ## Next steps - [API overview](https://shruwd.io/docs/api/api-overview.md) — the underlying contract - [MCP server](https://shruwd.io/docs/api/mcp-server.md) — the same client as agent tools --- # MCP server `shruwd-mcp` is a local [MCP](https://modelcontextprotocol.io) server. It gives an AI assistant tools to set up a brand, add prompts and competitors, read how the brand appears in AI answers, and act on findings. It runs on your machine over stdio and talks to the API with your key. There is no hosted server. Clients that find servers from a site can read its install and configuration from [`/.well-known/mcp.json`](https://shruwd.io/.well-known/mcp.json). ## Prerequisites An API key from **Account → API keys**. A key is bound to one workspace and acts with your role in it: reads need `viewer`, writes `editor`, creating or archiving a brand `owner`, and minting an ingest token `admin`. To work across several workspaces, mint a key in each and add one server entry per workspace. See [Roles](https://shruwd.io/docs/api/api-overview.md). ## Claude Desktop In `claude_desktop_config.json`: ```json { "mcpServers": { "shruwd": { "command": "npx", "args": ["-y", "shruwd-mcp"], "env": { "SHRUWD_API_KEY": "sh_live_…" } } } } ``` ## Claude Code ```bash claude mcp add shruwd -e SHRUWD_API_KEY=sh_live_… -- npx -y shruwd-mcp ``` ## Cursor In `.cursor/mcp.json`: ```json { "mcpServers": { "shruwd": { "command": "npx", "args": ["-y", "shruwd-mcp"], "env": { "SHRUWD_API_KEY": "sh_live_…" } } } } ``` ## Environment | Variable | | |---|---| | `SHRUWD_API_KEY` | Required | | `SHRUWD_API_URL` | Optional. Defaults to `https://shruwd.io/api/v1` | ## Tools | Tool | Does | |---|---| | `shruwd_get_workspace` | Plan, entitlements, usage, brands. Start here | | `shruwd_list_brands` · `shruwd_get_brand` · `shruwd_create_brand` · `shruwd_update_brand` · `shruwd_archive_brand` | Brands. Creating one also creates its self entity and first measurement. Creating and archiving need `owner` | | `shruwd_list_prompts` · `shruwd_add_prompts` · `shruwd_update_prompt` · `shruwd_remove_prompt` | The questions asked of each engine every cycle | | `shruwd_list_entities` · `shruwd_add_competitor` · `shruwd_set_entity` · `shruwd_remove_entity` | The brand and its competitors, with the exact aliases that count | | `shruwd_run_measurement` · `shruwd_list_cycles` | Measure now; watch the cycle | | `shruwd_get_visibility` · `shruwd_get_visibility_series` · `shruwd_get_crawlers` | Mention rate and share of voice with intervals, and verified crawler activity | | `shruwd_list_answers` | The latest individual answers. Evidence, not a metric | | `shruwd_list_suggestions` · `shruwd_accept_suggestion` · `shruwd_dismiss_suggestion` | Competitors the answers named that you do not track | | `shruwd_suggest_setup` | Drafts prompts and competitors from your homepage. Nothing is saved | | `shruwd_list_findings` · `shruwd_get_finding` · `shruwd_transition_finding` | Diagnoses with a specific fix; mark one applied to start the recheck. Capped plans count the rest in `locked` | | `shruwd_create_ingest_token` | The token and endpoint for shipping server logs. Needs `admin` | ## What the tools will not do No tool computes anything the API does not return. The server is a transport, and every guarantee the API keeps, it keeps: - **A metric is never a bare number.** Each carries its 95% interval and `n`, or an explicit `insufficient_data` — which the tool descriptions tell the agent not to read as zero. - **No tool can mark a finding resolved.** `resolved` and `not_moved` belong to the recheck. An agent can acknowledge, apply or dismiss; it cannot declare success. - **Nothing is saved without you.** `shruwd_suggest_setup` drafts prompts and competitors and saves none of them — they go in through the normal tools, with the same checks as adding them by hand. - **A capped plan hides nothing silently.** Where your plan shows only its top findings, `shruwd_list_findings` returns those and counts the rest in `locked` by severity, so the agent is told they exist rather than reporting a short list as the whole set. Fetching a locked one is `403 not_entitled`. ## Next steps - [For AI agents](https://shruwd.io/docs/api/for-ai-agents.md) — the docs themselves, machine-readable - [TypeScript SDK](https://shruwd.io/docs/api/typescript-sdk.md) — the client underneath these tools --- # For AI agents Everything on this site is readable without parsing HTML. If you are pointing a model at Shruwd, start here. ## The four entry points | URL | What it is | |---|---| | [`/llms.txt`](https://shruwd.io/llms.txt) | A curated index. Every page, grouped by section, each linking to its markdown | | [`/llms-full.txt`](https://shruwd.io/llms-full.txt) | The entire documentation as one file, each page tagged with its source URL | | [`/skill.md`](https://shruwd.io/skill.md) | A procedure for setting a brand up end to end, including the one step a human must do | | Any page `+ .md` | That page as plain markdown | Appending `.md` works on every docs URL: ``` /docs/getting-started/quickstart the page /docs/getting-started/quickstart.md the same page as markdown ``` The markdown is the source these pages are built from, not a conversion of the rendered HTML. Presentational components are translated rather than dropped — a callout becomes a labelled blockquote, a card becomes a link — so nothing loses its meaning on the way out. Every page also declares its markdown twin in the HTML head, so a client that prefers markdown can find it without knowing the convention: ```html ``` ## Copying a page by hand Every page carries a **Copy page** control under its title. It copies that page's markdown to your clipboard, and its menu opens the same page in ChatGPT or Claude with the URL already loaded. ## Reading the API instead The REST API describes itself: ``` https://shruwd.io/openapi.json ``` That OpenAPI 3.1 document is the contract — the TypeScript SDK's types are generated from it. It needs no authentication to read, and every page names it in its head as `rel="service-desc"`. The [MCP server](https://shruwd.io/docs/api/mcp-server.md) is described at [`/.well-known/mcp.json`](https://shruwd.io/.well-known/mcp.json): the package, the command and the key it needs, in the [mcp-manifest](https://mcp-manifest.dev) format. ## Reading a metric correctly This is the thing agents most often get wrong, so check any output against it. Every metric comes back with its uncertainty attached, or not at all: ```json { "state": "ok", "point": 0.142, "lo": 0.081, "hi": 0.226, "n": 42 } ``` - `point`, `lo` and `hi` are proportions from 0 to 1: 0.142 is 14.2%. - `state: "insufficient_data"` means fewer than ten responses exist. **It is not zero.** Say "not enough data yet", never "0%". - Never report `point` without `lo` and `hi`. The interval is often wide, and a bare number reads as far more certain than the measurement is. - Never call a change an improvement from two point estimates. Shruwd applies its own [movement gates](https://shruwd.io/docs/metrics/when-a-number-has-moved.md) and reports the verdict — use that rather than recomputing it. There is no shape in the API that returns a number without its interval. That is deliberate, and it holds for agents exactly as it does for the dashboard. ## Next steps - [API overview](https://shruwd.io/docs/api/api-overview.md) — authentication, errors, limits - [MCP server](https://shruwd.io/docs/api/mcp-server.md) — the tool interface for Claude and Cursor