Scrunch and SitecoreAI logos connected by flowing data streams and analytics cards, representing AI visibility measurement flowing into SitecoreAI agent grounding.
Back to home

When AI Visibility Becomes Agent Grounding: Scrunch's AI Discovery Files in SitecoreAI

Miguel Minoldo's picture
Miguel Minoldo

Since 2 October, Scrunch's AI visibility measurements sit inside the same retrieval corpus SitecoreAI agents use for grounding. Connect a brand context to Scrunch and it gains a read-only AI Discovery folder, refreshed daily, reporting how selected AI platforms answer the prompts you track about your brand. Sitecore shipped it alongside the general availability of its AI discoverability features, and while the GA entry is the bigger commercial news, this is the change I would study as an architect, because the system that measures your AI visibility can now shape the content written to improve it.

A month ago I described brand context as the descriptive half of the brand, the knowledge an agent reasons from, sitting beside the brand kit that holds the rules (From Brand Kit to Brand Context). At the time all of that knowledge was self-description, generated from your website and public information and then curated by an admin. The AI Discovery folder adds a different kind of content, a vendor-derived measurement of your brand, scoped to a prompt set you configured in another product, and overwritten every day.

Blog post image
Click to expand
The generated folders and the AI Discovery folder sit in the same corpus and reach the agent the same way, but they come from different places and are corrected in different places. An admin can fix a wrong persona profile in SitecoreAI. A wrong competitor in the AI Discovery files has to be fixed in the Scrunch configuration.

What shipped on 2 October

Two entries landed together. The first moves AI discoverability out of early access, so an organization with a Scrunch subscription now gets signals, AEO opportunities and Knowledge Studio, with what the changelog describes as improved onboarding. Signals was in early access when I wrote about it in September, so that piece now describes a GA feature.

The second is the brand context connection. The prerequisite is that your SitecoreAI CMS environment is already connected to Scrunch, and the Connect Scrunch button in the brand context file explorer only appears once that connection exists. You select a brand context, connect it to a configured brand in Scrunch, and SitecoreAI generates an AI Discovery folder with five files:

  • Overview and Platforms covers presence, position, sentiment, citation rates and share of voice, overall and per AI platform, split by search mode and by branded mode (whether the prompt names your brand).
  • Topics and Prompts reports performance by topic and tag, and lists zero-presence prompts, where responses were observed but your brand never appeared in them.
  • Competitors shows presence, share of voice, position and sentiment for the competitors you track in Scrunch.
  • Citations and Sources covers citation health, the top cited domains, and the brand URLs gaining or losing citations against the previous period.
  • Audience, Journey, and Geography cuts the same metrics by persona, funnel stage and country.

The folder is read-only and managed by SitecoreAI, so you can't edit, rename, move or delete the files, or add anything to the folder. It refreshes daily, the default reporting window is 30 days, and when comparison data exists the current period is compared with the previous one. Disconnecting deletes the folder and its contents, and the documentation is clear that this can't be undone. Scrunch's own announcement adds that access to individual features depends on your Scrunch plan, which is worth checking before you promise a client all five files.

Nothing changes in how agents consume it. In Agentic studio you select a brand context, and the agent decides which files are relevant to the request, nobody picks folders. Sitecore's example prompt for the new folder is "Create content for a topic where our competitors currently have greater AI visibility", which tells you the intended use.

Measurements are not facts

The generated folders (Audiences, Messaging, Product, Content Guidelines, Customer Evidence) hold claims about the brand. A persona profile or a value proposition is something you can read, disagree with and correct, and once an admin has edited it, it stays true until the brand changes. The AI Discovery files hold measurements, a share-of-voice figure for a topic is the result of sampling AI responses to a set of prompts over 30 days, and it will be a different number tomorrow. AI Discovery is evidence about model behavior, not ground truth about how the market sees you.

Both now sit in the same retrieval corpus, separated only by a folder name, and the agent decides what to read. The retrieval layer knows these are files, but nothing in it marks one file as durable brand knowledge and another as a noisy, time-bounded measurement, so that typing has to come from the file contents and from your instructions. The failure mode I would watch for is a measured figure used as if it were a brand fact, where "competitors dominate topic X" becomes the starting premise of a brief, when it is a reading taken over 30 days on a fixed set of platforms and prompts, and the sample behind it may be a handful of prompts.

Sitecore did the careful part in the file design. The files distinguish 0 (measured, and the value was zero), -- (no value available) and not observed (configured in Scrunch but absent from the period's data), and each file carries its reporting period and the data supporting it, prompts with volume and response counts. That is the right shape for a reader who pays attention, and whether the agent reads that metadata with the same care is the first thing I would test.

Two habits follow from this. When a task leans on AI Discovery data, ask the agent in the prompt to state the reporting period and the response counts behind any figure it uses, so a reviewer can tell a trend from noise. And if the reasoning behind a brief matters later (regulated content, or a campaign you will want to evaluate in three months), save the figures the agent used at the time, because the folder is overwritten daily with no history kept in SitecoreAI, and a disconnect removes even the current snapshot.

Brand context now has an upstream owner

In the brand context piece I argued that the edit step is where a generated corpus earns trust, and that it needs a named owner. For the AI Discovery folder that edit step doesn't exist in SitecoreAI. The files are read-only, so whatever is wrong in them has to be fixed at the source, which is the brand configuration in Scrunch, meaning the prompts, how they are organized by topics, tags, personas, funnel stages and countries, and the competitors you track.

That changes who owns brand context. Until now it was a SitecoreAI admin task, and with the connection on, part of the agent's grounding is defined by whoever maintains the Scrunch configuration, often a different person and sometimes a different team (SEO or digital marketing rather than the platform owner). A prompt set built to monitor visibility was never designed to brief content agents, and I doubt many teams reviewed theirs with that use in mind.

The overlap between the two sides is where I expect friction. The Messaging folder has a generated Competitive Positioning file and the AI Discovery folder has a Competitors file, and the Audiences folder has Persona Profiles while the Audience, Journey, and Geography file reports performance by the personas configured in Scrunch. Say the positioning file names competitor A as your main rival while Scrunch tracks B and C, both are retrieved as valid context, and the agent reconciles them on its own terms (the same goes for persona names that don't match). The documentation doesn't define a conflict policy for brand context files, and I would rather not create the conflict than find out how it gets resolved, which means aligning persona names and competitor lists on both sides before connecting, and making it one person's job to keep them aligned.

Two ways back into content

Signals acts after a visibility shift, on an existing page, through Page builder. The AI Discovery folder acts earlier, at brief and generation time, so it can influence content before the page exists.

The two also see different slices of the data. In the Signals post I described the Strategy page as an action queue, filtered to shifts that come with a page to edit, while the AI Discovery files carry the wider view, including zero-presence prompts and competitors that were never observed. The broader picture now goes to the agent and the actionable subset to the person, which is worth knowing even if it was never a deliberate design choice.

Blog post image
Click to expand
Signals routes a shift to an existing page, and the AI Discovery files route the same data into new content before the page exists. Both end in published content that the AI platforms read and cite, and the same Scrunch prompt set then measures the result.

The prompt set now steers and scores

Put the two halves together and the loop is tight. The Scrunch prompt set defines what gets measured, the measurements ground the agent, the agent writes content aimed at the gaps those prompts revealed, and the same prompt set then measures whether the content worked. Each step is reasonable on its own, and together they make the prompt set both the target and the scoreboard.

This is a Goodhart's law problem. Content gets optimized toward the queries you chose to track, the metrics on those queries improve, and nothing inside the loop tells you whether visibility on untracked queries moved, or whether the prompt set still reflects how your buyers ask. I raised the vendor-independence version of this concern in the Signals piece, and this release makes it more concrete, because measurement now feeds generation directly instead of passing through a person reading the Strategy page.

The mitigation is mostly process, starting with reviewing the prompt set on a schedule, as a content strategy artifact rather than a monitoring setting. The more important step is a holdout. If the same prompts decide what content gets created and whether it succeeded, you are evaluating the system against the questions it was optimized to answer, so keep a slice of prompts that agents are never pointed at and read it as a control group, alongside at least one measure of AI visibility from outside this stack. If the tracked prompts improve while the holdout stays flat, the gain is probably specific to the prompts you optimized for, and the loop on its own has no way to show you that.

Where I would not overstate this, the risk scales with how much weight agents give the AI Discovery files, and I have no visibility into that. Agentic studio decides relevance per request, so for a task like a tone rewrite the folder may never be retrieved. I also haven't confirmed whether the brand context tools in the Agent API and the Marketer MCP return the AI Discovery folder alongside the generated files, the documentation only describes its use in Agentic studio, so treat programmatic access as unverified until you see it in your own tenant.

Telemetry is now part of the grounding

The 2 October release makes AI discoverability a GA part of SitecoreAI, and the brand context connection is the piece that changes the architecture, because it turns vendor telemetry into agent grounding, and from then on the system measuring your AI visibility is part of the system deciding what content to create. The files are well specified, with explicit null semantics, reporting periods and sample sizes, which is more care than most retrieval corpora get. What the release doesn't settle is ownership, since the grounding now depends on a Scrunch configuration SitecoreAI can't edit, and on agents treating 30-day measurements as measurements.

Two questions for your own tenant. Who owns the Scrunch prompt set now that it shapes what your agents write, and do its personas and competitors match the ones in your brand context. And when an agent cites a visibility gap in a brief, can a reviewer see how many responses that gap rests on.

What's next

These five files provide most of the raw dimensions a governed share-of-model score needs, presence, share of voice and citation rates, cut by topic, persona and platform, with the period-over-period comparison already computed. The next post turns those dimensions into a score you can govern against, with a holdout set kept outside the optimization loop as part of the design.

References