How Envoyra measures AI recommendation presence

Envoyra runs as a three-stage pipeline. Each stage has a single responsibility and emits structured output the next stage consumes. We separate scanning from measurement, and measurement from insight, so that every claim in the weekly report traces to a specific row of data — and we can show our work when asked.


How we collect data

Every scan is a structured query to a specific AI engine. We do not paraphrase prompts. We do not interpret what an engine "meant." Each query is a literal string sent to a documented API endpoint, with the exact model and parameter set recorded alongside the response.

The current engines in our production scanner: Perplexity (sonar-pro), Claude (claude-sonnet-4-6), ChatGPT Search (gpt-4o), and Gemini (gemini-2.5-flash). Each engine is hit through its official, paid API. We do not scrape, we do not impersonate browsers, and we do not buy data from third parties. Every response in our database is one we paid the provider for.

Prompts are composed from a category-specific taxonomy that spans buyer stages: awareness, consideration, decision, and purchase. The set is curated, not generated by AI — we want the prompts a real buyer would type, not the prompts an AI thinks a buyer would type.

For every scan we record: the engine, the model version, the exact prompt, the full response, the response timestamp, and the latency. Nothing is dropped. Nothing is summarized at this stage.

What we measure

We measure outcomes, not algorithms. We do not claim to know how Perplexity, Claude, ChatGPT Search, or Gemini compose their answers. We measure what the answer was.

The primary metric is recommendation presence: in how many of N representative buying-stage prompts does the brand appear, by engine, by buying stage. We report this as a fraction (e.g., "3 of 25") not a percentage. The fraction communicates the sample size; the percentage hides it.

We also measure competitive displacement (which brands appear when the target brand does not), source influence (which publications and citations the engine references in its answer), engine divergence (where different engines name different vendors for the same buying question), and shortlist inclusion across the buying journey (a brand can be visible at awareness and absent at decision, or vice versa).

We do not measure product quality, customer satisfaction, deal-stage conversion, or anything an AI engine is not a credible source of opinion on. The signal is narrow and deliberate.

How we report

Every report ships with sample size disclosed, methodology linked, and confidence framing matched to the data. A 25-prompt single-run observation is directional, not conclusive — we say that in the report. A 5-week observation across three engines is more confident; we say that too. We do not present any signal stronger than the underlying data warrants.

When we cite a finding, we cite the row. When we make a recommendation, we cite the finding. The report is structured so a reader can trace any claim back to a specific scan response.

We use fractions throughout, not percentages. The exception is when we quote third-party research that originally used percentages — we preserve their notation and cite the source.

Scan health

Every scheduled scan run records its own status. We track expected prompt count vs received response count, errors per engine, response latency, and timestamp completion. A run that does not receive 100% of its expected responses is reported as PARTIAL or FAIL. Customers see this status alongside their data.

We do not silently drop failed responses. If a Gemini call returns HTTP 503, we record the error, retry per our backoff policy, and report the outcome. If after retries the call cannot complete, that prompt slot is marked as missing data — not as a brand absence. The distinction matters for the math.

Engine divergence

Different AI engines often return different brand sets for the same buying question. That divergence is a measurement, not a bug. We report it directly.

Engine divergence is a signal about which sources each engine weights. When Perplexity names a brand and Claude does not, the explanation is in the citation environment, not in the brand's quality. Our reports surface this divergence by engine, by prompt, by category — so customers can see the engine-specific patterns that matter for their buyer.

What we cannot know

We cannot know the exact training corpus of any engine. We cannot know the retrieval policy of any engine on any given query. We cannot reverse-engineer the weighting between citation count, recency, semantic similarity, or any other signal an engine may use to compose an answer.

When a customer asks "why did the engine name competitor X and not me," our honest answer is: we observed it, and here are the patterns that correlate with appearance — but the causal mechanism is not measured by this benchmark. We name that limit every time.

This is the discipline that separates measurement from speculation. We err toward the smaller claim.

Why we do not offer optimization

Envoyra is a measurement company. We sell measurement. We do not sell ranking improvement, citation engineering, AI SEO, or anything that promises score movement.

Two reasons. First: the underlying mechanisms are not stable enough for any honest vendor to make those promises with confidence. AI engine retrieval changes weekly. A practice that increases citation in one model can be neutral or negative in another. Second: incentive alignment. A vendor that sells measurement and optimization to the same customer has a conflict — incentives to find improvement where there is none. We choose to sell only the measurement.

Customers who want to act on our data are free to engage the agencies, consultants, or in-house teams who do that work. We will not refer business in either direction. The report is the product.