How the Diagnostic works

The methodology, in the open

Ten pillars. One perception gap. One composite score. Five tiers. We publish the conceptual model so you can decide whether the framework is sound before you trust the output. We reserve the algorithmic specifics — pillar importance weights, signal extraction, model prompts.

What we publish, what we reserve

A diagnostic earns trust through methodological transparency. It earns durability through algorithmic protection. We do both deliberately.

Published

  • The ten pillars — named, defined, scoped
  • The four-layer marketing stack model we sit at the top of
  • The perception-gap mechanic (math + philosophy)
  • The MIDAS composite (four sub-components + weights)
  • The five MIDAS tiers and what each one means
  • The signal sources the analysis draws from
  • The sub-score weights inside a pillar, and the agentic dimensions each one covers — published in full

Reserved

  • Per-pillar importance weights in the composite, and the curve applied to it
  • Business-type-conditional weighting is on the roadmap under evaluation; today's weights are deliberately uniform, so every score is comparable across the full scan corpus
  • The scoring bands behind each dimension
  • Signal extraction logic and routing rules
  • Model prompt structures and consensus formulas
  • Detected vs declared tie-breaking when signals disagree

The ten pillars

Each pillar is a surface of your marketing infrastructure that produces signal independently. We score them separately so a weak pillar can't be hidden behind a strong one.

01

Attribution

Do you measure what marketing dollars actually produce?

Analytics and tag-manager detection from page source and the third-party technology record, GA4-versus-legacy measurement evidence, server-side tagging signals, no-JS measurement fallback, and redundancy across the measurement tools we find. Presence, not firing — container contents are not externally observable, and a tag manager is never read as an absence of analytics.

02

Tech Stack

Are the tools you pay for doing the work you bought them for?

Tool inventory by category from three detection layers — HTML fingerprints, structured data, and a third-party technology record — with every row marked by how we know it: the vendor’s own fingerprint in your live page, or a historical record relayed and never corroborated. Plus functional redundancy inside a category, and an agentic-readiness sub-score (Conventional 0.7 / Agentic 0.3) covering llms.txt, MCP manifests, declarative agent attributes and agent-surface probes.

03

Conversion

Can a visitor actually become a customer without friction killing the path?

Form structure (label coverage, autocomplete tokens, placeholder-only anti-patterns, and whether a form posts to a real endpoint or fires into a mailto: link), CTA inventory classified self-serve versus sales-assisted, anti-bot stack, and payment infrastructure detected on commerce subdomains — plus an agent-compatibility read of whether an autonomous agent could complete the same path, weighted 80/20 against the human funnel. We never submit a form or complete a checkout.

04

Trust & Security

Does your stated posture match your detected behaviour?

Email authentication measured live from DNS — DMARC and SPF, plus a common-selector DKIM probe whose blank result is inconclusive rather than a finding — HTTPS reachability, privacy/cookie/terms page probes and the policy text itself, and consent-mechanism detection through three routes: vendor fingerprints, banner markup, and notice copy near an accept control, including whether a deployed banner ever renders for us. Also read: whether your privacy policy discloses AI usage and names AI vendors, crawler-governance clauses, and AI-content provenance markers. We describe posture; we do not certify compliance.

05

Brand

Is your voice the same person across every surface a prospect sees?

Three co-equal reads: voice consistency between your site and your public social surfaces, sentiment from public signals (review rating and volume, news-headline tone, public video and social metrics, trust markers), and content quality against an editorial rubric. No stylometric AI-writing detection — false-positive rates are too high to underwrite the accusation.

06

Marketing

Are the channels you run actually a system, or just running in parallel?

Detected email-platform, CRM and social-management tooling, content publishing cadence from RSS/Atom feeds and sitemap freshness, and social profile presence cross-validated against your site. Nothing inside a CRM or email platform is visible from outside, so campaign performance, list health and sequence logic are out of scope — and anything involving paid placement is scored under Advertising rather than counted twice.

07

Advertising

Do your ad-platform claims match what your site can actually receive?

Ad-platform pixel detection, a server-side tracking maturity ladder built from page signatures plus DNS and HEAD probes of candidate first-party tagging subdomains, public ad-library activity where a platform’s library discloses it, and the spend you state. No ad-account access — the score reads tracking infrastructure, so a pure server-to-server integration reads as undetected.

08

Competitors

Do you know who a buyer — or an AI assistant — would put beside you?

Competitor discovery through several lenses — jobs-to-be-done, substitutes, detected category, and which companies an AI assistant would name for the same job — with every competitor URL liveness-checked and non-resolving domains dropped, and a relationship type and confidence on each entry. Positioning is described from observable signals and labelled as estimated; where no genuine comparable exists the list comes back empty with the reason, because a confidently wrong peer costs more than a short list. Measured citation share across AI engines is a separate product — the $49 Citation Scan — not part of the Diagnostic.

09

Product

Can a first-time visitor tell what you do in under six seconds?

Offer legibility: offer-schema completeness (name, price, currency, availability, description), machine-readable pricing on the surfaces we fetch, whether the homepage exposes a single H1 a machine can lift as your value proposition, specification content in real semantic HTML rather than images, and pricing-page and CTA detection. We score whether the offer is legible to a human and a machine, not whether it is a good offer.

10

Discoverability

Web, SEO, AEO & GEO — are you discoverable by both Google and the systems replacing it?

Four sub-axes scored separately: Web (Lighthouse lab metrics on a mobile profile — LCP, CLS, first contentful paint and total blocking time — crawl directives, sitemap and canonical-tag hygiene, and DOM legibility: landmarks, alt coverage, link-text descriptiveness), SEO (scored from on-page evidence and content depth; no rank tracker and no backlink feed is wired), AEO (structured data, AI-crawler directives, llms.txt and answer-first structure, capped when retrieval bots are blocked), and GEO (off-site consensus driving whether AI systems name and recommend you). We also request your pages with each published AI-crawler user-agent string and diff the responses — which catches a robots.txt that welcomes a bot while the server refuses it. That measures how your server answers those strings from our infrastructure, not what any vendor’s crawler receives. Lab data only, no field metrics from your real users.

Where the Diagnostic sits in your stack

Most marketing stacks have three layers — execution, analytics, automation. Almost none have the fourth: infrastructure management. The Diagnostic operates at that fourth layer, testing whether the assumptions the first three layers depend on are actually true.

The four-layer model is published in full on a separate resource page, with the layer-by-layer breakdown and what each layer assumes about the layer beneath it.

Read the four-layer model

The perception-gap mechanic

The Diagnostic is two halves running in parallel. The gap between them is the most revealing thing it produces.

Side A

Your self-assessment

You rate your own marketing infrastructure across the ten pillars. One slider per pillar, takes about a minute. We assume you know your business — this is the operator's view.

Side B

Independently detected score

In parallel with your self-assessment, our system scrapes public signals, queries multiple LLMs, validates URLs, and produces a score per pillar from no opinion at all. This is the outside-observer view.

The math

Gap = Side A − Side B

Computed per pillar. A positive gap means you rated yourself higher than detected reality (the more common direction); a negative gap means detected reality is ahead of your perception (you're under-selling something that's working).

The philosophy

Most marketing dashboards measure outcomes (what happened) or activity (what we did). Almost none measure the gap between what you think is happening and what is actually happening. That gap is where surprise lives, where unbudgeted risk hides, and where the highest-leverage corrections sit.

The Yellowhead Diagnostic names that gap quantitatively, per pillar, and produces a roadmap from it. A pillar with a small gap and a low score is a known weakness. A pillar with a large gap is a blind spot.

The MIDAS composite

Ten pillar scores, plus a perception gap, plus a coherence read, plus a budget signal is a lot of numbers to hold at once. MIDAS rolls them into a single 0–100 score that lets you compare reports over time or against an industry baseline.

Pillar Performance

60%

Measures

Your detected pillar scores, weighted by business impact.

Why it's in the composite

Detects which specific surfaces of your marketing infrastructure are healthy and which are degraded.

Perception Accuracy

15%

Measures

How closely your self-assessment matches the independently detected reality.

Why it's in the composite

Detects how well you know your own infrastructure. Large gaps in either direction (over- or under-rating yourself) are themselves a finding.

Infrastructure Coherence

15%

Measures

Whether related systems reinforce each other, or operate as parallel silos.

Why it's in the composite

A well-funded stack of disconnected tools scores lower than a thoughtful stack of connected ones. Coherence is what compounds.

Budget Efficiency

10%

Measures

The ratio between detected investment signals and detected outcome signals.

Why it's in the composite

Detects when spend has outpaced strategy — when you're buying capability faster than you're using it.

The composite passes through a logarithmic curve so foundational gaps weigh more heavily than incremental gains.

The five MIDAS tiers

The composite score maps to one of five MIDAS tiers. Tier names communicate posture, not just rank. They tell you what state your infrastructure is in, not just how it compares to others.

Exceptional

Best-in-class infrastructure across pillars, with strong cross-system integration and accurate self-awareness. Rare — the bar this tier sets is real, and most stacks never reach it.

Strong

Well-integrated infrastructure with minor gaps. Growth-ready. The systems are talking to each other and the team understands their actual position.

Operational

Working infrastructure with optimisation headroom. Pillar scores are mixed and some blind spots exist, but the operator is broadly aware of them.

Exposed

Functional but blind spots create material risk. Perception gaps are large in at least one pillar. Decisions are being made on stale or incomplete information.

Critical

Foundational gaps. Infrastructure is actively leaking value — compliance, security, attribution, or brand surfaces are detectably broken. Either unknown to the operator or known and unaddressed.

Where the scans land: of the 58 sites scored under this methodology, 89.7% are Operational, 5.2% Strong and 5.2% Exposed. Nothing has reached Exceptional and nothing has fallen to Critical; the scores run 46 to 73. Those 58 sites are subjects we chose — weighted toward large software companies — so read this as an instrument record of what our scanner found as of Aug 2026, not as a rate for businesses at large.

Look under the hood

The mechanism, at the level a technical buyer needs it: what the scan reads, how a score is produced, and the inferences we refuse to make.

What a Diagnostic actually reads

We score ten pillars from what your infrastructure exposes to the public internet, plus what you tell us about your own spend and your own confidence. Nothing is read from inside your accounts. No login, no analytics access, no ad-account access, no tag-manager access — the Diagnostic needs a URL and your self-assessment.

The evidence divides into five classes, and they are not weighted equally:

  1. Direct technical evidence — what your pages, headers, DNS records and well-known files actually contain. Machine-verified, and it outranks everything below it.
  2. Direct site content — the text on the pages we fetch. First-party evidence of what you say you sell.
  3. Verified external signals — public business listings, review aggregates, public video and social metrics, public ad-library activity.
  4. Research synthesis — web-search-based market context. Useful for landscape, weak for facts about your site. Where it disagrees with class 1, class 1 wins, and the system is instructed to that effect.
  5. Inference — reasoning across gaps, flagged as inference wherever it appears.

That ordering is the methodology. A finding that reads as authoritative is one that came from class 1; a finding that came from class 4 or 5 says so in the report.

How the scoring works

Each pillar carries a detected score from 1 to 10, set against your own self-assessment of the same pillar — the perception gap above. Underneath most pillars sit named sub-scores, each scored on its own and combined by fixed weight. Discoverability separates whether a machine can load your site, whether you rank in organic search, whether your content can be lifted into an AI-generated answer, and whether AI systems name and recommend you off-site: four different failure modes with four different fixes, and a single "SEO score" hides which one you have. Tech Stack separates conventional tool efficiency from whether your systems can be operated by an autonomous agent. Conversion separates the human funnel from whether an agent could complete it.

Some scores are computed, not judged. Where a signal is hard — a DNS record, a measured ladder of tracking maturity, an arithmetic roll-up — the number is produced by code and the language model does not get a vote. Where the question is qualitative — offer clarity, editorial quality, competitive positioning — a model scores it against the evidence and the report says what evidence it used. We separate those two cases deliberately, because a model asked to do arithmetic will drift and a spreadsheet asked to judge a value proposition will not.

The composite score on the front of the report combines pillar performance, how accurate your self-assessment turned out to be, how well adjacent systems reinforce each other, and how your infrastructure quality relates to your stated spend. Those four component weights are published above. The per-pillar importance weights and the curve are not.

What we refuse to infer

This is the part worth reading twice, because it is where a diagnostic earns or loses its standing. Each line below is a constraint the system runs under — some enforced in code, the rest carried in the scoring instructions the model works from.

  • We do not treat "not detected" as "absent." A tool loaded through a tag manager, gated behind a login, or fired conditionally is invisible to an external scan. Our findings say "not detected in page source" and mean exactly that. Every report is required to name the analysis's primary limitation in its own executive summary.
  • We do not claim a tag fires. We detect that tracking infrastructure exists. Whether a given pixel fires, in what order, and under what consent state is a different measurement, and we do not have it from outside.
  • We do not claim to have tested your consent. We report what consent machinery is deployed and whether it renders — a banner that ships but never displays is reported as deployed-not-displayed, and a banner that never appeared for our scanner is reported as that, not as your visitors seeing none. Where load order is visible in the page itself — a chat widget loading ahead of the consent mechanism, say — we score it. Whether every tag then holds until consent is given needs network timing we do not have from outside, and we do not claim it.
  • We do not accuse your content of being AI-generated. Text-style AI detection is unreliable enough that a false positive would be an accusation we could not defend. We score AI-content governance only from concrete signals: detected tooling, and whether it is disclosed.
  • We do not claim to know a competitor's budget or strategy. Competitive figures are estimates from observable signals and are labelled as such. When we cannot name a genuine comparable, we return an empty list and say why — a padded competitor set is worse than a short one.
  • We do not evade your defences. Requests to your site identify themselves as YellowheadBot. The one check that has to send another crawler's user-agent string — the AI-crawler differential under Discoverability — appends a Yellowhead disclosure to it, so your logs name us even while your rules see the token they key on. Our render path is built to an identified-bot posture — no stealth automation, no fingerprint spoofing, no captcha solving — and when a site blocks us, the block is recorded as a block: a challenge page is never stored or scored as though it were your real page. A blocked scan produces a smaller report, not a worse score.
  • We do not read anything behind authentication, and we do not act on your site. Every request we make to your domain is an unauthenticated GET. We do not log in, submit a form, or complete a checkout, and we do not crawl your whole site — we request a fixed list of public paths plus the files that public standards put at known locations.

What we cannot see, and say so

Reports carry an Auditor's Notes block: the loose ends we could not confirm, each with the reason it stayed open. An early client told us that section was the most credible block in the document, and the reason is structural — a diagnostic that never says "we could not check this" is not reporting its own coverage.

The standing limits: no view inside tag containers, no analytics or ad-account data, no field performance data from your real users, no visibility into tools your site never mentions, and no observation of runtime behaviour that only a real user session would trigger.

What changes when you pay

Two different things, and they are worth separating.

First, depth inside the Diagnostic itself. The scan runs at full depth on every plan, including the free one — a paid plan does not start a better scan, it opens more of the one you already ran, with no re-scan. Free shows your ten pillar scores, the perception gap, the MIDAS composite, the executive summary, verdict cards and your first blind spot in full. Starter opens every blind spot, the evidence behind the brand-visibility and sentiment scores, the per-pillar budget breakdown, market intelligence with competitor profiles, and PDF export. Pro adds the investment roadmap, comparison of reports over time, action items, a machine-readable markdown export for agents and technical stakeholders, and an org account with team access. Enterprise adds scheduled re-runs, full action-item management, the journey canvas, and access to our MCP agent door on request.

Then, the part automation cannot do. Paid work adds three things: access (with your permission, we go inside the systems the scan could only see from outside), judgement (a human reading the findings against your commercial situation), and depth on one axis (a competitor set analysed properly, or your visibility inside AI answers measured across engines and questions rather than inferred).

Run the methodology on your own infrastructure

The Diagnostic is free and takes five to ten minutes end to end. The framework above is what it runs.

Start Your Free Diagnostic

No credit card. No commitment.