Methodology
We publish the rubric in full so anyone can check our working — and so a vendor who disputes a score
can point to the exact criterion. The Bench rates marketing-infrastructure tools across the
10 pillars and six
disciplines on a 0–4 scale — a 0 = absent gate plus four graded points — measuring
two separate things: how deep a capability runs, and how far an agent can operate it.
Rubric r1.4; scores as of 2026-08-05.
The three operator channels
- Self-serve — capability depth when you operate the product.
- Full-serve — capability depth when their experts operate it for you.
- Agentic — how autonomously an agent can operate it.
Every score is a cumulative gate. To reach level N, a tool must satisfy the criteria for every level up to N. That is the rule that stops score inflation, so each anchor states what counts and what explicitly does not.
The six disciplines
The practice of keeping a marketing stack honest — each column of the bench is one discipline, asked of every one of the 10 pillars. The four core disciplines (in violet) are the essential loop — detect, act, and stay accountable; Translate and Design are extended disciplines that go further, from findings to remediation to building the system.
- Verify
- Independently tests that the pillar is configured correctly — not just present?
- Diagnose
- Scores the pillar’s state with depth?
- Orchestrate
- Coordinates the tools and agents acting on the pillar?
- Govern
- Enforces consent, logging and accountability over the pillar’s actions?
- Translate
- Turns pillar findings into prioritized, specific remediation?
- Design
- Helps architect or build the pillar’s underlying system?
Capability depth — self-serve & full-serve
Levels 0–3 below are generic: the same four criteria apply to every discipline and every pillar, so a level 2 in Verify means the same amount of capability as a level 2 in Design.
No observable capability touching this pillar in this discipline.
A single generic signal or a checklist-level mention, with no method behind it.
Does not count: Naming a pillar in marketing copy is not capability — there must be an observable feature.
Real but partial: covers the common case, from self-reported or connected-account data (the tool reads what GA4 / the ad platforms / the user already report). Typically single-source.
Does not count: Summarising connected-account data is ALWAYS ≤ 2 — it is not independent detection. "Reads your GA4 and lists the events" is 2, never 3–4.
Thorough, multi-signal, edge-aware; specialist-grade in this pillar. Cross-references more than one independent source and handles non-trivial / edge cases.
Does not count: A single deep feature on one narrow sub-surface can be 3 (breadth is not required) — but multi-signal depth is.
Level 4 — "Top", by discipline family
Level 4 is the one rung that is not generic. Best-in-class means different things for different jobs — independent detection for a detector, production-grade execution for an actor — so a single "Top" wording would either flatter one family or be unreachable for the other. Each family gets its own, and a tool is measured against the one its discipline belongs to.
Must meet all of:
- Independent detection — observes the real artifact or behaviour itself (intercepts live traffic, fires synthetic probes, simulates the user journey, crawls the live surface), not the operator’s self-report or connected-account numbers.
- Cross-validation — checks the observed signal against an external ground truth or across sources.
- Surfaces the claimed-vs-actual gap (Verify) or yields a reproducible score from the independently-gathered signals (Diagnose). Verify additionally requires testing that the configuration BEHAVES correctly, not merely that a component is present.
Does not count: A scored dashboard built from connected-account data is ≤ 2 — forensic requires the signal be independently gathered.
Must meet all of:
- Operates at scale in production — not a template, demo, or single-shot.
- Multi-step / multi-target — coordinates across systems (Orchestrate), sequences prioritised remediation tied to impact (Translate), or builds a deployable artifact or configuration (Design).
- Reliable and repeatable — scheduled, stateful, or otherwise run repeatedly rather than ad hoc.
Does not count: A recommendation or a strategy document is not a 4 for an action discipline — the thing must actually be coordinated, sequenced, or built/deployed.
Must meet all of:
- Consent handling where personal data or tracking is involved.
- Immutable logging / audit trail of the actions and changes.
- A gate — human approval or a judge — before consequential actions.
- Accountability over time — regression or change detection across runs.
Does not count: Platform security certifications (SOC 2 / GDPR DPA) govern the vendor, not your marketing infrastructure — they do not count toward Govern on any pillar.
Agentic operability
One ladder, identical across all six disciplines. Each rung is a verifiable architectural fact, not a judgment of degree.
Output reaches the user only through human-facing surfaces — dashboards, PDF/email, a UI. No machine-consumable interface exists.
Does not count: An in-app AI chat/assistant is still a human surface; it does not make the tool agent-operable.
A structured, machine-parseable export or read endpoint exists that an external agent could ingest — a documented REST GET, JSON/CSV export, or a BI connector (Looker/Power BI/Tableau).
Does not count: A downloadable PDF or a screenshot is not "readable" — the output must be structured data.
An external agent can CALL the tool over a documented API or MCP server — to read data or trigger work (run a scan, fetch a report) — with real authentication (API key / OAuth / PAT).
Does not count: An internal pipeline a human triggers in the UI is not invokable. The tool ITSELF must be callable by an outside agent; merely reading the user’s connected data is Readable, not Invokable.
An agent can make the tool ACT on the user’s own systems (write / remediate) through an interface with explicit guardrails: scoped permissions AND (a human-approval gate OR documented rollback/undo).
Does not count: A write path WITHOUT a guardrail does not reach 3 — ungoverned action caps at agentic 2 (Invokable). An agent that writes with no approval and no rollback is a liability, not readiness.
The agent can run diagnose → act → verify, AND the platform governs those actions end to end: consent handling where personal data is touched, an immutable audit log of agent actions, a judge or human gate, and rollback.
Does not count: Platform-level security (SOC 2 / SSO / encryption) governs the VENDOR, not the agent’s writes to YOUR systems — it does not count toward agentic 4 (Closed-loop + governed).
Independence
A single 0–4 score per vendor, shown beside its name — higher means freer to leave. It measures portability and lock-in only: whether your data comes out, whether anything of theirs has to stay embedded in your properties, whether the output still means something once the tool is gone, and what the contract costs you on the way out.
It is not forensic independence. Whether a tool detects something for itself or just reads your connected accounts is a capability question and is scored on the ladders above. A vendor can be perfectly portable and still see nothing on its own.
Independence is scored per vendor rather than per cell, but it ladders the same way capability and agentic do: cumulative gates, each with what counts and what explicitly does not. Every score on the board was assigned against these criteria — the anchors and the re-scored column shipped together, because publishing a rule the numbers do not obey would be worse than publishing no rule at all.
No structured export of your own data. History lives and dies with the subscription.
A documented structured export of YOUR data exists — API, CSV, JSON — so the history survives departure.
Does not count: A PDF or a screenshot is not an export. A dashboard you can look at after cancelling is not portability.
Nothing of the vendor’s must stay embedded in your properties for the delivered value to persist. An outside-in scanner qualifies natively.
Does not count: A required snippet, tag, pixel or edge middleware caps at 1 — removing it ends the capability, it does not just stop new data.
Output is expressed in portable form — standards, config, plain findings — so replacing the tool does not mean redoing the analysis.
Does not count: A score that only means something inside the vendor’s own index is not transferable.
Leaving is a decision, not a project: month-to-month or short-term commitment, no minimum, no re-implementation to recover your own data.
Does not count: An annual or multi-year contract, a seat/site minimum, or a bespoke onboarding you would have to redo elsewhere caps at 3.
How well evidenced the number is
Independence is the one axis where an ambiguity cuts against the vendor: the lower rung is the harsher claim, so withholding a rung their own product visibly offers would not be caution. Where a vendor’s public surface indicates a rung we cannot find documented, we assign it and say so, rather than publishing the harsher number in silence. A small mark beside the score carries which of the three applies — and because the mark qualifies the evidence and not the number, the rung shown is the same in all three cases.
- Solid — every rung is met by a verbatim quote or a documented specification. No mark: this is the ordinary case, and annotating it would put a glyph on almost every row.
- Indicated — met by something seen rather than written: a control in a product screenshot, an in-product modal, a copy line with no specification behind it. Filled dot. The vendor has not documented the format, and we have not operated the product to find out.
- Inferred — no direct evidence on the vendor’s surface; extrapolated, with the basis stated in the row. Hollow ring.
Both marked tiers are invitations to correct us. A quote from the vendor’s own surface that refutes the reading drops the score; documentation of the thing we could only see upgrades it to Solid on the same rung.
What is deliberately not on this axis: the conflict of interest. Whether a vendor both diagnoses the problem and sells the remediation for it used to be one of the factors here. It came off, because counted across the whole field it capped roughly thirteen of seventeen vendors at the same rung — and the rung it left open was reachable most cleanly by the tools that simply act least. That measures abstention, not lock-in. It is now a plain disclosure shown against each vendor, including our own: Yellowhead sells remediation, holds Independence on portability grounds, and carries the flag. Draw your own conclusion from the combination — that is the point of showing it rather than scoring it.
Tier bands
Self-serve sub-band (from peak self-serve)
- 4 → Top
- 3 → Deep
- 2 → Functional
- 1 → Surface
Agentic sub-band (from peak agentic)
- 3–4 → Operable
- 2 → Invokable
- 1 → Readable
- 0 → Human-only
The bridge rule. Start at Tier 1 — Top capability and governed-agent-operable. Add half a tier for every rubric band you fall short on either axis. Half-steps only: each half is exactly one band, so there is no such thing as a quarter-tier.
| capability ↓ / agentic → | Operable (agentic 3–4) | Invokable (agentic 2) | Readable / none (≤1) |
|---|---|---|---|
| Top (self-serve 4) | Tier 1 | Tier 1.5 | Tier 2 |
| Deep (self-serve 3) | Tier 1.5 | Tier 2 | Tier 2.5 |
| Functional (self-serve 2) | Tier 2 | Tier 2.5 | Tier 3 |
| Surface (self-serve 1) | Tier 3 | Tier 3 | Tier 3 |
Top-right cells (high agentic over low self-serve) are crossovers — an agent reaching past the human channel. Empty or near-empty today. The first vendor to land there is the story.
How a score is assigned
The process that makes the rubric reproducible — two analysts applying it to the same evidence should land on the same number.
- 1 Cumulative gates. To score N, a tool must meet the criteria for every level up to N. Assign the highest level whose criteria are fully met.
- 2 Cite or cap. Every non-zero score carries at least one public source. Without a verifiable source, a score cannot exceed 2.
- 3 Directional tie-break. On the capability axes, ambiguity between two levels assigns the lower unless the higher is positively evidenced — inflation is the risk there. On Independence it runs the other way: if the vendor’s own surface indicates the higher rung (a visible affordance, an unlabelled control), we assign that rung at its floor and stamp it Indicated, because scoring down publishes the harsher claim about a named company.
- 4 Evidence-confidence tier. Every Independence score is stamped Solid (documented), Indicated (a seen affordance in the vendor’s own product imagery, format undocumented) or Inferred (no direct evidence). The tier changes how the number renders and how right of reply resolves — never the number.
- 5 Agentic caps (not a ceiling). The three operator channels are independent — an agent may exceed both human channels (a flagged crossover). The only caps are within the agentic ladder: an ungoverned write path caps agentic at 2; closed-loop (4) requires the full Govern set.
- 6 Three kinds of zero. Absent (checked, not present), Unknown (could not verify — never silently zeroed), and Out-of-scope-by-design (the tool deliberately doesn’t claim this pillar — neutral, not a failing).
- 7 Verification, not consensus. Every non-zero claim is anchored to a verbatim public quote with a retrieval date and, where the page permits, a third-party archive snapshot. Absence claims are re-tested against the raw served bytes with every control expanded. Each release then passes an adversarial review that re-resolves citations against their quotes before anything ships; disagreements resolve against the written anchors, with the rationale recorded.
- 8 As-of dating + refresh. Every competitor is stamped with the date its evidence was gathered and re-scored at least quarterly — sooner when a correction or a vendor’s shipped changes warrant it; momentum is computed against the last published snapshot.
- 9 We score ourselves the same way. Yellowhead is on the same rubric, shipped-state only (roadmap excluded), rounded down on ties.
- 10 Right of reply. Vendors may contest any score with evidence at hello@yellowhead.digital — acknowledged within 5 business days, assessed within 20. Corrections run in both directions and are logged, dated, in the public corrections log on the methodology page.
Worked examples
The rubric applied to real cells — including our own row, where the same rules score us down.
Self-serve 4 because it crawls and simulates multi-step journeys to observe whether tags actually fire and whether consent blocks cookies in accept/reject states (independent detection), cross-references against expected behaviour across 94+ providers (cross-validation), and reports the deviation. Agentic 2 because every feature is API-invokable, but there is no governed write-back — it can be called, not made to act.
Self-serve 2, not 3: its "GA4 event audit" reads the event list from the GA4 API and maps it to funnel stages — self-reported, connected-account data, single source. It never independently detects whether a tag fires, so by the anti-criterion it cannot reach Deep or Forensic. Agentic 2 because it is callable over MCP.
Agentic is held at 3, not 4: OTTO's fixes deploy — and undeploy — through published API operations, so the write path carries a documented guardrail (Actionable). Closed-loop + governed asks for the full govern set — an immutable audit log of agent actions, consent handling, a judge or human gate — and none of that is documented, so we score the highest governed rung the surface supports. (An earlier pre-publication draft of this example said the MCP path "bypasses" the dashboard gate; the vendor's own documentation contradicted that, and it was corrected before launch.)
We score ourselves down here on purpose. Our agent door is live, not roadmap: an MCP server with authenticated read tools and one write. But it exposes nothing that RUNS anything. An agent can fetch our findings and update an action item; it cannot trigger a scan or remediate a client system, which is what Orchestrate asks. Agent-driven remediation is roadmap, and the protocol scores shipped state only.
The closest we come to acting. An agent can update an action item through a scoped capability, behind a default-hold human approval gate, with every call audit-logged — which is exactly the guardrail Actionable asks for. It still scores 2, because Actionable requires the agent to act on the USER'S OWN systems, and this write lands on their record inside our platform, not on their marketing stack. Reading our own ladder generously on our own row is the one place we can least afford it.
Reading the chart
- ♛ Breadth leader — highest mean score across the 10 pillars in a discipline. A rougher, cross-pillar summary.
- ★ Agentic leader — the deepest governed agentic level reached on any pillar in a discipline. This is the rigorous metric: an ordinal maximum, no averaging.
- Cell tint — rank within the discipline column, so the leaderboard reads at a glance.
- Independence (0–4) — portability and lock-in only, not forensic independence. Full definition under Independence above.
On the scale’s limits. The levels are ordinal, so we lean on the parts that don’t require treating them as evenly spaced: scores drive ranking, the headline agentic metric is a pure maximum, and the field is re-sortable so no single composite is load-bearing. The cross-discipline "overall" average is the softest number on the page — read it as a rank, not a measurement.
Evidence boundaries
Every score is a public-surface assessment — marketing sites, docs, public API/MCP documentation, third-party reviews and press. We grade observable capability and architecture, never a test-run through a vendor’s product, so the Bench measures none of their output and never signs up where a vendor’s terms forbid competitive use. In full transparency: Yellowhead holds ordinary free-tier accounts with four vendors on this board, opened in May 2026 during early research, and a fifth account one vendor auto-created when we registered for its free public conference in June 2026 — none was ever used for any scoring input. No in-product observation is a scoring input on any row; every quote comes from an anonymous read of public pages. Confidence varies with what a vendor publishes; where a page was unreachable we say so.
Which way an ambiguity cuts depends on the axis. On the six capability axes we assign the lower rung unless the higher one is positively evidenced — scoring up would credit a vendor with something we cannot show. Independence is the exception, because there the lower rung is the harsher claim: calling a company captive is an accusation, and withholding a rung their own product visibly offers is not caution but unfairness. So where a vendor’s own public surface indicates a higher rung — a control in a product screenshot, a copy line without a specification — we assign that rung and mark the score Indicated, which says the number rests on a seen affordance rather than documented copy. A refuting quote from the same surface drops it; a vendor pointing us at documentation upgrades it.
Limitations — what this comparison does not cover
Every comparison has a shape, and its shape is an argument. Here is ours, and where it runs out.
- We did not operate anyone’s product
- Not theirs, and not ours in a way that would flatter us. Scores describe documented, observable capability — what a vendor’s architecture can be shown to do — never how good the output is when you actually use it. A tool can score well here and disappoint you in month two. No score or commentary is a quality judgement or a recommendation to buy.
- No verified client outcomes
- No vendor supplied independently audited results, and we did not validate the claims on anyone’s case-study page — including our own. Capability is not results.
- How this roster was chosen — and what that biases
- The field here was assembled the honest, unglamorous way: our own searches over time, plus the tools that surfaced in our own Diagnostic runs. It is not a market census, and it is not the output of a repeatable selection procedure. That biases the roster toward what we encountered — so absence from this page is not evidence of absence from the market, or of weakness. A capable vendor we never crossed paths with is simply missing, and we would rather say that than dress the sample up as something it isn’t. Tell us who we’ve missed.
- Pricing is not compared
- We requested no quotes. Pricing in this category is rarely standardised and often unpublished, so a price column would compare a list price against a negotiated one and call it a finding.
- The scale is ordinal
- A step from 2 to 3 is not the same size as 1 to 2, so the levels rank rather than measure. See On the scale’s limits above for how the page is built to lean only on the parts that survive that constraint.
- It is a snapshot, and this category moves fast
- Every row is stamped with the date its evidence was gathered and re-scored on the published cadence. A vendor may have shipped something the morning after we looked. Where we could not verify a capability claim we scored conservatively rather than generously, which means a fast-moving vendor is more likely to be under-scored here than over-scored on the six capability axes. Independence runs the other way by design — see Evidence boundaries above — so an unverified Independence score is more likely to be generous than harsh.
- We are on the page, and we score ourselves down
- Yellowhead is graded on the same rubric, from shipped state only — roadmap excluded — and rounded down on every tie. Where our own capability is ambiguous against an anchor, we take the lower level. That is a rule we can be held to: if you think a Yellowhead cell is generous, contest it the same way you would anyone else’s, and we will log the correction at the next refresh.
- One thing that will change, and how we will handle it
- Everything scored here — ours and everyone else’s — is assessed from the outside, with no access to anyone’s systems. We intend to offer opt-in first-party telemetry to clients on named engagements, which would mean holding data we do not hold today. When that ships, it changes nothing on this page: the Bench stays an outside-in assessment, and no client’s first-party data will ever be a scoring input to it. We are stating that before the capability exists rather than after, so the commitment is checkable.
Corrections & release log
The bench is refreshed at least quarterly — sooner when a correction or a vendor's shipped changes warrant it. Corrections run in both directions and are logged here, dated. Vendors: contest any score with evidence at hello@yellowhead.digital — acknowledged within 5 business days, assessed within 20.
- 2026-08-05 — First publication (release 1.19, rubric r1.4). The bench publishes after a full 17-row evidence pass: every non-zero claim anchored to a verbatim public quote with a retrieval date and, where the page permits, a third-party archive snapshot. Pre-publication drafting included corrections in both directions — including upward, against our own interest — under the same rules that govern this log. Post-publication corrections will be logged here, dated, newest first. A snapshot of each published release is saved to the Internet Archive.
All product names, trademarks and registered trademarks are the property of their respective owners. Vendor names are used for identification and comparison only; their use implies no affiliation with, sponsorship by, or endorsement from any vendor.
The 10 pillars
Attribution · Tech-Stack · Conversion · Trust & Security · Brand · Marketing · Advertising · Competitors · Product · Discoverability. Full definitions on The 10 Pillars.