We score the scorers.

The Readiness Bench

A readiness audit of the tools that sell you readiness.

A forensic, evidence-anchored look at the tools that diagnose, manage and operate marketing infrastructure — how deeply each one actually does the job, across the 10 pillars of the stack and the six disciplines of running it. Marketing infrastructure has needed honest scoring for years; the rise of AI agents just raises the stakes. Yellowhead is scored on the same rubric, no thumb on the scale.

No score or commentary is a quality judgement or a recommendation to buy — the bench grades documented, observable capability, never how good a product's output is in use. Every number is anchored: the full rubric is readable beside the matrix in the Rubric anchors panel, and the complete method — evidence rules, limits, right of reply — is on the methodology page.

How to read this bench

The bench is a matrix. Every tool is a row; each of the six disciplines is a column. Each cell is a radar plotting all 10 pillars, and every pillar is scored three ways — by operator. Brighter cells sit in a higher tier; marks the breadth leader and the agentic leader in each discipline. Tap any cell for its full per-pillar breakdown. Three things to know before you read one:

The six disciplines — the columns

The practice of keeping a marketing stack honest — each column of the bench is one discipline, asked of every one of the 10 pillars. The four core disciplines (in violet) are the essential loop — detect, act, and stay accountable; Translate and Design are extended disciplines that go further, from findings to remediation to building the system.

Verify
Independently tests that the pillar is configured correctly — not just present?
Diagnose
Scores the pillar’s state with depth?
Orchestrate
Coordinates the tools and agents acting on the pillar?
Govern
Enforces consent, logging and accountability over the pillar’s actions?
Translate
Turns pillar findings into prioritized, specific remediation?
Design
Helps architect or build the pillar’s underlying system?

The 10 pillars — the radar axes

The ten surfaces every marketing stack runs on. Each cell’s radar plots all ten as its axes (clockwise from the top): Attribution, Tech-Stack, Conversion, Trust & Security, Brand, Marketing, Advertising, Competitors, Product, Discoverability. Full definitions

The three rings — how you get the value

On top of how good a tool is, we score how you get that value — by operator channel. Every pillar carries three rings: self-serve (you run it), full-serve (their experts run it) and agentic (an agent runs it). Buy the mode you actually want.

Radar axes — the 10 pillars · tap to filter (pick one or more)
:self-servefull-serveagenticcell brightness = (self-serve × agentic) — lighter is higheryellow icons = the 10 pillar axes (key above)breadth leaderagentic leaderIndependence = portability / freedom from lock-in (0–4), laddered like the two above — tap any chip for its rung; sells the fix = does the vendor also sell the remediation it diagnoses — a disclosure, deliberately not scored

Tap any cell to expand its full per-pillar breakdown.

Verify Yellowhead
Diagnose Yellowhead
Orchestrate Viktor
Govern Yellowhead
Translate Yellowhead
Design Yellowhead
Yellowhead
Toffu
Marqea
Profound
Search Atlas
Stack Moxie
Boostify
Martechbase
Boomi
Viktor
ObservePoint
Trackingplan
BrightEdge
Obsero
Wordflow
Otterly.AI
Searchable

The field, ranked

by

Each row shows the competitor’s tier profile across the six disciplines — tap any chip for the breakdown. Within a shared rank, order is not claimed.

Reading the chips: T1 T3 are — Tier 1 is top capability and governed-agent-operable, and every half-step is one rubric band short of it. They are not the 0–4 pillar scores (those live inside each cell, where higher is better). Rank is set by the Overall sort above.

  1. 1
    YellowheadForensic infrastructure diagnosticsells the fix: yes

    Broad forensic depth from public signals alone — no credentials asked for, nothing embedded, reports exporting per report to PDF and Markdown, and a work product you own and may redistribute. The agent door is live and documented — twelve capability-gated tools over MCP, every write held for a human — and access to it is vendor-provisioned rather than self-serve. What it does not do is act on your stack: the one write updates an action item inside our platform, which is the same reason three competitors on this board cap where they do.

  2. =2
    ToffuAI marketing agentsells the fix: yes

    Broadest execution agent — functional, not forensic, and the most agent-native door in the field: an MCP server, a REST API, and a signup endpoint that onboards an agent with no human at all. Takes the Orchestrate lead, and the asymmetry is the story — campaign-change tools ship off by default with a propose → confirm → undo cycle, while the content-publishing path documents no equivalent gate.

  3. =2
    ViktorAI employee / ad-ops agentsells the fix: yes

    AI "employee" in Slack, and since July 2026 an agent door too — broad ad-ops orchestration an outside agent can drive over a documented MCP server or REST API, on scoped keys, with sensitive actions gated behind human approval at the tool layer. Nothing embeds in your properties; the working record is Slack-native, and the operating configuration you would rebuild elsewhere has no documented export.

  4. =2
    Search AtlasAgentic SEO executionsells the fix: yes

    Runs SEO, ads, local, social, content and page-building from one platform and acts fast — on the SEO path behind a deploy-or-undeploy step the vendor documents on the write path itself, while the ads path documents approval as an optional mode only — and it documents fixes staying live after cancellation, including writes made straight into the customer’s own CMS. The agent door is real and large, a hosted MCP server over 21 product scopes alongside a public API reference; what it does not document is a scope model it enforces on that door.

  5. =5
    MarqeaAI CMO / operating systemsells the fix: yes

    "AI CMO" that plans and ships — six continuous outside-in audits feeding a vendor-scored Visibility Index, with the audits themselves named as triggerable over an API and MCP the vendor documents on one marketing page, publishes no reference for, and does not name in its own llms.txt. Work is drafted and queued for approval rather than applied, including on the one documented write path into a customer channel; Studio adds a done-for-you tier over the same disciplines. Work product is portable where it counts: weekly written reports sent to the customer's own recipients, drafted content shipped into their channels, and Studio-tier IP assigned to the customer — while the Index itself only means something inside Marqea.

  6. =5
    ObseroAEO/GEO + consultingsells the fix: yes

    VPSC AI-search monitoring plus an in-house GEO consulting arm — deep on brand-in-AI, and not pre-agentic: Obsero documents a read-only MCP at mcp.obsero.ai and asserts "MCP and API access, full data export" on every plan, so an external agent can pull the detection data but is not documented to operate anything. A second detection mode logs AI crawler traffic to your own site through a connector at your edge — Cloudflare today, other stacks on request. Engine coverage is gated by tier rather than published as a count — a £399 Pro seat chooses three models, and only Enterprise gets all of them. Portable on its own evidence: the raw AI-response log and the cited-source list travel, while the VPSC and out-of-100 scores only mean something inside Obsero.

  7. =5
    ProfoundAEO pure-playsells the fix: yes

    Forensic on the Discoverability sliver, on a genuinely documented agent stack — REST, a hosted MCP server, and a CMS write behind an approval step. Nothing outside that sliver is documented on the public product surface.

  8. =5
    WordflowAEO/GEO + content executionsells the fix: yes

    GEO content engine — deep on AI-search detection, and genuinely outside-in: live prompt simulations across seven answer engines rather than a read of the client’s own accounts. The published MCP connector carries that visibility data from the Pro tier up, and carries nothing else — the content tooling is not documented as callable, and no write path back to the client’s own site is documented at all, so the agentic ceiling is 2. A managed tier puts a human brand-and-compliance review in front of delivery. Portability rests on an unlabelled download control on the GEO Writer output pane in their own product image, with the format undocumented.

  9. =5
    Stack MoxieVerification specialistsells the fix: no

    The one genuine mechanism peer — pushing synthetic records through the customer’s own production systems is real independent detection, but only on Attribution / Conversion / Trust. Portable on every rung: month-to-month by its own terms, a documented REST API over your scenarios and results, and the test framework itself published open-source.

  10. =5
    BrightEdgeEnterprise AI-SEOsells the fix: yes

    The deepest Discoverability stack in the field, and the corpus under the index scores travels — 22 read-only MCP tools, a REST API and a Looker Studio connector all reach the rankings, URLs and click data underneath. What is not documented is any approval gate or rollback on Autopilot’s zero-touch writes, or any price, tier or contract term on the public surface.

  11. =5
    ObservePointTag & consent governancesells the fix: no

    Enterprise tag/consent governance — forensic on Attribution & Trust (Verify + Govern = 4), an external cloud scanner with nothing to install on the site it audits, and callable end to end over a documented REST API that their own docs say covers all of the product. What costs on the way out is the contract, not the code: fees are invoiced annually in advance and the term auto-renews for another year unless notice lands 30 days before the renewal date.

  12. =5
    BoostifyAI marketing operating systemsells the fix: yes

    "Operating system" branding on a paid-social core — the vendor documents guardrails on the budget automation and scopes it itself, "Current automated budget updates are Facebook-focused", while no approval gate, rollback or audit log is documented on the public product surface. Nothing here is agent-operable beyond reading: the automation is configured in the product, and no API, MCP or webhook an outside agent could call is documented anywhere. The report half is genuinely outside-in: a crawl from a pasted URL plus a competitor read off Meta and Google ad-transparency surfaces, landing as prioritised plain findings and an executive SWOT the customer keeps, with a managed-spend tier and a monthly specialist meeting behind it. Thinly published for a row we name — real pricing, the expert channel and every product screenshot live on the Shopify listing, while boostify.ai ships a homepage, terms and a privacy policy over four untouched template pages.

  13. =13
    TrackingplanTracking & consent governancesells the fix: no

    Single-pillar deep — forensic tracking/consent monitoring on Attribution & Trust that rivals us there, read off real user traffic rather than source code. The lock-in reading does not survive their own docs: nothing of theirs alters the tracking it watches, the JS snippet is not even the only install path, and the plan exports as JSON Schema — what costs on the way out is the contract, an Order Form term that auto-renews yearly unless notice lands 30 days early.

  14. =13
    SearchableAEO/GEO platformsells the fix: yes

    AEO platform — independent crawl plus multi-LLM visibility scoring, reaching Brand and Competitors as well as Discoverability, over three documented programmatic surfaces whose keys carry named read/write scopes. Portable, and the write path stops at publishing content with no approval step or undo documented over it.

  15. =13
    BoomiEnterprise iPaaS (altitude reference)sells the fix: yes

    Right capability, wrong layer — production-grade governed agent orchestration, but on the IT integration layer, not marketing infrastructure. Independence turns on where the work runs rather than on what it costs: the runtime is Boomi software executing in the customer’s own cloud or behind their firewall, and on termination every installation of it comes out — so the integrations that were the delivered value stop with it.

  16. =13
    Otterly.AIAEO/GEO surface scannersells the fix: yes

    AI-search monitor with a per-URL auditor under it — genuine crawlability and citation scoring, plus brand sentiment, competitive benchmarking and ad-appearance tracking across the answer engines, all reachable over a documented REST API and an OAuth-gated, write-capable MCP server. Lightweight and portable: nothing of theirs need sit on your site for any of it to run, the data exports, and every self-serve tier is monthly-cancellable in the agreement itself.

  17. 17
    MartechbaseStack inventorysells the fix: no

    Inventory, not diagnosis — documents auto-detection of the apps you use, then tracks owners, spend and renewals; no test of whether any of it works and no agent surface is documented. Portability rests on a download control visible in their own product image, with the format undocumented. The narrowest row on the board.

How this board is dated

Two axes, gathered by two separate sweeps, each carrying its own date. Capability — the self-serve, full-serve and agentic cells — was gathered across the field as of 2026-08-05. Independence was gathered across the field as of 2026-07-27. A pass that re-gathers one axis does not restamp the other, so the two dates differ whenever the sweeps do.

Both are field-wide stamps, and neither is the provenance of any single claim. That is the retrieval date on the individual evidence entry, shown in every cell’s detail view beside the quote it dates: the exact day that page was read, which may fall before, between or on the two field stamps. Each row also states when its own evidence was last refreshed. Where a stamp and an entry disagree, the entry is the one to test — it is what a vendor contesting a score would check.

Release 1.19 · rubric r1.4 · public-surface analyst assessments — not a product test-run of any vendor. Each cell scores capability via self-serve, full-serve and agentic operators, independently (no enforced ceiling — an agent exceeding both human channels is a crossover, flagged not forbidden). Tier bridges self-serve × agentic. Independence measures portability / lock-in only, and “sells the fix” is disclosed rather than scored. The full-serve ring is scored across the field — vendors with no expert-delivered channel read zero. Full method: methodology. No score or commentary is a quality judgement or a recommendation to buy. Vendors may contest a score with evidence: hello@yellowhead.digital (acknowledged within 5 business days, assessed within 20; corrections logged, dated, in the public corrections log).

How this is scored

Scores are public-surface analyst assessments against a published rubric — drawn from marketing sites, docs, public API/MCP documentation, third-party reviews and press. We grade observable capability and architecture, never a test-run through a vendor's product, so the bench measures none of their output. The three operator channels are scored independently — and an agent that acts without a consent/audit/rollback gate is a liability, not "readiness," so ungoverned autonomy is capped within the agentic ladder, never rewarded. An agent reaching past both human channels is a crossover — flagged as the leading edge of the shift, not forbidden.

Read the full method on the methodology page, including what this comparison does not cover — how the roster was chosen, why no product was operated, and where the scale runs out. Scores reflect shipped state as of 2026-08-05 (rubric r1.4) and are refreshed at least quarterly — sooner when a correction or a vendor's shipped changes warrant it. Think a score is wrong? Contest it with evidence at hello@yellowhead.digital — we acknowledge within 5 business days, complete the assessment within 20, and log every correction, dated, in the public corrections log.

All product names, trademarks and registered trademarks are the property of their respective owners. Vendor names are used for identification and comparison only; their use implies no affiliation with, sponsorship by, or endorsement from any vendor.