Rank4AI · Internal Methodology · v1.3 · July 2026

The Authority Asset Framework

8 Asset Classes  •  7 Primary Jobs  •  7 Selection Filters  •  3 Research Routes
AI doesn't cite useful pages. It cites believable sources. Citation is the product of asset quality, retrieval eligibility, entity confidence, external corroboration, freshness and passage extractability. A framework that only manufactures useful pages manufactures content. This framework manufactures a source.

What an Authority Asset is

The strength assets of a website and the entity behind it: the indexes, datasets, calculators, downloads, research and reference structures that give a site substance worth citing — plus the external validation that makes engines willing to believe it. The things a brand owns that competitors cannot copy overnight, and that AI engines, journalists and other sites reference as a source. This is deliberately not structured data: schema is the packaging, these are the substance underneath it.

The six things citation is actually made of

Asset quality

Is there real substance here — data, method, evidence — or a prettier restatement?

Retrieval eligibility

Can the engine's crawler actually reach and index the answer at all?

Entity confidence

Is the publisher machine-verifiably real — company number, regulator, register links?

External corroboration

Does anyone else say this brand and author are credible? Engines read consensus.

Freshness

Is it dated, year-tagged and genuinely maintained? Undated loses to dated.

Extractability

Does the answer exist as a self-contained static passage a machine can lift?

The seven primary jobs (new in v1.3)

Citation is one payoff of six, not the definition of value. Every asset type now carries a primary job, and the governing rule is: an asset is judged, and cancelled, only against its primary job. An asset that fails at its job is cut; an asset that succeeds at it is never cut for performing poorly at a job it was not built for.

RANK — win organic positions and pull search traffic EARN LINKS — give other sites an editorial reason to link or embed RETAIN — repeat visits, bookmarks, subscriptions, brand CONVERT — capture, qualify and route leads GET CITED — be the extracted, attributed answer SUPPORT TRUST — raise everything else's credibility INFRASTRUCTURE — plumbing, justified by what it carries

The Eight Asset Classes

Starred = the four disproportionate-strength types for citation (indexes, trackers, calculators-with-companions, original research — each earning it only through its static, extractable layer). The v1.3 re-audit rescued a second tier for their non-citation jobs: spreadsheets & models (CONVERT), pipeline-fed widgets (EARN LINKS), quizzes (CONVERT), hand-built segment pages (RANK), calendar-driven expert commentary (EARN LINKS). Confirmed weak on every payoff: llms.txt, un-adopted APIs, literal mirrors, consumer-less feeds, generic curated libraries.
1

Reference assets

The workhorses. Evergreen, high citation value, and the hardest class for competitors to replicate once established.

Answers: "Who holds the definitive organised view of this niche?"

Types

Indexes (criteria · providers · rules · prices)Comparison tables (dated + versioned)Glossaries (niche terms only)Directories (published inclusion rules)Timelines & chronologiesCurated libraries
Easy win: timelines are rare in most niches, extract cleanly, and win date-shaped queries. Glossaries only pay where definitions carry niche-specific information gain — institutions own the generic terms.
2

Data assets

Numbers you own, or present better than the body that publishes them.

Answers: "Where does the canonical number for this live?"

Types

Stat pages & roundupsOriginal researchTrackers (pipeline-fed = gold standard)Benchmarks ("average X for firms like yours")Public data mirrors (must add resolution)Annual-refresh assets (year in the H1)Evidence assets (first-hand proof)
Warning: a stale tracker is worse than no tracker, and an unimproved mirror of an official source is redundant. Evidence assets (test logs, FOI responses, collection logs) are now a primary selection signal for experience-weighted queries, not trust decoration.
3

Interactive assets

Tools that answer a personal question. High engagement, strong snippets, natural link targets — and the highest lead capture in the taxonomy.

Answers: "What does this mean for MY numbers?"

Types

Calculators + static companionCheckers & eligibility toolsQuizzes & diagnosticsConfigurators & estimators
The one rule that governs the class: the engine cannot operate the tool. No AI crawler executes JavaScript or reads a computed output. A bare tool gets cited as "a tool exists here"; a tool with its formula, worked examples and results table in static HTML gets cited as the answer. Companion mandatory, always.
4

Downloadable assets

Things people take away and keep. Consistently the thinnest class on most sites — which makes it a reliable quick win.

Answers: "What can I walk away with right now?"

Types

Templates (letters · invoices · scripts)Checklists (real documents)Guides as PDFs (HTML twin mandatory)Spreadsheets & models (legit capture gates)Sample documentsDatasets as CSV/JSON (licence + version)
Who cites what: journalists and researchers cite the dataset file; engines cite the explanatory HTML around it. Both must be first-class. A PDF-only asset is invisible to most answer engines.
5

Editorial structures

Content shaped as an asset rather than an article.

Answers: "Is this a maintained resource or just another blog post?"

Types

Pillar hubs + spoke clustersQ&A libraries (from real demand)How-to / process guidesExplainersCase studies & worked examplesExpert commentary streamsTroubleshooting / emergency clustersSegment & vertical pages (highest risk)Multimedia with transcripts
Two under-built, two dangerous: case studies and troubleshooting clusters are rarely built and stand out (distress queries are snippet-shaped, high-anxiety, high-loyalty). Template-generated matrix pages and demand-harvesting Q&A are the types most likely to trip scaled-content enforcement.
6

Technical & machine-facing assets

Assets whose primary audience is a machine: crawler, engine or integrator. Moves the site from content to infrastructure.

Answers: "Can other systems build on this?"

Types

APIs (the documentation is cited)Embeddable widgets (backlink value only)Structured feeds (Bing/Copilot plumbing)Code repositories (reproducibility proof)llms.txt / AI pages (facts only, never instructions)DOI-anchored publications
Honesty note: llms.txt is an experimental discovery aid, nothing more — no engine has committed to it, and instruction-like content can be filtered as prompt injection. A DOI makes work formally citable; it does not make it authoritative without the Class 8 distribution behind it.
7

Proof & trust assets

Strength signals rather than traffic assets. These make every other class more citable — and on YMYL they gate citation eligibility outright.

Answers: "Who is behind this, how are the numbers made, and when were they last checked?"

Types

Methodology pagesData Sources & Updates pageEditorial / update / correction / funding policiesPublisher identity (legal entity, plainly)Author hubsVersioned knowledge products (changelogs)How-to-cite blocks
Reframed in v1.2: these are not enablers, they are prerequisites. Machine trust anchors to legal identity, not to an about page. A visible correction log is a trust signal, not an admission.
8

External validation & entity assets

The class the engines weigh most. Everything in classes 1–7 is self-asserted; engines resolve the brand and author as entities and test whether the wider web corroborates them before citing. Mandatory for every flagship asset and every YMYL site.

Answers: "Does anyone I already trust vouch for this entity?"

Types

The entity package (schema + sameAs → Companies House · regulators · ICO · ORCID)Register & body evidence (deep-linked)Third-party corroboration (mentions count, not just links)Adoption of your work (the market treating you as a source)Distribution assets (press page · journalist data packs)Review & reputation presenceAuthor external footprintCommunity presence (participation, not spam)
The v1.2 finding: research without distribution is a page, not a publication. Every flagship launches with a distribution plan and a named list of who will corroborate, use or cite it.
The mechanical reality

What the engine actually sees

AI crawlers execute no JavaScript, click nothing, fill nothing in, and see aggregate assets only in their default state. Four consequences drive most build decisions:

Tools are invisible

A calculator without a static companion is "a tool exists here" — never the answer. The companion carries the whole citation load.

Only reachable rows exist

Tables are cited at row and cell level. Anything behind load-more, filters or pagination needs its own static URL or it cannot be cited.

HTML is the citation surface

PDFs are rarely cited, charts are discarded, widgets render as an empty div, gated content is worth zero. The HTML twin does the work.

Q&A extracts cleanest

Heading-question + answer pairs are the cleanest extraction of any format. Timelines win date-shaped queries. Evidence pages prove experience.

The full 18-row reference table — every asset type, what the engine sees, what it cites the page as — is in the Full Framework tab.

Freshness & decay by channel

Decay is channel-specific and always paired with refresh cost — when an asset rots matters less than what it costs to stop it. A pipeline-fed tracker (fast decay, near-zero cost) is a build; a manually-maintained fast-decay guide is a consolidation candidate.

Perplexity — highest sensitivity; punishes stale data pages hardest AI Overviews / AI Mode — high; undated = stale on YMYL Bing / Copilot — IndexNow-driven; fastest propagation Gemini — entity confidence beats recency Google organic — most forgiving of evergreen ChatGPT — two channels: live search vs cached/trained

Selection: the seven filters

Never build the full taxonomy on one site. Every candidate asset runs through seven filters, in order — fail one, and it's redesigned, re-scoped or not built.

Authority ceiling

Can this entity plausibly be cited for this query class? Credentials are half; corroboration is the other half. Self-asserted credibility does not pass.

Honest hook

Does the claimed source or angle genuinely exist for this vertical? "Built on official data" collapses where the dataset doesn't exist.

Information gain

What does this reveal that current top sources and AI answers do not? If the gain can't be stated in a sentence, it's redundant before it's built.

Demand & demand shape

Evergreen heads → indexes and hubs. Situational long-tail → Q&A and checkers. Annual spikes → refresh assets. Some of the best citation wins have no keyword volume at all.

Query fan-out mapping

Document the 5–10 sub-questions engines generate from the head query; confirm each is answered as a self-contained, heading-matched passage. A build requirement.

Competitor weakness

Incumbents' pattern: thin lists, gated PDFs, undated data. Whole-class white space — timelines, models, case studies, widgets — is the cheapest differentiation there is.

Channel fit

Name the target channel before build: tables → Bing/Copilot, companions → snippets and AI answers, datasets + DOIs → journalists, transcripts → video surfaces, trackers → repeat visits.

Sequencing rule for a new or thin site: trust template + publisher identity + Data Sources page first (with the entity package) → one flagship index or tracker → the cheap wins (glossary, templates, CSVs of data already held) → the Q&A library from real demand → original research last, and never without a distribution plan. Research published before the trust and entity layers exist is wasted — nobody can verify who stands behind it.

Original research: the highest-value class, in one view

The bar is not "new data". The bar is a question with real demand that nobody currently owns the answer to, answered with a method you can defend, and distributed so third parties corroborate it.

Route 1 · Primary observation

Generate data that doesn't exist: operational data, surveys, FOI requests, systematic collection. Highest effort — the only route that makes you the primary source outright.

Route 2 · Derivative narrowing

Cut public data by a dimension nobody has published: region, sector, size, time. The fastest route to ownable research. The cut must reveal, not merely subdivide.

Route 3 · Recombination

Join datasets never put together: insolvency × payment terms, price × effective dose. The joined view is the original work.

The seven steps: find the unowned question → choose the route → verify nobody owns it → source & document (evidence trail captured during the work) → publish in citation format (capsule → tables → method, CSV + licence, DOI, named author) → distribute (journalist data pack, trade press, dataset deposits — third-party pickup is the point) → build the refresh cycle (annual, year-tagged, versioned; compounding begins at the second release).

The build standards that decide citations

Answer first, fan-out covered

Capsule above the tool, tool above the explanation, every sub-question a self-contained passage. Never bury the answer.

Show the real working

Every figure displays its method — and the displayed method is the one actually used. On money and health, anything else is disqualifying.

Claim-level sourcing & honest dating

Source, date and context beside every material claim. Year in the H1 on annual assets. Fake date-refreshing is maintenance theatre and engines detect it.

Named, verifiable people

Every asset carries a named author linked to their hub; YMYL credentials externally verifiable and deep-linked. Voices, not just pages.

Answer first, convert second

Nothing between a tool and its result. The conversion bridge comes after value: post-answer capture, take-away gates, a defined lead path from every high-intent asset.

Maintenance scheduled at build

Owner + refresh trigger + decommissioning rule, or it doesn't ship. An asset without a maintenance plan is a liability with a launch date.

The loop

Measurement & lifecycle

Are you actually cited?

Monthly tracking per flagship across Google AI surfaces, ChatGPT, Perplexity, Copilot and Claude. Listed-among-sources vs cited-for-the-key-claim is the metric that matters.

Does citation convert?

AI-referred vs organic conversion, assisted conversions via later brand search, lead quality fed back into the roadmap — so citation volume is never mistaken for commercial value.

"We publish content."   "We build citable pages."
Build, prove, distribute, measure, decide.

The framework is a loop, not a production line. Each asset is periodically refreshed, consolidated or retired on the evidence — and the whole thing is applied per brand as a gap map: score the taxonomy against what each site holds, find the white space, filter the candidates, sequence by the ratings layer, attach Class 8 work to every flagship.