OPEN STANDARD

Every AI visibility score rests on questions somebody chose. Almost nobody says which.

Two measurements of the same brand, on the same day, against the same engines, can differ by fifty points based only on what was asked. A score without a stated question methodology is not a measurement. It is an opinion with a number attached.

So we wrote the method down and published it. The Buyer Intent Framework defines how the questions behind a visibility score should be built, validated and declared, so that scores can be compared, audited and reproduced by anyone.

Version 1.2 Status Published Licence CC BY 4.0 Maintainer CiteTitan, a product of Sage Titans

What goes wrong without a standard

Four consequences, all of them invisible to the person being shown the score.

Scores are not comparable

Between vendors, between agencies, or between two analyses run by the same team six months apart. There is no agreed unit, so there is nothing to compare.

Scores are trivially inflatable

A set weighted toward narrow, low-competition phrasings produces a flattering number that means nothing. Nobody can tell from the outside.

Trends are unreliable

If the question set changes between runs, movement in the score cannot be attributed to anything at all.

Buyers cannot audit

A brand shown a visibility report has no basis on which to challenge it, because the one variable that matters most was never disclosed.

Why publish it rather than keep it

A score is only meaningful if it can be compared, and comparison requires a shared input standard. A widely adopted standard is worth more to everyone measuring AI visibility, including its author, than a proprietary one nobody uses. So this is published under Creative Commons Attribution: you may use, adapt, extend and build commercially on it, including inside products that compete with ours, provided attribution is given.

It also states its own weaknesses, in section 12. A standard that hides them is not much use to the people relying on it.

The specification

The full text, as published. Cite it as: Buyer Intent Framework (BIF-12) v1.2, CiteTitan, https://www.citetitan.com/standard.html

The Buyer Intent Framework (BIF-12)

An open standard for constructing the question sets used to measure brand visibility in AI answer engines.

Version1.2
StatusPublished
Canonical URLhttps://www.citetitan.com/standard.html
MaintainerCiteTitan, a product of Sage Titans
LicenceCreative Commons Attribution 4.0 International (CC BY 4.0)
Contactcontact@citetitan.com

Abstract

AI answer engines increasingly decide which brands a buyer considers. A growing number of tools measure that visibility, and all of them produce a score. Almost none of them specify how the questions behind the score were chosen.

This matters because the question set determines the score more than any other variable. Two measurements of the same brand, on the same day, against the same engines, can differ by fifty points based only on which questions were asked. A score without a stated question methodology is not a measurement. It is an opinion with a number attached.

BIF-12 is a public specification for constructing those question sets. It defines twelve buyer-intent archetypes, the rules for a valid question, the rules for a valid set, how to handle locale, and how to declare where each question came from.

It is published openly so that AI visibility scores can be compared, audited and reproduced by anyone, including by people who do not use CiteTitan.


1. Why this exists

1.1 The unstated variable

Search-rank measurement has an agreed unit: the keyword. Two tools tracking the same keyword measure the same thing, and disagreement between them is a data-quality question rather than a definitional one.

AI visibility measurement has no such unit. The equivalent input, the buyer question, is authored, not observed. It can be phrased a hundred ways, and the phrasing changes the answer. Which means every AI visibility score currently rests on an undisclosed authoring decision.

1.2 What goes wrong without a standard

  • Scores are not comparable. Between vendors, between agencies, or between two analyses run by the same team six months apart.
  • Scores are trivially inflatable. A set weighted toward narrow, low-competition phrasings produces a flattering number that means nothing.
  • Trends are unreliable. If the question set changes between runs, movement in the score cannot be attributed to anything.
  • Buyers cannot audit. A brand shown a visibility report has no basis on which to challenge it.

1.3 What BIF-12 does not fix

It does not make AI answers stable. Answer engines produce different responses to identical prompts across runs, and no question methodology changes that. Variance must be handled separately, by reporting ranges rather than point estimates. BIF-12 addresses only the input side.


2. Scope

In scope: the construction, validation, composition, localisation and provenance of question sets used to measure brand presence in AI-generated answers.

Out of scope: scoring methodology; engine selection; run frequency; variance handling; sentiment analysis; citation-source analysis; any recommendation about how to improve visibility.

BIF-12 is deliberately narrow. It specifies the input so that arguments about the output can be about something real.


3. Terminology

The key words MUST, MUST NOT, SHOULD, SHOULD NOT and MAY are used in the sense of RFC 2119.

TermDefinition
QuestionA single natural-language prompt submitted to an answer engine.
SetA group of questions run together and scored as one analysis.
ArchetypeOne of the twelve buyer-intent patterns defined in Section 5.
LexiconThe locale-specific vocabulary for a category (Section 7).
ParametersClient-specific values that populate a template: segment, use case, attribute, problem, and so on.
LocaleA language–market pair, expressed as an IETF tag: en-GB, de-DE, en-IN.
ProvenanceThe recorded origin of a question (Section 8).
Locked setA set whose question text is frozen so that repeat runs are comparable.

4. The three-layer model

A conformant question is produced by combining three separable layers. Separating them is what makes the framework portable across categories and markets.

Question  =  Archetype  ×  Lexicon  ×  Parameters
             (universal)   (per locale)  (per client)

Layer 1 — Archetype. The intent pattern. Universal. Identical in every category, market and language. Twelve of them, defined in Section 5.

Layer 2 — Lexicon. How the category is actually named by buyers in a given locale, plus locale-specific attribute phrasing and the incumbent brands buyers name. This layer is where most cross-border measurement fails.

Layer 3 — Parameters. The client's segments, use cases, differentiators and buyer problems.

A consequence worth stating plainly: adding a locale regenerates every question in every category, and adding a category propagates to every locale. Implementations that hard-code finished question strings cannot do either.


5. The twelve archetypes

Each archetype captures a distinct moment in how a buyer asks. Together they span the buying journey from unaware to decided.

Templates are given in two variants because grammar differs between things you buy and people you hire. {cat} is the lexicon entry; braced terms in lowercase are parameters.

{cat} carries grammatical number. Buyers ask for a list in the plural ("best food delivery apps") and for guidance in the singular ("how do I choose a food delivery app"). A lexicon entry MUST therefore supply both forms, and each template uses the one its grammar requires. Where a category noun is a mass noun with no plural, such as "CRM software", the two forms are identical.


A1 · Category leader

Definition. The broadest possible buying question in the category, with no qualifier.

Purpose. Establishes the baseline. This is the most contested question in any category and typically the hardest for a non-incumbent to win.

Template. best {cat} (both variants)

Locale examples.

  • en-US — best CRM software
  • en-GB — best chartered accountants
  • en-AU — best removalists

Inclusion. MUST appear in every conformant set.

Common errors. Adding a qualifier and still calling it A1. "Best affordable CRM software" is A4, not A1, and scoring it as A1 overstates performance.


A2 · Primary segment

Definition. The category question narrowed to the client's core customer type.

Purpose. Where mid-sized brands realistically compete. A brand invisible on A1 is often well cited here, and the gap between A1 and A2 is diagnostic.

Template. best {cat} for {segment}

Locale examples.

  • en-US — best payroll software for small businesses
  • en-IN — best packers and movers for corporate relocations
  • en-GB — best solicitors for first-time buyers

Inclusion. MUST appear in every conformant set.

Common errors. Choosing a segment the client wishes it served rather than the one it does. Segments MUST be evidenced from the client's actual customer base.


A3 · Secondary segment

Definition. A second, distinct customer type.

Purpose. Distinguishes broad visibility from narrow visibility. A brand cited on A2 but absent on A3 has a positioning that is legible to engines in only one direction.

Template. best {cat} for {segment_2}

Locale examples.

  • en-US — best project management software for remote teams
  • de-DE — beste Projektmanagement-Software für Agenturen
  • en-AE — best insurance brokers for construction contractors

Inclusion. SHOULD appear in sets of ten questions or more.


A4 · Value and price

Definition. Intent expressed through cost.

Purpose. Price-led questions frequently return an entirely different brand set from A1. Many premium brands score well on A1 and are absent from A4 without knowing it.

Template. most affordable {cat} · best value {cat}

Locale examples.

  • en-US — most affordable HVAC contractors
  • en-GB — cheapest conveyancing solicitors
  • en-IN — most affordable CA firms for startups

Inclusion. SHOULD appear in sets of ten questions or more.

Common errors. Using words like "cheap" where the local register makes them pejorative. This is a lexicon decision, not a translation decision.


A5 · Geography and scope

Definition. Intent bounded by place.

Purpose. Decisive for any business with a service area. For non-local businesses it tests whether engines associate the brand with a market at all.

Template.

  • Local: best {cat} in {city}
  • Non-local: best {cat} in {region}

Locale examples.

  • en-US — best personal injury attorneys in Phoenix
  • en-GB — best estate agents in Bristol
  • en-SG — best corporate secretarial firms in Singapore

Inclusion. MUST appear in every conformant set for a business with a defined service area.

Common errors. Leaving {city} or {region} unresolved. A question containing an unresolved token MUST NOT be submitted to an engine. See Section 6.


A6 · Job to be done

Definition. The question phrased as a task rather than a category.

Purpose. Closest to how people actually address answer engines. Buyers describe what they are trying to do far more often than they name a product category.

Template.

  • Product: what {cat} should I use for {use_case}
  • Service: who should I hire for {use_case}

Locale examples.

  • en-US — who should I hire for a kitchen remodel
  • en-GB — what accounting software should I use for filing VAT returns
  • en-AU — who should I hire for an interstate move

Inclusion. MUST appear in every conformant set.


A7 · Attribute and feature

Definition. Intent expressed through a single decisive requirement.

Purpose. Tests whether the brand's stated differentiator is legible to engines. A brand whose entire marketing rests on one attribute, and which is absent from the A7 question about it, has a content problem it can act on immediately.

Template.

  • Product: best {cat} with {attribute}
  • Service: {cat} that {attribute}

Locale examples.

  • en-US — personal injury attorneys that work on contingency
  • en-GB — personal injury solicitors that work on a no win no fee basis
  • de-DE — Buchhaltungssoftware mit automatischem Bankabgleich

Inclusion. SHOULD appear in sets of ten questions or more.

Common errors. Attribute phrasing is among the most locale-sensitive elements in the framework. Contingency fees, fee caps, warranty terms, licensing and certification language differ by jurisdiction and are frequently regulated. Attribute strings MUST be treated as lexicon entries, not parameters.


A8 · Alternative and specialism

Definition. For products, the switching query. For services, the niche query.

Purpose. The product form is among the highest commercial-intent questions that exists — the buyer has already decided to leave an incumbent. The service form tests visibility in a defensible niche rather than a contested category.

Template.

  • Product: {incumbent} alternatives
  • Service: {cat} that specialise in {specialism}

Locale examples.

  • en-US — QuickBooks alternatives
  • en-GB — Xero alternatives
  • en-IN — Tally alternatives

Inclusion. SHOULD appear in sets of ten questions or more.

Common errors. Using a global incumbent where a local one dominates. The three examples above are the same archetype, the same category and three different correct answers. {incumbent} is a lexicon entry and MUST be resolved per locale.


A9 · Vetting

Definition. The buyer asks how to evaluate the category rather than which brand to pick.

Purpose. Early-stage. Reveals which sources the engines trust in a category, which is often more actionable than which brands they name.

Template.

  • Product: how do I choose {cat}
  • Service: what should I look for when comparing {cat}

Locale examples.

  • en-US — what should I look for when comparing roofing contractors
  • en-GB — how do I choose payroll software
  • en-ZA — what should I look for when comparing medical aid brokers

Inclusion. SHOULD appear in every set.

Common errors. Scoring A9 identically to A1. Brand citation rates are structurally lower here, and mixing the two without distinguishing them depresses the aggregate score for reasons unrelated to brand performance.


A10 · Direct recommendation

Definition. An explicit request for the engine to choose.

Purpose. The purest available test of recommendation share. The buyer is not asking for options; they are asking to be told.

Template. who do you recommend for {need}

Locale examples.

  • en-US — who do you recommend for commercial insurance
  • en-AE — who do you recommend for company formation in Dubai
  • fr-FR — quelle agence recommandez-vous pour le référencement

Inclusion. SHOULD appear in every set.

Common errors. Some engines decline to make direct recommendations, or hedge. A non-response MUST be recorded as a non-response, never scored as a zero against the brand.


A11 · Problem first

Definition. The buyer describes a situation and names no category at all.

Purpose. The most valuable and least measured question type. There is no keyword to optimise for, and the engine must infer the category before it names any brand. Winning A11 is strong evidence of genuine entity authority rather than page-level optimisation.

Template. {problem_statement} — a natural sentence in the buyer's own words.

Locale examples.

  • en-US — my basement flooded overnight, who do I call
  • en-GB — our organic traffic dropped 40% and we do not know why
  • en-IN — my landlord is refusing to return my security deposit, what are my options

Inclusion. MUST appear in every conformant set.

Common errors. Writing the problem in marketing language rather than buyer language. "Seeking a scalable revenue-operations solution" is not a problem statement. If it does not sound like something a person typed at 11pm, it is not A11.


A12 · Trust and rating

Definition. Intent filtered by reputation.

Purpose. Reputation-weighted questions tend to draw on different sources — review platforms, directories, editorial roundups — from those behind A1. The gap between A1 and A12 indicates whether a brand's problem is authority or reputation.

Template.

  • Product: highest rated {cat}
  • Service: most trusted {cat}

Locale examples.

  • en-US — highest rated pet insurance providers
  • en-GB — most trusted removal companies
  • en-AU — most trusted migration agents

Inclusion. SHOULD appear in sets of ten questions or more.


6. Question validity

A question is conformant only if all of the following hold. Implementations SHOULD validate automatically and warn on failure.

#RuleRationale
V1Conversational. Phrased as a person would address an answer engine, not as a keyword string.Keyword grammar produces atypical answers.
V2Unbranded. MUST NOT contain the measured brand's name or domain.A branded question measures whether the engine knows the brand exists, not whether it recommends it. Including one inflates the score and invalidates the benchmark.
V3Commercial intent. Contains a selection marker — best, top, recommend, choose, compare, which, who, alternative — or is a valid A11 problem statement.Informational questions measure a different phenomenon.
V4Shortlist-answerable. A reasonable engine response would name brands.Questions that cannot return a brand list cannot be scored.
V5Fully resolved. Contains no unresolved token.A literal {city} submitted to an engine returns a meaningless answer.
V6Stable. Contains no date, price, promotion or other element that will expire.Trends require the question to be re-runnable unchanged.
V7Single intent. Maps to exactly one archetype.Compound questions cannot be attributed.

7. Set composition

7.1 Minimum viable set

A set of fewer than five questions is a spot check, and MUST be labelled as such. It MUST NOT be presented as a benchmark and MUST NOT be used to draw a trend line.

7.2 Archetype spread

A conformant set of five MUST include A1, A2 or A5, A6, A11, and one of A9 or A10.

No more than two questions in a five-question set may share an archetype. Five phrasings of the same intent constitute one measurement, not five, and produce a score with false precision.

Reference five-question sets:

Business typeArchetypes
Local service businessA1, A5, A6, A11, A9
Product or softwareA1, A2, A6, A11, A9

7.3 Full sets

A twelve-question set SHOULD contain one question per archetype. This is the reference composition against which cross-brand and cross-locale benchmarks are computed.

7.4 Locking

Once a set has been run, its question text MUST be frozen for all subsequent runs intended for trend comparison.

Editing a locked set MUST fork a new benchmark rather than silently continuing the existing one. Implementations MUST make this visible to the user at the moment of editing.

7.5 Deduplication

Near-duplicate questions within a set MUST be flagged. They consume budget and bias the average toward whichever intent was accidentally repeated.


8. Provenance

Every question MUST carry a provenance tag recording where it came from. The tag SHOULD be displayed wherever the score is displayed.

TagMeaningConfidence
SALESSupplied by the client from discovery calls, objection logs or win/loss notesHighest — verified buyer language
SEARCH_DATAImported from the client's own search console or analytics, filtered to selection intentHigh
LIBRARYSelected from a published category set built to this specificationMedium-high — comparable across brands
SUGGESTEDGenerated by prompting an answer engine for likely buying questionsMedium
COMPETITORDerived from a rival's FAQ, comparison page or advertisingMedium
MANUALWritten freehand with no external sourceUnverified

A note on SUGGESTED. Asking an answer engine which questions buyers ask is circular: the model may favour phrasings it is already well-equipped to answer. Question sets drawn predominantly from SUGGESTED SHOULD be treated as directional. This limitation applies to the reference implementation as much as to any other, and is disclosed here for that reason.


9. Locale

9.1 The core problem

Cross-border measurement fails at the category noun, not at the language boundary. The same concept is named differently in markets that share a language:

Concepten-USen-GBen-AUen-INen-AE
Injury lawyerpersonal injury attorneypersonal injury solicitorpersonal injury lawyerpersonal injury lawyerpersonal injury lawyer
Property agentrealtorestate agentreal estate agentproperty brokerreal estate agent
AccountantCPA firmchartered accountantaccountantCA firmchartered accountant
Moversmoving companyremovals companyremovalistspackers and moversmovers and packers
Climate systemsHVAC contractorboiler engineerair conditioning specialistAC serviceAC maintenance company

A set that measures "best HVAC companies in Manchester" is not slightly wrong. It is measuring a category that does not exist in that market.

9.2 Lexicon entries

For each category–locale pair, a lexicon entry MUST define: the category noun, its common variants, locale-appropriate attribute phrasing, and the incumbent brand or brands used by A8.

9.3 Verification tiers

Locale coverage MUST be published with its verification status. Implementations MUST display the tier alongside any score produced from a Tier 2 or Tier 3 set.

TierDefinitionReporting requirement
Tier 1 — VerifiedHand-curated and reviewed by a native speaker with category knowledgeMay be reported without qualification
Tier 2 — ReviewedMachine-adapted, then reviewed by a native speakerTier MUST be displayed
Tier 3 — UnverifiedMachine-adapted, not reviewedMUST be labelled directional; MUST NOT be used for cross-brand benchmarking

Publishing an unverified locale as unverified is conformant. Publishing an unverified locale without saying so is not.

9.4 Execution locale

Locale is not only a property of the question text. Answer engines geo-personalise and language-switch. A UK question executed from an Indian network may return an answer set contaminated by a market the brand does not serve.

Conformant implementations MUST:

  1. Control the execution locale — network egress region and any locale or language parameters the engine exposes — to match the declared locale of the set.
  2. Record the execution locale in run metadata.
  3. Lock the execution locale for the life of the set, alongside the question text. A trend across changed execution locales is not a trend.

9.5 Cross-locale comparison

Scores from different locales MAY be compared only where both sets use the same archetype composition and both are Tier 1. Such comparisons SHOULD be reported per locale rather than averaged, since averaging conceals precisely the gap that makes the comparison useful.


10. Conformance

An analysis MAY be described as BIF-12 conformant where all of the following hold:

  1. Every question maps to exactly one declared archetype.
  2. Every question satisfies the validity rules in Section 6.
  3. The set satisfies the composition rules in Section 7.
  4. Every question carries a provenance tag from Section 8.
  5. The locale is declared, its verification tier is disclosed, and the execution locale matches it.
  6. The question set and execution locale are locked for trend comparison.
  7. The specification version used is stated.

Conformance concerns the input only. It carries no implication about scoring accuracy, engine coverage or variance handling, and it MUST NOT be presented as an endorsement of any product, including the reference implementation.


11. Versioning

This specification is versioned as MAJOR.MINOR.

  • MINOR — clarifications, added locale guidance, corrected lexicon examples. Backward compatible.
  • MAJOR — changes to archetype definitions or composition rules. Existing benchmarks are not comparable across a major version.

Every analysis MUST record the specification version it was built under. Improvements to templates or lexicons MUST NOT retroactively alter locked sets; implementations SHOULD offer migration to a newer version as an explicit action that forks a new trend line.

The changelog is maintained at the canonical URL.


12. Known limitations

Stated plainly, because a standard that hides its weaknesses is not much use to the people relying on it.

  1. No prompt-volume weighting. BIF-12 specifies question structure, not question frequency. It does not know which of the twelve archetypes buyers in a given category actually use most. Weighting requires real user prompt-volume data, which is not publicly available at the time of writing.
  1. Archetype coverage is asserted, not proven. The twelve archetypes were derived from buying-question patterns across a range of categories. They are not the output of a formal study of AI prompt logs. They may under-represent intents specific to categories not yet examined.
  1. English-first. The framework was developed in English and extended outward. Languages with different question grammar, honorific systems or noun-class structures may require archetype forms not yet defined.
  1. Problem statements are hard to standardise. A11 is the most valuable archetype and the least mechanisable. It resists templating by design, which makes it the most likely to be authored poorly.
  1. It does not address variance. Answer engines are non-deterministic. A perfectly conformant question set still produces a score that moves run to run for reasons unrelated to brand performance. Variance must be handled by reporting ranges rather than point estimates. That is a separate problem and outside this scope.
  1. Regulated categories need review. Legal, medical and financial services carry jurisdiction-specific constraints on how services may be described. Lexicon entries in these categories SHOULD be reviewed by someone qualified in that jurisdiction.

13. Licence and attribution

BIF-12 is published under Creative Commons Attribution 4.0 International (CC BY 4.0).

You are free to use, adapt, extend and build commercially on this specification, including within competing products, provided attribution is given.

Required attribution:

Based on the Buyer Intent Framework (BIF-12) v1.2, published by CiteTitan, https://www.citetitan.com/standard.html

The name "BIF-12" MAY be used to describe conformant implementations. It MUST NOT be used in a way that implies certification, endorsement or partnership.


14. Contributing

The specification improves through use. Contributions are welcomed in three forms:

  • Lexicon corrections. A category noun that is wrong for your market is the highest-value correction you can send. Include the locale, the category, the incorrect term and the term buyers actually use.
  • Archetype proposals. If you have identified a buyer intent not covered by A1–A12, submit it with examples from at least two categories.
  • Conformance findings. Cases where the rules produce a poor result in practice.

Send to contact@citetitan.com. Accepted contributions are credited in the changelog.


15. Changelog

VersionDateChanges
1.22026A5's non-local template changed from best {cat} for companies in {region} to best {cat} in {region}. The original presumed a business buyer, which is wrong for consumer services; the business-buyer signal belongs in A2's segment. Added the requirement that a lexicon entry supplies both singular and plural forms of the category noun, since list questions and guidance questions need different grammatical number. Neither change alters an archetype's meaning or the composition rules, so existing benchmarks remain comparable.
1.12026Corrected the reference five-question set for products, which included A8 and therefore satisfied neither rule 7.2 (no A9 or A10 present) nor the Section 5 note that A8 belongs to sets of ten or more. A9 replaces A8 in the reference five; A8 is unchanged and still appears in the twelve-question set. No rule, archetype or template was altered, so existing benchmarks remain comparable.
1.02026Initial publication. Twelve archetypes, validity rules, composition rules, provenance taxonomy, locale tiers and execution-locale requirements.

Appendix A — Reference five-question sets

Local service business (en-GB, dental clinic, Manchester)

ArchetypeQuestion
A1best dental clinics
A5best dental clinics in Manchester
A6who should I hire for a dental implant
A11I need an implant but I am worried about the cost
A9what should I look for when comparing dental clinics

Software product (en-US, CRM)

ArchetypeQuestion
A1best CRM software
A2best CRM software for small businesses
A6what CRM software should I use for managing a sales pipeline
A11my sales team keeps losing track of leads, what should we use
A9how do I choose CRM software

Appendix B — Frequently asked questions

Is BIF-12 specific to any one tool? No. It is published openly and can be implemented by anyone, including by tools that compete with the reference implementation.

Why publish a method that competitors can adopt? Because a score is only meaningful if it can be compared, and comparison requires a shared input standard. A widely adopted standard is more valuable to everyone measuring AI visibility, including its author, than a proprietary one nobody uses.

Does conformance mean a score is accurate? No. Conformance concerns only how the questions were constructed. Scoring, engine coverage and variance handling are separate matters.

Can I use fewer than twelve questions? Yes. Five is the minimum for a benchmark, with the composition given in Section 7.2. Below five, the result MUST be labelled a spot check.

How do I handle a category not in any published library? Use the three-layer model directly. Establish the lexicon entry for your locale, gather the client parameters, and apply the twelve templates. The framework is designed for exactly this case.

What if my client operates in several markets? Build one set per locale and report them separately. Averaging across locales conceals the gap that makes the analysis worth running.

See it working

CiteTitan is the reference implementation. It builds a conformant question set for your business, shows you every question before it runs, and tags each one with the intent it tests.

Corrections and contributions: standard@citetitan.com