The public method · Instrument v1.0.0 · 2026-06-09

Show your workings.

We publish every question, with its rationale and evidence. You see the scoring model, the weights, the thresholds and every change. Trust in a benchmark comes from being more open about method than anyone else. So we are.

Scoring model

How the number is made

Every scored question maps to 0 to 4 points. Likert items run from strongly disagree (0) to strongly agree (4). Behavioural items run from anchor (a), 0, to anchor (e), 4. We store your raw answers, never just the computed scores. That is what lets us re-score you under a future instrument version, without asking you to sit the assessment again.

Readiness: 18 questions across six dimensions, maximum 72 points, normalised to 0 to 100. Each dimension produces a 0 to 100 sub-score from its three questions. Realisation: 6 questions, maximum 24 points, normalised to 0 to 100.

The headline AiR Bench is a weighted blend: 60% readiness, 40% realisation. The split reflects today's evidence. Realisation is the goal. But for the mid-market majority, the score has to stay diagnostic of what to fix, and readiness is where the fixes live. As the population matures, we expect the weighting to shift towards realisation. That shift will be a published finding in its own right.

The AiR Grid splits at 50 on both axes. That is the midpoint of the scale, not a band boundary, so the grid and the bands stay independent readings.

Emerging

0 to 39

Foundations need to be laid before a meaningful pilot is viable. Strategic clarity and leadership alignment come first.

Developing

40 to 59

Some foundations in place with one or two critical gaps. The dimension detail shows exactly which.

Ready

60 to 79

Strong foundations across most dimensions. Positioned to move into structured production work and start generating measured return.

Leading

80 to 100

The conditions for AI to compound value, and evidence that it is doing so. The question is sequencing, not whether to act.

We calibrate the bands so the top band stays genuinely rare. The research is consistent: real high performers are 5 to 13% of the population. A benchmark where a third of respondents are Leading is flattery, not measurement.

Foundations

Low readiness, low realisation

Early. The work is to define one quantified opportunity and build the conditions around it. Honest, common, fixable.

Prepared, Not Proving

High readiness, low realisation

The GenAI Divide quadrant, and the most populated in the research: conditions exist, value does not. The gap is execution design: workflow redesign, production discipline, measurement.

Running Hot

Low readiness, high realisation

Value is being created ahead of foundations. Often a few brilliant individuals or one heroic team. The value is real and fragile: governance, data and process debt will tax it.

Compounding

High readiness, high realisation

Conditions and evidence. The work is sequencing and protecting focus. This is where compounding competitive advantage lives.

The instrument · 24 scored questions · Why we ask each one

Every question, with its evidence

Behavioural anchors are the credibility engine: five defined states, hard to flatter, easy to verify. We use Likert agreement scales only where the construct is genuinely attitudinal. Questions marked ✳ form the short-form subset.

D1. Strategic Clarity

Whether AI ambition has been converted into named, owned, quantified opportunities.

D1.1 · Agreement (Likert 0 to 4)

We have named specific AI opportunities, each tied to a measurable business outcome with a single accountable owner.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Named, owned, quantified opportunities are the single clearest marker of strategy converted into work. McKinsey's high performers are distinguished by outcome-based objectives, not by ambition statements.

Evidence: McKinsey, The State of AI (2025 edition)

D1.2 · Behavioural anchors (0 to 4)

Which best describes your AI priorities today?

  1. (a)=0No defined priorities.
  2. (b)=1A general ambition without named opportunities.
  3. (c)=2A shortlist of opportunities, not yet quantified.
  4. (d)=3Quantified opportunities with value targets and owners.
  5. (e)=4A quantified portfolio reviewed at board level.

Why we ask · The ladder from ambition to board-reviewed portfolio is observable and hard to flatter. Each anchor is a state you either are in or are not.

Evidence: McKinsey, The State of AI (2025 edition) · MIT NANDA, The GenAI Divide: State of AI in Business 2025

D1.3 · Agreement (Likert 0 to 4)

Senior leadership can explain why AI matters to this organisation in the language of our strategy, not in general terms.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · When leaders can only describe AI in general terms, prioritisation defaults to whoever shouts loudest. Strategy-specific language is a proxy for genuine strategic integration.

Evidence: McKinsey, The State of AI (2025 edition)

D2. People and Change

Leadership engagement, frontline involvement, and actual weekly use across the organisation.

D2.1 · Agreement (Likert 0 to 4)

Our leadership team actively champions AI adoption and dedicates meaningful time to it.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · High performers show roughly three times the leadership engagement of everyone else. It is the strongest single differentiator in the McKinsey data.

Evidence: McKinsey, The State of AI (2025 edition)

D2.2 · Behavioural anchors (0 to 4)

What share of your people use AI tools in their actual work at least weekly?

  1. (a)=0Effectively none.
  2. (b)=1Under 10%.
  3. (c)=210 to 30%.
  4. (d)=330 to 60%.
  5. (e)=4Over 60%.

Why we ask · Weekly use in real work is the adoption measure that survives scrutiny. Licence counts and pilot enrolment do not.

Evidence: Stanford HAI AI Index 2026 · Anthropic Economic Index (January 2026)

D2.3 · Agreement (Likert 0 to 4)

The teams who will use AI day to day have been involved in shaping how it is deployed.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · The MIT research locates the failure of most pilots in a learning gap, not a technology gap. Frontline involvement in deployment design is the observable counter to that gap.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

D3. Process Readiness

Whether the work AI will touch is understood, documented and being redesigned, not just augmented.

D3.1 · Agreement (Likert 0 to 4)

The processes we want AI to improve are documented, well understood and currently measurable.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · You cannot redesign what you cannot describe, and you cannot prove improvement without a baseline. Documented, measurable processes are the precondition for measured value.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

D3.2 · Behavioural anchors (0 to 4)

Has any workflow been redesigned end to end around AI, rather than AI being added to the existing way of working?

  1. (a)=0No, and not discussed.
  2. (b)=1Discussed, nothing started.
  3. (c)=2One redesign in progress.
  4. (d)=3One redesigned workflow shipped.
  5. (e)=4Several redesigned workflows shipped.

Why we ask · End-to-end workflow redesign is the highest-signal single behaviour in the current research: only around 21% of organisations have done it once, and it is where AI's value concentrates.

Evidence: McKinsey, The State of AI (2025 edition)

D3.3 · Agreement (Likert 0 to 4)

For our target use cases, we know which steps AI will replace, which it will augment, and which it will leave unchanged.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Replace, augment, leave alone: the three-way split is the working vocabulary of process redesign. Knowing it per use case marks the difference between a plan and a hope.

Evidence: Anthropic Economic Index (January 2026)

D4. Data Foundations

Whether the data the use cases need is accessible, usable and governed.

D4.1 · Agreement (Likert 0 to 4)

We have access to the data our priority use cases require, and it is in a usable, organised state.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Anchored to priority use cases deliberately: data readiness in the abstract is unanswerable, data readiness for the three things you intend to do is not.

Evidence: Cisco AI Readiness Index 2025

D4.2 · Behavioural anchors (0 to 4)

If you adopted a new AI tool tomorrow, how quickly could it be connected to the business data it needs, with appropriate controls?

  1. (a)=0We could not.
  2. (b)=1Months.
  3. (c)=2Weeks.
  4. (d)=3Days.
  5. (e)=4Days, through established, governed patterns we have used before.

Why we ask · Time-to-connect is the practical measure of data readiness, and the repeated, governed pattern at the top anchor is what Cisco calls the antidote to AI Infrastructure Debt.

Evidence: Cisco AI Readiness Index 2025

D4.3 · Agreement (Likert 0 to 4)

Security, privacy and quality constraints on our data are understood and managed; they shape our AI work rather than blocking it.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Constraints that are understood become design inputs. Constraints that are vague become vetoes. The difference shows up directly in pilot-to-production conversion.

Evidence: Cisco AI Readiness Index 2025 · MIT NANDA, The GenAI Divide: State of AI in Business 2025

D5. Governance and Trust

Whether use is governed by policy people actually know, with defined human review and explainability.

D5.1 · Agreement (Likert 0 to 4)

We have a clear policy on acceptable AI use, and the people doing the work actually know what it says.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · The second clause is the question. Most organisations have a policy; far fewer have a policy anyone could recite. Governance that exists only as a document governs nothing.

Evidence: Cisco AI Readiness Index 2025

D5.2 · Behavioural anchors (0 to 4)

Where AI contributes to work that matters, is there a defined human review point?

  1. (a)=0AI is not in meaningful use.
  2. (b)=1Review is informal and inconsistent.
  3. (c)=2Defined for some workflows.
  4. (d)=3Defined for all critical workflows.
  5. (e)=4Defined, monitored, and refined as autonomy increases.

Why we ask · As deployment becomes agentic, the defined human review point is the control that matters most. The top anchor describes review as a living system, which is what increasing autonomy requires.

Evidence: Cisco AI Readiness Index 2025 · Anthropic Economic Index (January 2026)

D5.3 · Agreement (Likert 0 to 4)

We could explain to a customer or a regulator, today, where and how AI is used in our business.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Explainability on demand is the practical test of governance maturity. 'Today' is in the question because a capability that needs six weeks of preparation is not a capability.

Evidence: Cisco AI Readiness Index 2025

D6. Execution Track Record

Whether the organisation ships digital change, and how fast.

D6.1 · Agreement (Likert 0 to 4)

We have launched new digital tools or ways of working successfully in the past 24 months.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · Organisations that ship digital change ship AI change. A 24-month window keeps the evidence recent enough to mean something.

Evidence: DORA, State of DevOps (2014 to present)

D6.2 · Behavioural anchors (0 to 4)

From decision to working in production, how long does digital change typically take here?

  1. (a)=0More than 12 months.
  2. (b)=16 to 12 months.
  3. (c)=23 to 6 months.
  4. (d)=3Under 3 months.
  5. (e)=4Under a month for changes of this kind.

Why we ask · Lead time from decision to production is DORA's most durable metric, carried into the AI context. Speed of change is itself a readiness asset, because the frontier will move again.

Evidence: DORA, State of DevOps (2014 to present)

D6.3 · Agreement (Likert 0 to 4)

When an initiative is stalling, we find out quickly and act, rather than letting it drift.

Strongly disagree = 0 → Strongly agree = 4

Why we ask · The 95% pilot failure figure is mostly a story of drift. Fast detection and decisive action on stalling work is the organisational muscle that prevents it.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R. Value Realisation

The realisation axis

Production deployments, pilot-to-production conversion, measured value, where value concentrates, frontier responsiveness.

R1 · Behavioural anchors (0 to 4)

How many AI use cases are in production use today, meaning relied upon in real work, not piloted?

  1. (a)=0None.
  2. (b)=1One.
  3. (c)=2Two to four.
  4. (d)=3Five to ten.
  5. (e)=4More than ten.

Why we ask · Production use, relied upon in real work, is the bar. The definition is in the question because the adoption-to-value gap lives precisely in the space between a pilot and a dependency.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R2 · Behavioural anchors (0 to 4)

Of the AI initiatives you have started in the last 18 months, what share reached production?

  1. (a)=0We have not started any.
  2. (b)=1Under 10%.
  3. (c)=210 to 30%.
  4. (d)=330 to 60%.
  5. (e)=4Over 60%.

Why we ask · Pilot-to-production conversion is the rate that separates the 5% from the 95% in the MIT data. It measures execution design, not enthusiasm.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R3 · Behavioural anchors (0 to 4)

Has any AI deployment delivered value you have measured against a baseline?

  1. (a)=0No.
  2. (b)=1We believe so, but have not measured it.
  3. (c)=2One, measured.
  4. (d)=3Several, measured.
  5. (e)=4Several, measured and reported at board level.

Why we ask · Measured against a baseline is the only claim of value this instrument accepts. Anchor (b), we believe so, is where most of the market sits, and the instrument names it without judgement.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025 · Anthropic Economic Index (January 2026)

R4 · Behavioural anchors (0 to 4)

Where is the value you have seen concentrated?

  1. (a)=0We have not seen value yet.
  2. (b)=1Individual productivity.
  3. (c)=2Team-level efficiency.
  4. (d)=3Function-level results visible in the P&L.
  5. (e)=4Business-model level: new offers, new pricing, new capacity.

Why we ask · Value climbs a ladder: individual, team, function, business model. The MIT finding that budgets concentrate where ROI does not makes the location of value a first-class measurement.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R5 · Behavioural anchors (0 to 4)

When a significant new model capability is released, how long before it reaches your workflows?

  1. (a)=0We do not track releases.
  2. (b)=1More than 12 months, if at all.
  3. (c)=2Within 6 to 12 months.
  4. (d)=3Within 3 months.
  5. (e)=4Within weeks, through an established evaluation route.

Why we ask · The frontier moves twice a year. Frontier responsiveness measures whether your organisation compounds those moves or pays for them later. This is the question most exposed to re-anchoring as the benchmark evolves.

Evidence: Anthropic Economic Index (January 2026) · Stanford HAI AI Index 2026

R6 · Behavioural anchors (0 to 4)

Have you deployed any agentic systems, meaning AI carrying out multi-step work with defined autonomy and oversight?

  1. (a)=0No, and no plans.
  2. (b)=1Exploring.
  3. (c)=2Piloting.
  4. (d)=3One in production.
  5. (e)=4Several in production.

Why we ask · 83% of organisations plan agentic deployment while foundations lag. The definition (defined autonomy and oversight) is in the question so that an unsupervised script cannot score as an agentic system.

Evidence: Cisco AI Readiness Index 2025

Benchmarks and percentiles

Honest about sample size

Every readout shows your percentile against the full UK dataset and against your sector, on the headline score and each dimension. We publish sector percentiles only when the sector cell reaches the minimum sample of 25 respondents. Below that threshold, the readout says so plainly.

Until the dataset reaches critical mass, we seed benchmarks from the published research distributions (McKinsey, Cisco) and label them research-calibrated wherever they appear. The calibration targets are clear: high performers held to 5 to 13% of the population, the median in the Developing band, and realisation trailing readiness (the GenAI Divide). We will announce the switch to live data in the changelog. The instrument never pretends seeded data is its own.

On anti-gaming: behavioural anchors are the scoring backbone, the verified stratum is the calibration reference, and we flag straight-line response patterns internally. We never scold anyone in public.

The Evidence Register · Founding set

What the instrument is built on

MIT NANDA, The GenAI Divide: State of AI in Business 2025

active · scope: 150 interviews, 350 surveyed employees, 300 analysed deployments

95% of enterprise AI pilots delivered no measurable P&L impact. The barrier is a learning gap, not the technology. Vendor partnerships succeed roughly twice as often as internal builds. Budgets concentrate in sales and marketing while ROI concentrates in operations.

McKinsey, The State of AI (2025 edition)

active · scope: Global survey of 1,491 participants across 101 countries

88% of organisations use AI in at least one function, but only around 21% have redesigned any workflow end to end, and high performers are roughly 5% of the population. High performers show 3x leadership engagement and outcome-based objectives.

Cisco AI Readiness Index 2025

active · scope: Annual global survey, around 8,000 respondents

Around 13% of organisations qualify as Pacesetters and the share is static year on year. 83% plan agentic deployment while foundations lag. Introduces the concept of AI Infrastructure Debt.

Anthropic Economic Index (January 2026)

active · scope: Analysis of anonymised usage across Claude.ai and the API

Five economic primitives for measuring AI's workplace impact. Reliability-adjusted productivity contribution falls from 1.8 to roughly 1.0 percentage points: real but more modest than claimed, and dependent on integration quality.

Stanford HAI AI Index 2026

active · scope: Annual index, multi-source

Generative AI reached 53% population adoption within three years, faster than the PC or the internet. Adoption pace varies sharply by geography and correlates with GDP.

DORA, State of DevOps (2014 to present)

active · scope: Longitudinal research programme, 36,000+ professionals over a decade

Four outcome metrics plus a capabilities model, published methodology, annual re-benchmarking, and tier classification turned a research programme into the industry's shared vocabulary.

The register is living. A monthly research review decides what counts as signal, and revisions to the instrument follow from it, never from the calendar alone. Every citation carries the study's scope, so no claim can be accused of overreach.

Changelog · Semantic versioning

Every change, published

Patch versions clarify wording, with no scoring change. Minor versions recalibrate anchors or weights, with a published adjustment. Major versions change questions: a new benchmark era, with a bridge analysis. When the instrument re-anchors, your stored answers can score differently under the new version. Your answers have not changed; the frontier has. We show you that drift directly.

v1.1 · expected Q4 2026

Agentic deployment anchors (R6, D5.2) and frontier-responsiveness recalibration (R5). An agentic pilot that scores near the top of R6 today is expected to score lower under v1.1, not because the answer changed, but because the frontier did. Next research review: 6 July 2026.

v1.0.0 · 2026-06-09

Founding version. Six readiness dimensions, the realisation block, behavioural anchors as the scoring backbone. Weights set at 60/40 readiness to realisation, reflecting the current evidence: realisation is the goal, but for the mid-market majority the score must remain diagnostic of what to fix.

The covenant

The data promise, in plain language

Part of the product, not the small print

Individual results are private. Always. No individual or identifiable organisational result is ever published, shared or sold.

Public

Aggregated, anonymised data is public: sector cuts, band distributions, dimension patterns, the perception gap, the readiness-to-realisation relationship.

Minimum cell

Minimum cell size of 25 respondents for any published cut. Below that, the cut does not publish, full stop.

Deletion

You can delete your data at any time, and deletion propagates to aggregates at the next publication cycle.

No trackers

No third-party advertising trackers anywhere in the product. It would be incoherent next to this covenant.

Statement of interest

Naming the conflict

Projject builds this instrument, and Projject sells services. Here is how we keep the two honest. The instrument does authority work, not vendor work. The methodology is public, the scoring is deterministic and published, and recommendations come from an evidence-mapped library. They cite research, not Projject offerings. The readout has one call to action: a no-cost conversation, clearly labelled.

The benchmark only has value if you trust it. And you will only trust it if it survives scrutiny from someone who never intends to buy anything. That is the standard we hold every release to.

AiR Advisory · Protocol summary · Versioned with the instrument

How verification works

Our consultants run a structured 60-minute evidence session with you. We check a defined sample of answers, weighted toward the behavioural anchors and the realisation block, against real artefacts: production systems shown live, measurement baselines, governance documents and deployment records.

Where the evidence does not support an answer, we correct it and recalculate the score. You keep the corrected result, a badge valid for 12 months, and a one-page note recording what was evidenced. Advisory scores form a separate stratum in the public dataset. The gap between self-reported and verified results becomes one of the most interesting findings in the annual report.

The public method · Instrument v1.0.0

AiR Bench

The method is public. The workings are yours to check.

Benchmark your business →