The public method · Version 1.0.0

A transparent approach

We publish every question, its rationale and the evidence behind it. You see the scoring model, the weights, the thresholds and every change. Trust in a benchmark comes from showing how the number is made. We do.

Scoring model

How the number is made

Every scored question maps to 0 to 4 points. Likert items run from strongly disagree (0) to strongly agree (4). Behavioural items run from anchor (a), 0, to anchor (e), 4. We store your raw answers, never just the computed scores. That is what lets us recalculate your score when the assessment changes, without asking you to sit it again.

Readiness: 18 questions across six dimensions, maximum 72 points, normalised to 0 to 100. Each dimension produces a 0 to 100 sub-score from its three questions. Realisation: 6 questions, maximum 24 points, normalised to 0 to 100.

The headline AiR Bench is a weighted blend: 60% readiness, 40% realisation. The split reflects today's evidence. Realisation is the goal. But for the mid-market majority, the score has to stay diagnostic of what to fix, and readiness is where the fixes live. As the population matures, we expect the weighting to shift towards realisation. That shift will be a published finding in its own right.

The AiR Grid splits at 50 on both axes. That is the midpoint of the scale, not a band boundary, so the grid and the bands stay independent readings.

Emerging

0 to 39

Foundations need to be laid before a meaningful pilot is viable. Strategic clarity and leadership alignment come first.

Developing

40 to 59

Some foundations in place with one or two critical gaps. The dimension detail shows exactly which.

Ready

60 to 79

Strong foundations across most dimensions. Positioned to move into structured production work and start generating measured return.

Leading

80 to 100

The conditions for AI to compound value, and evidence that it is doing so. The question is sequencing, not whether to act.

We calibrate the bands so the top band stays genuinely rare. The research is consistent: real high performers are 5 to 13% of the population. A benchmark where a third of respondents are Leading is flattery, not measurement.

Foundations

Low readiness, low realisation

Early. The work is to define one quantified opportunity and build the conditions around it. Honest, common, fixable.

Prepared, Not Proving

High readiness, low realisation

The GenAI Divide quadrant, and the most populated in the research: conditions exist, value does not. The gap is execution design: workflow redesign, production discipline, measurement.

Running Hot

Low readiness, high realisation

Value is being created ahead of foundations. Often a few brilliant individuals or one heroic team. The value is real and fragile: governance, data and process debt will tax it.

Compounding

High readiness, high realisation

Conditions and evidence. The work is sequencing and protecting focus. This is where compounding competitive advantage lives.

The questions · 24 scored · Why we ask each one

Every question, with its evidence

Behavioural anchors give us five clear states to score against, so answers are hard to game and easy to verify. We use Likert agreement scales only when the question is about attitude, not behaviour. Questions marked ✳ form the short-form subset.

D1. Strategic Clarity

Whether AI ambition has been converted into named, owned, quantified opportunities.

D1.1 · Agreement (Likert 0 to 4)

We have named specific AI opportunities, each tied to a measurable business outcome with a single accountable owner.

Why we ask · Named, owned, quantified opportunities are the single clearest marker of strategy converted into work. McKinsey's high performers are distinguished by outcome-based objectives, not by ambition statements.

Evidence: McKinsey, The State of AI (2025 edition)

D1.3 · Agreement (Likert 0 to 4)

Senior leadership can explain why AI matters to this organisation in the language of our strategy, not in general terms.

Why we ask · When leaders can only describe AI in general terms, prioritisation defaults to whoever shouts loudest. Strategy-specific language is a proxy for genuine strategic integration.

Evidence: McKinsey, The State of AI (2025 edition)

D2. People and Change

Leadership engagement, frontline involvement, and actual weekly use across the organisation.

D2.1 · Agreement (Likert 0 to 4)

Our leadership team actively champions AI adoption and dedicates meaningful time to it.

Why we ask · High performers show roughly three times the leadership engagement of everyone else. It is the strongest single differentiator in the McKinsey data.

Evidence: McKinsey, The State of AI (2025 edition)

D2.3 · Agreement (Likert 0 to 4)

The teams who will use AI day to day have been involved in shaping how it is deployed.

Why we ask · The MIT research locates the failure of most pilots in a learning gap, not a technology gap. Frontline involvement in deployment design is the observable counter to that gap.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

D3. Process Readiness

Whether the work AI will touch is understood, documented and being redesigned, not just augmented.

D3.1 · Agreement (Likert 0 to 4)

The processes we want AI to improve are documented, well understood and currently measurable.

Why we ask · You cannot redesign what you cannot describe, and you cannot prove improvement without a baseline. Documented, measurable processes are the precondition for measured value.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

D3.2 · Behavioural anchors (0 to 4)

Has any workflow been redesigned end to end around AI, rather than AI being added to the existing way of working?

Why we ask · End-to-end workflow redesign is the highest-signal single behaviour in the current research: only around 21% of organisations have done it once, and it is where AI's value concentrates.

Evidence: McKinsey, The State of AI (2025 edition)

D3.3 · Agreement (Likert 0 to 4)

For our target use cases, we know which steps AI will replace, which it will augment, and which it will leave unchanged.

Why we ask · Replace, augment, leave alone: the three-way split is the working vocabulary of process redesign. Knowing it per use case marks the difference between a plan and a hope.

Evidence: Anthropic Economic Index (January 2026)

D4. Data Foundations

Whether the data the use cases need is accessible, usable and governed.

D4.1 · Agreement (Likert 0 to 4)

We have access to the data our priority use cases require, and it is in a usable, organised state.

Why we ask · Anchored to priority use cases deliberately: data readiness in the abstract is unanswerable, data readiness for the three things you intend to do is not.

Evidence: Cisco AI Readiness Index 2025

D4.2 · Behavioural anchors (0 to 4)

If you adopted a new AI tool tomorrow, how quickly could it be connected to the business data it needs, with appropriate controls?

Why we ask · Time-to-connect is the practical measure of data readiness, and the repeated, governed pattern at the top anchor is what Cisco calls the antidote to AI Infrastructure Debt.

Evidence: Cisco AI Readiness Index 2025

D4.3 · Agreement (Likert 0 to 4)

Security, privacy and quality constraints on our data are understood and managed; they shape our AI work rather than blocking it.

Why we ask · Constraints that are understood become design inputs. Constraints that are vague become vetoes. The difference shows up directly in pilot-to-production conversion.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025 · Cisco AI Readiness Index 2025

D5. Governance and Trust

Whether use is governed by policy people actually know, with defined human review and explainability.

D5.1 · Agreement (Likert 0 to 4)

We have a clear policy on acceptable AI use, and the people doing the work actually know what it says.

Why we ask · The second clause is the question. Most organisations have a policy; far fewer have a policy anyone could recite. Governance that exists only as a document governs nothing.

Evidence: Cisco AI Readiness Index 2025

D5.2 · Behavioural anchors (0 to 4)

Where AI contributes to work that matters, is there a defined human review point?

Why we ask · As deployment becomes agentic, the defined human review point is the control that matters most. The top anchor describes review as a living system, which is what increasing autonomy requires.

Evidence: Cisco AI Readiness Index 2025 · Anthropic Economic Index (January 2026)

D5.3 · Agreement (Likert 0 to 4)

We could explain to a customer or a regulator, today, where and how AI is used in our business.

Why we ask · Explainability on demand is the practical test of governance maturity. 'Today' is in the question because a capability that needs six weeks of preparation is not a capability.

Evidence: Cisco AI Readiness Index 2025

D6. Execution Track Record

Whether the organisation ships digital change, and how fast.

D6.1 · Agreement (Likert 0 to 4)

We have launched new digital tools or ways of working successfully in the past 24 months.

Why we ask · Organisations that ship digital change ship AI change. A 24-month window keeps the evidence recent enough to mean something.

Evidence: DORA, State of DevOps (2014 to present)

D6.2 · Behavioural anchors (0 to 4)

From decision to working in production, how long does digital change typically take here?

Why we ask · Lead time from decision to production is DORA's most durable metric, carried into the AI context. Speed of change is itself a readiness asset, because the frontier will move again.

Evidence: DORA, State of DevOps (2014 to present)

D6.3 · Agreement (Likert 0 to 4)

When an initiative is stalling, we find out quickly and act, rather than letting it drift.

Why we ask · The 95% pilot failure figure is mostly a story of drift. Fast detection and decisive action on stalling work is the organisational muscle that prevents it.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R. Value Realisation

The realisation axis

Production deployments, pilot-to-production conversion, measured value, where value concentrates, frontier responsiveness.

R1 · Behavioural anchors (0 to 4)

How many AI use cases are in production use today, meaning relied upon in real work, not piloted?

Why we ask · Production use, relied upon in real work, is the bar. The definition is in the question because the adoption-to-value gap lives precisely in the space between a pilot and a dependency.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R2 · Behavioural anchors (0 to 4)

Of the AI initiatives you have started in the last 18 months, what share reached production?

Why we ask · Pilot-to-production conversion is the rate that separates the 5% from the 95% in the MIT data. It measures execution design, not enthusiasm.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R4 · Behavioural anchors (0 to 4)

Where is the value you have seen concentrated?

Why we ask · Value climbs a ladder: individual, team, function, business model. The MIT finding that budgets concentrate where ROI does not makes the location of value a first-class measurement.

Evidence: MIT NANDA, The GenAI Divide: State of AI in Business 2025

R5 · Behavioural anchors (0 to 4)

When a significant new model capability is released, how long before it reaches your workflows?

Why we ask · The frontier moves twice a year. Frontier responsiveness measures whether your organisation compounds those moves or pays for them later. This is the question most exposed to re-anchoring as the benchmark evolves.

Evidence: Anthropic Economic Index (January 2026) · Stanford HAI AI Index 2026

R6 · Behavioural anchors (0 to 4)

Have you deployed any agentic systems, meaning AI carrying out multi-step work with defined autonomy and oversight?

Why we ask · 83% of organisations plan agentic deployment while foundations lag. The definition (defined autonomy and oversight) is in the question so that an unsupervised script cannot score as an agentic system.

Evidence: Cisco AI Readiness Index 2025

Benchmarks and percentiles

Honest about sample size

Every readout shows your percentile against the full UK dataset and against your sector, on the headline score and each dimension. We publish sector percentiles only when the sector cell reaches the minimum sample of 25 respondents. Below that threshold, the readout says so plainly.

Until the dataset reaches critical mass, we seed benchmarks from the published research distributions (McKinsey, Cisco) and label them research-calibrated wherever they appear. The calibration targets are clear: high performers held to 5 to 13% of the population, the median in the Developing band, and realisation trailing readiness (the GenAI Divide). We will announce the switch to live data in the changelog. The assessment never pretends seeded data is its own.

Behavioural anchors are the backbone of scoring, verified advisory scores act as a calibration reference, and we flag repeated answer patterns internally. That is how we discourage gaming. We never call anyone out in public.

The evidence base · Founding sources

What the assessment is built on

MIT NANDA, The GenAI Divide: State of AI in Business 2025

active · scope: 150 interviews, 350 surveyed employees, 300 analysed deployments

95% of enterprise AI pilots delivered no measurable P&L impact. The barrier is a learning gap, not the technology. Vendor partnerships succeed roughly twice as often as internal builds. Budgets concentrate in sales and marketing while ROI concentrates in operations.

McKinsey, The State of AI (2025 edition)

active · scope: Global survey of 1,491 participants across 101 countries

88% of organisations use AI in at least one function, but only around 21% have redesigned any workflow end to end, and high performers are roughly 5% of the population. High performers show 3x leadership engagement and outcome-based objectives.

Cisco AI Readiness Index 2025

active · scope: Annual global survey, around 8,000 respondents

Around 13% of organisations qualify as Pacesetters and the share is static year on year. 83% plan agentic deployment while foundations lag. Introduces the concept of AI Infrastructure Debt.

Anthropic Economic Index (January 2026)

active · scope: Analysis of anonymised usage across Claude.ai and the API

Five economic primitives for measuring AI's workplace impact. Reliability-adjusted productivity contribution falls from 1.8 to roughly 1.0 percentage points: real but more modest than claimed, and dependent on integration quality.

Stanford HAI AI Index 2026

active · scope: Annual index, multi-source

Generative AI reached 53% population adoption within three years, faster than the PC or the internet. Adoption pace varies sharply by geography and correlates with GDP.

DORA, State of DevOps (2014 to present)

active · scope: Longitudinal research programme, 36,000+ professionals over a decade

Four outcome metrics plus a capabilities model, published methodology, annual re-benchmarking, and tier classification turned a research programme into the industry's shared vocabulary.

The evidence base is updated every month. A research review decides what findings are strong enough to change the assessment, and revisions follow from that, not from the calendar alone. Every citation carries the study's scope, so claims cannot overreach.

Changelog · Semantic versioning

Every change, published

Patch versions clarify wording, with no scoring change. Minor versions adjust anchors or weights, with a published note. Major versions change questions: a new phase for the benchmark, with a bridge analysis. When the assessment is updated, your stored answers can score differently under the new version. Your answers have not changed; the criteria have. We show you that movement directly.

v1.1 · expected Q4 2026

Agentic deployment anchors (R6, D5.2) and frontier-responsiveness recalibration (R5). An agentic pilot that scores near the top of R6 today is expected to score lower under v1.1, not because the answer changed, but because the frontier did. Next research review: Research review overdue.

v1.0.0 · 2026-06-09

Founding version. Six readiness dimensions, the realisation block, behavioural anchors as the scoring backbone. Weights set at 60/40 readiness to realisation, reflecting the current evidence: realisation is the goal, but for the mid-market majority the score must remain diagnostic of what to fix.

The covenant

The data promise, in plain language

Part of the product, not the small print

Individual results are private. Always. We never publish, share or sell results that identify you or your organisation.

Public

Aggregated, anonymised data is public: sector comparisons, score distributions, dimension patterns and the relationship between readiness and real value.

Minimum cell

We only publish a cut when it reaches 25 respondents. Below that, the cut does not publish, full stop.

Deletion

You can delete your data at any time. Deletion flows through to the aggregates at the next publication cycle.

No trackers

No third-party advertising trackers anywhere in the product. It would be incoherent next to this covenant.

Statement of interest

Naming the conflict

Projject builds this assessment, and Projject sells services. Here is how we keep the two honest. The assessment is independent work, not a sales pitch. The methodology is public, the scoring is deterministic and published, and recommendations come from an evidence-mapped library. They cite research, not Projject offerings. The readout has one call to action: a no-cost conversation, clearly labelled.

The benchmark only has value if you trust it. And you will only trust it if it survives scrutiny from someone who never intends to buy anything. That is the standard we hold every release to.

AiR Advisory · Protocol summary · Versioned with the assessment

How verification works

Our consultants run a structured 60-minute evidence session with you. We check a defined sample of answers, weighted toward the behavioural anchors and the realisation questions, against real artefacts: production systems shown live, measurement baselines, governance documents and deployment records.

Where the evidence does not support an answer, we correct it and recalculate the score. You keep the corrected result, a badge valid for 12 months, and a one-page note recording what was evidenced. Advisory scores sit in a separate verified set in the public dataset. The gap between self-reported and verified results becomes one of the most interesting findings in the annual report.

The public method · Version 1.0.0

AiR Bench

The method is public. You can inspect every step.

Benchmark your business →