6 data sources · 24 measurements · Updated monthly
Are you ready for the Ai Revolution?
Compare your business against the latest frontier understanding of AI’s impact, and see exactly how to optimise for it.
Grounded in frontier research from
The case
AiR Bench measures what your organisation does with AI, not how it feels about it. We re-score you every month as the frontier moves, so your standing always stays current.
01
We ask behavioural questions with observable anchors. The benchmarks that last, like DORA, measure what teams do, not how confident they feel.
02
It is versioned like software. Every score is timestamped to an instrument version. When the frontier moves, your standing can drift even if nothing you do changes. We show you that drift.
03
We publish every question with its rationale and evidence. Weights, thresholds, the changelog, the conflicts of interest: all public. We earn trust by being more open about method than anyone else.
One instrument · Two axes · 24 scored measurements
D1
Whether AI ambition has been converted into named, owned, quantified opportunities.
D2
Leadership engagement, frontline involvement, and actual weekly use across the organisation.
D3
Whether the work AI will touch is understood, documented and being redesigned, not just augmented.
D4
Whether the data the use cases need is accessible, usable and governed.
D5
Whether use is governed by policy people actually know, with defined human review and explainability.
D6
Whether the organisation ships digital change, and how fast.
R · THE REALISATION AXIS
What no competitor measures
We look at production deployments, pilot-to-production conversion, measured value against baselines, where that value concentrates, and how fast the frontier reaches your workflows. Current research keeps finding the same gap between adoption and value. An instrument that cannot see that gap cannot claim accuracy.
The Curve · Updated monthly
See how AI adoption converts to value over time. Three lines track adoption, production and measured value, annotated with model releases, research findings and instrument versions. It is the visible heartbeat of the research layer.
Explore The Curve →The living research layer · The differentiator
Every month the research loop runs. New studies, model releases and market evidence enter the Evidence Register, and a standing review decides what counts as signal. Signal becomes the next instrument version. When the instrument re-anchors, we re-score your stored answers under it.
Your answers stay the same and the frontier moves. That gap is drift, and this is the only instrument that shows it to you.
Last review
2 June 2026
Next review
6 July 2026
v1.1 expected
Q4 2026
From the Evidence Register · Live feed
12 Jun
Moonshot's Kimi K2.7 Code extends the run of low-cost, open-weight, near-frontier coding/agentic models (alongside earlier DeepSeek V4), lowering the economic barrier to agentic build-out for cost-sensitive mid-market firms; sourced from release trackers rather than a single primary lab post, so treated as a developing signal. Informs R5 (frontier responsiveness) and R6 (agentic deployment cost).
08 Jun
Microsoft AI announced seven in-house multimodal models (incl. MAI-Thinking-1, ~35B active MoE, 256K context, 53% SWE-Bench Pro) and 'Frontier Tuning' — reinforcement learning on a customer's own workflow traces inside their environment. Signals a shift toward customer-data-as-moat and tunable near-frontier models; informs R5 (frontier responsiveness) and D4 (Data Foundations as the input to tuned models). Primary vendor source but partly product marketing, so monitoring independent benchmarks.
02 Jun
Establishes a voluntary framework for federal early access (up to 30 days) to 'covered frontier models' and a classified NSA/CISA/Treasury benchmarking process for cyber capability, explicitly disclaiming any mandatory licensing or preclearance; deliverables due 1 Aug 2026 (corroborated by Wiley law alert and multiple outlets). Sets the near-term US governance posture against which UK mid-market buyers judge vendor risk. Informs D5 (Governance & Trust) and R5 (frontier responsiveness).
02 Jun
Reliability-adjusted productivity contribution revised upward for integrated deployments. Informs the R5 frontier-responsiveness anchors for v1.1.
02 Jun
Watching. Explainability obligations may re-anchor D5.3 — what counts as 'could explain to a regulator, today' is about to get a legal definition.
01 Jun
Copilot shifted from flat per-seat pricing to token/credit metering across all tiers, with heavy agentic users reportedly paying $60–100+/month versus the $10 sticker (verified against GitHub's pricing page in trade coverage). Concrete evidence that agentic deployment carries variable, usage-scaled cost that reshapes pilot-to-production economics. Informs R6 (agentic deployment), R2 (pilot-to-production conversion) and D3 (Process Readiness).
Monthly research review. Minor versions roughly twice a year, driven by the register, never by the calendar. · Every change published in the changelog
Two ways in
10 to 12 minutes · No cost
Benchmark your business against the frontier. You get your band, the AiR Grid, seven dimension percentiles, the perception gap, and a clear Now, Next, Later.
Benchmark your business →The benchmark plus one session
We check your answers against real artefacts with our consultants. You get a badge valid for 12 months, and a separate stratum in the public dataset.
Talk to us →The covenant
We never publish, share or sell an individual or identifiable organisational result. We do publish aggregated, anonymised data as a public resource: sector cuts, band distributions, the perception gap, and the readiness-to-realisation relationship.
You can delete your data at any time. Deletion carries through to the aggregates at the next publication cycle.
Part of the product, not the small print. Read it in full.
6 data sources · 24 measurements · Updated monthly
Twelve minutes now. A score you can defend in the boardroom.
See your headline before any sign-up · Save and resume · v1.0.0