AI capability is rising fast — and very unevenly.
Stanford’s 2026 AI Index measures a frontier moving faster than any benchmark can absorb — and failing at tasks nobody thought to test. The practical consequence for companies is that capability has to be measured on their own work rather than read off a leaderboard.
Les Insights sont publiés en anglais.
Stanford’s Institute for Human-Centered Artificial Intelligence published the 2026 edition of its AI Index in April. One pair of figures has travelled further than the rest. On Humanity’s Last Exam — a benchmark built from questions that specialists in dozens of fields were asked to make as hard as they could — the leading model scored 8.8% in the 2025 edition. In the 2026 edition, the leading score is 38.3%.
That figure was already behind on publication: reviewing the Index on 13 April 2026, IEEE Spectrum noted that the best-scoring models then available, Claude Opus 4.6 and Gemini 3.1 Pro, were above 50%.
The same report records that leading models cannot reliably read an analogue clock. Both findings measure the same systems in the same year; for a company deciding what to hand over to software, the second is the more useful.
A benchmark designed to resist AI for years lost most of its difficulty in one
Humanity’s Last Exam was released in January 2025 precisely because earlier tests had stopped discriminating: models were already clearing 90% on standard examinations such as MMLU. Assembled by the Center for AI Safety and Scale AI from contributions by close to a thousand subject-matter experts, it collects 2,500 questions with verifiable answers that cannot be found on the open internet — “the final closed-ended academic benchmark of its kind”, in its authors’ words.
One year moved it by thirty percentage points, and agentic performance moved comparably. On OSWorld, which asks models to carry out real tasks on a computer, accuracy rose from roughly 12% to 66.3% — within six percentage points of human performance. Underneath sits an industrial expansion: the Index records world AI compute capacity growing more than threefold every year since 2022, and thirtyfold since 2021.
The same systems fail a task most ten-year-olds complete
In 2025 Google’s Gemini Deep Think reached a gold-medal score at the International Mathematical Olympiad. On ClockBench, which asks models to read analogue clock faces, the best model in the Index scored 50.6%, against 90.1% for people. ClockBench is not an exotic construction: 720 questions over 180 clocks, covering reading the time, adding intervals, rotating the hands and shifting time zones. In its author’s words, the task “sets a high bar for doing reasoning within the visual space (as opposed to text space)”.
The unevenness persists inside professional work, in a narrower band: the Index reports model performance ranging from 60% to 90% across evaluations in tax, mortgage processing, corporate finance and legal reasoning. A thirty-point spread within knowledge work is not a curiosity; it is the planning problem. Whatever its cause, the pattern is stable — how hard a task feels to a person predicts very little about how hard it is for a model.
Scores are becoming less informative as they are quoted more often
The Index is candid about its own instruments. Widely used benchmarks contain invalid questions at rates from 2% on MMLU Math to 42% on GSM8K. Standing on public arena leaderboards, it notes, “may partly reflect adaptation to the platform rather than general capability”. And, in the Index’s own words, evaluations “intended to be challenging for years are saturated in months”.
None of this makes benchmarks worthless: they remain the only comparable public record of a fast-moving field. It does make them the wrong document to buy from. A leaderboard reports what a model did on a public test set at one moment, with someone else’s data, formats and tolerance for error.
Two controlled experiments locate both the gain and the loss
The most useful evidence for managers is older than the benchmark race. In a pre-registered experiment with Boston Consulting Group, 758 consultants were randomly assigned to work with or without GPT-4. Across 18 realistic consulting tasks inside the model’s capability, those using AI completed 12.2% more tasks, worked 25.1% faster and produced work rated more than 40% higher in quality; consultants below the average performance threshold improved by 43%, those above it by 17%. On one task deliberately chosen to sit outside that capability, consultants using AI were 19 percentage points less likely to reach the correct solution. The study, which introduced the phrase “jagged technological frontier”, appeared as HBS working paper 24-013 in 2023 and in Organization Science in March 2026.
A second experiment addresses what managers rely on more than they admit: self-assessment. In a randomised trial run by METR in early 2025, 16 experienced open-source developers worked through 246 real issues in their own repositories, with AI assistance randomly permitted or withheld. They expected it to make them 24% faster. It made them 19% slower — and afterwards they still believed it had made them roughly 20% faster. METR is explicit that the result describes its own setting rather than software development at large — and that caution is the lesson: evidence about a company’s tasks has to be produced at company level.
Swiss firms have settled the adoption question and left the measurement question open
Adoption is no longer the open question. On 15 July 2026, SECO’s SME portal relayed an EY survey of 604 respondents in Swiss companies published on 27 May: 89% use AI in their daily work, 55% say their company deploys it deliberately in one or more business areas, 31% are still at pilot stage and 9% report a business model already changed by it. The obstacles named are organisational rather than technical: data quality and silos (20%), security and data protection (19%), scarce qualified staff (18%).
The same survey shows where Swiss constraints bind: data sovereignty is business-critical for 51% of respondents, who rate the importance of Swiss and EU data-protection standards at 8.7 out of 10, and 56% call for a sovereign Swiss AI infrastructure. The market they face is concentrated: according to Epoch AI, the United States holds about three quarters of global GPU-cluster performance and China about 15%, leaving roughly a tenth for everywhere else, Europe included.
Switzerland’s institutional answer is instructive precisely because it is not a leaderboard entry. Apertus, released on 2 September 2025 by EPFL, ETH Zurich and the Swiss National Supercomputing Centre, was trained on the Alps machine in Lugano across more than a thousand languages, including Swiss German and Romansh, and published under a permissive licence with its weights, training data and methods documented; it runs on Hugging Face and on Swisscom’s Swiss-hosted platform. Its proposition is auditability and location rather than a top score. Where the binding constraint is jurisdiction over customer data, that can be the right trade — provided the model clears the threshold the task requires. Establishing that is a measurement, and only the firm can take it.
What a management team can decide this month
None of this requires new technology, only a small body of internal evidence, produced once and maintained.
- Build an evaluation set from your own work: twenty to thirty real tasks — quotations, technical answers, supplier correspondence, order classification — each with an output your team accepts as correct.
- Measure against a control: have the same work done with and without the tool, recording time and quality separately. Reported speed-ups are unreliable, as METR’s developers showed on themselves.
- Set the passing threshold by the cost of an error: a first draft of marketing copy and a customer-specific price calculation do not need the same score.
- Re-run the set whenever the model changes. Providers update continuously, and movement on public benchmarks does not transfer automatically to your tasks, in either direction.
- Deploy where the measured gain is largest rather than where enthusiasm is: in the BCG experiment the largest improvement, 43%, went to the group performing below average.
Two developments are worth following: whether occupational, task-based evaluations displace academic benchmarks as the reference buyers use, and the next edition of the AI Index in spring 2027, where a change in the shape of the frontier would first be visible.
Capability is likely to keep rising. Whether its unevenness recedes is not settled by present measurements: jaggedness has appeared at every level of capability recorded so far. Companies that know the shape of their own edge — which tasks their systems clear and which they do not — can keep deciding what to automate on evidence rather than on a headline figure already out of date the day it was published.
