← The Score CardClaim 05 of 06

Claim 05 of 06

Genius for everyone, on demand.

Progress is measured by a constructed index: frontier capability against human expert benchmarks, multiplied by practical daily reach. The index has risen from 0.5 to 19.4 since 2020, against a completion threshold of 100.

Progress since 2020
0%
Share of the distance from the 2020 value of the primary measure to its completion threshold.
Estimated completion
2034
expert parity combined with majority reach
Trend
Rising fastest on the board
Direction of the primary measure over the six years since the baseline.

Primary measure

Basis of the progress score

The single objective measure assigned to this claim, tracked from the 2020 baseline to the current value. This measure alone determines the progress percentage.

capability parity multiplied by practical daily reach

Expert-Parity Access Index

A constructed composite, not a published statistic. The capability component measures frontier model performance against human expert benchmarks and currently stands near 62 percent of expert parity. The reach component measures practical daily availability by internet access, price, and language coverage, and currently stands near 30 percent. The index is their product. Completion requires expert parity available to the entire population.

2020 baseline0.5%
Current value19.4%
Completion threshold100%
0%25%50%75%100%
Progress = (19.4% now − 0.5% in 2020) ÷ (100% at completion − 0.5% in 2020) = 19%
0.5%20200.8%20211.5%20224.0%20238.0%202413.5%202519.4%2026
Annual values of the primary measure, 2020 to 2026. The progress score above is derived from the 2020 and 2026 values only.

Supporting measures

Related indicators

Reported alongside the primary measure to show whether it is consistent with the wider evidence. These do not enter the progress score. Bars show relative movement only.

Humanity's Last Exam, best model
near 0%55.5%
Claude Fable 5 scored 55.5 percent in August 2026, Claude Opus 5 54.9 percent, and GPT-5.6 Sol 49.5 percent. Human domain experts average approximately 90 percent. Frontier scores rose roughly 30 percentage points over twelve months.
GPQA Diamond
about 30%above 90%
Graduate-level science questions. Frontier models score above 90 percent. Domain PhDs answering within their own field score approximately 65 to 75 percent. Expert parity on this benchmark has been passed.
Practical reach
near zeroabout 30%
ChatGPT crossed one billion monthly active users in June 2026. Microsoft estimates approximately one in six people worldwide used a generative AI tool in the second half of 2025. Global internet penetration near 68 percent is the current ceiling.

Recent developments

Events since 2025

Developments relevant to this claim, and their effect on the primary measure where there is one.

Aug 2026
The Humanity's Last Exam leaderboard stood at 55.5 percent (Claude Fable 5), 54.9 percent (Claude Opus 5), and 49.5 percent (GPT-5.6 Sol). Twelve months earlier the leading score was in the low twenties.
Jun 2026
ChatGPT crossed one billion monthly active users, the fastest application in history to reach that scale.
Feb 2026
OpenAI confirmed 900 million weekly active users, up from 400 million a year earlier. Frontier-class inference cost has fallen more than two orders of magnitude since 2022.
Jan 2026
Humanity's Last Exam was published in Nature. It comprises approximately 2,500 questions contributed by roughly 1,000 experts across 500 institutions in 50 countries, designed to resist saturation.
Feb 2026
The Metaculus community median for AGI stood at November 2033, with 25 percent probability by 2029. In 2020 the same community median was approximately fifty years out.

Completion estimate

Basis for 2034

Starting point, adjustments applied, and the resulting estimate.

This measure is a constructed composite rather than a published statistic. Its two components behave differently and are reported separately.

The capability component stands near 62 percent of expert parity. Frontier models score above 90 percent on GPQA Diamond, where domain PhDs score approximately 65 to 75 percent within their own field. On Humanity's Last Exam the leading model scored 55.5 percent in August 2026 against approximately 90 percent for human domain experts, having gained roughly 30 percentage points in twelve months.

The reach component stands near 30 percent, derived from global internet penetration near 68 percent and the share of that population with affordable frontier-tier access in a supported language.

Capability is not expected to remain the binding constraint. The Metaculus community median for AGI is November 2033, and the remaining capability gap is primarily long-horizon reliability rather than domain knowledge.

The date estimate is therefore governed by reach. Internet penetration is projected near 90 percent in the late 2030s and inference cost has fallen more than two orders of magnitude since 2022. The estimate used here is 2034 for expert parity combined with majority reach. Coverage approaching the full population falls in the 2040s.

Assessment

Arguments for and against

The strongest available case on each side, stated without weighting.

Arguments for

Reasons to expect completion

  • Benchmark improvement has been rapid and sustained.Frontier scores on Humanity's Last Exam rose approximately 30 percentage points in twelve months, on a benchmark constructed specifically to resist saturation.
  • Expert parity has already been exceeded in several domains.GPQA Diamond, competition mathematics, most competitive programming, and protein structure prediction are all above expert human performance.
  • Marginal distribution cost is near zero.Serving an additional user of a trained model is inexpensive relative to training. ChatGPT alone reached approximately one billion monthly active users within four years of launch.
  • Language coverage has broadened.Frontier models perform competently across dozens of languages, including several with limited digital training corpora.
  • Forecast timelines have shortened substantially.The Metaculus community median for AGI moved from approximately fifty years out in 2020 to November 2033 as of February 2026.
Arguments against

Reasons to expect failure

  • Reach is the binding constraint and improves slowly.Approximately one third of world population lacks reliable internet access, and connectivity growth is limited by infrastructure investment rather than by software.
  • Availability does not imply use or benefit.Access statistics measure account activity, not the application of expert-level capability to consequential tasks.
  • Reliability remains unresolved.Models produce confident errors, and the domains where expert capability is most valuable are those in which a non-expert user cannot verify the output.
  • Benchmark saturation limits measurement.Humanity's Last Exam was constructed to be difficult and is being solved rapidly, which constrains how much the score indicates about general capability.
  • Frontier capability is concentrated.A small number of laboratories control frontier training. Population-scale access depends on their pricing and deployment decisions.
  • Cognitive capability has not historically been the constraint on outcomes.Coordination, institutional capacity, and implementation rather than availability of expertise have limited progress on most large-scale problems.

Scenarios

Most likely paths

Two sequences: the most probable route to completion and the most probable route to failure. These are structured projections, not forecasts, and neither is assigned a probability.

Completion scenario

Most likely path to completion

The most probable path requires continuation of current trends rather than a discontinuity.

Capability reaches expert parity on general benchmarks around 2029 at the current rate of improvement. Long-horizon agentic reliability, the principal remaining gap, is addressed through training and evaluation methods over a similar period.

Inference cost continues to fall. By approximately 2031 frontier-class capability runs locally on consumer hardware, removing per-query cost and network dependency as barriers.

Reach then expands faster than capability. The marginal user is added at negligible cost, and distribution proceeds through mobile device penetration, which substantially exceeds fixed broadband penetration in low-income regions.

The largest measurable effects appear in education and in professional services in regions with low practitioner density, where the counterfactual is absence of expertise rather than substitution for an existing professional.

Majority reach combined with expert parity is reached near 2034. Coverage approaching the full population follows in the 2040s as connectivity infrastructure completes.

Failure scenario

Most likely path to failure

The most probable failure is in the reach component rather than the capability component.

Capability continues to improve on the observed trajectory. Expert parity on general benchmarks is achieved on approximately the projected schedule.

Access stratifies by price tier. Free and low-cost tiers remain substantially less capable than frontier deployments, and the performance gap between tiers widens rather than narrows, because competitive pressure at the frontier is limited to a small number of providers.

Connectivity growth in low-income regions slows as the remaining unconnected population becomes more expensive to reach, being predominantly rural, low-density, and low-income.

Among users with access, the dominant application is consumption rather than production, following the pattern of prior consumer information technologies. Measured effects on educational attainment and output remain smaller than capability improvements alone would suggest.

Available cognitive capability increases by orders of magnitude while measured outcomes on the other five claims improve at rates closer to their historical trends.

The other claims

Open another claim page

Each claim carries its own colour, applied as the background of its page.

← Previous claim
Work becomes a choice.
← Back to the full score card
Jesse Walker
Jesse Walker
Jesse Walker is a philosopher, a meditation teacher, a business founder and a father. He is optimistic about humanity’s ability to shape AI into a force for global good.