Claim 05 of 06
Progress is measured by a constructed index: frontier capability against human expert benchmarks, multiplied by practical daily reach. The index has risen from 0.5 to 19.4 since 2020, against a completion threshold of 100.
The single objective measure assigned to this claim, tracked from the 2020 baseline to the current value. This measure alone determines the progress percentage.
A constructed composite, not a published statistic. The capability component measures frontier model performance against human expert benchmarks and currently stands near 62 percent of expert parity. The reach component measures practical daily availability by internet access, price, and language coverage, and currently stands near 30 percent. The index is their product. Completion requires expert parity available to the entire population.
Reported alongside the primary measure to show whether it is consistent with the wider evidence. These do not enter the progress score. Bars show relative movement only.
Developments relevant to this claim, and their effect on the primary measure where there is one.
Starting point, adjustments applied, and the resulting estimate.
This measure is a constructed composite rather than a published statistic. Its two components behave differently and are reported separately.
The capability component stands near 62 percent of expert parity. Frontier models score above 90 percent on GPQA Diamond, where domain PhDs score approximately 65 to 75 percent within their own field. On Humanity's Last Exam the leading model scored 55.5 percent in August 2026 against approximately 90 percent for human domain experts, having gained roughly 30 percentage points in twelve months.
The reach component stands near 30 percent, derived from global internet penetration near 68 percent and the share of that population with affordable frontier-tier access in a supported language.
Capability is not expected to remain the binding constraint. The Metaculus community median for AGI is November 2033, and the remaining capability gap is primarily long-horizon reliability rather than domain knowledge.
The date estimate is therefore governed by reach. Internet penetration is projected near 90 percent in the late 2030s and inference cost has fallen more than two orders of magnitude since 2022. The estimate used here is 2034 for expert parity combined with majority reach. Coverage approaching the full population falls in the 2040s.
The strongest available case on each side, stated without weighting.
Two sequences: the most probable route to completion and the most probable route to failure. These are structured projections, not forecasts, and neither is assigned a probability.
The most probable path requires continuation of current trends rather than a discontinuity.
Capability reaches expert parity on general benchmarks around 2029 at the current rate of improvement. Long-horizon agentic reliability, the principal remaining gap, is addressed through training and evaluation methods over a similar period.
Inference cost continues to fall. By approximately 2031 frontier-class capability runs locally on consumer hardware, removing per-query cost and network dependency as barriers.
Reach then expands faster than capability. The marginal user is added at negligible cost, and distribution proceeds through mobile device penetration, which substantially exceeds fixed broadband penetration in low-income regions.
The largest measurable effects appear in education and in professional services in regions with low practitioner density, where the counterfactual is absence of expertise rather than substitution for an existing professional.
Majority reach combined with expert parity is reached near 2034. Coverage approaching the full population follows in the 2040s as connectivity infrastructure completes.
The most probable failure is in the reach component rather than the capability component.
Capability continues to improve on the observed trajectory. Expert parity on general benchmarks is achieved on approximately the projected schedule.
Access stratifies by price tier. Free and low-cost tiers remain substantially less capable than frontier deployments, and the performance gap between tiers widens rather than narrows, because competitive pressure at the frontier is limited to a small number of providers.
Connectivity growth in low-income regions slows as the remaining unconnected population becomes more expensive to reach, being predominantly rural, low-density, and low-income.
Among users with access, the dominant application is consumption rather than production, following the pattern of prior consumer information technologies. Measured effects on educational attainment and output remain smaller than capability improvements alone would suggest.
Available cognitive capability increases by orders of magnitude while measured outcomes on the other five claims improve at rates closer to their historical trends.
Each claim carries its own colour, applied as the background of its page.
