Trend watch

Our probabilities are judgments. These are measurements. Here we extend real trends forward to see when they would cross the thresholds that matter for the hypotheses, if they hold. A projection is not a prediction: trends bend and break, and the bands show how quickly the uncertainty grows.

How long a task AI can finish on its own

METR's time horizon: the length of task, measured in skilled-human time, that frontier models complete with 50% or 80% success, on software, ML and cyber tasks. Log scale, so a straight line means steady doubling. Dots are measured frontier models. The dashed line and bands are the projection.

What this means for our numbers

The 80% horizon reaching a month of work is a necessary condition for our strict AGI bar, not the bar itself: it measures reliability on software, ML and cyber tasks only, and METR says measurements above 16 hours are unreliable, so the crossing date is an extrapolation no current test can check. If the trend holds, that crossing lands around –. The breadth half of the bar, whether AI can do most kinds of real remote work, is tracked below. We put strict AGI at – by the end of 2030, more cautious than the straight line, for three reasons:

The lighter band on the chart is where a single new frontier model should land if the trend holds. If new models keep landing above that prediction band, our 2030 number should rise. If they land below it, it should fall.

How much AI research AI is already doing

Other trends we watch

Get Hidden AGI watch by email. The latest reading, what changed, and the news that matters. Free.
Subscribe