Trend watch
Our probabilities are judgments. These are measurements. Here we extend real trends forward to see when they would cross the thresholds that matter for the hypotheses, if they hold. A projection is not a prediction: trends bend and break, and the bands show how quickly the uncertainty grows.
How long a task AI can finish on its own
METR's time horizon: the length of task, measured in skilled-human time, that frontier models complete with 50% or 80% success, on software, ML and cyber tasks. Log scale, so a straight line means steady doubling. Dots are measured frontier models. The dashed line and bands are the projection.
What this means for our numbers
The 80% horizon reaching a month of work is a necessary condition for our strict AGI bar, not the bar itself: it measures reliability on software, ML and cyber tasks only, and METR says measurements above 16 hours are unreliable, so the crossing date is an extrapolation no current test can check. If the trend holds, that crossing lands around –. The breadth half of the bar, whether AI can do most kinds of real remote work, is tracked below. We put strict AGI at – by the end of 2030, more cautious than the straight line, for three reasons:
- Benchmark tasks are cleaner than real work.
- METR's suite can't measure horizons above about 16 hours, and the work-month threshold is ten times longer.
- Trends like this can bend with compute, energy, data or safety pauses. OpenAI paused RL training for two weeks (Fortune dates it to late July 2026; OpenAI announced it in mid-August) and paused training again on Sept 25, 2026 (Fortune).
The lighter band on the chart is where a single new frontier model should land if the trend holds. If new models keep landing above that prediction band, our 2030 number should rise. If they land below it, it should fall.
How much real remote work AI can already do
The Remote Labor Index (RLI) is the closest public measure of the breadth half of our bar: the share of real freelance projects an AI delivers at a quality a client would accept, judged by people. Our bar asks for at least 80% of remote professional tasks; the RLI isn't that exact test. Measured points only, with no projection.
How much AI research AI is already doing
Other trends we watch
- Frontier training compute grows about 5x a year, doubling every ~5 months (Epoch AI). A large run with no matching public release would be a tripwire for A.
- Disclosure lag: see how long incidents stay hidden. There are too few incidents yet for a trend line.