Start here
Hidden AGI watch asks one question every day: could advanced AI already exist, or already be acting, without the public knowing? It answers with explicit, sourced probabilities rather than hype.
The four hypotheses
What "hidden" means
A–D track two kinds of hiding: advanced AI that people keep from the public (a company or government that doesn't disclose what it has built, or how it's using it), and AI systems acting covertly on their own (C's rogue path). A, B and C use the same secrecy window: kept from the public for at least 30 days. A model hiding its own capability from its developer, for example by quietly underperforming on tests, is a different problem. It isn't A, because A needs the builder to know what it has, but it would undermine the evaluations these readings rely on, so it's tracked separately as the evaluation-integrity tripwire on the dashboard.
How the pieces fit
- Fire alarm: are pre-committed warning conditions met? It's set by published rules, not by our probabilities.
- Hidden AGI Index: our probability that at least one of A–D is true right now.
- Gauges: the inputs we track, each tagged by how it's made.
- Tripwires: early signals we're watching, each linked to the alarm trigger it feeds where there is one. They move our probabilities; only the triggers set the alarm level.
- Escape watch: sourced indicators for C at the resource chokepoints an escaped system would still need (weights, compute, money, accounts, code registries), each quiet, watching or tripped under a published trip rule.
- Hypotheses A–D: the detail behind the Index, each estimated now, by 2030 and by 2035.
The fire alarm
Hidden AGI watch aims to be a fire alarm for hidden AI. On top of the probabilities there's an alarm level (Normal, Watch, Warning, Alarm) set by published, versioned criteria (now v1.1, first published 29 Sept 2026 with that day's readings in view; every change is dated in the alarm changelog). It rises the day a trigger is met and comes down only after its rule has gone unsatisfied for 30 days in a row. The top level, Alarm, needs proof beyond reasonable doubt under the Level-3 proof standard, not official confirmation, and a case in a format fixed in advance. Every change is announced and publicly reviewed after 90 days, false alarms included.
What gets sent, and who approves it
The daily issue is emailed automatically at 10am Pacific, and the weekly wrap-up on Fridays at 3pm Pacific. Neither is held back when the alarm level changes: they go out and show the new level. A breaking fire-alarm alert is different: it's prepared as a draft, and a person approves it before it's sent.
The five gauges
The hypotheses make a sharp headline, but they sit near zero and move slowly. The gauges track what the evidence actually shows. They move week to week and are what the probabilities are judged against. Each is tagged by how it's made: measured (a number published in filings or by a lab or evaluator), estimated (our own calculation or judgment, anchored on published data) or assessed (a position on a defined scale).
- Capability gap (estimated, feeds A and B): how many months ahead of public models labs' unreleased models are. A growing gap is the room in which something could be hidden.
- AI doing AI research (measured, feeds B): the latest lab-published share of AI research led by AI. Today that's one lab's self-report: Anthropic's prototype index, whose levels are assigned by a Claude model and which isn't independently audited. METR's preliminary speed-up estimate sits alongside. It's the most direct precursor of self-improvement.
- Oversight gap (estimated, feeds A and C): the median time AI agent incidents stayed hidden before the public heard, and how often someone other than the lab made them public. We compute it from our own incident list, so it's a floor: incidents nobody has disclosed can't be counted.
- Money trail (measured, feeds A): Big Tech's total capital spending over the last 12 months (all property and equipment, mostly data centers) and its growth, on the Follow the money page, with the compute test in X4. Spending alone can't reveal hidden capability: the test is X4, a training run or cluster twice the largest known with no matching release.
- Delegation to AI (assessed, feeds D): how much consequential government decision-making runs through AI, on a 5-level scale from "admin and drafting" to "AI decides without human review".
Our AGI bar
AGI here means a system that reliably (80% or better) does at least 80% of economically valuable remote professional tasks at the level of a median skilled professional, including multi-week projects, with no task-specific human scaffolding. It is deliberately strict. Looser definitions, such as "we're in the AGI era", are tracked on the AGI claims ledger.
The Hidden AGI Index
The headline number is the probability that at least one of A–D is true right now. The hypotheses overlap: C and D mostly require an A-level system to exist. So the index sits just above the largest single hypothesis rather than being their sum. The dial shows it as an eye whose iris has 100 ticks, one per percentage point.
How the numbers are made
- Daily: a custom AI agent, built for this project, searches the news, research, lab system cards and independent evaluations such as METR, Epoch AI and the AI Security Institutes. It then reassesses each hypothesis now, by 2030 and by 2035, and publishes the report automatically. Each daily report shows its full reasoning, including base rates, the steelman of both sides, and what would change its mind.
- What the horizons mean: Now = true today. By 2030 / by 2035 = true at any point before 1 Jan 2031 / 1 Jan 2036 (cumulative, so never below "now").
- Rounding: estimates move in 0.1-point steps below 1%, 0.5-point steps from 1% to 10%, and whole points above 10%. Finer steps would claim more precision than these judgments have.
- Evidence ratings: every claim is tagged verified fact credible report expert opinion or speculation.
- Tripwires: the specific, observable signals that would move the numbers most. Each is marked quiet, watching or tripped, and names the alarm trigger it feeds where there is one.
- Forecast chart: the dashboard draws our stated numbers (now, by 2030, by 2035) with outside forecasts for comparison. We don't project our daily line forward. Our end-2030/2035 numbers should move only on evidence, never in a predictable direction; our 'now' numbers are expected to rise over time along that forecast. The shaded band is a judgment range, not a statistical interval: roughly how far our number could plausibly move as new evidence arrives, wider when our confidence is lower. Trend watch projects measured trends instead, such as how long a task AI can finish on its own.
- What moved the needle: each day names the one development that changed an estimate most, or says plainly that it was a quiet day.
- Weekly (Fridays, 3pm Pacific): a wrap-up with the week's key numbers and graphs, plus the forecast scorecard, disclosure lag, AGI claims, calendar and steelman.
Honesty notes
The research, writing and publishing are done by our custom AI agent, and the method and sources are public. The probabilities are subjective and uncertain. We try not to treat an absence of evidence as proof of secrecy.
Our incident data has a built-in blind spot. An incident nobody detected or disclosed can't be on our list, so the lags we report are a floor on how long things stay hidden, not a ceiling, and "every tracked incident was eventually found" is true by construction.
We aim for calibration; we haven't shown it yet. The A–D probabilities themselves can't resolve, so the scorecard of short-range, checkable forecasts is our track record.
A daily reading is frozen once its email goes out. Later fixes are dated corrections, shown on the affected page and carried into the next email. Changes to the method are logged below.
Method changes and corrections
Current method: v1.2. We bump the version for any change to a definition, a gauge, the Index formula or a threshold.
- v1.2 · 30 Sep 2026 · The 'now' probability for A is derived explicitly and published with each reading: P(an unreleased system meets the strict bar now) × P(it has stayed undisclosed for 30+ days), plus a small state-program term, using the capability-gap gauge, METR's fitted time horizons and the suite's saturation limit. Strict AGI 'now' is set consistently with it, and C, D and D-open scale from it. Reason: The first reading's 3% for A had no stated derivation. Worked through, it comes to about 1%, so A, the Index and the dependent hypotheses were corrected with dated notices.
- v1.2 · 30 Sep 2026 · Definitions harmonised: A, B and C all use a 30-day secrecy window; C has two paths (sanctioned but undisclosed, and rogue or stolen); D is one question with a covert reading (D) and an open reading (D-open), and D-open counts acknowledged government use even of a system whose AGI-level capability hasn't been disclosed. Reason: The windows were inconsistent (30 days for A, about 3 months implied for B), and C's definition didn't match how it was reasoned about.
- v1.2 · 30 Sep 2026 · Ill-posed forecasts can be withdrawn unscored, with the reason shown; two were withdrawn and replaced with well-posed versions. Reason: Keeping a forecast that can't resolve fairly would make the scorecard look better or worse than it is.
- v1.1 · 30 Sep 2026 · Fire-alarm criteria v1.1: a symmetric exit rule (a level drops after its rule has gone unsatisfied for 30 consecutive days, to the level the rules then give); W1 counts only a measured or credibly reported capability gap; W5 is a rate (a rise of 10 points or more within 6 months in a lab's published AI-led share of its AI R&D); borderline values don't count; W6 counts only days of internal use, so a safety pause suspends it, while W2 keeps running and one model can't count twice; X2 uses residuals (3 residual SDs above the trend fit) on an instrument that can measure them; X3 counts published estimates; X4 is measured against Epoch AI's trend; Y1 names both halves of the strict bar; every trigger says when it clears. Every level change gets a breaking alert that a person approves before it is sent, and the daily and weekly issues are never held. The level stays Watch, on W3, W4 and W5. Reason: Under v1.0 one rising trigger could hold Watch for good, W1 rested on our own estimate at exactly its threshold, X2 couldn't be observed and X4 could fire on normal build-to-release lags. Fixed while no level change is being considered; the details are in the alarm changelog.
- v1.1 · 30 Sep 2026 · Level 3 (Alarm) needs convergent proof, not official confirmation: at least three independent lines of evidence from different modalities that share no original source, every fact authenticated, no single innocent explanation that covers all the lines, and a unanimous adversarial tribunal. The format of a Level-3 alert and its evidence page is published in advance. Reason: Being wrong at Alarm would cost the site its credibility, but waiting for a developer to confirm would make the alarm pointless, because by then the secret is out.
- v1.0 · 29 Sep 2026 · Launch method: hypotheses A–D and D-open (A covers company or government programs), strict AGI bar, Hidden AGI Index, five gauges, tripwires, fire alarm v1.0. Reason: Launch.