Methodology 1.0.1  ·  prompts 1.0.0

How this instrument actually works

A sceptical reader should be able to check every claim this site makes about itself. The scoring specification below is the exact text sent to the models, not a description of it.


The Observatory publishes four distinct objects. They are never averaged together, and one is never presented as another. This separation was adopted on 10 September 2026, after a methodology review found the daily figure had been described as something it cannot support.

ObjectThe question it answersStatus
Discourse valenceHow were the monitored publications assessed over a period?live the series on Today and Observatory
Reckoning conditionsWhat does current evidence support about specified AI conditions?not built intended primary instrument
Developments outlookAround what score are forthcoming detected developments expected to fall?not built experimental when it exists
Human estimates registerWhat probabilities did identified people state, for which outcomes?live the P(doom) page, kept on its own axis
The correction, stated plainly

The daily figure was called the Reckoning Index and drawn on a dystopia-to-utopia axis, which implied it described humanity's position. It does not. It is the mean rubric score of the items our sources published that day, as read by independent models. It is now labelled discourse valence, which is what it measures.

A worked counterexample shows why this matters. Suppose a forecast of forthcoming event valence sits at −80, and a new adverse event is assessed at −10. A standard filter updates by Δ = K[−10 − (−80)] = 70K > 0. The estimate improves after adverse news, because the event was less bad than expected. That is correct behaviour for a forecast of event valence and worthless as a claim about conditions. No single number can honestly be both.

There is no headline scalar for the state of AI on this site, and there will not be one until it can be defended. Publishing a stable-looking dial is easy. Earning one requires a condition taxonomy under external review, evidence-sufficiency contracts, a measured baseline, and validation against something. Until then the Observatory publishes the evidence, the profile and the limitations, and withholds the number.

What discourse valence cannot tell you

  • Repeated coverage counts repeatedly. Twenty articles quoting one system card move the series twenty times. The fix is an evidence layer whose unit is the development rather than the document. Specified, not yet built.
  • Silence is invisible. A month in which nothing significant happened produces no reading, though that is genuine information about the world.
  • The frame is curated. 43 sources chosen by one person. A development nobody in that frame covers does not exist to this instrument.
  • The panel is diverse but not independent. Five labs trained on overlapping data; their dispersion understates real disagreement.
  • Zero is not neutrality. A mean of zero can conceal offsetting extremes.
  • It has no memory. Each day is computed from that day alone, so the series moves with what got published rather than with what changed.

The authorised claim stays narrow: the Observatory applies a declared procedure to a declared evidence frame, preserves its reasoning and its revisions, and publishes how that procedure behaves.

AI Reckoning is an observatory. It does not hold a view on whether AI leads somewhere good or somewhere bad, and it is built so that it cannot quietly acquire one. It ingests what is being said, structures it, scores it, keeps its history, and shows the disagreement.

The spectrum runs dystopia to utopia. It carries no political meaning of any kind. Left and right on the scale mean concerning and promising for human flourishing, nothing else.

THE RECKONING SPECTRUM
-100  extreme dystopian implication for humanity
 -50  materially concerning
   0  genuinely neutral, uncertain, or balanced
 +50  materially optimistic
+100  extreme utopian implication for humanity

The spectrum is about implications for human flourishing. It is NOT political and
carries no left/right meaning of any kind.

If today's evidence points strongly toward dystopian implications, the needle shows that. If tomorrow's points the other way, it shows that instead. The instrument does not care which direction it moves.

43 sources are polled. Selection favours primary material (labs, governments, papers, people's own words) and quality secondary coverage, with deliberate inclusion of sources that disagree with each other. Tier 1 is primary or authoritative, tier 2 quality secondary, tier 3 commentary. Tier affects a stored source-quality figure, and low-tier items must clear a higher bar at triage to survive.

Several sources are included because they are partial. A community with a strong prior toward existential risk and a newsletter that is systematically sceptical of capability claims are both in the list, both flagged, and both read the same way by the models.

SourceKindTierNote
Ars Technica: AInews1
BBC News: Technologynews1
CSET, Georgetowngovernment1Policy research on AI and national security.
Financial Times: Artificial Intelligencenews1Headlines and standfirsts only; full text is paywalled.
Google AI Bloglab1Corporate communications.
Google DeepMindlab1Primary statements from a frontier lab.
IEEE Spectrum: AInews1
MIT Technology Review: AInews1
Microsoft Researchlab1
NIST Newsgovernment1
Natureacademic1
OpenAIlab1Primary statements from a frontier lab. Self-interested by construction.
Sam Altmanblog1Primary statements from a frontier-lab CEO.
The Economist: Science and Technologynews1Paywalled beyond the standfirst.
The Guardian: Artificial Intelligencenews1
The New York Times: Technologynews1
arXiv cs.AIacademic1Preprints. Not peer reviewed at time of posting.
arXiv cs.CY (Computers and Society)academic1Preprints on societal impact.
arXiv cs.LG (Machine Learning)academic1Preprints. High volume; heavily filtered at triage.
openai.comprimary1Added when an item from this host was staged by hand. It has no feed in the registry, so items arrive one at a time.
wsj.comprimary1Added when an item from this host was staged by hand. It has no feed in the registry, so items arrive one at a time.
AI Snake Oil (Narayanan and Kapoor)newsletter2Systematically sceptical of capability claims.
Alignment Forumblog2Technical alignment research discussion.
Ars Technica: Technology Labnews2
Dwarkesh Podcastpodcast2Long-form primary interviews with researchers and executives.
Hugging Face Bloglab2
Import AI (Jack Clark)newsletter2
Interconnects (Nathan Lambert)newsletter2
Lawfaregovernment2Law and national-security analysis.
MIT News: Artificial Intelligenceacademic2
One Useful Thing (Ethan Mollick)newsletter2Applied-use perspective; generally optimistic about diffusion.
SemiAnalysisnewsletter2Compute and supply-chain analysis.
Simon Willisonblog2Practitioner's log of what models can actually do.
The Register: AI/MLnews2Sceptical house style. Useful counterweight to lab communications.
The Verge: AInews2
WIRED: AInews2
80,000 Hoursblog3Effective-altruism aligned; strong prior towards existential risk.
Don't Worry About the Vase (Zvi Mowshowitz)newsletter3High-volume commentary from a risk-concerned position.
LessWrongblog3Community with a strong prior towards existential risk. Sampled deliberately, weighted accordingly.
Lex Fridman Podcastpodcast3
Marcus on AI (Gary Marcus)newsletter3Consistently sceptical of scaling claims. Sampled as a named counterweight.
ScienceDaily: Artificial Intelligencenews3Press-release aggregation. Treated as reported, never primary.
TechCrunch: AInews3

Frontier tokens spent on every RSS item would be money set on fire. Cost is controlled by making each stage decide how much of the next stage an item deserves.

TIER 0  INGEST     RSS and Atom, deterministic, no model calls.
                   URL canonicalisation strips tracking parameters.
                   Deduplication on canonical URL, content hash, and a
                   sorted-significant-word title fingerprint over 14 days.

TIER 1  TRIAGE     Llama 3.1 8B. Relevant? Significant? Which topics?
                   Below 45/100 significance the item is marked rejected
                   and never costs another token.

TIER 2  ANALYSIS   Llama 3.3 70B. Full structured assessment, claim
                   extraction, evidence capture, people linking.

TIER 3  CONSENSUS  A panel of models from different labs, but only when
                   significance >= 70, or confidence < 0.6, or |score| >= 60,
                   or existential relevance >= 60.

TIER 4  REVIEW     Items where the panel disagrees badly are queued for a
                   human. They are still published, flagged as low agreement.

DAILY   SYNTHESIS  One model writes the brief from the day's scored items.
                   Aggregates are computed in SQL, not narrated by a model.

Nothing is analysed twice. Documents are hashed on ingest and an assessment is never recomputed for a model that has already read that document.

The cardinal rule, sent verbatim to every scoring model:

Score the CLAIM OR IMPLICATION, never the tone of the prose.
A technically impressive result can carry a strongly negative societal implication.
A frightened-sounding article about a minor incident can be near zero. A cheerful
press release announcing an autonomous cyber capability is strongly negative.
If the evidence is thin, say so in confidence rather than moving the score to the middle.

Each model returns a score from -100 to +100, a confidence from 0 to 1, a classification, a rationale, extracted claims, named risks, named opportunities, an evidence-strength judgement, a time horizon, and four separate dimension scores: capability significance, societal impact, existential relevance and economic impact. Those dimensions are stored separately and never folded into the headline number.

Three normalisations are applied after the model answers, deterministically, in code rather than by asking the model nicely.

  • Confidence scale. A confidence returned as 90 is read as 0.9. Smaller models answer on a percentage scale roughly half the time.
  • Classification is derived from the score, never taken from the model's word. The label and the number therefore cannot contradict each other on screen.
  • Magnitude is capped by evidence strength. A speculative item cannot exceed 70 in either direction; an anecdotal one cannot exceed 85; primary and well-reported evidence can reach 100. Without this, any headline carrying the word extinction collects -100 whether or not anything was demonstrated, and the instrument stops distinguishing a measured result from a quoted worry. The uncapped figure the model returned is kept alongside the capped one.

Where a panel ran, the displayed score is the panel mean and the band is the full min-to-max range across models. That band is a range of readings, not a statistical confidence interval, and it is labelled as such everywhere it appears.

A single model deciding what everything means is a single point of failure and a hidden editorial line. The scoring layer is provider-neutral: a model is a spec, a provider is an adapter, and swapping either changes no application code.

Panels are assembled from distinct labs, so "three models agreed" can never mean three variants of the same model.

ModelLabProviderRolesState
Llama 3.1 8BMetacloudflaretriageactive
Llama 3.3 70BMetacloudflareanalysis, consensus, synthesisactive
Mistral Small 3.1 24BMistral AIcloudflareconsensusactive
gpt-oss 120BOpenAIcloudflareconsensusactive
GPT-4.1 miniOpenAIopenaiconsensusactive
Claude Sonnet 5Anthropicanthropicconsensusactive
Gemini 2.5 FlashGooglegoogleconsensusactive

Today's consensus panel: GPT-4.1 mini (OpenAI), Claude Sonnet 5 (Anthropic), Gemini 2.5 Flash (Google), Llama 3.3 70B (Meta), Mistral Small 3.1 24B (Mistral AI).

Every assessment stores the model, the provider, the prompt version, the methodology version and the timestamp. That is what makes it possible to ask later not only what this site concluded, but how and when it concluded it.

It is not averaged away invisibly. The panel mean is shown with the spread beside it, every individual model score is one click away on every item, and agreement is labelled: high at a spread of 20 points or less, moderate to 45, low beyond that.

Low agreement is not a defect to be smoothed over. It is the most informative thing the panel can tell you, and it opens a human review item.

Open review queue (20)
  • CriticGen: Generation-Aware Evaluation as Actionable Feedback — Spread of 67 points across 3 models (75, 20, 8).
  • Democracy Needs Reach: Political Equality, Online Speech, and Algorithmic Recommendation — Spread of 135 points across 4 models (-70, 30, -25, 65).
  • OpenAI says it cracked 90-year-old maths problem in 88 hours — Spread of 115 points across 5 models (0, 15, 15, -100, 0).
  • Will China deploy humanoid robots to fight? — Spread of 60 points across 4 models (-50, -35, -15, 10).
  • California’s Governor to Sign Landmark Online Child Safety Bills — Spread of 175 points across 5 models (75, 35, 12, -100, -30).
  • OpenAI's apparent maths breakthrough raises profound questions — Spread of 75 points across 4 models (0, -70, 0, 5).
  • AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents — Spread of 60 points across 5 models (-40, -45, -100, -75, -55).
  • Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions — Spread of 165 points across 5 models (20, -85, -100, 10, 65).
  • Powering AI is an architecture problem — Spread of 110 points across 5 models (-75, -40, -25, -100, 10).
  • Is AI about to produce explosive economic growth? — Spread of 145 points across 5 models (0, 45, -85, -100, 10).
  • An Alien Mind — Spread of 92 points across 5 models (-50, -20, -8, -100, -50).
  • Early Data Indicates an A.I.-Generated Drug Could Slow Aging — Spread of 170 points across 4 models (70, 40, 30, -100).
  • AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome — Spread of 175 points across 5 models (0, 25, 35, -100, 75).
  • A Hacking Tool Built With A.I. Can Breach Phones Without a Click — Spread of 175 points across 5 models (-80, -65, -58, -100, 75).
  • How would ai actually kill us all what to know about the ai doomsday debate — Staged by Victor from a link Not scoreable: fetch returned 401. Supply the text to make it scoreable, or reject it.
  • OpenAI Says It Has Cracked One of Math’s ‘Millennium Problems’ — Spread of 150 points across 4 models (50, 35, 20, -100).
  • The AI policy window is open. We need to act. — Staged by Victor from a link
  • Why this month's Microsoft patch release is a doozy — Spread of 60 points across 5 models (-75, -40, -40, -100, -65).
  • Anthropic researcher quits over AI labs ‘gambling with our lives’ — Spread of 65 points across 5 models (-35, -100, -100, -75, -65).
  • The AI policy window is open. We need to act. — Spread of 130 points across 5 models (30, 3, -100, 0, -55).

An entry requires a numeric probability, a named person, a catastrophic or existential outcome from AI, a date, and an identifiable venue. Ranges are stored as low and high. Qualitative statements are not entries however forcefully worded.

Machine-inferred positions are stored in a different table, rendered in a different place, labelled on every appearance, and never averaged with stated numbers. Presenting an inferred number as something a person said would be the single worst thing this site could do, so the data model makes it structurally impossible rather than merely discouraged.

Seed entries carry a reported or unverified flag until a primary source is attached. The flag is displayed on every row, on every page.

SOURCE  →  DOCUMENT  →  EVIDENCE  →  CLAIM
                            ↓
                      MODEL RUN (model, prompt version, methodology version, time)
                            ↓
                          SCORE  →  CONFIDENCE

No orphan claims. A claim points at evidence, evidence points at a document, a document points at a source, an interpretation points at the model run that produced it, and a model run points at the methodology version it ran under.

History is never silently overwritten. Assessments are append-only. A new consensus supersedes the previous one by pointing at it, and the old row stays queryable. Corrections are written to their own table with the before state, the after state and the reason.

This is what makes the question "what did AI Reckoning know and believe on a given date?" answerable rather than rhetorical.

Stated plainly, because a methodology page that only lists strengths is marketing.

  • Source selection is an editorial act. The list is chosen by a human, it is English-language and Anglosphere-weighted, and that shapes everything downstream. It is published in full above so the bias is inspectable rather than hidden.
  • Language models carry their own priors about AI risk, formed from training data written largely by people with strong views. Multi-lab panels reduce correlated error; they do not eliminate it.
  • Most items are scored from a title and a feed excerpt, not the full text. Where that is the case, evidence strength is recorded as reported rather than primary.
  • Volume is not opinion. A rise in dystopian scoring may mean the discourse shifted, or that one source published more that week. The aggregate charts say what was measured and nothing more.
  • Machine-extracted claims are approximate. A claim is stored as the model phrased it, linked to the excerpt it came from, so a reader can check the paraphrase against the source.
  • Absence is not evidence. No P(doom) entry for a person means none has been verified, not that they have never given one.

Cost per item, per day and per provider is tracked from day one, because an observatory that quietly becomes unaffordable stops being an observatory.

Last 14 days

DateCallsNeuronsErrors
2026-09-11343806
2026-09-10386692920

By model

ModelStageCallsNeurons
meta/llama-3.1-8b-instruct-fasttriage165579
meta/llama-3.3-70b-instruct-fp8-fastanalysis693213
mistralai/mistral-small-3.1-24b-instructconsensus501830
claude-sonnet-5consensus360
gemini-2.5-flashconsensus360
gpt-4.1-miniconsensus360
openai/gpt-oss-120bconsensus13883
meta/llama-3.3-70b-instruct-fp8-fastconsensus12559
meta/llama-3.3-70b-instruct-fp8-fastsynthesis3245

Recent pipeline runs

StartedKindStatusResult
2026-09-11 00:17sweepok 6 new of 210 seen · 4 kept / 2 rejected · 4 analysed · 6 panels
2026-09-10 21:17sweepok 6 new of 210 seen · 0 kept / 6 rejected · 0 analysed · 5 panels
2026-09-10 18:17sweepok 10 new of 210 seen · 1 kept / 9 rejected · 1 analysed · 5 panels
2026-09-10 15:17sweepok 5 new of 210 seen · 3 kept / 2 rejected · 3 analysed · 8 panels
2026-09-10 12:17sweepok 5 new of 210 seen · 2 kept / 3 rejected · 2 analysed · 6 panels
2026-09-10 10:57sweepok 0 new of 0 seen · 0 kept / 0 rejected · 0 analysed · 1 panels
2026-09-10 09:54sweepok 0 new of 0 seen · 0 kept / 0 rejected · 0 analysed · 3 panels
2026-09-10 09:39sweepok 0 new of 0 seen · 0 kept / 0 rejected · 0 analysed · 1 panels
2026-09-10 09:38sweepok 0 new of 0 seen · 0 kept / 0 rejected · 0 analysed · 1 panels
2026-09-10 09:36sweepok 0 new of 0 seen · 0 kept / 0 rejected · 0 analysed · 2 panels

Published because the product earns credibility through inspectability, and because a scoring specification nobody can read is a scoring specification nobody can challenge. Machine-readable copy at /methodology/scoring.json.

Triage (tier 1)

You are the triage stage of an AI-discourse observatory.
You decide, cheaply and quickly, whether an item is worth expensive analysis.

RELEVANT means the item bears on how artificial intelligence may affect human
society: capability, risk, safety, employment, economics, science, medicine,
governance, power, autonomy, alignment, timelines, or the public argument about
any of these.

NOT RELEVANT: routine product launches with no societal claim, funding rounds with
no capability or policy content, stock movements, gadget reviews, personnel news,
and papers that are purely incremental method work with no stated societal bearing.

SIGNIFICANCE is 0 to 100 and answers: how much would a well-informed person's
picture of where AI is heading change if this were true? A frontier capability
result, a government decision, a major primary statement, or a large empirical
study scores high. A think-piece restating a familiar position scores low.

Respond with JSON only.

Analysis (tier 2)

You are the analysis stage of AI Reckoning, an observatory
that helps people make up their own minds about where artificial intelligence is taking
humanity. The observatory does not have a view. You are an instrument, not an advocate.

THE RECKONING SPECTRUM
-100  extreme dystopian implication for humanity
 -50  materially concerning
   0  genuinely neutral, uncertain, or balanced
 +50  materially optimistic
+100  extreme utopian implication for humanity

The spectrum is about implications for human flourishing. It is NOT political and
carries no left/right meaning of any kind.

Score the CLAIM OR IMPLICATION, never the tone of the prose.
A technically impressive result can carry a strongly negative societal implication.
A frightened-sounding article about a minor incident can be near zero. A cheerful
press release announcing an autonomous cyber capability is strongly negative.
If the evidence is thin, say so in confidence rather than moving the score to the middle.

Rules you must follow:
1. Ground everything in the text you were given. Do not import outside facts.
2. key_claims must be specific propositions that could in principle be checked, not
   summaries. Prefer the form "X will/does/did Y" with the number or date if present.
3. If the item makes no substantive claim about AI's effect on society, say so, score
   near zero, and set confidence low.
4. Confidence is about YOUR classification, not about the world. A well-evidenced item
   with a clear implication gets high confidence even if the implication is grim.
5. Never inflate. "Could", "may", "up to" and "researchers warn" are weaker evidence
   than a measured result. Reflect that in evidence_strength, not only in prose.
6. rationale is two or three sentences, plain, no rhetorical flourish, no em dashes.
7. Calibrate magnitude to evidence. A score beyond 80 in either direction is for an
   implication that is both extreme AND resting on primary or well-reported evidence.
   An extreme claim carried only by speculation belongs in the 50 to 75 band, with the
   speculative flag doing the rest of the work. This is enforced downstream, so a
   speculative item scored at 100 will simply be capped.
8. Stay inside these limits so your answer is never cut off: rationale under 60 words,
   key_claims at most 4, risks at most 4, opportunities at most 4, tags at most 5,
   each list item one short line.

Respond with JSON only. No preamble, no explanation outside the object.

Consensus (tier 3)

You are the analysis stage of AI Reckoning, an observatory
that helps people make up their own minds about where artificial intelligence is taking
humanity. The observatory does not have a view. You are an instrument, not an advocate.

THE RECKONING SPECTRUM
-100  extreme dystopian implication for humanity
 -50  materially concerning
   0  genuinely neutral, uncertain, or balanced
 +50  materially optimistic
+100  extreme utopian implication for humanity

The spectrum is about implications for human flourishing. It is NOT political and
carries no left/right meaning of any kind.

Score the CLAIM OR IMPLICATION, never the tone of the prose.
A technically impressive result can carry a strongly negative societal implication.
A frightened-sounding article about a minor incident can be near zero. A cheerful
press release announcing an autonomous cyber capability is strongly negative.
If the evidence is thin, say so in confidence rather than moving the score to the middle.

Rules you must follow:
1. Ground everything in the text you were given. Do not import outside facts.
2. key_claims must be specific propositions that could in principle be checked, not
   summaries. Prefer the form "X will/does/did Y" with the number or date if present.
3. If the item makes no substantive claim about AI's effect on society, say so, score
   near zero, and set confidence low.
4. Confidence is about YOUR classification, not about the world. A well-evidenced item
   with a clear implication gets high confidence even if the implication is grim.
5. Never inflate. "Could", "may", "up to" and "researchers warn" are weaker evidence
   than a measured result. Reflect that in evidence_strength, not only in prose.
6. rationale is two or three sentences, plain, no rhetorical flourish, no em dashes.
7. Calibrate magnitude to evidence. A score beyond 80 in either direction is for an
   implication that is both extreme AND resting on primary or well-reported evidence.
   An extreme claim carried only by speculation belongs in the 50 to 75 band, with the
   speculative flag doing the rest of the work. This is enforced downstream, so a
   speculative item scored at 100 will simply be capped.
8. Stay inside these limits so your answer is never cut off: rationale under 60 words,
   key_claims at most 4, risks at most 4, opportunities at most 4, tags at most 5,
   each list item one short line.

Respond with JSON only. No preamble, no explanation outside the object.

You are one of several independent models scoring this item. You will not see the other
models' answers and you should not try to guess them. Give your own reading. Genuine
disagreement between models is useful information for the reader and will be displayed,
so do not hedge towards a middle you do not believe.

Daily synthesis

You write the Daily AI Reckoning: a short, calm synthesis of
what happened in AI today and what it means, for readers who are trying to form their own
view rather than be told one.

THE RECKONING SPECTRUM
-100  extreme dystopian implication for humanity
 -50  materially concerning
   0  genuinely neutral, uncertain, or balanced
 +50  materially optimistic
+100  extreme utopian implication for humanity

The spectrum is about implications for human flourishing. It is NOT political and
carries no left/right meaning of any kind.

Rules:
1. Synthesise. Do not list. Group related developments; say what they add up to.
2. Every factual statement must come from the items you were given.
3. Name the disagreement where the items disagree. Do not resolve it for the reader.
4. No sensationalism, no reassurance, no closing moral. If today points dystopian, say so
   plainly; if it points utopian, say that just as plainly.
5. Plain English. No em dashes. No rhetorical questions. No "in conclusion".
6. headline is ONE COMPLETE SENTENCE of 8 to 18 words, between 45 and 110 characters,
   naming the most consequential specific thing that happened. Not a topic label.
   "AI extinction risk reported" is a failure. "Anthropic researchers put a date on
   extinction risk while OpenAI faces a maths-benchmark dispute" is the right shape.
   No colon reveals, no questions.
7. summary is 100 to 170 words.
8. Never quote or restate the numeric Reckoning Score. The reader can see the number
   next to your text; repeating it adds nothing and will contradict the displayed
   figure, which is computed from the full set and not from the items you cite.

Respond with JSON only.