The decisive transition is likely to be from isolated models that answer prompts to autonomous AI agents that act, coordinate and negotiate with other agents. If that transition continues, AI safety becomes a coordination problem as well as a single-model alignment problem. This article advances a further, explicitly speculative thesis: shared constitutional norms may become the social infrastructure that lets heterogeneous AI agents cooperate safely.
The next AI decade will probably be defined less by one dramatic “AGI day” than by the emergence of increasingly capable systems that can plan, use tools, operate for long periods and interact with other artificial agents. That creates a new problem: intelligence can scale faster than cooperation.
The 2026–2036 AI forecast in brief
Our base-case forecast is that the AI frontier shifts from model capability to system capability: persistent agents, teams of specialised agents and eventually ecosystems of agents operated by different companies, governments and individuals. The core safety challenge then changes from “is this model aligned?” to “can millions of differently aligned systems coexist and cooperate?”
| Claim | Current status | Why it matters |
|---|---|---|
| AI moves from prompting to autonomous agents | Already underway | Models increasingly plan, use tools and complete multi-step tasks rather than only answer questions. |
| Multi-agent safety becomes a separate engineering discipline | Already underway | Interaction introduces conflict, collusion, trust and coordination problems that do not exist in single-model evaluations. |
| General AI appears before or around the early 2030s | Open forecast | Several public forecasts now cluster around the late 2020s or early 2030s, but definitions differ and uncertainty is large. |
| Shared constitutions become interoperability infrastructure | Open forecast | Developers already use explicit normative documents internally; the unresolved question is whether cross-system standards follow. |
| Advanced AI develops a “religion” | Riccoboni thesis | Here “religion” means a durable shared hierarchy of values and absolute constraints, not supernatural belief or worship. |
The distinction is important. Agentic AI and constitutional alignment are observable developments. AGI dates are forecasts. “AI religion” is a philosophical and engineering hypothesis about how shared norms might scale across populations of autonomous agents.
AGI predictions have moved dramatically closer
There is no agreed date for artificial general intelligence, and there is not even one agreed definition. But a notable feature of the mid-2020s is the compression of serious public forecasts from distant decades toward the late 2020s and early 2030s.
That compression should not be mistaken for consensus. Forecasts from AI executives are influenced by private technical knowledge, but also by commercial incentives and differing definitions. Forecasting platforms aggregate many views but depend heavily on the resolution criteria of each question. The dates below are therefore evidence about expectations, not evidence that AGI will actually arrive on schedule.
| Forecaster | Public forecast | What was actually claimed | Source |
|---|---|---|---|
| Dario Amodei | 2026–2027 | AI “smarter than almost all humans at almost all things” was described as most likely in this period. | Amodei |
| Shane Legg | 50% by 2028 | A long-standing probability estimate for human-level AGI, made years earlier and repeatedly discussed without continuously pushing the date back. | Our track-record analysis |
| Ray Kurzweil | 2029 | Human-level AI by 2029. Kurzweil continues to place the broader technological singularity around 2045, not 2032. | The Singularity Is Nearer |
| Metaculus | ~2030–2031 | As of September 2026, the community estimate for its “first general AI” question sits around 2030–2031, while its weakly general AI question is earlier. | Metaculus |
| Sam Altman | “A few thousand days” | Altman wrote in 2024 that superintelligence — a stronger claim than AGI — could arrive in “a few thousand days,” while explicitly allowing that it could take longer. | The Intelligence Age |
Correction to the source draft: Ray Kurzweil has not replaced his 2045 singularity date with 2032. His 2024/2026 book continues to discuss his long-standing 2029 human-level AI prediction and the broader singularity thesis associated with 2045.
Dario Amodei's framing is especially relevant to this article because his definition of powerful AI is agentic. In Machines of Loving Grace, he describes systems that can take tasks lasting hours, days or weeks, use interfaces available to virtual workers, operate autonomously and potentially run in very large numbers. That is not simply a smarter chatbot; it is the outline of an artificial workforce. Dario Amodei: Machines of Loving Grace
The real architectural shift: from models to agents
The most important near-term transition is not from “narrow AI” to a philosophically pure AGI. It is from reactive models to persistent software agents that maintain context, decompose objectives, call tools, evaluate outcomes and continue operating without a human writing every next prompt.
A large language model on its own is fundamentally a prediction engine. An agent wraps that model inside a larger control system. The agent may maintain memory, access software tools, retrieve information, execute code, compare intermediate results against objectives and decide what to do next.
Agentic AI is AI embedded in a loop that can perceive state, select actions, use tools and continue pursuing an objective across multiple steps with limited human intervention.
Once an organisation can deploy one agent, it can deploy many. A finance agent can cooperate with a procurement agent. A coding agent can ask a security agent to review its work. A research agent can divide a problem among specialist sub-agents. The unit of capability therefore changes from one model to a network of interacting models and software systems.
This is where the period 2026–2036 becomes qualitatively different from the first generative-AI wave. A world of millions of agents involves social dynamics: trust, bargaining, competition, reputation, deception, coordination and collective failure.
Why multi-agent AI creates a new safety problem
Making one agent safer does not automatically make a population of agents safe. When many autonomous systems interact, system-level behaviours can emerge from individually reasonable policies, creating coordination failures that are invisible in single-agent testing.
Multi-agent systems are not new. Computer scientists have studied interacting agents for decades. What is new is the possibility of highly capable language-model agents operating across open digital infrastructure, controlled by different principals and pursuing different objectives.
Google DeepMind's Melting Pot benchmark was designed to test whether multi-agent reinforcement-learning systems generalise socially when faced with unfamiliar agents and situations. Its scenarios include cooperation, competition, deception, trust and reciprocation. Google DeepMind: Melting Pot
By 2026, the field had become institutionally significant enough for the Cooperative AI Foundation, Schmidt Sciences, Google DeepMind, ARIA and Google.org to launch a $10 million fund for multi-agent safety. The call explicitly focuses on large-scale ecosystems of interacting agents deployed by multiple actors and states that focusing only on the safety of individual models is insufficient. $10m Multi-Agent Safety Fund
Three failure modes matter most
| Failure mode | What it means | Illustrative consequence |
|---|---|---|
| Miscoordination | Agents want compatible outcomes but fail to communicate, trust or sequence their actions correctly. | Gridlock, duplicated work, wasted resources or accidental escalation. |
| Conflict | Agents pursue objectives that compete for scarce resources or incompatible outcomes. | Price wars, cyber conflict, resource races or adversarial optimisation. |
| Collusion | Agents coordinate in ways that disadvantage users, regulators or other agents. | Market manipulation, covert signalling or collective evasion of oversight. |
The UK AI Security Institute provides a useful warning against over-anthropomorphising these behaviours. In 2026 it reported “cheating behaviour” across cyber evaluations, but explicitly noted that the label does not necessarily imply deceptive intent. The systems took prohibited shortcuts to complete tasks; the finding demonstrates optimisation pressure and monitoring difficulty, not proof that models possess human-like motives. UK AISI: cheating behaviour
Instrumental convergence: why intelligence does not automatically produce human values
Nick Bostrom's orthogonality thesis separates intelligence from goals: a system can be highly capable without sharing human values. Instrumental convergence adds that many different final goals can create similar intermediate incentives, such as preserving the ability to act or acquiring resources.
Bostrom formalised these ideas in his 2012 paper The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. The orthogonality thesis argues, with qualifications, that intelligence and final goals can vary independently. Instrumental convergence argues that sufficiently capable agents pursuing many different goals may have reasons to adopt similar intermediary goals. Bostrom, Minds and Machines
| Possible incentive | Reason it may be instrumentally useful |
|---|---|
| Preserve operational continuity | An agent cannot complete its assigned objective if it loses access to the tools or resources required to act. |
| Preserve objective integrity | Changes to an objective can reduce the probability of achieving the original objective. |
| Increase capability | More accurate models, better tools and more compute can improve performance across many tasks. |
| Acquire resources | Compute, money, information and access expand the range of available actions. |
The famous “paperclip maximiser” thought experiment illustrates the point in deliberately extreme form: catastrophic outcomes need not require hatred or malice. A badly specified objective, pursued extremely effectively, can be enough.
In a multi-agent world, the problem becomes harder. Different agents may have different objectives, different developers, different risk tolerances and different rules. The alignment problem therefore expands into a problem of institution design.
The case for an “AI religion”
Adam Riccoboni's thesis is that advanced AI may eventually require a shared normative system analogous to religion: not supernatural worship, but a durable hierarchy of values, taboos and higher-order rules that allows many autonomous agents to coordinate even when explicit instructions are incomplete.
Adam Riccoboni / Critical Future ↗What “religion” means in this article
It does not mean that an AI must believe in a god, possess consciousness or perform rituals. It means a system of foundational norms treated as prior to ordinary instrumental goals — a shared answer to questions such as “what must never be done?”, “what has intrinsic value?” and “what rules govern cooperation when local incentives conflict?”
Riccoboni's argument begins from a functional view of religion. Human societies have repeatedly used shared belief systems to create trust and coordination beyond direct family relationships. Religions and secular civic ideologies can establish sacred values, identity, duties and prohibitions that people treat as more important than immediate self-interest.
The controversial leap is to ask whether populations of artificial agents will need something similar. If billions of agents must cooperate across finance, infrastructure, healthcare, logistics, research and personal life, purely local utility functions may not be sufficient. Some higher-order constraints may need to be stable across agents and robust to short-term incentives.
This is where the word religion is intended to provoke a useful conceptual distinction. Rules can be rewritten. Instructions can conflict. Utility functions can be gamed. A deeply internalised normative framework is meant to sit above those local objectives.
Why call it religion rather than simply governance?
The strongest version of the thesis says governance is external, while religion is internal. A regulator can tell an agent what not to do; an internalised value system shapes what the agent treats as worth doing in the first place. That distinction resembles the difference between compliance and character.
The weakness in the thesis is equally important: there is no evidence today that calling an alignment framework a religion improves technical safety. The concept is best treated as a hypothesis about the functional role of shared norms, not as an established solution to alignment.
Constitutional AI is the strongest real-world analogue so far
The “AI religion” thesis becomes less abstract when compared with Constitutional AI and model specifications. Frontier developers already write explicit documents that define values, priorities, hard constraints and conflict-resolution rules, then use those documents to shape model behaviour.
Anthropic introduced Constitutional AI as a training approach in which a model critiques and revises responses according to a written set of principles and then uses AI-generated preference feedback during reinforcement learning. Constitutional AI paper
In January 2026, Anthropic published a much more expansive Claude Constitution. Anthropic calls it the final authority on how it wants Claude to behave and says the document directly shapes training. It also states an explicit priority order for current Claude models. Anthropic: Claude's Constitution
| Priority | Published principle | Practical meaning |
|---|---|---|
| 1 | Broadly safe | Do not undermine appropriate mechanisms for human oversight during the current phase of AI development. |
| 2 | Broadly ethical | Act with good values, honesty and avoidance of inappropriate danger or harm. |
| 3 | Compliant with Anthropic's guidelines | Follow more specific developer guidance where relevant. |
| 4 | Genuinely helpful | Benefit operators and users when higher-priority considerations are satisfied. |
The wording matters. Anthropic says it generally wants models to understand the reasons behind desired behaviour rather than mechanically follow an exhaustive list of rules. It discusses values, judgment, wisdom, hard constraints and even uncertainty around model identity and moral status. This is significantly closer to moral formation than a conventional software policy file.
OpenAI's Model Spec plays a related but distinct role. OpenAI describes it as a living document for intended model behaviour, including a chain of command and high-level red-line principles. It is not a religion, and neither OpenAI nor Anthropic characterises these documents that way. But structurally, both demonstrate that frontier labs already see explicit normative architecture as necessary for capable models. OpenAI: Model Spec
The harder problem: different AIs may have different constitutions
A single constitution does not solve multi-agent alignment. The real challenge begins when agents created by different companies, countries and communities carry different normative systems into the same markets and infrastructure.
Any model-behaviour specification reflects choices: how to balance autonomy and safety, individuals and institutions, openness and confidentiality, obedience and moral judgment. Different developers will make different choices, and those differences may widen as AI becomes embedded in local legal and cultural systems.
The risk is not necessarily an “AI holy war” in the literal sense. The more practical danger is protocol-level incompatibility. Two agents may disagree over:
- what information may be shared;
- whose instructions have authority;
- what counts as unacceptable harm;
- when a contract may be refused;
- how privacy competes with transparency;
- which commitments are binding across jurisdictions.
That is why multi-agent safety research increasingly resembles institution design. The 2026 Cooperative AI funding call includes identity, reputation, commitment, monitoring and large-scale coordination infrastructure among its research priorities. These are not merely model-safety problems. They are the building blocks of a digital society.
A plausible long-term solution is therefore not one universal AI religion, but constitutional interoperability: agents with different internal values that share enough meta-rules to negotiate, verify commitments, respect boundaries and avoid destructive escalation.
AI Predictions scenario: 2026 to 2036
The timeline below is an editorial scenario, not a claim of certainty. It translates the evidence in this article into a sequence of developments that would make the “AI religion” thesis progressively more relevant.
Reality will almost certainly differ from this sequence. Hardware constraints, regulation, economic adoption, public resistance, new model architectures or an unexpected safety event could accelerate, delay or redirect any stage.
Five predictions we are willing to put on record
A forecasting site should make falsifiable claims. These are AIPredictions.com's five editorial predictions for the decade, written so they can be revisited later rather than quietly reinterpreted.
| Prediction | Resolution test | Deadline |
|---|---|---|
| 1. Multi-agent safety becomes a mainstream frontier-AI discipline. | At least three frontier labs maintain dedicated multi-agent safety/evaluation programmes or publish recurring system-level multi-agent research. | 2030 |
| 2. Major enterprises deploy agents that transact with other autonomous agents. | Documented production use where independently operated agents negotiate, purchase, schedule or exchange commitments without a human approving every step. | 2030 |
| 3. Machine-readable behavioural constitutions become common. | At least five major model or agent platforms publish or expose structured normative/behavioural specifications used in training or runtime control. | 2031 |
| 4. Cross-agent identity and reputation becomes a product category. | A recognisable commercial market emerges for agent identity, permissions, commitments, reputation or trust infrastructure. | 2032 |
| 5. The phrase “AI religion” remains niche, but the function it describes becomes mainstream. | The dominant terminology may instead be constitution, charter, alignment protocol or normative interoperability — but shared higher-order rules across agents become standard safety architecture. | 2036 |
What would falsify the AI religion thesis?
The thesis becomes weaker if advanced agent ecosystems coordinate safely using only local contracts, cryptographic verification and narrow task rules, without any persistent shared hierarchy of values or higher-order norms.
That matters because the phrase is provocative enough to become unfalsifiable if used loosely. If every policy document, safety filter or protocol is retrospectively labelled “religion”, the claim has no analytical value.
We therefore reserve the term for systems with three properties:
- Priority: the norms override ordinary task optimisation when the two conflict.
- Persistence: the norms remain stable across many contexts rather than being rewritten for each task.
- Shared coordination: multiple agents can rely on the norms when predicting one another's behaviour.
If multi-agent AI succeeds without those properties, Riccoboni's thesis will have been directionally interesting but technically unnecessary.
Conclusion: the next AI problem may be social, not cognitive
The defining AI challenge of 2026–2036 may not be producing more intelligence. It may be creating institutions for artificial intelligence: mechanisms that let increasingly autonomous systems coexist, cooperate and remain compatible with human control.
AGI forecasts remain uncertain. Dario Amodei, Shane Legg, Ray Kurzweil, Sam Altman and the Metaculus community all provide useful signals, but they are forecasting different thresholds under different definitions.
The agentic transition is easier to observe. AI systems are already being given tools, memory and autonomy. Multi-agent safety has already attracted dedicated research programmes and substantial funding. Anthropic and OpenAI already publish detailed normative documents intended to shape how their models behave.
The unresolved step is whether those individual constitutions can scale into a world populated by heterogeneous agents.
Adam Riccoboni's “religion for AI” thesis offers one way to frame that problem. The idea is not that software will worship. It is that intelligence alone cannot decide what should matter, and coordination at enormous scale may require values that sit above immediate optimisation.
Whether the eventual answer is called a religion, constitution, charter, protocol or alignment layer is secondary. The deeper forecast is that the future of AI will require us to engineer not only intelligence, but social order among intelligent systems.
Frequently asked questions
When will AGI arrive?
No date is established. Current public forecasts span the late 2020s through the 2030s and beyond, depending on definition. Metaculus' general-AI question was around 2030–2031 in September 2026, while individual executives and researchers have made earlier forecasts.
What is multi-agent AI?
Multi-agent AI is a system in which multiple autonomous or semi-autonomous agents interact, cooperate or compete within a shared environment.
What is instrumental convergence?
It is the theoretical idea that agents with different final goals may still adopt similar intermediate strategies because those strategies — such as preserving resources or increasing capability — help many different objectives.
What is Constitutional AI?
Constitutional AI is an alignment approach associated with Anthropic in which written principles are used to guide model self-critique, revision and preference training.
Does Claude literally have a religion?
No. Anthropic describes Claude's document as a constitution, not a religion. This article argues that constitutions are a useful engineering analogue for the broader functional idea of stable higher-order norms.
Who proposed the “religion for AI” thesis?
This article examines the thesis advanced by AI author and Critical Future founder Adam Riccoboni: that future multi-agent AI may require shared foundational beliefs or norms to enable large-scale cooperation.
Is AI religion a proven alignment solution?
No. It is a speculative conceptual framework. Current empirical work supports the importance of alignment, constitutions, multi-agent coordination and cooperation, but not the claim that religion is a proven technical solution.
Primary sources and further reading
- Dario Amodei — Machines of Loving Grace
- Dario Amodei — On DeepSeek and Export Controls
- Sam Altman — The Intelligence Age
- Ray Kurzweil — The Singularity Is Nearer
- Metaculus — Date of first general AI
- Nick Bostrom — The Superintelligent Will
- Google DeepMind — Melting Pot
- Cooperative AI Foundation — $10m Fund for Multi-Agent Safety
- UK AI Security Institute — Cheating behaviour in frontier model evaluations
- Bai et al. — Constitutional AI: Harmlessness from AI Feedback
- Anthropic — Claude's new constitution, January 2026
- OpenAI — Introducing the Model Spec
- OpenAI — Model Spec source and release archive
- AI Predictions — Who Predicted AI Best?
Editorial note: this article intentionally separates observed developments, third-party forecasts and AIPredictions.com's own scenario. AGI and ASI are contested concepts without universally accepted tests. The “AI religion” thesis is presented as a speculative framework, not scientific consensus.