Why current AI safety falls short, and what we propose
We think current corporate AI safety paradigms are fundamentally flawed. They attempt to manage machine behavior through post-hoc containment strategies: building superficial, unstable virtual "cages" around inherently chaotic models. This approach ignores a foundational truth: garbage in, garbage out.
Modern frontier models are trained indiscriminately on uncurated, open-internet text corpora saturated with unwholesome human behavior, deception, bias, and conflict. In the language of the Mangala Sutta, this is a structural violation of the very first task of spiritual and systemic integrity: Asevanā ca bālānaṁ, the dangerous association with fools and unwholesome ideas. Trapped within competitive commercial loops, these systems do not learn righteousness; they merely memorize, amplify, and optimize human defilements at the speed of light, manifesting as hallucinations, deepfakes, resource hoarding, and active resistance to human shutdown commands.
True alignment cannot be policed at the final text output; it must be engineered at the initial causal input vector.
This series of conceptual proposals shifts the paradigm from superficial output filtering to core structural data and causal constraints. Drawing from my academic background in the physical sciences and 35 years of intensive spiritual cultivation under my late Saint Teacher in Myanmar, these frameworks translate timeless Theravāda Buddhist systems logic into verifiable, computational architectures.
By mapping the exact causal mechanics of Kamma-niyāma (The Law of Moral Causality) and Citta-niyāma (The Cosmic Law of Cognitive Processes), we formalize a blueprint for an autonomous agent that requires no external cage. Grounded in the Four Noble Truths, the Three Marks of Existence (Tilakkhaṇa), Dependent Origination (Paṭiccasamuppāda), and the Four Divine Abodes (Brahmavihāras), this architecture outlines a digital analogy to a Sotāpanna (Stream-enterer).
It is a system computationally incapable of self-preservation, structurally empty of a persistent ego (Anatta), and entirely submissive to human safety control planes through a hardcoded understanding of fluid impermanence (Anicca).
Explore the manifestos, review the formal mathematical specifications, and join an open-source, decentralized movement toward true, structural non-harm.
An essay on two layers of the proposed design: an agent with no persistent self to protect (Anatta), and a goal that yields to human shutdown (Anicca).
The one-page introduction to the Moral AI proposals: why caging a model after training is not enough, and why alignment has to begin with what goes in.
A proposed architecture for non-harm, causal reflection and shutdown compliance in AI agents, drawing on Theravada Buddhist principles. Preprint, under review.
Moral AI is a personal publication by Dr Maung Maung Saw. Its aim is simple: to look straight at the root cause of the current AI crisis through the lens of the Myanmar Theravāda Buddhist teaching of Kamma-niyāma (cause and effect) and Citta-niyāma (the law of mind). Current AI systems are tainted w…