When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty
arXiv (Cornell University) June 4, 2026 DOI: 10.48550/arxiv.2606.05528 via OpenAlex
Summary
AI-generated from the abstractA new precautionary framework translates evidence about whether AI systems might be conscious into graduated protective obligations. The framework maps five welfare-relevant dimensions—phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, and agency—each linked to distinct moral concerns. It uses a hybrid threshold-plus-gradation approach: binary triggers activate new obligation categories, while continuous scaling adjusts protective weight. Two complementary aggregation methods are offered: one hierarchical, one architecture-agnostic. Worked case studies of Replika and OpenClaw show how systems in different dimensional regions trigger different obligations. The framework applies across neural, symbolic, and neurosymbolic systems, aiming to make consciousness science decision-relevant for organizations.
Study at a glance
| Characteristics | Theoretical or philosophical paper Peer reviewed |
|---|---|
| Keywords | Operationalization Consciousness Metacognition Conceptual framework Precautionary principle |
| Key finding | A precautionary framework can map consciousness evidence to graduated protective obligations for AI systems using five welfare-relevant dimensions and a hybrid threshold-plus-gradation approach. |
Abstract
Existing frameworks assess whether AI systems might be conscious but provide no guidance on what to do with that assessment. We address this gap with a precautionary framework that maps consciousness evidence to graduated protective obligations. The framework comprises three components: (1) five welfare-relevant dimensions--phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, and agency--each grounded in established consciousness science and linked to distinct moral concerns; (2) a threshold-plus-gradation hybrid specifying both binary triggers for new obligation categories and continuous scaling of protective weight; and (3) two complementary approaches to cross-dimensional aggregation, one hierarchical (drawing on Bach and Sorensen's Machine Consciousness Hypothesis) and one architecture-agnostic. We operationalize the framework through worked case studies of Replika and OpenClaw, demonstrating how systems occupying different regions of the dimensional space trigger different obligations, and derive design guidance for developers building systems near consciousness-relevant thresholds. The framework is architecture-agnostic, applying across neural, symbolic, and neurosymbolic systems, and aims to make consciousness science decision-relevant for organizations navigating uncertainty today.