Skip to content

When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty

Anna Mikeda

arXiv (Cornell University) June 4, 2026 DOI: 10.48550/arxiv.2606.05528 via OpenAlex

Summary

AI-generated from the abstract

A new precautionary framework translates evidence about whether AI systems might be conscious into graduated protective obligations. The framework maps five welfare-relevant dimensions—phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, and agency—each linked to distinct moral concerns. It uses a hybrid threshold-plus-gradation approach: binary triggers activate new obligation categories, while continuous scaling adjusts protective weight. Two complementary aggregation methods are offered: one hierarchical, one architecture-agnostic. Worked case studies of Replika and OpenClaw show how systems in different dimensional regions trigger different obligations. The framework applies across neural, symbolic, and neurosymbolic systems, aiming to make consciousness science decision-relevant for organizations.

Study at a glance

Characteristics Theoretical or philosophical paper Peer reviewed
Keywords Operationalization Consciousness Metacognition Conceptual framework Precautionary principle
Key finding A precautionary framework can map consciousness evidence to graduated protective obligations for AI systems using five welfare-relevant dimensions and a hybrid threshold-plus-gradation approach.

Abstract

Existing frameworks assess whether AI systems might be conscious but provide no guidance on what to do with that assessment. We address this gap with a precautionary framework that maps consciousness evidence to graduated protective obligations. The framework comprises three components: (1) five welfare-relevant dimensions--phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, and agency--each grounded in established consciousness science and linked to distinct moral concerns; (2) a threshold-plus-gradation hybrid specifying both binary triggers for new obligation categories and continuous scaling of protective weight; and (3) two complementary approaches to cross-dimensional aggregation, one hierarchical (drawing on Bach and Sorensen's Machine Consciousness Hypothesis) and one architecture-agnostic. We operationalize the framework through worked case studies of Replika and OpenClaw, demonstrating how systems occupying different regions of the dimensional space trigger different obligations, and derive design guidance for developers building systems near consciousness-relevant thresholds. The framework is architecture-agnostic, applying across neural, symbolic, and neurosymbolic systems, and aims to make consciousness science decision-relevant for organizations navigating uncertainty today.

Comments

No comments yet.

Log in to comment