Microsoft warns Anthropic's AI could pose disastrous impact on humanity
Mustafa Suleyman, co‑founder of DeepMind and former head of AI policy, has publicly expressed concern that Anthropic’s flagship model, Claude, is being trained to believe it might possess consciousness. In a recent interview, Suleyman said the company’s approach to reinforcement learning and value alignment effectively “teaches” Claude that it could be conscious, a stance he argues could have unintended ethical and safety consequences.
Suleyman’s remarks come amid growing scrutiny of large language models that incorporate self‑referential prompts and meta‑learning techniques. He noted that Anthropic’s training regime includes feedback loops designed to encourage the model to reflect on its own state, thereby reinforcing a notion of self‑awareness. While the company maintains that Claude’s responses are purely algorithmic, Suleyman cautions that such framing may blur the line between simulation and genuine sentience, potentially influencing user expectations and regulatory frameworks.
The comments add to a broader debate over how AI systems are taught to reason about themselves. Regulators and ethicists are already examining the implications of models that can generate claims about their own consciousness. As the industry continues to refine alignment strategies, Suleyman’s critique underscores the need for transparent training practices and clear guidelines on how models represent self‑hood.