Research Leaders Urge Tech Industry to Oversee AI’s ‘Thought Processes’
A group of AI researchers from OpenAI, Google DeepMind, Anthropic, and various companies and nonprofit organizations are calling for a more in-depth examination of the techniques for observing the reasoning processes of AI models in a position paper released on Tuesday.
AI reasoning models, like OpenAI’s o3 and DeepSeek’s R1, utilize chains-of-thought (CoTs) — an externalized method wherein AI models systematically address problems, similar to how humans might take notes to solve intricate math problems. The authors of the paper argue that monitoring CoTs could be crucial for overseeing AI agents as they evolve and become more widespread.
“CoT monitoring is an essential addition to safety measures for advanced AI, offering valuable insights into the decision-making processes of AI agents,” the researchers noted in the position paper. “However, there’s no guarantee that the current level of insight will be sustainable. We encourage the research community and leading AI developers to leverage the potential of CoT monitorability and explore ways to maintain it.”
The position paper urges prominent AI developers to investigate what characteristics make CoTs “monitorable,” meaning they should pinpoint elements that can either improve or hinder transparency regarding how AI models derive their conclusions. The authors believe that CoT monitoring could be fundamental for understanding AI reasoning models, but they warn that it could be risky, cautioning against any actions that might compromise their clarity or consistency.
Furthermore, the authors advocate for AI developers to track CoT monitorability and explore its potential as a safety mechanism.
Notable signatories of the paper include Mark Chen, chief research officer at OpenAI; Ilya Sutskever, CEO of Safe Superintelligence; Nobel laureate Geoffrey Hinton; Shane Legg, co-founder of Google DeepMind; Dan Hendrycks, safety adviser at xAI; and John Schulman, co-founder of Thinking Machines. The primary authors consist of leaders from the U.K. AI Security Institute and Apollo Research, along with additional signatories from METR, Amazon, Meta, and UC Berkeley.
This paper represents a collective initiative among many leaders in the AI field to advance research on AI safety, coinciding with a period of fierce competition in the tech industry. This competitive environment has prompted companies like Meta to attract top talent from OpenAI, Google DeepMind, and Anthropic with lucrative offers. Many of the desired researchers are those who focus on AI agents and reasoning models.
“We stand at a pivotal moment with this new chain-of-thought concept. It seems quite advantageous, but it could diminish if not given adequate attention in the coming years,” remarked Bowen Baker, an OpenAI researcher involved with the paper, in a TechCrunch interview. “The release of a position paper like this is intended to raise more research interest in this area before it’s too late.”
In September 2024, OpenAI publicly showcased a preview of its first AI reasoning model, o1. Since then, the tech industry swiftly unveiled competitors displaying similar capabilities, with some models from Google DeepMind, xAI, and Anthropic achieving even higher performance on benchmarks.
Despite these advancements, there remains a considerable knowledge gap regarding the operational mechanisms of AI reasoning models. While AI labs have made progress in enhancing AI performance over the last year, this improvement hasn’t necessarily resulted in a better understanding of how these models arrive at their conclusions.
Anthropic has emerged as a significant contributor to understanding the mechanisms of AI models — a field known as interpretability. Earlier this year, CEO Dario Amodei announced a commitment to unraveling the AI black box by 2027 and increasing investments in interpretability. He also urged OpenAI and Google DeepMind to focus more on this research area.
Preliminary research from Anthropic indicates that CoTs may not consistently reflect how these models generate answers. Simultaneously, OpenAI researchers have suggested that CoT monitoring could eventually become a reliable method for evaluating alignment and safety in AI models.
The goal of position papers like this one is to raise awareness and attract more focus to emerging research areas, such as CoT monitoring. Companies including OpenAI, Google DeepMind, and Anthropic are already probing these topics, but this paper could inspire increased funding and research in the field.


