OTHER

Discovering OpenAI’s Mission: Enabling AI to Support Your Success

Soon after Hunter Lightman began his research position at OpenAI in 2022, he witnessed the rapid rise of ChatGPT, which quickly became one of the most popular products ever. At the same time, Lightman was contributing to a team dedicated to enhancing OpenAI’s models for high school math competitions.

This team, known as MathGen, plays a crucial role in OpenAI’s mission to develop advanced AI reasoning models, the technology that allows AI agents to perform tasks akin to human capabilities.

“Our aim was to enhance the models’ mathematical reasoning skills, which were quite limited back then,” Lightman remarked in a TechCrunch interview regarding MathGen’s early efforts.

While OpenAI’s models still encounter difficulties—frequently producing incorrect information and sometimes struggling with intricate tasks—they have made notable advancements in mathematical reasoning. Recently, one of OpenAI’s models secured a gold medal at the International Math Olympiad, a prestigious event for top high school students. The company anticipates that these reasoning skills will extend to various subjects, ultimately paving the way for the development of general-purpose agents.

The introduction of ChatGPT was serendipitous—initially just a research preview that unexpectedly went viral—but the advancement of OpenAI’s agents is the product of years of extensive internal research.

“Before long, you’ll simply ask the computer for what you need, and it will handle all those tasks for you,” stated OpenAI CEO Sam Altman at the company’s inaugural developer conference in 2023. “In AI, these capabilities are often referred to as agents, and their benefits will be enormous.”

TechCrunch event

San Francisco
|
October 27-29, 2025

OpenAI CEO Sam Altman discusses developments during the OpenAI DevDay event on November 06, 2023, in San Francisco, California.
OpenAI CEO Sam Altman discusses developments during the OpenAI DevDay event on November 06, 2023, in San Francisco, California. (Photo by Justin Sullivan/Getty Images)Image Credits:Justin Sullivan / Getty Images

Although Altman’s future objectives are still uncertain, OpenAI stirred excitement in the tech community with the unveiling of its groundbreaking AI reasoning model, o1, in the fall of 2024. Within a year, the 21 innovative researchers driving this project became highly sought after assets in Silicon Valley.

Mark Zuckerberg successfully recruited five members from the o1 team for Meta’s new superintelligence division, reportedly offering over $100 million in compensation. One of these researchers, Shengjia Zhao, presently fills the role of chief scientist at Meta Superintelligence Labs.

The Reinforcement Learning Renaissance

The origin of OpenAI’s reasoning models and agents is linked to a machine learning method known as reinforcement learning (RL), which provides feedback to AI models regarding the accuracy of their decisions in simulated environments.

RL has a noteworthy history; for example, shortly after OpenAI’s founding in 2016, Google’s DeepMind launched AlphaGo, an AI that captivated public interest by defeating a Go champion.

Lee Se-Dol, a professional Go player from South Korea, prepares for his fourth match against Google’s AlphaGo during the Google DeepMind Challenge Match on March 13, 2016, in Seoul, South Korea. Lee competed in a five-game match against the computer program AlphaGo. (Photo by Google via Getty Images)

During this period, one of OpenAI’s early employees, Andrej Karpathy, began exploring the potential of using RL to develop an AI agent capable of controlling a computer. However, the creation of the necessary models and training frameworks took several years.

By 2018, OpenAI had released its first large language model in the GPT series, trained on extensive datasets from the internet using powerful GPU clusters. While GPT models excelled in generating text, they struggled with basic math.

The breakthrough transpired in 2023 with a model initially branded “Q*,” later renamed “Strawberry,” which combined LLMs, RL, and a novel technique called test-time computation. This innovation provided models with additional planning and verification time prior to reaching a conclusion.

This advancement birthed a new methodology termed “chain-of-thought” (CoT), which greatly enhanced AI performance on unfamiliar mathematical challenges.

“I noticed the model beginning to reason,” El Kishky commented. “It became aware of errors and retraced its steps, displaying frustrations reminiscent of human emotions. It truly mirrored a human cognitive process.”

While none of these isolated techniques were revolutionary, OpenAI successfully integrated them to create Strawberry, which directly contributed to o1’s development. The company quickly acknowledged that the planning and verification strengths of AI reasoning models could empower AI agents.

“We confronted a challenge I had grappled with for years,” Lightman conveyed. “That was among the most exhilarating moments of my research career.”

Scaling Reasoning

With the advent of AI reasoning models, OpenAI identified two new strategies to improve its AI capabilities: deploying more computational resources during the post-training phase and allowing models more time and resources when tackling queries.

“OpenAI is committed not only to the present landscape but also to exploring avenues for expansion,” Lightman asserted.

Shortly following the 2023 Strawberry milestone, OpenAI established an “Agents” team, led by researcher Daniel Selsam, to advance this new paradigm, according to insights from two sources who spoke with TechCrunch. Initially, the team blurred the distinctions between reasoning models and the agents of today; their goal was to develop AI systems capable of handling intricate tasks.

Ultimately, the efforts of Selsam’s Agents team contributed to the broader initiative to shape the o1 reasoning model, involving pivotal figures such as OpenAI co-founder Ilya Sutskever, chief research officer Mark Chen, and chief scientist Jakub Pachocki.

Ilya Sutskever, co-founder and Chief Scientist of OpenAI.
Ilya Sutskever, co-founder and Chief Scientist of OpenAI, speaks at Tel Aviv University on June 5, 2023. (Photo by JACK GUEZ / AFP)Image Credits:Getty Images

Creating o1 required OpenAI to invest significant resources, particularly concerning talent and GPUs. Throughout this phase, researchers frequently had to negotiate for resource allocation with leadership; demonstrating breakthroughs was an effective means of securing necessary support.

“A fundamental principle at OpenAI is our collaborative research approach,” Lightman stressed. “When we presented our findings regarding o1, the organization responded favorably, saying, ‘This makes sense; let’s explore further.’”

Some former employees assert that OpenAI’s ambition for AGI was pivotal in advancing AI reasoning models. By emphasizing the development of the most complex AI rather than just focusing on products, OpenAI could invest in creating o1—a luxury that many competing labs couldn’t afford.

This decision to investigate innovative training methods proved beneficial. By late 2024, numerous AI labs began facing the limitations of traditional pretraining scaling. Presently, much of the progress in the AI field originates from advancements in reasoning models.

What Does It Mean for an AI to “Reason?”

AI research strives to replicate human intelligence within machines through various methods. Since the launch of o1, the user experience of ChatGPT has incorporated increasingly human-like attributes such as “thinking” and “reasoning.”

When queried about whether OpenAI’s models genuinely demonstrate reasoning capabilities, El Kishky took a cautious stance, clarifying that he views the concept through a computer science lens.

“We train the model to efficiently utilize computing resources to arrive at answers. Therefore, if you frame reasoning in that context, then yes, it exhibits reasoning,” El Kishky explained.

Lightman adopts a more results-oriented perspective, focusing on the outcomes produced by the model rather than its methods or their connection to human thought processes.

The OpenAI logo displayed on their developer day stage.
The OpenAI logo displayed on their developer day stage. (Credit: Devin Coldeway)Image Credits:Devin Coldewey

“If the model excels at difficult questions, then it is employing whatever reasoning is necessary for that situation,” Lightman pointed out. “We may classify it as reasoning because it reflects reasoning patterns, but ultimately, it’s about developing powerful and effective AI tools for users.”

OpenAI researchers recognize that definitions and terms surrounding reasoning can differ, and while many critics have arisen, they believe the capabilities of their models take precedence over semantic discussions. Other AI experts share this sentiment.

Nathan Lambert, an AI researcher at the non-profit AI2, likened AI reasoning models to airplanes in a blog post, indicating that while both are human-designed systems inspired by nature, they function through entirely different mechanisms. This distinction does not diminish their utility or capability to yield comparable outcomes.

Recently, a coalition of AI researchers from OpenAI, Anthropic, and Google DeepMind issued a position paper acknowledging that AI reasoning models remain only partially understood and necessitate further exploration. It may be premature to make definitive conclusions about their internal processes.

The Next Frontier: AI Agents for Subjective Tasks

Currently, AI agents excel in clearly defined and easily verifiable areas, such as programming. OpenAI’s Codex agent is tailored to aid software developers with straightforward programming tasks. Likewise, Anthropic’s models have made strides in AI coding tools like Cursor and Claude Code—innovative AI agents that attract eager users.

However, general-purpose AI agents like OpenAI’s ChatGPT Agent and Perplexity’s Comet struggle to automate many complex and subjective tasks that users wish to simplify. During attempts to use these tools for online shopping or securing long-term parking, I found they frequently took longer than anticipated and committed basic errors.

These agents are still in their nascent stages and are certain to evolve. Nevertheless, researchers must first pinpoint more effective training methodologies for foundational models in order to effectively address subjective tasks.

AI applications (Photo by Jonathan Raa/NurPhoto via Getty Images)

“As with many challenges in machine learning, this ultimately boils down to a data issue,” Lightman noted concerning the limitations of agents in subjective tasks. “Some research I’m currently enthusiastic about involves discovering training techniques for less verifiable tasks. We’ve seen some promising advancements in this area.”

Noam Brown, an OpenAI researcher involved in the IMO model and o1, informed TechCrunch that the organization has made headway in innovative general-purpose RL methods, enabling them to train AI models in ways that are not easily verifiable. This methodology was crucial for the model that achieved gold at the IMO.

OpenAI’s IMO model illustrates a more sophisticated AI system that employs multiple agents exploring various concepts concurrently, ultimately selecting the most effective response. This model architecture is becoming more prevalent, with Google and xAI recently launching advanced models that adopt this strategy.

“I believe these models will not only refine their mathematical capabilities but will also boost performance across other reasoning domains,” Brown affirmed. “The rate of progress has been astounding, and I see no reason for it to slow down.”

These advancements could empower OpenAI’s models to achieve remarkable outcomes, with expectations set for the forthcoming GPT-5 model. OpenAI aims to establish a competitive advantage over rivals by introducing GPT-5, which they believe will deliver premier AI models that energize agents for developers and users alike.

Additionally, the company plans to enhance its products for users. El Kishky emphasized that OpenAI aspires to develop AI agents that inherently grasp user intentions, alleviating the need for specific configuration choices. He envisions systems that know when to activate certain tools and how long to deliberate.

These ambitions outline a vision for the ultimate iteration of ChatGPT: an agent capable of executing any online task on your behalf while understanding your preferences. This idea significantly departs from ChatGPT’s current capacities, yet the company’s research is steadfastly oriented toward this aim.

While OpenAI has undoubtedly spearheaded the AI industry in recent years, the organization now grapples with intense competition. The focus has shifted from whether OpenAI will achieve its envisioned agentic future to whether it can do so ahead of competitors like Google, Anthropic, xAI, or Meta.