Back to Timeline

Event Summary

On September 12, 2024, OpenAI released o1 (codenamed 'Strawberry'), a large language model trained to spend more time 'thinking' before generating responses — using a chain-of-thought reasoning process internally. o1 scored in the 89th percentile on the International Mathematics Olympiad qualifier and reached PhD-level accuracy on physics, chemistry, and biology benchmarks. It demonstrated that scaling compute at inference time — not just training time — could unlock new capabilities, opening a second axis of AI progress.

Context & Narrative

o1 represented the most significant conceptual shift in LLM architecture since the Transformer itself. The core insight was simple: if you let a model 'think' longer by generating an internal chain-of-thought before answering, it performs dramatically better on complex reasoning tasks. The model was trained using reinforcement learning to improve its chain-of-thought process — effectively learning how to think, not just what to output. The results were eye-opening. On the AIME (American Invitational Mathematics Examination), o1 solved 12/15 problems (vs. 1.8/15 for GPT-4o). On GPQA Diamond (PhD-level science), it exceeded expert accuracy. The model's reasoning ability generalized across chemistry, physics, biology, math, and coding. OpenAI framed this as a new paradigm: scaling inference-time compute (more 'thinking' = better answers) was now an axis of improvement alongside scaling training-time compute (more parameters + more data = better accuracy). The release also came with significant safety implications. o1's chain-of-thought reasoning made it more resistant to jailbreaking — it could 'think through' whether a request was harmful even when cleverly phrased. However, the internal reasoning being hidden from users raised new questions about AI transparency. The model was quickly followed by o1-mini (a smaller, cheaper version), o3 (December 2024, a further improvement in reasoning capability), and Google's Gemini 2.0 Flash Thinking. o1 established 'reasoning models' as a distinct category that all major AI labs would race to compete in throughout 2024 and 2025.

Key Findings

  • Fact Grade A

    OpenAI released o1 on September 12, 2024 — a reasoning model that uses internal chain-of-thought to spend more time 'thinking' before answering, achieving PhD-level accuracy on science benchmarks.

    Sources [1]
  • Impact Grade A

    o1 established test-time compute scaling as a second axis of AI progress alongside training-time scaling, creating the 'reasoning model' category that all major AI labs would adopt.

    Sources [1][2]

Impact Assessment

  • Capability Leap +3 · Long-term

    Introduced test-time compute scaling as a new axis of AI progress. o1 achieved PhD-level accuracy on science benchmarks (GPQA Diamond) and solved 12/15 AIME math problems vs. GPT-4o's 1.8/15. The chain-of-thought reasoning approach generalized across chemistry, physics, biology, math, and coding — matching or exceeding expert human performance on hard reasoning tasks.

    Affected Groups: AI researchers, mathematicians, scientists, software engineers

  • Paradigm Shift +3 · Long-term

    Changed the AI field's understanding of scaling from a single axis (training compute) to two axes (training + inference compute). The concept of 'reasoning models' became a new category. Labs including Google, Anthropic, and DeepSeek all launched their own reasoning models in response. The idea that giving a model more time to 'think' improves performance had profound implications for AI system design, pricing, and capability forecasting.

    Affected Groups: entire AI field, researchers, AI companies, policymakers

  • Risk Creation -1 · Medium-term

    Hidden chain-of-thought reasoning raised transparency concerns — users could see the model's output but not its reasoning process. The model's improved resistance to jailbreaking was positive for safety, but the potential for deceptive reasoning (the model could 'think' harmful things without revealing them) created new AI governance challenges.

    Affected Groups: AI safety researchers, policymakers, ethicists, general public

Consensus & Sources

Significance L2
Category Capability Breakthrough
Consensus Broad Consensus
Impact Index 7/10
  • 1

    URL: https://openai.com/index/introducing-openai-o1-preview/

    We've developed a new series of AI models designed to spend more time thinking before they respond. They can reason through complex tasks and solve harder problems than previous models.
    Reference Evidence Citation logged Live source
  • 2

    URL: https://www.theverge.com/2024/9/12/24242439/openai-o1-model-reasoning-strawberry-chatgpt

    OpenAI releases o1, its first model with 'reasoning' abilities.
    News Report Citation logged Live source