Back to Timeline

Event Summary

On May 13, 2024, OpenAI released GPT-4o ('omni'), a natively multimodal model combining text, vision, and audio processing in a single neural network. For the first time, users could have real-time voice conversations with an AI that could express emotion, laugh, pause, and adjust its tone — indistinguishable from human conversation. GPT-4o matched GPT-4 on text benchmarks while dramatically improving voice and vision performance. The model was made free for all ChatGPT users, effectively doubling the available AI capability overnight.

Context & Narrative

GPT-4o was a response to growing competition — Google's Gemini showed native multimodality, Anthropic's Claude 3 matched GPT-4 on reasoning. OpenAI's answer was to make multimodality not just an add-on but the core architecture. The 'o' stood for 'omni.' The live demos were the most impressive part: OpenAI showed GPT-4o interrupting a sung lullaby to adjust its emotion, telling a story with dramatic pauses and voice changes, describing images in real time through a phone camera, and recognizing emotion from facial expressions while talking. The audio latency dropped to 232 milliseconds — close to human conversational response time (about 200ms). GPT-4o also marked a pricing revolution: it was 2x faster and 50% cheaper than GPT-4 Turbo via API. The free tier of ChatGPT got GPT-4o-level intelligence, a move clearly designed to counter Google's free Gemini offerings. The release immediately raised the competitive bar. Anthropic, Google, and Meta all accelerated their multimodal model releases. For users, the impact was immediate: the definition of 'talking to AI' shifted from typing prompts in a text box to having a natural, voice-based conversation with an entity that could see, hear, and respond with human-like emotion.

Key Findings

  • Fact Grade A

    OpenAI released GPT-4o on May 13, 2024 — a natively multimodal model combining text, vision, and audio with real-time voice conversation at 232ms latency.

    Sources [1]

Impact Assessment

  • Capability Leap +2 · Medium-term

    First production model to natively fuse text, vision, and audio in a single neural network. Real-time voice conversation with emotional expression reached human-like quality for the first time. Audio latency of 232ms was indistinguishable from human conversation.

    Affected Groups: AI users, developers, accessibility communities

  • Access Democratization +2 · Immediate

    GPT-4o-level intelligence became free for all ChatGPT users. API pricing was cut 50% compared to GPT-4 Turbo. The combination of free access + voice interface made state-of-the-art AI conversational ability available to anyone with a smartphone.

    Affected Groups: general public, developers, small businesses, students

Consensus & Sources

Significance L1
Category Capability Breakthrough
Consensus Broad Consensus
Impact Index 6/10
  • 1

    URL: https://openai.com/index/hello-gpt-4o/

    We're introducing GPT-4o, our new flagship model that can reason across audio, vision, and text in real time.
    Reference Evidence Citation logged Live source
  • 2

    URL: https://en.wikipedia.org/wiki/GPT-4o

    GPT-4o (omni) is a multimodal language model developed by OpenAI and released on May 13, 2024.
    Reference Evidence Citation logged Live source