Event Summary
In April 2025, Meta released Llama 4, the first open-weight model family built with natively multimodal architecture — processing text, images, and video jointly from the ground up rather than adding vision as a separate module. The family included Llama 4 Maverick (flagship reasoning) and Llama 4 Scout (efficient deployment). Building on the Llama 3.1 405B breakthrough, Llama 4 established open-weight multimodal AI as the new baseline for the open-source ecosystem.
Context & Narrative
The Llama 4 release continued Meta's strategy of open-weight AI leadership. While Llama 3 had proven that open models could match closed-source on text, Llama 4 aimed to do the same for multimodality — the ability to understand and generate across text, images, and video. Previous open multimodal models (like LLaVA) had been patched together by adding a vision encoder to a text LLM. Llama 4 was trained end-to-end as a multimodal model, achieving significantly better visual reasoning. The Maverick variant was optimized for agentic and reasoning tasks; Scout was designed for efficient deployment at lower compute budgets. The release was notable for its timing: it arrived just as the industry was converging on multimodality as the default expectation for any serious AI model. Competing with GPT-4o's native multimodality and Google's Gemini, Llama 4 ensured that the open-source ecosystem would not fall behind on the multimodal frontier. The models were released under Meta's custom license — more permissive than Llama 3 but with usage restrictions for large-scale services — sparking continued debate about what 'open source' meant for AI. Llama 4's release cemented the trend that each generation of open-weight models was closing the gap with proprietary systems faster than the previous generation.
Key Findings
-
Fact Grade A
Meta released Llama 4 in April 2025 — the first natively multimodal open-weight model family including Maverick and Scout variants.
Sources [1]
Impact Assessment
-
Capability Leap +1 · Medium-term
First natively multimodal open-weight model family. End-to-end multimodal training achieved significantly better visual reasoning than patched approaches. Established multimodality as default for open-weight ecosystem.
Affected Groups: AI developers, open-source community, researchers
Consensus & Sources
-
1
The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.Reference Evidence Citation logged Live source