
In its July 2025 deep dive on Meta's AI troubles, the research firm SemiAnalysis reported that unlike all other leading AI labs, including OpenAI and DeepSeek, Meta does not use YouTube data, and suggested this may be why Meta struggled to build a strong multimodal model (one that understands images, video and sound, not just text). Google, which owns YouTube, confirmed to CNBC in June 2025 that it trains its Gemini models and its Veo 3 video generator on a subset of YouTube's roughly 20 billion videos. The New York Times reported in 2024 that OpenAI transcribed more than a million hours of YouTube video to help train GPT-4, and a 2024 Proof News investigation found subtitles from 173,536 YouTube videos inside a dataset used by Apple, Nvidia, Anthropic and others. YouTube says scraping its videos breaks its rules, so this is legally and ethically contested ground. The lesson for everyone else: in AI, data you own outright is a moat that money cannot instantly buy, which is partly why Meta paid around $14 billion for 49% of the data company Scale AI.
What happened
When people explain who is winning in AI, they usually talk about chips and talent. Who has the most Nvidia GPUs (the graphics chips AI runs on)? Who poached the best researchers? Those matter. But buried in a long SemiAnalysis report on Meta's AI struggles is a quieter reason one of the richest companies in the world fell behind: it did not have the right homework to learn from.
SemiAnalysis, an independent research firm that tracks the AI industry closely, wrote that "unlike all other leading AI labs including OpenAI and Deepseek, Meta does not utilize YouTube data". It added that YouTube lecture transcripts and other videos are an incredible source of data, and that Meta "may have struggled to produce a multimodal model without the data". Multimodal means a model that understands more than text: pictures, video and sound too.
The facts so far
Start with who has the video. Google owns YouTube, and in June 2025 it confirmed to CNBC that it trains its Gemini models and its Veo 3 video generator on YouTube content. Google said it uses only a subset of the roughly 20 billion videos on the platform and honours its agreements with creators and media companies. Creators CNBC spoke to said they had not known, and there is no way for them to opt out of Google's own training.
Others got at it from outside. In April 2024 The New York Times reported that OpenAI built its speech-to-text tool Whisper partly to transcribe more than a million hours of YouTube video into text for training GPT-4. In July 2024 a Proof News investigation, co-published with Wired, found a dataset called YouTube Subtitles, holding transcripts from 173,536 videos across more than 48,000 channels, used by companies including Apple, Nvidia and Anthropic. YouTube's chief executive Neal Mohan has said that using its videos to train AI without permission would be a "clear violation" of its rules.
Meta, by SemiAnalysis's account, stayed out. It relied on public web text, then switched partway through training its giant Llama 4 Behemoth model to a web crawler it had built itself, and struggled to clean the new data. Behemoth was never released.
"YouTube lecture transcripts and other videos are an incredible source for data." (SemiAnalysis)
Why video is so valuable
Text tells a model what people say about the world. Video shows it what the world actually does: how a glass tips over, how a hand ties a knot, how a voice rises when someone is surprised, how a lecturer walks through a proof on a whiteboard. If you want an AI that can watch, listen and generate realistic footage, you need an enormous amount of real footage to learn from.
And YouTube is unusually good footage. Much of it is people explaining things on purpose, with speech that lines up with what is on screen, plus titles, chapters and captions that act as free labels. That combination of picture, sound and description is exactly what multimodal models are hungry for.
The everyday version
Imagine two cooking students with the same talent and the same expensive kitchen. One has spent years watching thousands of chefs on video, seeing every chop and every pan flip. The other has only read cookbooks. Give them both a new dish and the first one will move like a cook; the second will know the words but fumble the knife. Meta built a world-class kitchen. SemiAnalysis's point is that it was learning mostly from the cookbooks.
Is this actually new?
Data moats are an old idea. Google search got better because it saw more searches than anyone else, and that loop was hard to copy. What has changed is that the most valuable data is no longer only clicks and text but video, and video is far harder to collect at scale without owning a platform.
It also explains a few moves that looked odd at the time. In June 2025 Meta paid around $14 billion for 49% of Scale AI, a company that organises and labels training data, and hired its chief executive, Alexandr Wang. SemiAnalysis read that as a direct attempt to fix Meta's data problems, rather than a consolation prize.
What it means
Two honest caveats. First, this is one well-sourced analyst report, not something Meta has confirmed, and it describes Meta in mid-2025; Meta may have changed course since. Second, scraping YouTube is contested ground. Some of what gave rivals an edge may have broken YouTube's rules, and lawsuits over training data are still working their way through courts.
The bigger lesson holds either way. In AI, money can buy chips within months and buy researchers within weeks, but a decade of the world filming itself is not for sale. The companies that own that kind of data, Google above all, have an advantage that rarely makes the headlines. Next time someone ranks the AI race only by chips and salaries, ask who owns the video.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.