Facebook has ~10x the proprietary language data on database as was used to train the LLaMa models. In images Facebook have 20x more than that.
Instagram and Youtube have 2x more that in uploaded video.
Tesla’s data capture-ability dwarfs all (at 20x more again) than Youtube.
Size matters
Facebook has ~10x the proprietary language data on database as was used to train the LLaMa models
In images they have 20x more than that
Instagram and Youtube have 2x more that in uploaded video
And yet Tesla's data capture-ability dwarfs all (at 20x more again) pic.twitter.com/SeSRNK6HtM
— Brett Winton (@wintonARK) September 3, 2024

Brian Wang is a Futurist Thought Leader and a popular Science blogger with 1 million readers per month. His blog Nextbigfuture.com is ranked #1 Science News Blog. It covers many disruptive technology and trends including Space, Robotics, Artificial Intelligence, Medicine, Anti-aging Biotechnology, and Nanotechnology.
Known for identifying cutting edge technologies, he is currently a Co-Founder of a startup and fundraiser for high potential early-stage companies. He is the Head of Research for Allocations for deep technology investments and an Angel Investor at Space Angels.
A frequent speaker at corporations, he has been a TEDx speaker, a Singularity University speaker and guest at numerous interviews for radio and podcasts. He is open to public speaking and advising engagements.
Size matters, but Tesla’s data is near useless to train a general intelligence model, since it’s so limited in variation. You cannot get a good generalisation when the data only covers a microscopic part that f reality. It that respect, the data from facebook and YouTube is probably much more varied.