
106. Yang Gao - Sample-efficient AI
Towards Data Science
12/08/21
•49m
About
Comments
Featured In
Historically, AI systems have been slow learners. For example, a computer vision model often needs to see tens of thousands of hand-written digits before it can tell a 1 apart from a 3. Even game-playing AIs like DeepMind’s AlphaGo, or its more recent descendant MuZero, need far more experience than humans do to master a given game.
So when someone develops an algorithm that can reach human-level performance at anything as fast as a human can, it’s a big deal. And that’s exactly why I asked Yang Gao to join me on this episode of the podcast. Yang is an AI researcher with affiliations at Berkeley and Tsinghua University, who recently co-authored a paper introducing EfficientZero: a reinforcement learning system that learned to play Atari games at the human-level after just two hours of in-game experience. It’s a tremendous breakthrough in sample-efficiency, and a major milestone in the development of more general and flexible AI systems.
---
Intro music:
➞ Artist: Ron Gelinas
➞ Track Title: Daybreak Chill Blend (original mix)
➞ Link to Track: https://youtu.be/d8Y2sKIgFWc
---
Chapters:
0:00 Intro
1:50 Yang’s background
6:00 MuZero’s activity
13:25 MuZero to EfficiantZero
19:00 Sample efficiency comparison
23:40 Leveraging algorithmic tweaks
27:10 Importance of evolution to human brains and AI systems
35:10 Human-level sample efficiency
38:28 Existential risk from AI in China
47:30 Evolution and language
49:40 Wrap-up
Previous Episode

105. Yannic Kilcher - A 10,000-foot view of AI
December 1, 2021
•63m
There once was a time when AI researchers could expect to read every new paper published in the field on the arXiv, but today, that’s no longer the case. The recent explosion of research activity in AI has turned keeping up to date with new developments into a full-time job.
Fortunately, people like YouTuber, ML PhD and sunglasses enthusiast Yannic Kilcher make it their business to distill ML news and papers into a digestible form for mortals like you and me to consume. I highly recommend his channel to any TDS podcast listeners who are interested in ML research — it’s a fantastic resource, and literally the way I finally managed to understand the Attention is All You Need paper back in the day.
Yannic is joined me to talk about what he’s learned from years of following, reporting and doing AI research, including the trends, the challenges and the opportunities that he expects are going to shape the course of AI history in coming years.
---
Intro music:
➞ Artist: Ron Gelinas
➞ Track Title: Daybreak Chill Blend (original mix)
➞ Link to Track: https://youtu.be/d8Y2sKIgFWc
---
Chapters:
0:00 Intro
1:20 Yannic’s path into ML
7:25 Selecting ML news
11:45 AI ethics → political discourse
17:30 AI alignment
24:15 Malicious uses
32:10 Impacts on persona
39:50 Bringing in human thought
46:45 Math with big numbers
51:05 Metrics for generalization
58:05 The future of AI
1:02:58 Wrap-up
Next Episode

107. Kevin Hu - Data observability and why it matters
December 15, 2021
•49m
Imagine for a minute that you’re running a profitable business, and that part of your sales strategy is to send the occasional mass email to people who’ve signed up to be on your mailing list. For a while, this approach leads to a reliable flow of new sales, but then one day, that abruptly stops. What happened?
You pour over logs, looking for an explanation, but it turns out that the problem wasn’t with your software; it was with your data. Maybe the new intern accidentally added a character to every email address in your dataset, or shuffled the names on your mailing list so that Christina got a message addressed to “John”, or vice-versa. Versions of this story happen surprisingly often, and when they happen, the cost can be significant: lost revenue, disappointed customers, or worse — an irreversible loss of trust.
Today, entire products are being built on top of datasets that aren’t monitored properly for critical failures — and an increasing number of those products are operating in high-stakes situations. That’s why data observability is so important: the ability to track the origin, transformations and characteristics of mission-critical data to detect problems before they lead to downstream harm.
And it’s also why we’ll be talking to Kevin Hu, the co-founder and CEO of Metaplane, one of the world’s first data observability startups. Kevin has a deep understanding of data pipelines, and the problems that cap pop up if you they aren’t properly monitored. He joined me to talk about data observability, why it matters, and how it might be connected to responsible AI on this episode of the TDS podcast.
Intro music:
➞ Artist: Ron Gelinas
➞ Track Title: Daybreak Chill Blend (original mix)
➞ Link to Track: https://youtu.be/d8Y2sKIgFWc 0:00
Chapters:
- 0:00 Intro
- 2:00 What is data observability?
- 8:20 Difference between a dataset’s internal and external characteristics
- 12:20 Why is data so difficult to log?
- 17:15 Tracing back models
- 22:00 Algorithmic analyzation of a date
- 26:30 Data ops in five years
- 33:20 Relation to cutting-edge AI work
- 39:25 Software engineering and startup funding
- 42:05 Problems on a smaller scale
- 46:40 Future data ops problems to solve
- 48:45 Wrap-up
If you like this episode you’ll love
Promoted




