
54. Tim Rocktäschel - Deep reinforcement learning, symbolic learning and the road to AGI
Towards Data Science
10/15/20
•53m
About
Comments
Featured In
Reinforcement learning can do some pretty impressive things. It can optimize ad targeting, help run self-driving cars, and even win StarCraft games. But current RL systems are still highly task-specific. Tesla’s self-driving car algorithm can’t win at StarCraft, and DeepMind’s AlphaZero algorithm can with Go matches against grandmasters, but can’t optimize your company’s ad spend.
So how do we make the leap from narrow AI systems that leverage reinforcement learning to solve specific problems, to more general systems that can orient themselves in the world? Enter Tim Rocktäschel, a Research Scientist at Facebook AI Research London and a Lecturer in the Department of Computer Science at University College London. Much of Tim’s work has been focused on ways to make RL agents learn with relatively little data, using strategies known as sample efficient learning, in the hopes of improving their ability to solve more general problems. Tim joined me for this episode of the podcast.
Previous Episode

53. Edouard Harris - Emerging problems in machine learning: making AI "good"
October 8, 2020
•66m
Where do we want our technology to lead us? How are we falling short of that target? What risks might advanced AI systems pose to us in the future, and what potential do they hold? And what does it mean to build ethical, safe, interpretable, and accountable AI that’s aligned with human values?
That’s what this year is going to be about for the Towards Data Science podcast. I hope you join us for that journey, which starts today with an interview with my brother Ed, who apart from being a colleague who’s worked with me as part of a small team to build the SharpestMinds data science mentorship program, is also collaborating with me on a number of AI safety, alignment and policy projects. I thought he’d be a perfect guest to kick off this new year for the podcast.
Next Episode

If you walked into a room filled with objects that were scattered around somewhat randomly, how important or expensive would you assume those objects were?
What if you walked into the same room, and instead found those objects carefully arranged in a very specific configuration that was unlikely to happen by chance?
These two scenarios hint at something important: human beings have shaped our environments in ways that reflect what we value. You might just learn more about what I value by taking a 10 minute stroll through my apartment than by spending 30 minutes talking to me as I try to put my life philosophy into words.
And that’s a pretty important idea, because as it turns out, one of the most important challenges in advanced AI today is finding ways to communicate our values to machines. If our environments implicitly encode part of our value system, then we might be able to teach machines to observe it, and learn about our preferences without our having to express them explicitly.
The idea of leveraging deriving human values from the state of an human-inhabited environment was first developed in a paper co-authored by Berkeley PhD and incoming DeepMind researcher Rohin Shah. Rohin has spent the last several years working on AI safety, and publishes the widely read AI alignment newsletter — and he was kind enough to join us for this episode of the Towards Data Science podcast, where we discussed his approach to AI safety, and his thoughts on risk mitigation strategies for advanced AI systems.
If you like this episode you’ll love
Promoted




