
64. David Krueger - Managing the incentives of AI
Towards Data Science
12/30/20
•50m
About
Comments
Featured In
What does a neural network system want to do?
That might seem like a straightforward question. You might imagine that the answer is “whatever the loss function says it should do.” But when you dig into it, you quickly find that the answer is much more complicated than that might imply.
In order to accomplish their primary goal of optimizing a loss function, algorithms often develop secondary objectives (known as instrumental goals) that are tactically useful for that main goal. For example, a computer vision algorithm designed to tell faces apart might find it beneficial to develop the ability to detect noses with high fidelity. Or in a more extreme case, a very advanced AI might find it useful to monopolize the Earth’s resources in order to accomplish its primary goal — and it’s been suggested that this might actually be the default behavior of powerful AI systems in the future.
So, what does an AI want to do? Optimize its loss function — perhaps. But a sufficiently complex system is likely to also manifest instrumental goals. And if we don’t develop a deep understanding of AI incentives, and reliable strategies to manage those incentives, we may be in for an unpleasant surprise when unexpected and highly strategic behavior emerges from systems with simple and desirable primary goals. Which is why it’s a good thing that my guest today, David Krueger, has been working on exactly that problem. David studies deep learning and AI alignment at MILA, and joined me to discuss his thoughts on AI safety, and his work on managing the incentives of AI systems.
Previous Episode

63. Geordie Rose - Will AGI need to be embodied?
December 23, 2020
•72m
The leap from today’s narrow AI to a more general kind of intelligence seems likely to happen at some point in the next century. But no one knows exactly how: at the moment, AGI remains a significant technical and theoretical challenge, and expert opinion about what it will take to achieve it varies widely. Some think that scaling up existing paradigms — like deep learning and reinforcement learning — will be enough, but others think these approaches are going to fall short.
Geordie Rose is one of them, and his voice is one that’s worth listening to: he has deep experience with hard tech, from founding D-Wave (the world’s first quantum computing company), to building Kindred Systems, a company pioneering applications of reinforcement learning in industry that was recently acquired for $350 million dollars.
Geordie is now focused entirely on AGI. Through his current company, Sanctuary AI, he’s working on an exciting and unusual thesis. At the core of this thesis is the idea is that one of the easiest paths to AGI will be to build embodied systems: AIs with physical structures that can move around in the real world and interact directly with objects. Geordie joined me for this episode of the podcast to discuss his AGI thesis, as well as broader questions about AI safety and AI alignment.
Next Episode

65. Helen Toner - The strategic and security implications of AI
January 6, 2021
•43m
With every new technology comes the potential for abuse. And while AI is clearly starting to deliver an awful lot of value, it’s also creating new systemic vulnerabilities that governments now have to worry about and address. Self-driving cars can be hacked. Speech synthesis can make traditional ways of verifying someone’s identity less reliable. AI can be used to build weapons systems that are less predictable.
As AI technology continues to develop and become more powerful, we’ll have to worry more about safety and security. But competitive pressures risk encouraging companies and countries to focus on capabilities research rather than responsible AI development. Solving this problem will be a big challenge, and it’s probably going to require new national AI policies, and international norms and standards that don’t currently exist.
Helen Toner is Director of Strategy at the Center for Security and Emerging Technology (CSET), a US policy think tank that connects policymakers to experts on the security implications of new technologies like AI. Her work spans national security and technology policy, and international AI competition, and she’s become an expert on AI in China, in particular. Helen joined me for a special AI policy-themed episode of the podcast.
If you like this episode you’ll love
Promoted




