
74. Ethan Perez - Making AI safe through debate
Towards Data Science
03/10/21
•52m
About
Comments
Featured In
Most AI researchers are confident that we will one day create superintelligent systems — machines that can significantly outperform humans across a wide variety of tasks.
If this ends up happening, it will pose some potentially serious problems. Specifically: if a system is superintelligent, how can we maintain control over it? That’s the core of the AI alignment problem — the problem of aligning advanced AI systems with human values.
A full solution to the alignment problem will have to involve at least two things. First, we’ll have to know exactly what we want superintelligent systems to do, and make sure they don’t misinterpret us when we ask them to do it (the “outer alignment” problem). But second, we’ll have to make sure that those systems are genuinely trying to optimize for what we’ve asked them to do, and that they aren’t trying to deceive us (the “inner alignment” problem).
Creating systems that are inner-aligned and superintelligent might seem like different problems — and many think that they are. But in the last few years, AI researchers have been exploring a new family of strategies that some hope will allow us to achieve both superintelligence and inner alignment at the same time. Today’s guest, Ethan Perez, is using these approaches to build language models that he hopes will form an important part of the superintelligent systems of the future. Ethan has done frontier research at Google, Facebook, and MILA, and is now working full-time on developing learning systems with generalization abilities that could one day exceed those of human beings.
Previous Episode

73. David Roodman - Economic history and the road to the singularity
March 3, 2021
•68m
There’s a minor mystery in economics that may suggest that things are about to get really, really weird for humanity.
And that mystery is this: many economic models predict that, at some point, human economic output will become infinite.
Now, infinities really don’t tend to happen in the real world. But when they’re predicted by otherwise sound theories, they tend to indicate a point at which the assumptions of these theories break down in some fundamental way. Often, that’s because of things like phase transitions: when gases condense or liquids evaporate, some of their thermodynamic parameters go to infinity — not because anything “infinite” is really happening, but because the equations that define a gas cease to apply when those gases become liquids and vice-versa.
So how should we think of economic models that tell us that human economic output will one day reach infinity? Is it reasonable to interpret them as predicting a phase transition in the human economy — and if so, what might that transition look like? These are hard questions to answer, but they’re questions that my guest David Roodman, a Senior Advisor at Open Philanthropy, has thought about a lot.
David has centered his investigations on what he considers to be a plausible culprit for a potential economic phase transition: the rise of transformative AI technology. His work explores a powerful way to think about how, and even when, transformative AI may change how the economy works in a fundamental way.
Next Episode

75. Georg Northoff - Consciousness and AI
March 17, 2021
•59m
For the past decade, progress in AI has mostly been driven by deep learning — a field of research that draws inspiration directly from the structure and function of the human brain. By drawing an analogy between brains and computers, we’ve been able to build computer vision, natural language and other predictive systems that would have been inconceivable just ten years ago.
But analogies work two ways. Now that we have self-driving cars and AI systems that regularly outperform humans at increasingly complex tasks, some are wondering whether reversing the usual approach — and drawing inspiration from AI to inform out approach to neuroscience — might be a promising strategy. This more mathematical approach to neuroscience is exactly what today’s guest, Georg Nortoff, is working on. Georg is a professor of neuroscience, psychiatry, and philosophy at the University of Ottawa, and as part of his work developing a more mathematical foundation for neuroscience, he’s explored a unique and intriguing theory of consciousness that he thinks might serve as a useful framework for developing more advanced AI systems that will benefit human beings.
If you like this episode you’ll love
Promoted




