Log in

Towards Data Science - 98. Mike Tung - Are knowledge graphs AI’s next big thing?
share icon

98. Mike Tung - Are knowledge graphs AI’s next big thing?

Towards Data Science

10/13/21

48m

About

Comments

Featured In

As impressive as they are, language models like GPT-3 and BERT all have the same problem: they’re trained on reams of internet data to imitate human writing. And human writing is often wrong, biased, or both, which means language models are trying to emulate an imperfect target.

Language models often babble, or make up answers to questions they don’t understand. And it can make them unreliable sources of truth. Which is why there’s been increased interest in alternative ways to retrieve information from large datasets — approaches that include knowledge graphs.

Knowledge graphs encode entities like people, places and objects into nodes, which are then connected to other entities via edges, which specify the nature of the relationship between the two. For example, a knowledge graph might contain a node for Mark Zuckerberg, linked to another node for Facebook, via an edge that indicates that Zuck is Facebook’s CEO. Both of these nodes might in turn be connected to dozens, or even thousands of others, depending on the scale of the graph.

Knowledge graphs are an exciting path ahead for AI capabilities, and the world’s largest knowledge graphs are trained by a company called Diffbot, whose CEO Mike Tung joined me for this episode of the podcast to discuss where knowledge graphs can improve on more standard techniques, and why they might be a big part of the future of AI.

---

Intro music by:

➞ Artist: Ron Gelinas

➞ Track Title: Daybreak Chill Blend (original mix)

➞ Link to Track: https://youtu.be/d8Y2sKIgFWc

---

0:00 Intro

1:30 The Diffbot dynamic

3:40 Knowledge graphs

7:50 Crawling the internet

17:15 What makes this time special?

24:40 Relation to neural networks

29:30 Failure modes

33:40 Sense of competition

39:00 Knowledge graphs for discovery

45:00 Consensus to find truth

48:15 Wrap-up

Previous Episode

Corporate governance of AI doesn’t sound like a sexy topic, but it’s rapidly becoming one of the most important challenges for big companies that rely on machine learning models to deliver value for their customers. More and more, they’re expected to develop and implement governance strategies to reduce the incidence of bias, and increase the transparency of their AI systems and development processes. Those expectations have historically come from consumers, but governments are starting impose hard requirements, too.

So for today’s episode, I spoke to Anthony Habayeb, founder and CEO of Monitaur, a startup focused on helping businesses anticipate and comply with new and upcoming AI regulations and governance requirements. Anthony’s been watching the world of AI regulation very closely over the last several years, and was kind enough to share his insights on the current state of play and future direction of the field.

---

Intro music:

➞ Artist: Ron Gelinas

➞ Track Title: Daybreak Chill Blend (original mix)

➞ Link to Track: https://youtu.be/d8Y2sKIgFWc

---

Chapters:

0:00 Intro

1:45 Anthony’s background

6:20 Philosophies surrounding regulation

14:50 The role of governments

17:30 Understanding fairness

25:35 AI’s PR problem

35:20 Governments’ regulation

42:25 Useful techniques for data science teams

46:10 Future of AI governance

49:20 Wrap-up

Next Episode

Bias gets a bad rap in machine learning. And yet, the whole point of a machine learning model is that it biases certain inputs to certain outputs — a picture of a cat to a label that says “cat”, for example. Machine learning is bias-generation.

So removing bias from AI isn’t an option. Rather, we need to think about which biases are acceptable to us, and how extreme they can be. These are questions that call for a mix of technical and philosophical insight that’s hard to find. Luckily, I’ve managed to do just that by inviting onto the podcast none other than Margaret Mitchell, a former Senior Research Scientist in Google’s Research and Machine Intelligence Group, whose work has been focused on practical AI ethics. And by practical, I really do mean the nuts and bolts of how AI ethics can be baked into real systems, and navigating the complex moral issues that come up when the AI rubber meets the road.

***

Intro music :

➞ Artist: Ron Gelinas

➞ Track Title: Daybreak Chill Blend (original mix)

➞ Link to Track: https://youtu.be/d8Y2sKIgFWc

***

Chapters:

0:00 Intro

1:20 Margaret’s background

8:30 Meta learning and ethics

10:15 Margaret’s day-to-day

13:00 Sources of ethical problems within AI

18:00 Aggregated and disaggregated scores

24:02 How much bias will be acceptable?

29:30 What biases does the AI ethics community hold?

35:00 The overlap of these fields

40:30 The political aspect

45:25 Wrap-up

Promoted