Log in

Towards Data Science - 37. Sean Knapp - The brave new world of data engineering
share icon

37. Sean Knapp - The brave new world of data engineering

Towards Data Science

06/10/20

44m

About

Comments

Featured In

There’s been a lot of talk in data science circles about techniques like AutoML, which are dramatically reducing the time it takes for data scientists to train and tune models, and create reliable experiments. But that trend towards increased automation, greater robustness and reliability doesn’t end with machine learning: increasingly, companies are focusing their attention on automating earlier parts of the data lifecycle, including the critical task of data engineering.

Today, many data engineers are unicorns: they not only have to understand the needs of their customers, but also how to work with data, and what software engineering tools and best practices to use to set up and monitor their pipelines. Pipeline monitoring in particular is time-consuming, and just as important, isn’t a particularly fun thing to do. Luckily, people like Sean Knapp — a former Googler turned founder of data engineering startup Ascend.io — are leading the charge to make automated data pipeline monitoring a reality.

We had Sean on this latest episode of the Towards Data Science podcast to talk about data engineering: where it’s at, where it’s going, and what data scientists should really know about it to be prepared for the future.

Previous Episode

For the last decade, advances in machine learning have come from two things: improved compute power and better algorithms. These two areas have become somewhat siloed in most people’s thinking: we tend to imagine that there are people who build hardware, and people who make algorithms, and that there isn’t much overlap between the two.

But this picture is wrong. Hardware constraints can and do inform algorithm design, and algorithms can be used to optimize hardware. Increasingly, compute and modelling are being optimized together, by people with expertise in both areas.

My guest today is one of the world’s leading experts on hardware/software integration for machine learning applications. Max Welling is a former physicist and currently works as VP Technologies at Qualcomm, a world-leading chip manufacturer, in addition to which he’s also a machine learning researcher with affiliations at UC Irvine, CIFAR and the University of Amsterdam.

Next Episode

One Thursday afternoon in 2015, I got a spontaneous notification on my phone telling me how long it would take to drive to my favourite restaurant under current traffic conditions. This was alarming, not only because it implied that my phone had figured out what my favourite restaurant was without ever asking explicitly, but also because it suggested that my phone knew enough about my eating habits to realize that I liked to go out to dinner on Thursdays specifically.

As our phones, our laptops and our Amazon Echos collect increasing amounts of data about us — and impute even more — data privacy is becoming a greater and greater concern for research as well as government and industry applications. That’s why I wanted to speak to Harvard PhD student and frequent Towards Data Science contributor Matthew Stewart about to get an introduction to some of the key principles behind data privacy. Matthew is a prolific blogger, and his research work at Harvard is focused on applications of machine learning to environmental sciences, a topic we also discuss during this episode.

Promoted