Uploaded December 2022 | Updated September 2026, 4 hours ago
In this video, we're going to learn how to use the corr() method to create data showing correlation. In the first part of this tutorial, we discussed how to prepare the data for modeling. In this part, we'll get acquainted with collinear features and the importance of removing redundancy in our data. By removing redundancy, we'll be able to improve our data's overall accuracy and make it easier to understand. This is an essential step in data analysis, and you'll want to pay attention to it when working with data!
Watch our previous video:
π Part 1: Data Preparation for Modeling: youtu.be/Sf6jn8QZHhc
π§βπ» Go to the project through the link below and follow along with me: platform.stratascratch.com/data-projects/delivery-duration-prediction?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link&utm_content=collinearity+%26+removing+redundancies
______________________________________________________________________
π Subscribe to my channel: bit.ly/2GsFxmA
π Playlist for more data science interview questions and answers: bit.ly/3jifw81
π Playlist for data science interview tips: bit.ly/2G5hNoJ
π Practice more real data science interview questions: platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link
______________________________________________________________________
Timeline:
Intro: (0:00βββ)
Data project: (0:25 )
The approach: (1:51)
Creating a mask: (2:40)
Functions to test the correlations: (4:00)
Feature engineering: (7:15)
Conclusion: (β8:00)
______________________________________________________________________
About The Platform:
I'm using StrataScratch (platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link), a platform that allows you to practice real data science interview questions. There are over 1000+ interview questions that cover coding (SQL and python), statistics, probability, product sense, and business cases.
So, if you want more interview practice with real data science interview questions, visit platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link. All questions are free and you can even execute SQL and python code in the IDE, but if you want to check out the solutions from me or from other users, you can use ss15 for a 15% discount on the premium plans.
______________________________________________________________________
Contact:
If you have any questions, comments, or feedback, please leave them here!
Feel free to also email me at nathan@stratascratch.com
______________________________________________________________________
#StrataScratch #DoordashDataProject #DataModeling
In this video, we're going to learn how to use the corr() method to create data showing correlation. In the first part of this tutorial, we discussed how to prepare the data for modeling. In this part, we'll get acquainted with collinear features and the importance of removing redundancy in our data. By removing redundancy, we'll be able to improve our data's overall accuracy and make it easier to understand. This is an essential step in data analysis, and you'll want to pay attention to it when working with data!
Watch our previous video:
π Part 1: Data Preparation for Modeling: youtu.be/Sf6jn8QZHhc
π§βπ» Go to the project through the link below and follow along with me: platform.stratascratch.com/data-projects/delivery-duration-prediction?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link&utm_content=collinearity+%26+removing+redundancies
______________________________________________________________________
π Subscribe to my channel: bit.ly/2GsFxmA
π Playlist for more data science interview questions and answers: bit.ly/3jifw81
π Playlist for data science interview tips: bit.ly/2G5hNoJ
π Practice more real data science interview questions: platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link
______________________________________________________________________
Timeline:
Intro: (0:00βββ)
Data project: (0:25 )
The approach: (1:51)
Creating a mask: (2:40)
Functions to test the correlations: (4:00)
Feature engineering: (7:15)
Conclusion: (β8:00)
______________________________________________________________________
About The Platform:
I'm using StrataScratch (platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link), a platform that allows you to practice real data science interview questions. There are over 1000+ interview questions that cover coding (SQL and python), statistics, probability, product sense, and business cases.
So, if you want more interview practice with real data science interview questions, visit platform.stratascratch.com/coding?utm_source=youtube&utm_medium=click&utm_campaign=YT+description+link. All questions are free and you can even execute SQL and python code in the IDE, but if you want to check out the solutions from me or from other users, you can use ss15 for a 15% discount on the premium plans.
______________________________________________________________________
Contact:
If you have any questions, comments, or feedback, please leave them here!
Feel free to also email me at nathan@stratascratch.com
______________________________________________________________________
#StrataScratch #DoordashDataProject #DataModeling










