Posts

Predicting Customer Conversion: A Data-Driven Approach

Image
The following blog post reflects the efforts of my team - Daria Herasymova, Hirak Bhayani, and I - during Duke University's Datathon competition. As Business Analytics graduate students of Duke University’s Fuqua School of Business, we have a plethora of resources to explore & have access to many opportunities to build our data science skills and experience. We were excited to put to good use all that we learned in our classroom sessions conducted by Prof. Alex Belloni at an annual datathon that took place on campus. Organized by Duke Machine Learning, the Duke Datathon was a weekend-long event where teams solved an existent business challenge faced by a marketing technology & consumer engagement company that works with over 60,000 brands across a wide array of industries. The firm predicts customer churn by examining content across 80 billion web pages each day and deriving meaning from words and sentences to determine likely purchase intent through page...

Investigating House Prices

Image
Having now learnt a variety of machine learning models, as well as evaluative methods, it was time for some light practice using an old Kaggle dataset on house prices. Predicting house prices have a variety of applications from prospective buyers estimating the price of a house based on desired features to homeowners who wish to value their house before selling. While the entire analysis was carried out in R, I did a few cleaning bits in Python too to gain further practice. Data Exploration The training dataset is quite small, with only 1460 observations and about 79 variables. These variables included a variety of salient features about houses, including the neighbourhood, square footage, details on the number of rooms, whether the house has a fireplace, pool, garage and so on, as well as the age, and so on. Let's first begin by exploring the data. Here is the distribution of house prices. We can see that the distribution is skewed to the right, with most houses around t...

Modelling Probability of Loan Default

Image
When banks and other financial institutions evaluate a loan request, they must ascertain the risk of default. Defaults can be very costly to banks and predictive analytics can be deployed to help banks make the right decisions. This project involves using a past Kaggle competition to model default of a customer based on many different features. Data Exploration The dataset comprises over 300,000 customers with over 100 features; it is quite a large dataset packed with data. Let's try to make sense of it. It gives demographic and financial information of a customer and then says if he/she defaulted on a loan. Let's first get a broad-level overview of the fraction of defaults. The bank issues both cash loans and revolving loans. The latter are flexible loans and are rather open-ended. Let us see how default rates vary based on the type of loan. It seems that a smaller proportion of revolving loans,...