Jan 28, 2016

Kaggle Rossmann Sales Forecasting Challenge

This write-up is about the recently completed Kaggle Rossmann Sales Forecasting Challenge.

Results
I am happy to report that I came at Rank 39 / 3303 on the final leaderboard.

Kaggle Rossmann Leaderboard





Source Code
The source code for my submission is available at the below github repo:
https://github.com/ajerome9/kaggle-time-series-forecasting

Problem description:
Forecast sales for 6 weeks across 3000 stores using store, promotion and competitor data for those stores during the previous 2.5 years.

Evaluation metric: 
Root Mean Square Percentage Error (RMSPE). My final model had a RMSPE of 0.11146


Example of sales predictions from one of the models in the ensemble (days with sales = 0 were not scored)

Model description:
The final model consisted of an ensemble of:
       (a) Gradient Boosted Decision Tree models (xgboost)
       (b) Linear models - Lasso/Ridge using glmnet with various values of alpha
       (c) tslm for removing trend/seasonality + boosted decision tree model fitted on residuals.

The models could be categorized as
1. Store-wise models 
2. Models using merged-data
3. Hybrid Models

Store-wise models:
These models were fit individually on each store. On the positive side, this model could focus on store specific behavior alone. However, at the same time, it was also missing out on information that could be derived from patterns across other similar stores.
  1. linear model to fit trend, and
      a. residuals fit using xgboost
      b. residuals fit using xgboost with feature interactions
      c. residuals fit using glmnet with various alpha values
  2. tslm (from Rob Hyndman's excellent 'forecast' package) to fit trend+seasonality. Residuals fit using xgboost
  3. xgboost per store

Models using merged-data: 
These models used a combination of store sales data and store metadata. With this, the model could now start looking at how similar stores performed.
  1. xgboost on merged train+store
      a. without feature engineering (depth=5, trees=300)
      b. without feature engineering (depth=10, trees=3000)
      c. with feature engineering (depth=10, trees=4000, min_child_weight=2)
      d. with feature engineering + feature interactions (depth=18, trees=3000). Although tree models can capture feature interactions by themselves, I thought it might help to provide the model with guesses on feature interactions that are likely to have an impact.

Hybrid Models: 
Combination of merged-data models to capture behaviors shared across stores. Residuals from this model were fit using storewise models so as to capture aspects unique to each store (eg: Store-A has high sales on the day before Valentine's day)

250 feats were constructed from merged data, including binary features indicating the type of holiday on a given day in a particular state. 50 principal components were extracted from this using the prcomp function, and fit using xgboost. Per-store residuals of this model were fit using separate xgboost models.

Other models that were explored and dropped:
1. Using instance weights to emphasize recent data in xgboost.
2. Using stl for decomposing trend+seasonality, and fitting residuals using xgboost

Ensembling approaches:
1. Simple average of predictions
2. Weighted average (weighted by performance of model for each store during cross validation). Two approaches were explored.
      (a) light penalty for poor predictions 
      (b) high penalty for poor predictions
I went with approach-a because approach-b was resulting in overfitting to the cross validation set.

Feature engineering
1. Categorical variables were transformed to {mean, sd of sales in category}. This was done for store_id, store_type, assortment, and few other combinations like store_id:day_of_week, store_id:day_of_week:promo
2. Feature that spikes up on the day a competitor opens up, and then gradually decreases to a steady value.
3. Pay-day effect: feature that takes a value of 3 on the first working day in the month, and {2,1} on the subsequent days. If the first day is a Saturday/Sunday, then Friday is assumed to be the payday. Plotting the values also showed a small spike around the start of the month.
4. Feature that cycles through values 1,2,3 through the Promo2 interval (eg: for Promo2 interval=Jan-Apr-Jul-Oct, days in January would have 1, days in February would have 2, days in March would have 3, and then back to 1 in April)
5. Large number of Leading/Lagging variables. eg: was yesterday a holiday, is tomorrow a holiday, was store open yesterday, was promo=1 yesterday etc
6. External data: This competition allowed the usage of external data that was publicly declared on the forums. I tried some of the datasets posted there, however did not see a substantial performance improvement. The store-location and holidays dataset were somewhat useful. Other datasets that were explored: State metadata, Daily/Weekly Google trends at Country/State level.

Approaches that seemed promising, but were not implemented due to lack of time
1. Linear mixed effects/hierarchical model (lme4 package)
2. Arima: some contestants reported good results with arima models. I noticed that the autocorrelation plots showed residuals having autocorrelations at various lags (1,2,3,7,14). Rob Hyndman's forecast package has an arima implementation called auto.arima that finds lags/order of differencing automatically. Might be worth checking out for the next competition.
3. KNeighborsRegressor in scikit learn.
4. Ensemble strategy: Stacking instead of averaging.