May 12, 2013

Using KFold cross validation in scikit learn

The idea in KFold cross validation is to split the dataset into K subsets, and use K-1 of these for training and the remaining one for testing. In each iteration we cycle through the different subsets, thereby giving each subset a chance to be the test set.

The following is an illustration of KFold cross validation using scikit-learn.
from sklearn.cross_validation import KFold
import numpy as np


X=np.arange(16).reshape(2,8).T # to get the row numbers in order
y = np.arange(8)
kf = KFold(n=len(y), n_folds=4, indices=False, shuffle=True)

print "Entire dataset"
print X

for train, test in kf:
    X_train, X_test, y_train, y_test = X[train], X[test], y[train], y[test]
    print "------------------------------"
    print "Train using"
    print X_train
    print "Test using"
    print X_test
Output
Entire dataset
[[ 0  8]
 [ 1  9]
 [ 2 10]
 [ 3 11]
 [ 4 12]
 [ 5 13]
 [ 6 14]
 [ 7 15]]
------------------------------
Train using
[[ 0  8]
 [ 1  9]
 [ 2 10]
 [ 3 11]
 [ 6 14]
 [ 7 15]]
Test using
[[ 4 12]
 [ 5 13]]
------------------------------
Train using
[[ 2 10]
 [ 3 11]
 [ 4 12]
 [ 5 13]
 [ 6 14]
 [ 7 15]]
Test using
[[0 8]
 [1 9]]
------------------------------
Train using
[[ 0  8]
 [ 1  9]
 [ 2 10]
 [ 3 11]
 [ 4 12]
 [ 5 13]]
Test using
[[ 6 14]
 [ 7 15]]
------------------------------
Train using
[[ 0  8]
 [ 1  9]
 [ 4 12]
 [ 5 13]
 [ 6 14]
 [ 7 15]]
Test using
[[ 2 10]
 [ 3 11]]
8 rows in the original dataset were divided up into 4 subsets. Through the 4 iterations printed in the output, we can observe that each subset was used as the test set once.

Following is a use case with the Boston housing prices data set
d = pickle.load( open( "housing_prices_shuffled-cp.p", "rb" ) )
X = d[:, :13]
y = d[:, 13]
X = np.hstack((X, X**2))

train_x, test_x, train_y, test_y = cross_validation.train_test_split(
                            X, y, test_size=0.3, random_state=0)

print "Score with 70-30 split => "+str(linreg(train_x, train_y, test_x, test_y))

kf = KFold(n=len(y), n_folds=4, indices=False, shuffle=True)
for i, (train, test) in enumerate(kf):
    train_x, test_x, train_y, test_y = X[train], X[test], y[train], y[test]
    score = linreg(train_x, train_y, test_x, test_y)
    print ("Score in KFold round %s => %s" %( i+1, score))
Output
Score with 70-30 split => 0.788362712017
Score in KFold round 1 => 0.881987112096
Score in KFold round 2 => 0.754062158544
Score in KFold round 3 => 0.774394139433
Score in KFold round 4 => 0.767306134201