May 10, 2013

Linear Regression - ensemble approach

This is a continuation of the post : Linear Regression with Scikit-Learn.

Here we attempt to run linear regression with different combinations of input variables and compare the scores generated by these different models.

from sklearn.linear_model import LinearRegression
import cPickle as pickle
import itertools
import operator
from pprint import pprint

def linreg(train_x, train_y, test_x, test_y):
    reg = LinearRegression()
    reg.fit(train_x, train_y)
    reg.predict(test_x)
    return reg.score(test_x, test_y) 


k = pickle.load( open( "housing_prices_shuffled-cp.p", "rb" ) )
train_y = k[:-20, 13]
test_y  = k[-20:, 13]
train_x = k[:-20, :13]
test_x  = k[-20:, :13]

scoremap = {}
for numvars in xrange(4):
    comb = itertools.combinations(tuple(range(train_x.shape[1])),numvars+1)
    for i in comb:
        scoremap[i] = linreg(train_x[:, i], train_y, test_x[:, i], test_y)


pprint (scoremap)
m = max(scoremap.iteritems(), key=operator.itemgetter(1))[0]
print "Highest score was for the combination "+str(m) + " and the score was " + str( scoremap[m] )
print "Score with all parameters => "+str(linreg(train_x, train_y, test_x, test_y))
Output
{(0,): 0.094934881718741981,
 (0, 1): 0.15440076324231666,
 (0, 1, 2): 0.31263198950857518,
 (0, 1, 2, 3): 0.33078144141566046,
 (0, 1, 2, 4): 0.32265618841454957,
 (0, 1, 2, 5): 0.71474203629461419,
 (0, 1, 2, 6): 0.31654350182389657,
 (0, 1, 2, 7): 0.24129421943954743,
 (0, 1, 2, 8): 0.31354584652512074,
 (0, 1, 2, 9): 0.28242669419930766,
 (0, 1, 2, 10): 0.18453687711944267,
 (0, 1, 2, 11): 0.37156015631289618,
 (0, 1, 2, 12): 0.62662369683834229,
 (0, 1, 3): 0.14564590334510252,
 (0, 1, 3, 4): 0.30448864205509185,
 (0, 1, 3, 5): 0.65694580409701808,
 (0, 1, 3, 6): 0.29525760888290975,
 ...
 ...
 ...
 (10, 11): 0.091172739042896023,
 (10, 11, 12): 0.67117510695094174,
 (10, 12): 0.66456393824019266,
 (11,): 0.26293003704903495,
 (11, 12): 0.59626994121671273,
 (12,): 0.59380010075786283}
Highest score was for the combination (5, 10, 11, 12) and the score was 0.844977437779
Score with all parameters => 0.778321085183
In the above script, different combinations of the input vector was generated and regression models were trained using each of these. From the above crude attempt, we can see that the case where all the variables were blindly used resulted in a score of 0.7783, whereas the one where the dimensions (5,10,11,12) alone were chosen resulted in a higher score of 0.8449.

From the graphs plotted in the previous linear regression post, we can see that parameters 5,10,11,12 had a higher visual correlation when compared with the other parameters.