May 12, 2013

Linear Regression with non-linear features

The objective in this post is to add non-linear features to the input and see how the error is impacted. Here we apply a non-linear transform to the input vectors (x^2) and append this transformed set of input vectors to the original set. i.e.
train_x, test_x = np.hstack((train_x, train_x**2)), np.hstack((test_x, test_x**2))
With this modified set of input vectors, we attempt to train a linear regression model
scoremap = {}
for numvars in xrange(3):
    comb = itertools.combinations(tuple(range(train_x.shape[1])),numvars+1)
    for i in comb:
        scoremap[i] = linreg(train_x[:, i], train_y, test_x[:, i], test_y)

pprint (scoremap)
m = max(scoremap.iteritems(), key=operator.itemgetter(1))[0]
print "Highest score was for the combination "+str(m) + " and the score was " + str( scoremap[m] )
print "Score with all parameters => "+str(linreg(train_x, train_y, test_x, test_y))
Here we have an input with 26 features (13*2), and hence the ensemble approach would generate a larger number of combinations. For a quick test, numvars was restricted to 3.

Output
{(0,): 0.16362181986382918,
 (0, 1): 0.2684875948162504,
 (0, 1, 2): 0.31889908199852668,
 (0, 1, 3): 0.30721137411019506,
 ...
 ...
 ...
 (24, 25): 0.37072146513331328,
 (25,): 0.37395586902331879}
Highest score was for the combination (12, 18, 25) and the score was 0.704451741594
Score with all parameters => 0.788362712017
In the output, we can see that the score did improve over the previous value (0.68447). However the increased complexity of the model could also mean that we are potentially overfitting the data, and this might need further analysis.