train_x, test_x = np.hstack((train_x, train_x**2)), np.hstack((test_x, test_x**2))With this modified set of input vectors, we attempt to train a linear regression model
scoremap = {}
for numvars in xrange(3):
comb = itertools.combinations(tuple(range(train_x.shape[1])),numvars+1)
for i in comb:
scoremap[i] = linreg(train_x[:, i], train_y, test_x[:, i], test_y)
pprint (scoremap)
m = max(scoremap.iteritems(), key=operator.itemgetter(1))[0]
print "Highest score was for the combination "+str(m) + " and the score was " + str( scoremap[m] )
print "Score with all parameters => "+str(linreg(train_x, train_y, test_x, test_y))
Here we have an input with 26 features (13*2), and hence the ensemble approach would generate a larger number of combinations. For a quick test, numvars was restricted to 3.Output
{(0,): 0.16362181986382918,
(0, 1): 0.2684875948162504,
(0, 1, 2): 0.31889908199852668,
(0, 1, 3): 0.30721137411019506,
...
...
...
(24, 25): 0.37072146513331328,
(25,): 0.37395586902331879}
Highest score was for the combination (12, 18, 25) and the score was 0.704451741594
Score with all parameters => 0.788362712017
In the output, we can see that the score did improve over the previous value (0.68447). However the increased complexity of the model could also mean that we are potentially overfitting the data, and this might need further analysis.