May 8, 2013

Linear Regression with Scikit Learn

The Boston housing dataset consists of 506 house prices and a set of features that might have influenced these prices.

* Input Vector : 13 features (Number of rooms, Age of house, Demographics etc)
* Target Function : House prices in $1000s (More details on the dataset)

Visualizing the dataset

When the price is plotted against individual features, we can observe that some of the features have a correlation with the price.
Eg:
1. In the first plot (CRIM), as the crime rate goes up, the housing price goes down.
2. As the number of rooms (RM) increases, the price increases.
House Price -Vs- Individual Features

The following is a quick and dirty attempt at running linear regression using all 13 features. From the output, we can see that the weights learned in this case are not optimal.
import numpy as np
from sklearn import datasets
from matplotlib import pyplot as plt
from sklearn.linear_model import LinearRegression

d = datasets.load_boston()

train_x = d.data[:-20]
train_y = d.target[:-20]
test_x = d.data[-20:]
test_y = d.target[-20:]


reg = LinearRegression()
reg.fit(train_x, train_y)
 
print "Weights learned by linear regression => "
print reg.coef_

res = reg.predict(test_x)

print "Root mean square error"
print np.sqrt(np.mean((res - test_y)**2))

print "Score"
print reg.score(test_x, test_y)
Output
Weights learned by linear regression => 
[ -1.01939810e-01   4.94888480e-02   2.33211595e-02   2.59303688e+00
  -1.68214723e+01   3.76487853e+00   6.86236795e-03  -1.45762710e+00
   3.50148357e-01  -1.56920780e-02  -9.03727306e-01   9.04441274e-03
  -5.52819289e-01]
Root mean square error
4.22970052109
Score
0.234803679279

NOTE: The Score here is the Coefficient of Determination

Shuffling the dataset

If the dataset has some order within it (eg: training examples in increasing order of house prices), then a naive subset consisting of the first N rows might not be a good representative sample of the data we have. Lets see if shuffling the training examples would make any difference.

d = datasets.load_boston()
k = np.array([np.hstack((d.data[i],d.target[i])) for i in range(len(d.target))])
np.random.shuffle(k) 

train_y = k[:-20, 13]
test_y  = k[-20:, 13]
train_x = k[:-20, :13]
test_x  = k[-20:, :13]

reg = LinearRegression()
reg.fit(train_x, train_y)
 
print "Weights learned by linear regression => "
print reg.coef_

res = reg.predict(test_x)

print "Root mean square error"
print np.sqrt(np.mean((res - test_y)**2))

print "Score"
print reg.score(test_x, test_y)
With this change, the score has improved over the previous attempt. However we might need to further tweak the regression model to achieve better results.
Weights learned by linear regression => 
[ -1.05619886e-01   4.57828979e-02   2.07741453e-02   2.77093055e+00
  -1.77208901e+01   3.74134479e+00  -1.46239003e-03  -1.49920141e+00
   2.95667008e-01  -1.17928650e-02  -9.52054667e-01   9.86461510e-03
  -5.28157244e-01]
Root mean square error
4.00764038083
Score
0.776918740753