* Input Vector : 13 features (Number of rooms, Age of house, Demographics etc)
* Target Function : House prices in $1000s (More details on the dataset)
Visualizing the dataset
When the price is plotted against individual features, we can observe that some of the features have a correlation with the price.Eg:
1. In the first plot (CRIM), as the crime rate goes up, the housing price goes down.
2. As the number of rooms (RM) increases, the price increases.
![]() |
| House Price -Vs- Individual Features |
The following is a quick and dirty attempt at running linear regression using all 13 features. From the output, we can see that the weights learned in this case are not optimal.
import numpy as np from sklearn import datasets from matplotlib import pyplot as plt from sklearn.linear_model import LinearRegression d = datasets.load_boston() train_x = d.data[:-20] train_y = d.target[:-20] test_x = d.data[-20:] test_y = d.target[-20:] reg = LinearRegression() reg.fit(train_x, train_y) print "Weights learned by linear regression => " print reg.coef_ res = reg.predict(test_x) print "Root mean square error" print np.sqrt(np.mean((res - test_y)**2)) print "Score" print reg.score(test_x, test_y)Output
Weights learned by linear regression => [ -1.01939810e-01 4.94888480e-02 2.33211595e-02 2.59303688e+00 -1.68214723e+01 3.76487853e+00 6.86236795e-03 -1.45762710e+00 3.50148357e-01 -1.56920780e-02 -9.03727306e-01 9.04441274e-03 -5.52819289e-01] Root mean square error 4.22970052109 Score 0.234803679279
NOTE: The Score here is the Coefficient of Determination
Shuffling the dataset
If the dataset has some order within it (eg: training examples in increasing order of house prices), then a naive subset consisting of the first N rows might not be a good representative sample of the data we have. Lets see if shuffling the training examples would make any difference.d = datasets.load_boston() k = np.array([np.hstack((d.data[i],d.target[i])) for i in range(len(d.target))]) np.random.shuffle(k) train_y = k[:-20, 13] test_y = k[-20:, 13] train_x = k[:-20, :13] test_x = k[-20:, :13] reg = LinearRegression() reg.fit(train_x, train_y) print "Weights learned by linear regression => " print reg.coef_ res = reg.predict(test_x) print "Root mean square error" print np.sqrt(np.mean((res - test_y)**2)) print "Score" print reg.score(test_x, test_y)With this change, the score has improved over the previous attempt. However we might need to further tweak the regression model to achieve better results.
Weights learned by linear regression => [ -1.05619886e-01 4.57828979e-02 2.07741453e-02 2.77093055e+00 -1.77208901e+01 3.74134479e+00 -1.46239003e-03 -1.49920141e+00 2.95667008e-01 -1.17928650e-02 -9.52054667e-01 9.86461510e-03 -5.28157244e-01] Root mean square error 4.00764038083 Score 0.776918740753
