Multiple Regression Equations

\[ E(Y|X=\mathbf{x})=\beta_0+\beta_1x_1+\dots+\beta_px_p \tag{1} \] \[ Var(\mathbf{Y}|\mathbf{X})=\sigma^2 \tag{2} \]

\[ \hat{\boldsymbol{\beta}}=(\mathbf{X}'\mathbf{X})^{-1}\mathbf{X}'\mathbf{Y} \tag{3} \]

\[\text{RSS}=\sum(y_i-\mathbf{x}_i'\boldsymbol{\beta})^2=(\mathbf{Y}-\mathbf{X}\beta)'(\mathbf{Y}-\mathbf{X}\beta) \tag{4} \]

\[ \hat{\sigma}^2=\frac{\text{RSS}}{n-(p+1)} \tag{5} \] \[ E(\mathbf{e}|X)=0 \tag{6} \]

\[Var(\mathbf{e}|X)=\sigma^2\mathbf{I}_n \tag{7}\]

\[(\mathbf{e}|X)\sim \text{N}(\mathbf{0},\sigma^2\mathbf{I}_n) \tag{8}\]

Kroger Annual Spend

Kroger is interested in better understanding the purchasing behavior of its loyalty-card customers. The company has assembled a data set containing information from 250 loyalty-card customers. Each row in the data set represents one customer. Some information is available from customers’ purchasing activity, while other information would need to be provided by customers when they apply for a loyalty card. The goal is to model Annual_Spend based to understand what characteristics are related to a customer’s spend.

kroger<-read.csv("I:\\Classes\\ISA 391\\Exam 1\\Kroger.csv")
head(kroger)
##   Customer_ID Annual_Spend Weekly_Visits Household_Income Household_Size Age
## 1       K0001      4681.12          2.89           108086              5  46
## 2       K0002      3452.27          2.38           159514              1  48
## 3       K0003      1828.91          0.74            36549              4  64
## 4       K0004      3446.22          1.31            52048              3  33
## 5       K0005       300.00          0.85            31997              1  44
## 6       K0006      3966.60          3.14            72745              1  34
##   Miles_to_Store
## 1            1.9
## 2            4.8
## 3            2.2
## 4            0.9
## 5           12.7
## 6            0.7
library(dplyr)
library(skimr)

kroger %>% skim() %>% yank("numeric") %>% select(-hist) 

Variable type: numeric

skim_variable n_missing complete_rate mean sd p0 p25 p50 p75 p100
Annual_Spend 0 1 3173.53 1400.21 300.0 2223.88 3224.61 4118.42 7507.90
Weekly_Visits 0 1 2.05 0.73 0.3 1.55 2.07 2.44 4.05
Household_Income 0 1 80930.81 48724.50 18000.0 48567.75 68803.00 99929.25 260000.00
Household_Size 0 1 2.70 1.21 1.0 2.00 3.00 3.00 7.00
Age 0 1 45.30 15.11 18.0 34.00 44.00 56.75 82.00
Miles_to_Store 0 1 4.23 3.01 0.2 2.10 3.45 5.47 15.70

Model 1:

reg<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
summary(reg)
## 
## Call:
## lm(formula = Annual_Spend ~ Weekly_Visits + Household_Income + 
##     Household_Size + Miles_to_Store, data = kroger)
## 
## Residuals:
##      Min       1Q   Median       3Q      Max 
## -2712.05  -591.96    15.72   614.06  2883.44 
## 
## Coefficients:
##                    Estimate Std. Error t value Pr(>|t|)    
## (Intercept)      -4.229e+02  2.670e+02  -1.584  0.11448    
## Weekly_Visits     1.163e+03  8.283e+01  14.041  < 2e-16 ***
## Household_Income  1.031e-02  1.239e-03   8.319 6.21e-15 ***
## Household_Size    2.396e+02  5.029e+01   4.765 3.25e-06 ***
## Miles_to_Store   -6.265e+01  2.005e+01  -3.125  0.00199 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 951.8 on 245 degrees of freedom
## Multiple R-squared:  0.5453, Adjusted R-squared:  0.5379 
## F-statistic: 73.46 on 4 and 245 DF,  p-value: < 2.2e-16

Added Variable Plot for Age

#step 1

step1<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.1<-step1$residuals

#step 2
step2<-lm(Age~Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.2<-step2$residuals

#step 3
plot(
  residuals.1 ~ residuals.2,
  xlab = "e-hat from Age on Model",
  ylab = "e-hat from Annual Spend on MOdel",
  main = "Added Variable Plot: Age"
)

abline(lm(residuals.1~ residuals.2), lwd = 2)

OU student creates a new variable

library(dplyr)

kroger<- kroger %>% mutate(
  Kilometers_to_store=Miles_to_Store*1.60934
)

Added variable plot for the new variable

#step 1

step1<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.1<-step1$residuals

#step 2
step2<-lm(Kilometers_to_store~Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.2<-step2$residuals

#step 3
plot(
  residuals.1 ~ residuals.2,
  xlab = "e-hat from Kilometers_to_store on Model",
  ylab = "e-hat from Annual Spend on MOdel",
  main = "Added Variable Plot: Kilometers_to_store",
  xlim=c(0,0.1)
)

abline(lm(residuals.1~ residuals.2), lwd = 2)

Model 2: Using log(Annual_Spend)

reg<-lm(log(Annual_Spend)~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
summary(reg)
## 
## Call:
## lm(formula = log(Annual_Spend) ~ Weekly_Visits + Household_Income + 
##     Household_Size + Miles_to_Store, data = kroger)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -2.1664 -0.1929  0.0387  0.2853  0.8167 
## 
## Coefficients:
##                    Estimate Std. Error t value Pr(>|t|)    
## (Intercept)       6.439e+00  1.237e-01  52.054  < 2e-16 ***
## Weekly_Visits     5.087e-01  3.838e-02  13.256  < 2e-16 ***
## Household_Income  3.628e-06  5.741e-07   6.319 1.23e-09 ***
## Household_Size    9.338e-02  2.330e-02   4.008 8.14e-05 ***
## Miles_to_Store   -2.436e-02  9.288e-03  -2.623  0.00927 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 0.441 on 245 degrees of freedom
## Multiple R-squared:  0.4885, Adjusted R-squared:  0.4801 
## F-statistic: 58.49 on 4 and 245 DF,  p-value: < 2.2e-16

Hotel Performance

A national hotel company operates hotels of different sizes throughout the United States. Management would like to better understand the factors associated with room revenue performance and use this information to make decisions about hotel operations and marketing.

The company has collected data from 90 hotels for a particular month. Each row in the dataset represents one hotel during that month. The goal of the model building exercise is to understand what factors best predict room revenue. In the data each row represents a single hotel. Please note the hotels are of varying size.

hotel<-read.csv("I:\\Classes\\ISA 391\\Exam 1\\hotel.csv")
head(hotel)
##   Hotel_ID Number_of_Rooms Days_in_Month Rooms_Available Rooms_Sold
## 1     H001             136            30            4080       2277
## 2     H002              71            30            2130       1146
## 3     H003             108            31            3348       1607
## 4     H004             426            31           13206       6339
## 5     H005             239            31            7409       4496
## 6     H006             279            28            7812       4125
##   Room_Revenue Housekeeping_Hours Maintenance_Calls Guest_Rating Advertising
## 1       300959               1312                95         3.95        3773
## 2       155665                526                56         4.76        1372
## 3       176177                689                94         3.82        2695
## 4       639626               3615               336         3.54        5112
## 5       556433               1958               150         4.24        8230
## 6       421964               2012               134         3.98        6097