\[ E(Y|X=\mathbf{x})=\beta_0+\beta_1x_1+\dots+\beta_px_p \tag{1} \] \[ Var(\mathbf{Y}|\mathbf{X})=\sigma^2 \tag{2} \]
\[ \hat{\boldsymbol{\beta}}=(\mathbf{X}'\mathbf{X})^{-1}\mathbf{X}'\mathbf{Y} \tag{3} \]
\[\text{RSS}=\sum(y_i-\mathbf{x}_i'\boldsymbol{\beta})^2=(\mathbf{Y}-\mathbf{X}\beta)'(\mathbf{Y}-\mathbf{X}\beta) \tag{4} \]
\[ \hat{\sigma}^2=\frac{\text{RSS}}{n-(p+1)} \tag{5} \] \[ E(\mathbf{e}|X)=0 \tag{6} \]
\[Var(\mathbf{e}|X)=\sigma^2\mathbf{I}_n \tag{7}\]
\[(\mathbf{e}|X)\sim \text{N}(\mathbf{0},\sigma^2\mathbf{I}_n) \tag{8}\]
Kroger is interested in better understanding the purchasing behavior of its loyalty-card customers. The company has assembled a data set containing information from 250 loyalty-card customers. Each row in the data set represents one customer. Some information is available from customers’ purchasing activity, while other information would need to be provided by customers when they apply for a loyalty card. The goal is to model Annual_Spend based to understand what characteristics are related to a customer’s spend.
kroger<-read.csv("I:\\Classes\\ISA 391\\Exam 1\\Kroger.csv")
head(kroger)
## Customer_ID Annual_Spend Weekly_Visits Household_Income Household_Size Age
## 1 K0001 4681.12 2.89 108086 5 46
## 2 K0002 3452.27 2.38 159514 1 48
## 3 K0003 1828.91 0.74 36549 4 64
## 4 K0004 3446.22 1.31 52048 3 33
## 5 K0005 300.00 0.85 31997 1 44
## 6 K0006 3966.60 3.14 72745 1 34
## Miles_to_Store
## 1 1.9
## 2 4.8
## 3 2.2
## 4 0.9
## 5 12.7
## 6 0.7
library(dplyr)
library(skimr)
kroger %>% skim() %>% yank("numeric") %>% select(-hist)
Variable type: numeric
| skim_variable | n_missing | complete_rate | mean | sd | p0 | p25 | p50 | p75 | p100 |
|---|---|---|---|---|---|---|---|---|---|
| Annual_Spend | 0 | 1 | 3173.53 | 1400.21 | 300.0 | 2223.88 | 3224.61 | 4118.42 | 7507.90 |
| Weekly_Visits | 0 | 1 | 2.05 | 0.73 | 0.3 | 1.55 | 2.07 | 2.44 | 4.05 |
| Household_Income | 0 | 1 | 80930.81 | 48724.50 | 18000.0 | 48567.75 | 68803.00 | 99929.25 | 260000.00 |
| Household_Size | 0 | 1 | 2.70 | 1.21 | 1.0 | 2.00 | 3.00 | 3.00 | 7.00 |
| Age | 0 | 1 | 45.30 | 15.11 | 18.0 | 34.00 | 44.00 | 56.75 | 82.00 |
| Miles_to_Store | 0 | 1 | 4.23 | 3.01 | 0.2 | 2.10 | 3.45 | 5.47 | 15.70 |
reg<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
summary(reg)
##
## Call:
## lm(formula = Annual_Spend ~ Weekly_Visits + Household_Income +
## Household_Size + Miles_to_Store, data = kroger)
##
## Residuals:
## Min 1Q Median 3Q Max
## -2712.05 -591.96 15.72 614.06 2883.44
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -4.229e+02 2.670e+02 -1.584 0.11448
## Weekly_Visits 1.163e+03 8.283e+01 14.041 < 2e-16 ***
## Household_Income 1.031e-02 1.239e-03 8.319 6.21e-15 ***
## Household_Size 2.396e+02 5.029e+01 4.765 3.25e-06 ***
## Miles_to_Store -6.265e+01 2.005e+01 -3.125 0.00199 **
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 951.8 on 245 degrees of freedom
## Multiple R-squared: 0.5453, Adjusted R-squared: 0.5379
## F-statistic: 73.46 on 4 and 245 DF, p-value: < 2.2e-16
#step 1
step1<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.1<-step1$residuals
#step 2
step2<-lm(Age~Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.2<-step2$residuals
#step 3
plot(
residuals.1 ~ residuals.2,
xlab = "e-hat from Age on Model",
ylab = "e-hat from Annual Spend on MOdel",
main = "Added Variable Plot: Age"
)
abline(lm(residuals.1~ residuals.2), lwd = 2)
library(dplyr)
kroger<- kroger %>% mutate(
Kilometers_to_store=Miles_to_Store*1.60934
)
#step 1
step1<-lm(Annual_Spend~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.1<-step1$residuals
#step 2
step2<-lm(Kilometers_to_store~Household_Income+Household_Size+Miles_to_Store, data=kroger)
residuals.2<-step2$residuals
#step 3
plot(
residuals.1 ~ residuals.2,
xlab = "e-hat from Kilometers_to_store on Model",
ylab = "e-hat from Annual Spend on MOdel",
main = "Added Variable Plot: Kilometers_to_store",
xlim=c(0,0.1)
)
abline(lm(residuals.1~ residuals.2), lwd = 2)
reg<-lm(log(Annual_Spend)~Weekly_Visits+Household_Income+Household_Size+Miles_to_Store, data=kroger)
summary(reg)
##
## Call:
## lm(formula = log(Annual_Spend) ~ Weekly_Visits + Household_Income +
## Household_Size + Miles_to_Store, data = kroger)
##
## Residuals:
## Min 1Q Median 3Q Max
## -2.1664 -0.1929 0.0387 0.2853 0.8167
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 6.439e+00 1.237e-01 52.054 < 2e-16 ***
## Weekly_Visits 5.087e-01 3.838e-02 13.256 < 2e-16 ***
## Household_Income 3.628e-06 5.741e-07 6.319 1.23e-09 ***
## Household_Size 9.338e-02 2.330e-02 4.008 8.14e-05 ***
## Miles_to_Store -2.436e-02 9.288e-03 -2.623 0.00927 **
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.441 on 245 degrees of freedom
## Multiple R-squared: 0.4885, Adjusted R-squared: 0.4801
## F-statistic: 58.49 on 4 and 245 DF, p-value: < 2.2e-16
A national hotel company operates hotels of different sizes throughout the United States. Management would like to better understand the factors associated with room revenue performance and use this information to make decisions about hotel operations and marketing.
The company has collected data from 90 hotels for a particular month. Each row in the dataset represents one hotel during that month. The goal of the model building exercise is to understand what factors best predict room revenue. In the data each row represents a single hotel. Please note the hotels are of varying size.
hotel<-read.csv("I:\\Classes\\ISA 391\\Exam 1\\hotel.csv")
head(hotel)
## Hotel_ID Number_of_Rooms Days_in_Month Rooms_Available Rooms_Sold
## 1 H001 136 30 4080 2277
## 2 H002 71 30 2130 1146
## 3 H003 108 31 3348 1607
## 4 H004 426 31 13206 6339
## 5 H005 239 31 7409 4496
## 6 H006 279 28 7812 4125
## Room_Revenue Housekeeping_Hours Maintenance_Calls Guest_Rating Advertising
## 1 300959 1312 95 3.95 3773
## 2 155665 526 56 4.76 1372
## 3 176177 689 94 3.82 2695
## 4 639626 3615 336 3.54 5112
## 5 556433 1958 150 4.24 8230
## 6 421964 2012 134 3.98 6097