Anda di halaman 1dari 24

CHAPTER 9

INTRODUCTION TO LINEAR REGRESSION


AND
CORRELATION ANALYSIS
Goals
After this, you should be able
to:
Calculate and interpret the
simple correlation between two
variables
Calculate and interpret the
simple linear regression
equation for a set of data
Understand the assumptions
Regression and Correlation
a) Income –consumption
b) Sales of ice-cream –with temperature
of the day
c) Industrial production and
consumption of electricity
d) The yield of crops, amount of rainfall,
type of fertilizer, humidity.
e) Weight and height, age and strength,
blood pressure and time of exercise
Cont…
-Income and expenditure
- Number of hours spent in
studying and the score
obtained
Distance covered and fuel
consumed by car.
Simple Linear Regression

Derive an equation for the line that


best models the relationship between
variables. This equation has the
mathematical form:
Y = a + bX
where, Y is the value of the
dependent variable, X is the value of
the independent variable, a is the
intercept of the regression line on the
Y axis when X = 0, and b is the slope
Scatter Plots and Correlation

A scatter plot (or scatter diagram) is


used to show the relationship between
two variables
Correlation analysis is used to measure
strength of the association (linear
relationship) between two variables
 Only concerned with strength of the
relationship
 No causal effect is implied
Scatter Plot Examples
Linear Curvilinear
relationships relationships
y y

x x

y y

x x
Scatter Plot Examples
(continued)
Strong Weak
relationships relationships
y y

x x

y y

x x
Scatter Plot Examples
(continued)
No
relationship
y

x
Correlation Coefficient (continued)
The population correlation coefficient
ρ (rho) measures the strength of the
association between the variables

The sample correlation coefficient


r
is an estimate of ρ and is used to
measure the strength of the linear
relationship in the sample
observations
Features of ρ and r
Unit free
Range between -1 and 1
The closer to -1, the stronger
the negative linear
relationship
The closer to 1, the stronger
the positive linear relationship
The closer to 0, the weaker
Examples of Approximate
r Values

y y y

x x x
r = -1 r = -.6 r=0
y y

x x
r = +.3 r = +1
Calculating the
Correlation Coefficient
Simple correlation
coefficient:
r
 ( x  x )( y  y )
[ ( x  x ) ][  ( y  y ) ]
2 2

or the algebraic
equivalent:
n xy   x  y
r
[n(  x 2 )  (  x )2 ][n(  y 2 )  (  y )2 ]
where:
r = Sample correlation coefficient
n = Sample size
x = Value of the independent variable
y = Value of the dependent variable
2.Rank Correlation for qualitative
data

Given by:
The Least Squares Equation
The formulas for b1 and b0 are:

b1 
 ( x  x )( y  y )
 (x  x) 2

algebraic
equivalent: and
 xy   x y
b1  n b0  y  b1 x
 x 2

(  x ) 2

n
Interpretation of the
Slope and the Intercept
b0 is the estimated average value
of y when the value of x is zero

b1 is the estimated change in the


average value of y as a result of a
one-unit change in x
Coefficient of Determination,
R2
The coefficient of determination is
the portion of the total variation in
the dependent variable that is
explained by variation in the
independent variable

The coefficient of determination is


also called R-squared where
and is0 denoted
R 1
2

as R2
Coefficient of Determination,
R2
(continued)
Coefficient of determination
Note: In the single independent variable
case, the coefficient of determination is
R r
2 2

where:
R2 = Coefficient of determination
r = Simple correlation coefficient
Examples of Approximate
R2 Values

y
R2 = 1

Perfect linear
relationship between x
x
R2 = 1 and y:
y
100% of the variation
in y is explained by
variation in x
x
R = +1
2
Examples of Approximate
R2 Values

y
0 < R2 < 1

Weaker linear
relationship between x
x and y:
y
Some but not all of the
variation in y is
explained by variation
in x
x
Examples of Approximate
R2 Values

R2 = 0
y
No linear relationship
between x and y:

The value of Y does


x not depend on x.
R2 = 0
(None of the variation
in y is explained by
variation in x)
Summary
Introduced correlation analysis
Discussed correlation to measure the
strength of a linear association
Introduced simple linear regression
analysis
Calculated the coefficients for the
simple linear regression equation
measures of variation (R2 and sε)
Addressed assumptions of regression
and correlation
EXERCISE
 Explain Regression & correlation analysis
Given The summary statistics for Crop yield
and Amount of fertilizer used, be able to  Fit
linear regression line & Interprete.
Calculate the correlation coefficient and
comment on it.
Calculate the coefficient of determination and
comment on it.
Find the fitted value corresponding to x= 10
Thank
You!

Anda mungkin juga menyukai