A line of best fit is a straight line drawn through the points of a scatter plot to show the overall trend. It turns a cloud of dots into something you can describe in one sentence (“scores rise about four points for every extra hour of study”) and use to make predictions. This guide shows how to draw a line of best fit by eye, how the exact least squares line is calculated, how to read its equation and R², and the cases where a best fit line on a scatter plot misleads.
What a Line of Best Fit Is
Data rarely falls on a perfect line. Twelve students who studied different hours got different scores, and the scores do not rise in neat steps. A line of best fit, also called a trend line or least squares line, is the single straight line that comes closest to all the points at once.
Start with the scatter plot on its own. Each dot is one student: hours studied along the bottom, test score up the side.
Twelve students before any line is drawn. The dots drift upward from left to right. Sample data.
Show the data
| Hours studied | Test score | Label |
|---|---|---|
| 1 | 58 | Ava |
| 2 | 62 | Ben |
| 2.5 | 66 | Cam |
| 3 | 61 | Dee |
| 4 | 72 | Eli |
| 4.5 | 70 | Fay |
| 5 | 78 | Gus |
| 6 | 75 | Hal |
| 6.5 | 84 | Ivy |
| 7 | 81 | Jon |
| 8 | 90 | Kim |
| 9 | 88 | Lou |
Now add the line. The tool computes it with least squares, which we explain below, and prints its equation and R² above the plot.
The same students with the least squares line of best fit. The equation reads y = 4.137x + 53.58 and R² = 0.924. Sample data.
Show the data
| Hours studied | Test score | Label |
|---|---|---|
| 1 | 58 | Ava |
| 2 | 62 | Ben |
| 2.5 | 66 | Cam |
| 3 | 61 | Dee |
| 4 | 72 | Eli |
| 4.5 | 70 | Fay |
| 5 | 78 | Gus |
| 6 | 75 | Hal |
| 6.5 | 84 | Ivy |
| 7 | 81 | Jon |
| 8 | 90 | Kim |
| 9 | 88 | Lou |
The line answers two questions the dots only hint at. How fast does the score rise? About 4.1 points per hour. Where does it start? Near 54 points for a student who studied very little. OpenStax’s Introductory Statistics (2023 edition) calls this line the line of best fit or least-squares line, and describes the whole process as linear regression.
Open this line of best fit in the scatter plot maker
How to Draw a Line of Best Fit by Eye
Before calculators, and still in many classrooms, the line is drawn by hand. OpenStax describes the exercise in its chapter on the regression equation: plot the points, draw a line that appears to fit, pick two convenient points on it, and work out its equation.
Step 1: Plot the Points Carefully
Use graph paper and a ruler. Put the variable you think drives the other on the x-axis (hours studied) and the one that responds on the y-axis (score). A sloppy plot gives a sloppy line.
Step 2: Place the Ruler So the Points Balance
Lay a clear ruler across the cloud and rotate it until the points are balanced around it. Good rules of thumb:
- About as many points above the line as below it.
- The points above and below are spread along the whole length, not all above on the left and all below on the right.
- The line follows the direction of the cloud, not the direction of one or two extreme points.
The line does not have to touch any point. It often touches none.
Step 3: Read Two Points on Your Line
Pick two points on the line you drew, not two data points, and choose them far apart so small reading errors matter less. Say your line passes through (1, 57) and (9, 91).
Step 4: Work Out the Slope and Intercept
The slope is the rise over the run: (91 minus 57) divided by (9 minus 1), which is 34 divided by 8, or 4.25. For the intercept, substitute one point into y = mx + b: 57 = 4.25 times 1 + b, so b = 52.75. Your hand-drawn line is y = 4.25x + 52.75.
Compare that with the calculated line, y = 4.137x + 53.58. Close, but not the same. OpenStax makes exactly this point: if each person in a class fits a line by eye, each draws a different line. That is why statistics uses one agreed rule, least squares, to pick a single best line.
How the Least Squares Line Is Calculated
Residuals: How Far Each Point Is From the Line
For any line, each point has a residual: its actual y value minus the y value the line predicts at that x. OpenStax defines it as the vertical distance between the data point and the line. A point above the line has a positive residual, because the line underestimates it. A point below has a negative residual.
Take Dee, who studied 3 hours and scored 61. The line predicts 4.137 times 3 + 53.58, about 66.0. Her residual is 61 minus 66.0, about minus 5. Gus studied 5 hours and scored 78; the line predicts about 74.3, so his residual is about plus 3.7.
The Rule: Make the Squared Residuals as Small as Possible
Square every residual and add them up. OpenStax calls the total the sum of squared errors, or SSE. The least squares line is the one line with the smallest possible SSE: any other line you could draw has a larger total. Squaring does two jobs. It stops positive and negative residuals from cancelling out, and it punishes big misses more than small ones.
The idea is old. The NIST/SEMATECH e-Handbook of Statistical Methods says the method of least squares was developed independently in the late 1700s and early 1800s by Carl Friedrich Gauss, Adrien Marie Legendre and possibly Robert Adrain, and calls linear least squares regression by far the most widely used modeling method.
The Formulas for Slope and Intercept
You do not need calculus to use the result. With x̄ as the mean of the x values and ȳ as the mean of the y values:
- Slope: b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²
- Intercept: a = ȳ − b × x̄
Two facts follow from the formulas and make good checks. First, the line always passes through the point (x̄, ȳ). For the twelve students that point is (4.875, 73.75). Second, the slope can also be written as r times the standard deviation of y divided by the standard deviation of x, so the slope and the correlation r always have the same sign.
In practice nobody does this by hand for more than a handful of points. A spreadsheet, a TI-84 or the scatter plot maker does it instantly, and all of them use the same least squares rule, so they give the same answer for the same data.
How to Read the Equation y = mx + b
US schools usually write the line as y = mx + b. OpenStax writes the same line as y = a + bx. Only the letters change: one letter is the slope, the other the y-intercept.
The Slope Is the Average Change Per Unit
OpenStax’s rule for the slope is to say it in plain words, in the context of the data: it is how much y changes, on average, for every one unit increase in x. For the students, a slope of 4.137 means each extra hour of study goes with about 4.1 more points, on average. For a line that falls, the slope is negative.
A falling line of best fit, y = −2.382x + 31.8. Each extra year of age goes with a price about $2,380 lower, on average. Sample data.
Show the data
| Age (years) | Price ($ thousands) | Label |
|---|---|---|
| 1 | 31 | |
| 2 | 27 | |
| 3 | 25 | |
| 4 | 21 | |
| 5 | 19 | |
| 6 | 17 | |
| 7 | 14 | |
| 8 | 13 | |
| 9 | 11 | |
| 10 | 9 |
The Intercept Is the Value of y When x Is Zero
The intercept is where the line crosses the y-axis. Sometimes it means something: a student who studied zero hours would be predicted to score about 54. Often it does not. The car line crosses at 31.8, a price of $31,800 for a car aged zero, but no car in the data was newer than one year, so that number is a mathematical anchor, not a fact about new cars.
Using the Line to Predict
To predict, put an x value into the equation. A student who studies 5.5 hours: 4.137 times 5.5 + 53.58 is about 76.3, so the prediction is a score of about 76.
OpenStax adds the rule that matters most: predict only for x values inside the range of your data. The students studied between 1 and 9 hours, so 5.5 is safe. Twenty hours of study would predict a score over 136 on a test out of 100, which shows why extrapolating past the data is a guess, not a result.
What R² Tells You About the Fit
R² Is the Share of Variation the Line Explains
R², the coefficient of determination, runs from 0 to 1. OpenStax defines it as the square of the correlation coefficient r and reads it as a percent: the share of the variation in y that the line explains through x. For the students, R² = 0.924, so about 92 percent of the differences in scores follow the line, and about 8 percent is scatter around it.
Excel’s documentation describes R² the same way, as a number from 0 to 1 that shows how closely the trendline’s estimates match the actual data. Google Sheets shows it as an optional label on a trendline.
r Adds the Direction
The correlation coefficient r runs from minus 1 to plus 1, and its sign always matches the sign of the slope. OpenStax notes that r was developed by Karl Pearson in the early 1900s. For the students, r is about 0.961; for the cars it is about minus 0.991. Both are strong. One rises and one falls.
OpenStax gives no official cut-off for “strong” or “weak”, so be careful with any rule that says a fixed value of r always means a strong relationship. Look at the plot and at what the numbers mean in your subject.
When a Line of Best Fit Misleads
A line of best fit can be calculated for any set of points. That does not mean it should be.
The Pattern Is Curved
Here is the temperature through one day. It rises in the morning, peaks mid-afternoon and falls in the evening.
A curved pattern with a straight line forced through it. R² is only 0.101, yet the relationship between hour and temperature is very strong. Sample data.
Show the data
| Hour of the day (24 h clock) | Temperature (°F) | Label |
|---|---|---|
| 6 | 52 | |
| 8 | 60 | |
| 10 | 68 | |
| 12 | 75 | |
| 13 | 78 | |
| 14 | 79 | |
| 15 | 78 | |
| 16 | 75 | |
| 18 | 67 | |
| 20 | 58 |
The straight line is almost flat and R² is about 0.10, which could fool you into saying the hour and the temperature are unrelated. They are strongly related, just not in a straight line. OpenStax warns that curved patterns can give a correlation near 0, and NIST’s handbook says that when no straight line can describe the points, a curved function is needed, with a quadratic as the simplest one to try. Always look at the picture before trusting the number.
One Point Pulls the Line
Add one student to the class: Max studied 8.5 hours but was ill on test day and scored 52.
One unusual point changes the line. The slope drops from 4.137 to 2.613 and R² falls from 0.924 to 0.323. Sample data.
Show the data
| Hours studied | Test score | Label |
|---|---|---|
| 1 | 58 | Ava |
| 2 | 62 | Ben |
| 2.5 | 66 | Cam |
| 3 | 61 | Dee |
| 4 | 72 | Eli |
| 4.5 | 70 | Fay |
| 5 | 78 | Gus |
| 6 | 75 | Hal |
| 6.5 | 84 | Ivy |
| 7 | 81 | Jon |
| 8 | 90 | Kim |
| 9 | 88 | Lou |
| 8.5 | 52 | Max |
One point out of thirteen cut the slope by more than a third and R² from 0.92 to 0.32. OpenStax suggests a rough rule: flag any point more than two standard deviations of the residuals away from the line as a possible outlier. It also warns about influential points, far from the others along the x-axis, which can swing the slope; remove one and refit to see how much it matters.
What to do next is a judgment call. NIST’s handbook says an outlier that comes from a different process should be left out of the fit, while its section on outliers and OpenStax both say to investigate the cause first, because some outliers are errors and some carry the most important information in the data. A fair practice is to report both lines and say which one you used.
The Spread Changes Along the Line
Sometimes the points hug the line at one end and fan out at the other. The line can still be drawn, but predictions at the wide end are much less certain. OpenStax recommends a residual plot, with x along the bottom and each residual up the side: for a good straight-line fit it should look random, with no pattern and roughly even spread.
The Line Is Read as a Cause
A line of best fit shows that two things move together. It does not show that one causes the other. OpenStax says it plainly: correlation does not imply causation. More study hours going with higher scores is a plausible cause and effect, but the line alone cannot prove it. If a third factor, such as the season or the age of the people measured, drives both variables, the line will look just as convincing and mean nothing about cause.
Line of Best Fit in Excel, Google Sheets and on a Calculator
Every common tool draws the same least squares line.
- Excel: add a linear trendline to a scatter chart and turn on the options that show the equation and R² on the chart. Excel’s documentation says its linear trendline uses least squares with y = mx + b.
- Google Sheets: double-click the chart, then Customize, Series, Trendline. Google documents the linear trendline as y = mx + b and shows R² as an option.
- Online: the scatter plot maker draws the line, prints the equation and R², and lists the slope, intercept and r for each group.
The step-by-step clicks for both spreadsheets are in our guide on how to make a scatter plot in Excel and Google Sheets. For practice with answer keys, use the scatter plot and line of best fit worksheets, and for the basics of reading the plot itself, see what a scatter plot is.
A Checklist Before You Use the Line
- Plot first. Is the pattern roughly straight? If it curves, do not force a line.
- Check for outliers. Does one point sit far from the others? Refit without it and compare.
- Read the slope in words. “Each extra unit of x goes with about m more units of y, on average.”
- Report R² with the line. A line with R² near 0 describes very little.
- Predict only inside your data. Stay within the smallest and largest x you measured.
- Say “goes with”, not “causes”. The line shows association, nothing more.
Try It With Your Own Data
Paste two columns from a spreadsheet, or type the pairs, and the line of best fit, its equation and R² appear as you type. Turn on groups to fit a separate line for each group. Your data stays in your browser.
Open the example with the outlier and try removing it
Questions people ask
What is a best fit line on a scatter plot?
It is one straight line that summarizes the trend in a cloud of points, so you can describe the relationship with a slope and a starting value instead of a list of pairs. Most calculators and spreadsheets mean the least squares line when they say best fit, trendline or linear regression.
How do you calculate a best fit line?
Find the mean of x and the mean of y. For every point, multiply its x distance from the mean by its y distance, add those products, and divide by the sum of squared x distances. That is the slope. The intercept is the mean of y minus the slope times the mean of x.
When should a line of best fit be used?
Use one when the scatter plot shows a roughly straight pattern and one variable plausibly helps explain or predict the other. Skip it when the points curve, when the cloud has no shape, or when a single odd point drives the result. Predictions are only safe inside the range of x you measured.
What is the most reliable trend line to use on a scatter plot?
The one whose shape matches the pattern you see. For points that follow a straight path, the linear least squares line is the standard choice. If the points bend, a straight line gives a poor fit however it is computed, so a curve type such as polynomial or logarithmic may describe the data better.
Which two points should the line of best fit go through to best represent?
It does not have to pass through any data point. The least squares line always passes through the point made of the two averages, mean x and mean y. When you draw by eye, aim for about as many points above the line as below it, spread along its whole length.
What type of graph has a line of best fit?
A scatter plot, where every dot is one pair of measurements. The line is laid over the dots to show the overall trend. Excel and Google Sheets also let you add a trendline to line, bar and column charts, but for two numeric measurements the scatter plot is the natural home.