Skip to content

Guide

Scatter Plot: What It Is, How to Read One, and Examples

A scatter plot shows how two numbers relate. Learn to read direction, strength and outliers, spot correlation and when not to use one, with live examples.

Published September 30, 2026

In this guide
  1. What Is a Scatter Plot?
  2. How to Read a Scatter Plot
  3. Correlation: Positive, Negative and None
  4. Patterns Beyond a Straight Line
  5. Correlation Does Not Mean Causation
  6. Adding a Line of Best Fit
  7. Scatter Plot Examples
  8. When to Use a Scatter Plot, and When Not To
  9. Scatter Plot vs Line Graph vs Histogram
  10. Scatter Plots With More Than Two Variables
  11. How to Make a Scatter Plot
  12. Common Mistakes
  13. Try It With Your Own Data

A scatter plot is a graph of dots that shows how two numeric variables relate. Each dot is one case, such as one student, one day or one car, placed by two measurements: one along the horizontal axis and one up the vertical axis. Look at the whole cloud and you can see whether the two measurements move together, how closely, and which cases do not fit. This guide explains what a scatter plot is, how to read one, what correlation looks like, when a scatter plot is the wrong choice, and how to make one from your own numbers.

What Is a Scatter Plot?

The NIST/SEMATECH e-Handbook of Statistical Methods defines a scatter plot as a plot of the values of one variable, Y, against the corresponding values of another, X. Y goes on the vertical axis and is usually the response; X goes on the horizontal axis and is the variable you suspect is related to it. The handbook says the plot reveals relationships or association between the two, and lists the questions it answers:

  • Are X and Y related?
  • Are they related in a straight line?
  • Are they related in a curve?
  • Does the spread of Y change as X changes?
  • Are there outliers?

The same handbook calls the scatter plot the cornerstone of exploratory data analysis, and the most important and most heavily used graph for uncovering relationships between variables. You will also see it called a scatter graph, scatter chart, scatter diagram or XY graph: Microsoft’s Excel documentation says a scatter chart is often referred to as an xy chart, and Google Sheets lists it as a scatter chart.

Here is a simple one. Twelve students reported how many hours they studied, and each dot shows one student’s hours and test score.

Hours studied and test scoreScatter plot with 12 points. (1, 58), (2, 62), (2.5, 66), (3, 61), (4, 72), (4.5, 70), (5, 78), (6, 75), (6.5, 84), (7, 81), (8, 90), (9, 88).Hours studied and test scoreEach dot is one student556065707580859095012345678910Ava: (1, 58)Ben: (2, 62)Cam: (2.5, 66)Dee: (3, 61)Eli: (4, 72)Fay: (4.5, 70)Gus: (5, 78)Hal: (6, 75)Ivy: (6.5, 84)Jon: (7, 81)Kim: (8, 90)Lou: (9, 88)Test scoreHours studied

A scatter plot of hours studied and test score. Each dot is one student. Sample data.

Show the data
Hours studiedTest scoreLabel
158Ava
262Ben
2.566Cam
361Dee
472Eli
4.570Fay
578Gus
675Hal
6.584Ivy
781Jon
890Kim
988Lou

Notice what is missing: there are no bars, no connecting line and no categories. NIST notes that the most common way to draw a scatter plot is a mark at each point with nothing joining them, so that nothing reaches the screen except the data.

Open this scatter plot in the scatter plot maker

How to Read a Scatter Plot

OpenStax’s Introductory Statistics (2023 edition) advises looking for the overall pattern first and then for any deviations from it. In practice, read a scatter plot in five steps.

1. Read the Axes

What is measured along the bottom, and what up the side? In what units? Is either axis starting far from zero? A scatter plot’s axes do not have to start at zero, but you should know where they start before you judge distances.

2. Find the Direction

Do the dots drift up from left to right, down, or neither? OpenStax describes the two clear directions: high values of one variable with high values of the other (and low with low), or high values of one with low values of the other.

3. Name the Form

Does the cloud follow a straight path, a curve, or no shape at all? Straight patterns are the most common in textbooks, but a curve is just as much a relationship.

4. Judge the Strength

How tightly do the dots hug the path? OpenStax says the strength is judged by how close the points are to a line or to another function. A narrow band is strong; a loose cloud is weak.

5. Look for Exceptions

Is there a dot far from the others? Are there separate clumps? Does the spread widen at one end? Exceptions are often the most interesting part of the plot.

Put together, the students’ plot reads like this: as hours of study rise, scores rise; the pattern is roughly straight; the dots stay fairly close to it; and there are no obvious outliers.

Correlation: Positive, Negative and None

Correlation is the word for how two variables move together. A scatter plot is the fastest way to see it.

Positive Correlation

When one variable goes up, the other tends to go up. NIST’s example of a strong positive linear relationship is one where a straight line fits comfortably through the points, the scatter about the line is small, and small X goes with small Y while large X goes with large Y. The students’ plot above is a positive correlation.

Negative Correlation

When one variable goes up, the other tends to go down. The dots drift from the top left to the bottom right.

Age of a used car and its priceScatter plot with 10 points. (1, 31), (2, 27), (3, 25), (4, 21), (5, 19), (6, 17), (7, 14), (8, 13), (9, 11), (10, 9).Age of a used car and its priceNegative correlation: older cars sell for less510152025303501234567891011(1, 31)(2, 27)(3, 25)(4, 21)(5, 19)(6, 17)(7, 14)(8, 13)(9, 11)(10, 9)Price ($ thousands)Age (years)

Negative correlation. Older cars of the same model sell for less, and the dots fall from left to right. Sample data.

Show the data
Age (years)Price ($ thousands)Label
131
227
325
421
519
617
714
813
911
109

No Correlation

When knowing one variable tells you nothing about the other, the dots form a shapeless cloud. NIST describes it as a plot where, for a given X, the Y values range all over the place, leading to the conclusion of no relationship.

Day of the month born and quiz scoreScatter plot with 14 points. (2, 71), (4, 88), (5, 62), (7, 79), (9, 70), (11, 91), (13, 66), (15, 84), (17, 73), (20, 60), (22, 86), (24, 75), (26, 68), (28, 90).Day of the month born and quiz scoreNo correlation: the dots form a shapeless cloud556065707580859095051015202530(2, 71)(4, 88)(5, 62)(7, 79)(9, 70)(11, 91)(13, 66)(15, 84)(17, 73)(20, 60)(22, 86)(24, 75)(26, 68)(28, 90)Quiz scoreDay of the month born

No correlation. The day of the month a student was born tells you nothing about the quiz score. Sample data.

Show the data
Day of the month bornQuiz scoreLabel
271
488
562
779
970
1191
1366
1584
1773
2060
2286
2475
2668
2890

Strong and Weak, and the Number r

Strength is how tightly the dots follow the pattern. It is often summarized by the correlation coefficient r, which OpenStax says was developed by Karl Pearson in the early 1900s. It runs from minus 1 to plus 1:

r What it means
Close to +1 Strong positive straight-line relationship
Close to 0 Little or no straight-line relationship
Close to −1 Strong negative straight-line relationship
Exactly +1 or −1 Every point lies on one straight line

For the students r is about 0.96, for the cars about minus 0.99, and for the birthdays about 0.10. The textbook sets no official cut-off between strong and weak, so be wary of any rule that treats one value as the boundary. A curious special case, also from OpenStax: if every point lies on a horizontal line, the fit is perfect but there is no relationship, because Y does not change at all.

Patterns Beyond a Straight Line

Curved Relationships

Some strong relationships are not straight. Temperature through a day rises in the morning, peaks in the afternoon and falls in the evening.

Temperature through one dayScatter plot with 10 points. (6, 52), (8, 60), (10, 68), (12, 75), (13, 78), (14, 79), (15, 78), (16, 75), (18, 67), (20, 58).Temperature through one dayA strong relationship that is not a straight line505560657075808546810121416182022(6, 52)(8, 60)(10, 68)(12, 75)(13, 78)(14, 79)(15, 78)(16, 75)(18, 67)(20, 58)Temperature (°F)Hour of the day (24 h clock)

A curved relationship. The hour tells you a lot about the temperature, but not in a straight line. Sample data.

Show the data
Hour of the day (24 h clock)Temperature (°F)Label
652
860
1068
1275
1378
1479
1578
1675
1867
2058

This is where r misleads. OpenStax warns that curved data can give a correlation of 0, and for these ten readings r is only about 0.32, even though the hour predicts the temperature well. NIST’s handbook says that when no straight line can describe the points, a curved function is needed, and suggests trying a quadratic first. The lesson: always look at the plot, not just at r.

Outliers

An outlier is a point that sits far from the pattern of the rest. NIST defines an outlier more generally as an observation at an abnormal distance from the other values, and leaves the exact meaning of abnormal to the analyst.

Hours studied and test scoreScatter plot with 13 points. (1, 58), (2, 62), (2.5, 66), (3, 61), (4, 72), (4.5, 70), (5, 78), (6, 75), (6.5, 84), (7, 81), (8, 90), (9, 88), (8.5, 52).Hours studied and test scoreOne student does not fit the pattern5060708090100012345678910(1, 58)(2, 62)(2.5, 66)(3, 61)(4, 72)(4.5, 70)(5, 78)(6, 75)(6.5, 84)(7, 81)(8, 90)(9, 88)Max: (8.5, 52)MaxTest scoreHours studied

One student, Max, studied many hours but scored far below the pattern. Sample data.

Show the data
Hours studiedTest scoreLabel
158
262
2.566
361
472
4.570
578
675
6.584
781
890
988
8.552Max

Outliers are worth investigating, not just deleting. NIST advises understanding why they appeared before removing them, because they may be bad data or may carry valuable information. Here, a note that Max was ill on test day would explain the point. To see how one point can move a line of best fit, read our guide to the line of best fit.

Clusters

Sometimes the dots fall into separate groups. Microsoft’s Excel documentation names clusters, along with linear or non-linear trends and outliers, as the patterns scatter charts are good at showing.

Trunk width and height of two tree speciesScatter plot with 14 points in 2 groups. Species A: (6, 22), (7, 25), (8, 24), (9, 28), (10, 30), (11, 29), (12, 33). Species B: (14, 48), (15, 52), (16, 50), (17, 56), (18, 55), (19, 60), (20, 58).Trunk width and height of two tree speciesTwo groups form two separate clustersSpecies ASpecies B20304050607046810121416182022(6, 22), Species A(7, 25), Species A(8, 24), Species A(9, 28), Species A(10, 30), Species A(11, 29), Species A(12, 33), Species A(14, 48), Species B(15, 52), Species B(16, 50), Species B(17, 56), Species B(18, 55), Species B(19, 60), Species B(20, 58), Species BHeight (feet)Trunk diameter (inches)

Two clusters. Measurements of two tree species form two separate groups, each with its own trend. Sample data.

Show the data
Trunk diameter (inches)Height (feet)Label
622
725
824
928
1030
1129
1233
1448
1552
1650
1756
1855
1960
2058

A cluster usually means a hidden grouping variable, such as species, region or product type. Coloring the dots by group, as above, makes it visible. Fitting one line through both clusters would mix two different stories.

Spread That Changes

In some data the dots hug the trend at one end and fan out at the other. Statisticians call this heteroscedasticity. NIST’s example shows small X values with small scatter in Y and large X values with large scatter, and notes that an ordinary fit still gives unbiased coefficients but less precise ones. For a reader, the practical point is that predictions are much less certain where the fan is wide.

Correlation Does Not Mean Causation

This is the most important rule for any scatter plot. NIST puts it directly: causality implies association, but association does not imply causality. A scatter plot reveals association, which the handbook calls only step one toward cause and effect, and no statistical procedure, the scatter plot included, can prove cause and effect; that conclusion rests on the researcher’s knowledge of the subject.

OpenStax says the same in the short form most students learn: correlation does not imply causation. When two variables move together, there are always several possible explanations. X may affect Y, Y may affect X, a third factor may drive both, or the pattern may be a coincidence in a small sample. The plot cannot tell these apart. Only knowledge of how the data was produced, or a controlled experiment, can.

Adding a Line of Best Fit

When the pattern is roughly straight, a line of best fit summarizes it. NIST notes that a straight-line pattern suggests a linear regression model may be appropriate. The line gives you a slope (“scores rise about 4.1 points per hour”), a starting value and a way to predict.

Hours studied and test scoreScatter plot with 12 points. (1, 58), (2, 62), (2.5, 66), (3, 61), (4, 72), (4.5, 70), (5, 78), (6, 75), (6.5, 84), (7, 81), (8, 90), (9, 88). Line of best fit y = 4.137x + 53.58, R squared 0.924.Hours studied and test scoreThe same students with a line of best fity = 4.137x + 53.58R² = 0.9245060708090100012345678910Line of best fit: y = 4.137x + 53.58Ava: (1, 58)Ben: (2, 62)Cam: (2.5, 66)Dee: (3, 61)Eli: (4, 72)Fay: (4.5, 70)Gus: (5, 78)Hal: (6, 75)Ivy: (6.5, 84)Jon: (7, 81)Kim: (8, 90)Lou: (9, 88)Test scoreHours studied

The students' scatter plot with the least squares line, its equation and R². Sample data.

Show the data
Hours studiedTest scoreLabel
158Ava
262Ben
2.566Cam
361Dee
472Eli
4.570Fay
578Gus
675Hal
6.584Ivy
781Jon
890Kim
988Lou

OpenStax adds that a regression line only makes sense when one variable helps explain or predict the other. How to draw the line by eye, how least squares works and when the line misleads are covered in detail in our line of best fit guide.

Scatter Plot Examples

Scatter plots appear wherever two measurements are taken on the same things. A few common pairings, as illustrations:

  • School: hours studied and test score; arm span and height; shoe size and reading level in young children, where a third factor, age, could sit behind any pattern you find.
  • Science: the reading of an instrument against a known reference value, to check calibration; NIST’s handbook uses a load cell calibration case study to demonstrate the scatter plot.
  • Business: advertising spending and sales; price and units sold; delivery distance and delivery time.
  • Health: age and resting heart rate; hours of sleep and reaction time.
  • Everyday life: the age and price of used cars; the outdoor temperature and the heating bill.

For each one, the question is the same: when one number changes, what tends to happen to the other?

When to Use a Scatter Plot, and When Not To

Use One When

  • Both variables are numbers measured on the same cases.
  • You want to know whether and how they are related.
  • The order of the rows does not matter.
  • You want to spot outliers or groups before fitting a model.

Choose Something Else When

  • The horizontal variable is a category. Product names, countries or survey answers have no numeric position. Use a bar graph.
  • The horizontal variable is time in even steps. Microsoft’s rule of thumb is to use a line chart when your x values are not numbers or are evenly spaced labels such as months or years, and a scatter chart when the x values are numeric. A line graph connects the points in order, which is what a trend over time needs.
  • You care about one variable on its own. To see how a single measurement is spread out, use a histogram or a box plot.
  • You have thousands of overlapping points. The cloud turns into a solid blob. Smaller or transparent dots help, or a summary chart instead.

Scatter Plot vs Line Graph vs Histogram

These three are often confused because all of them have two axes.

Feature Scatter plot Line graph Histogram
Variables Two numeric One numeric over an ordered sequence One numeric
Each mark is One case One step in the sequence A count of cases in an interval
Points connected? No Yes, in order Bars touch
Answers Are X and Y related? How did it change over time? How are the values spread?

Microsoft’s documentation explains the key difference between scatter and line charts: a scatter chart has two value axes, while a line chart has a category axis along the bottom that spaces its points evenly whatever their values. Plot uneven numeric x values on a line chart and the picture is distorted. A histogram is a different kind of graph altogether: it summarizes one variable into bins and counts, so it cannot show a relationship between two.

Scatter Plots With More Than Two Variables

A plain scatter plot shows two variables. There are three common ways to add more.

Color or Shape for a Group

Give each group its own color, as in the tree species example. This is how Google Sheets treats extra columns too: its help page says each column of Y values shows up as a separate series of points.

Size for a Third Number

A bubble chart is a scatter plot where the size of each dot shows a third value. Microsoft describes it as a type of scatter chart where bubble size adds a third data dimension, with the values in the order x, y, then size.

Many Plots at Once

NIST’s handbook describes the scatterplot matrix, which draws every pair of variables on one page, and the conditioning plot, which draws Y against X separately for different values of a third variable. Both help with data sets that have more than two variables.

How to Make a Scatter Plot

You can make one in a minute with the scatter plot maker, with no sign-up.

  1. Enter the pairs. One row per case: X in the first column, Y in the second.
  2. Or paste from a spreadsheet. Copy two columns from Excel or Google Sheets and paste them. A header row becomes the axis titles.
  3. Title the axes. Always include the units.
  4. Add the line of best fit if the pattern is straight. The equation and R² appear on the chart.
  5. Turn on groups to color the dots by category, with one line per group.
  6. Export a PNG for slides or documents, or an SVG you can scale.

If you work in a spreadsheet, our step-by-step guide covers scatter plots in Excel and Google Sheets. For classroom practice, see the scatter plot worksheets with answer keys.

Common Mistakes

  1. Swapping the axes. Put the explanatory variable on X and the response on Y. Swapping them changes the line of best fit.
  2. Reading correlation as cause. Say “goes with”, not “causes”.
  3. Trusting r without looking. A curve or a single outlier can make r small or large for the wrong reason.
  4. Drawing a line through clusters. Split the groups first.
  5. Connecting the dots. Unless the order of the points means something, lines between dots suggest a sequence that is not there.
  6. Leaving out units. “Score” and “hours” are not enough if the reader cannot tell the scale.

Try It With Your Own Data

Paste two columns of numbers and the scatter plot draws itself, with an optional line of best fit, R² and groups. Your data stays in your browser.

Questions people ask

What is a scatter plot used for?

It is used to check whether two numeric measurements are related: whether one tends to rise or fall as the other rises, how tightly, and whether some cases break the pattern. Scientists, engineers, teachers and analysts use it to explore data before fitting a model or making a prediction.

How to explain a scatter plot?

Say what each axis measures, then describe the cloud in four words: direction, form, strength and exceptions. For example: as hours of study rise, test scores rise in a roughly straight, fairly tight pattern, with one student well below the rest. End by saying it shows association, not cause.

How do I make a scatter plot?

Collect pairs of numbers for the same cases, such as the height and arm span of each student. Put the explanatory variable on the horizontal axis and the response on the vertical axis, draw one dot per pair, and title both axes with their units. A spreadsheet or an online maker does the plotting for you.

What is the difference between a scatter plot and a histogram?

A scatter plot uses two variables and shows how they relate, with one dot per case. A histogram uses a single variable and shows how its values are spread, with bars counting how many values fall in each interval. Use the histogram to see a distribution and the scatter plot to see a relationship.

When not to use scatter plot?

Avoid it when the horizontal variable is a category or a label, such as product names, or when your values are evenly spaced time steps and the order matters. Bars suit categories and a line graph suits time. It also struggles when thousands of points overlap into a single blob.