Introduction
Creating a scatter plot in R is one of the most intuitive ways to visualize the relationship between two continuous variables. Whether you are exploring data for a school project, conducting scientific research, or preparing visualizations for a professional report, mastering this plot type will give you immediate insight into patterns, trends, and outliers. In this guide we will walk through the entire process—from installing necessary packages to customizing the appearance of your chart—so you can produce clear, publication‑ready scatter plots quickly and confidently.
Steps
1. Prepare Your Data
Before you can plot, you need a data frame that contains at least two numeric columns. R works best with vectors of equal length, so make sure your observations line up correctly.
# Example data
data <- data.frame(
x = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10),
y = c(2, 5, 3, 8, 7, 6, 9, 12, 11, 14)
)
Tip: Use the head() function to preview your data and str() to confirm column types. If any column is stored as a factor, coerce it to numeric with as.numeric() Turns out it matters..
2. Choose Your Plotting Method
R offers two popular approaches for scatter plots: base R graphics and the ggplot2 package from the tidyverse. Base R is built‑in and sufficient for simple plots, while ggplot2 provides a grammar of graphics that makes customization far more flexible.
Base R – The Quick Method
plot(x = data$x,
y = data$y,
main = "Scatter Plot – Base R",
xlab = "X‑axis Label",
ylab = "Y‑axis Label",
pch = 19, # solid points
col = "steelblue")
plot()is the fundamental function for scatter plots in base R.pchcontrols point shape;colsets the point color.
ggplot2 – The Flexible Method
library(ggplot2)
ggplot(data, aes(x = x, y = y)) +
geom_point(color = "darkgreen", size = 3) +
labs(title = "Scatter Plot – ggplot2",
x = "X‑axis Label",
y = "Y‑axis Label")
aes()defines the aesthetic mappings (x, y).geom_point()adds the points; you can tweak color, size, shape, and alpha here.
3. Enhance the Plot
Adding a Trend Line (Linear Regression)
Base R:
plot(data$x, data$y, pch = 19, col = "steelblue")
abline(lm(data$y ~ data$x), col = "red", lwd = 2)
ggplot2:
ggplot(data, aes(x = x, y = y)) +
geom_point() +
geom_smooth(method = "lm", se = FALSE, color = "red", size = 1) +
labs(title = "Scatter Plot with Regression Line")
The lm() function fits a linear model; geom_smooth() draws the line and, with se = FALSE, removes the confidence band for a cleaner look.
Customizing Colors and Shapes
In ggplot2 you can map additional variables to color, shape, or size, turning a simple scatter plot into a multivariate visualization.
ggplot(data, aes(x = x, y = y, color = factor(group), shape = factor(group))) +
geom_point(size = 4) +
labs(title = "Multi‑Variable Scatter Plot")
Here group is an extra column you add to data. This technique is especially useful for comparing categories within the same plot Worth knowing..
4. Save Your Plot
Once you’re satisfied with the visualization, export it for reports or presentations That's the part that actually makes a difference..
Base R:
pdf("scatter_plot_base.pdf") # or png("scatter_plot_base.png", width = 800, height = 600)
plot(data$x, data$y, pch = 19, col = "steelblue")
dev.off()
ggplot2:
ggsave("scatter_plot_ggplot.png", width = 8, height = 6, dpi = 300)
ggsave() automatically uses the current ggplot object, making file generation a one‑liner Turns out it matters..
Scientific Explanation
A scatter plot displays individual data points as coordinates on a two‑dimensional plane, where the horizontal axis represents one variable (often called the independent or explanatory variable) and the vertical axis represents the other (dependent variable). By inspecting the spatial distribution of points, analysts can quickly assess:
- Direction: Positive (upward trend) or negative (downward trend) correlation.
- Strength: How tightly points cluster around an imaginary line.
- Form: Linear, curvilinear, or no apparent relationship.
- Outliers: Points that deviate markedly from the overall pattern.
The underlying mathematics for a linear trend line is the least‑squares regression, which minimizes the sum of squared vertical distances between observed points and the line. In R, lm(y ~ x) computes the coefficients (intercept and slope) that define this optimal line. When you overlay the line on a scatter plot, you are essentially visualizing the fitted model, making it easier to predict values or discuss the strength of association Turns out it matters..
Scatter plots are not limited to simple bivariate displays. By encoding a third variable through color, shape, or size, you create a multivariate scatter plot, also known as a bubble plot when size encodes a fourth dimension. This extension follows the same principle: each point’s position still reflects two variables, while additional aesthetics reveal extra layers of information Nothing fancy..
FAQ
Q: Do I need to install anything to create a scatter plot?
A: No. Base R comes with the plot() function, so you can generate a basic scatter plot immediately. If you prefer the richer customization options of ggplot2, install it once with install.packages("ggplot2") and load it with library(ggplot2).
Q: What’s the difference between plot() and ggplot()?
A: plot() belongs to base R graphics and uses a procedural style—you set parameters directly in the function call. ggplot() follows the grammar of graphics paradigm, building a plot layer by layer. ggplot2 is generally more flexible for complex designs, while base R is quicker for simple, one‑off visualizations.
Q: Can I add a regression line without writing the lm() function manually?
A: Yes. Both base R and ggplot2 have built‑in shortcuts. In base R, `abline(l
In base R, you can add the regression line directly with
abline(lm(y ~ x), col = "steelblue", lwd = 2)
where lm(y ~ x) fits the least‑squares model and abline draws the resulting intercept and slope. Adjusting col and lwd lets you match the line to your plot’s theme.
Enhancing the ggplot2 scatter plot
Beyond geom_smooth(method = "lm"), ggplot2 offers several ways to enrich a scatter plot:
| Goal | ggplot2 code | What it adds |
|---|---|---|
| Confidence band | geom_smooth(method = "lm", se = TRUE, fill = "lightgray") |
Shaded area showing the uncertainty of the fitted line. |
| Color gradient | geom_point(aes(color = z)) + scale_color_viridis_c() |
Maps a fourth variable to a perceptually uniform color scale. That's why |
| Point size encoding | geom_point(aes(size = z)) |
A third continuous variable (z) visualized as bubble size. |
| Faceting | facet_wrap(~ category) |
Creates small multiples for each level of a categorical variable. |
| Custom theme | `theme_minimal(base_size = 12) + theme(panel. | |
| Non‑linear fit | geom_smooth(method = "loess") |
A locally weighted regression that captures curvature. grid = element_line(color = "grey80"))` |
Putting it all together, a sophisticated multivariate scatter plot might look like:
ggplot(df, aes(x = predictor, y = response,
color = temperature, size = pressure)) +
geom_point(alpha = 0.7) +
geom_smooth(method = "lm", se = TRUE, color = "black", linetype = "dashed") +
scale_color_viridis_c(option = "C") +
labs(title = "Response vs. Predictor",
subtitle = "Colored by temperature, sized by pressure",
x = "Predictor variable",
y = "Response variable",
color = "Temperature (°C)",
size = "Pressure (kPa)") +
theme_minimal(base_size = 11)
This single block produces a plot where:
- Position shows the core bivariate relationship,
- Color reveals how temperature modulates that relationship,
- Size highlights pressure variations,
- The dashed line with a gray confidence band summarizes the linear trend and its uncertainty.
Practical tips
- Check assumptions – Before interpreting a regression line, examine residual plots (
plot(lm(y ~ x))) to verify linearity, homoscedasticity, and normality. - Avoid overplotting – With large datasets, use
geom_jitter(),alphatransparency, orgeom_hex()/geom_bin2d()to reveal density. - Save with appropriate resolution – For manuscripts, 300 dpi is standard; for presentations, 150 dpi often suffices. Adjust
widthandheightinggsave()to match the target column or slide dimensions. - Version control – Keep the script that generates the plot alongside the data; this ensures reproducibility and makes it easy to tweak aesthetics later.
Conclusion
Scatter plots remain one of the most intuitive tools for exploring relationships between variables. Even so, whether you prefer the quick, procedural approach of base R’s plot() and abline() or the layered, grammar‑of‑graphics flexibility of ggplot2, both environments let you move from raw data to insightful visualizations in just a few lines of code. By adding regression lines, confidence bands, and additional aesthetic mappings, you transform a simple point cloud into a rich, multivariate story that highlights direction, strength, form, and outliers—all while keeping the workflow reproducible and publication‑ready. Embrace these techniques, and your data will speak clearly through every plot you create.
It sounds simple, but the gap is usually here.