14  Your First Data Wrangling

Learning Objectives: After this chapter you will be able to: - Use the five core dplyr verbs: filter(), select(), mutate(), arrange(), summarise(). - Chain steps with the native base pipe |>. - Know where to go next after this book.

library(dplyr)
# ggplot2 bundles the `mpg` dataset (MIT license, same as ggplot2 itself)
data(mpg, package = "ggplot2")

We use the native base pipe |> (standard since R 4.1.0; zero external dependencies). We do not use %>%.

14.1 The five verbs

mpg |>
  filter(manufacturer == "audi") |>
  select(manufacturer, model, displ, hwy) |>
  mutate(km_per_litre = hwy * 0.425144) |>
  arrange(desc(hwy))
# A tibble: 18 Γ— 5
   manufacturer model      displ   hwy km_per_litre
   <chr>        <chr>      <dbl> <int>        <dbl>
 1 audi         a4           2      31        13.2 
 2 audi         a4           2      30        12.8 
 3 audi         a4           1.8    29        12.3 
 4 audi         a4           1.8    29        12.3 
 5 audi         a4 quattro   2      28        11.9 
 6 audi         a4           3.1    27        11.5 
 7 audi         a4 quattro   2      27        11.5 
 8 audi         a4           2.8    26        11.1 
 9 audi         a4           2.8    26        11.1 
10 audi         a4 quattro   1.8    26        11.1 
11 audi         a4 quattro   1.8    25        10.6 
12 audi         a4 quattro   2.8    25        10.6 
13 audi         a4 quattro   2.8    25        10.6 
14 audi         a4 quattro   3.1    25        10.6 
15 audi         a4 quattro   3.1    25        10.6 
16 audi         a6 quattro   3.1    25        10.6 
17 audi         a6 quattro   2.8    24        10.2 
18 audi         a6 quattro   4.2    23         9.78
mpg |>
  summarise(mean_hwy = mean(hwy), n = dplyr::n())
# A tibble: 1 Γ— 2
  mean_hwy     n
     <dbl> <int>
1     23.4   234

Common Mistake: mean() on a vector with NA returns NA. Add na.rm = TRUE when missing values are present.

Try it in your browser

Edit the code below and press Run. It executes entirely in your browser via WebR β€” no R installation needed. The runtime downloads once, on this page only.

14.2 Practice Check

Knowledge Check: Filter then select What does mpg |> filter(manufacturer == "audi") |> select(model, hwy) return?

Click to reveal answer

Answer: A data frame with two columns (model, hwy) containing only the rows where manufacturer is "audi".

Explanation: filter() keeps matching rows; select() keeps matching columns. The |> pipe passes the left-hand result into the first argument of the right-hand function.

Knowledge Check: mutate() adds columns What does mpg |> mutate(km_per_litre = hwy * 0.425144) return, and how many rows does the result have compared to mpg?

Click to reveal answer

Answer: The full mpg data with one extra column (km_per_litre); the row count is unchanged.

Explanation: mutate() adds (or modifies) columns row-wise β€” it never adds or removes rows. Use filter() to change row counts.

Knowledge Check: arrange() ordering After mpg |> arrange(desc(hwy)), which kind of car sits in row 1?

Click to reveal answer

Answer: The car with the highest highway mileage (hwy).

Explanation: arrange() sorts ascending by default; wrapping in desc() flips it to descending, so the maximum comes first.

Knowledge Check: summarise() collapses rows How many rows does mpg |> summarise(mean_hwy = mean(hwy), n = dplyr::n()) return, and why?

Click to reveal answer

Answer: Exactly one row: the mean of hwy across the whole dataset and the total row count.

Explanation: Without group_by(), summarise() collapses the entire data frame into a single summary row.

Knowledge Check: Reading a pipeline Read this pipeline aloud in order: mpg |> filter(manufacturer == "audi") |> select(model, hwy) |> arrange(desc(hwy)). What question does it answer?

Click to reveal answer

Answer: β€œAmong Audi models, which has the best (and worst) highway mileage?” β€” it keeps only Audis, keeps the model and mileage columns, and sorts best-first.

Explanation: Pipelines read top-to-bottom as a sequence of verbs; narrating each step in plain words is the fastest way to check you built the right one.

14.3 Where to go next

Where to go next: - Book: R for Data Science (2e) β€” start at Chapter 1 - Courses: Rsquared Academy Data Wrangling series - More books: ebooks.rsquaredacademy.com