RWDIR assignment data

Goal of this notebook

The goal here is to create Billboard Hot 100 data for the Reporting with Data in R book. It is purposely mucked up for the assignment.

In June 2023 the script action_combine_charts.R was updated to create assignment data from “current” data. So this isn’t really needed anymore except for posterity.

Setup

library(tidyverse)
library(lubridate)

Import data

Import the most current Billboard Hot 100 data.

hot100_current <- read_csv("data-out/hot-100-current.csv")
Warning: One or more parsing issues, call `problems()` on your data frame for details,
e.g.:
  dat <- vroom(...)
  problems(dat)
Rows: 355600 Columns: 8
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr  (2): title, performer
dbl  (4): current_week, last_week, peak_pos, wks_on_chart
lgl  (1): wks_at_no1
date (1): chart_week

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Clip the year

For the book assignment, we want only full years so we remove rows from the most recent year.

This is bad … we don’t clip the year anymore

# This gets the current year.
hot100_clipped <- hot100_current |> 
  filter(year(chart_week) < year(Sys.Date()))

hot100_clipped$chart_week |> summary()
        Min.      1st Qu.       Median         Mean      3rd Qu.         Max. 
"1958-08-04" "1975-06-14" "1992-04-18" "1992-04-17" "2009-02-21" "2025-12-27" 

Muck the data

We change the names of the column headers and muck up the data to force students to deal with that in Reporting with Data in R Chapter 3.

hot100_assignment <- hot100_clipped %>%
  # muck the date
  mutate(
    chart_week = paste(
      month(chart_week) %>% as.character(),
      day(chart_week) %>% as.character(),
      year(chart_week) %>% as.character(),
      sep = "/"
      )
  ) |>   
  # muck the names
  select(
    `CHART WEEK` = chart_week,
    `THIS WEEK` = current_week,
    `TITLE` = title,
    `PERFORMER` = performer,
    `LAST WEEK` = last_week,
    `PEAK POS.` = peak_pos,
    `WKS ON CHART` = wks_on_chart
  )

hot100_assignment %>% glimpse()
Rows: 351,700
Columns: 7
$ `CHART WEEK`   <chr> "1/1/2022", "1/1/2022", "1/1/2022", "1/1/2022", "1/1/20…
$ `THIS WEEK`    <dbl> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, …
$ TITLE          <chr> "All I Want For Christmas Is You", "Rockin' Around The …
$ PERFORMER      <chr> "Mariah Carey", "Brenda Lee", "Bobby Helms", "Burl Ives…
$ `LAST WEEK`    <dbl> 1, 2, 4, 5, 3, 7, 9, 11, 6, 13, 15, 17, 18, 0, 8, 25, 1…
$ `PEAK POS.`    <dbl> 1, 2, 3, 4, 1, 5, 7, 6, 1, 10, 11, 8, 12, 14, 7, 16, 12…
$ `WKS ON CHART` <dbl> 50, 44, 41, 25, 11, 26, 24, 19, 24, 15, 31, 18, 14, 1, …

Export

Here we export the data for the assignment.

This is commented out so I don’t overwrite “current” data.

# hot100_assignment |> write_csv("data-out/hot100_assignment.csv")