library(tidyverse)
library(lubridate)RWDIR assignment data
Goal of this notebook
The goal here is to create Billboard Hot 100 data for the Reporting with Data in R book. It is purposely mucked up for the assignment.
In June 2023 the script
action_combine_charts.Rwas updated to create assignment data from “current” data. So this isn’t really needed anymore except for posterity.
Setup
Import data
Import the most current Billboard Hot 100 data.
hot100_current <- read_csv("data-out/hot-100-current.csv")Warning: One or more parsing issues, call `problems()` on your data frame for details,
e.g.:
dat <- vroom(...)
problems(dat)
Rows: 355600 Columns: 8
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): title, performer
dbl (4): current_week, last_week, peak_pos, wks_on_chart
lgl (1): wks_at_no1
date (1): chart_week
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
Clip the year
For the book assignment, we want only full years so we remove rows from the most recent year.
This is bad … we don’t clip the year anymore
# This gets the current year.
hot100_clipped <- hot100_current |>
filter(year(chart_week) < year(Sys.Date()))
hot100_clipped$chart_week |> summary() Min. 1st Qu. Median Mean 3rd Qu. Max.
"1958-08-04" "1975-06-14" "1992-04-18" "1992-04-17" "2009-02-21" "2025-12-27"
Muck the data
We change the names of the column headers and muck up the data to force students to deal with that in Reporting with Data in R Chapter 3.
hot100_assignment <- hot100_clipped %>%
# muck the date
mutate(
chart_week = paste(
month(chart_week) %>% as.character(),
day(chart_week) %>% as.character(),
year(chart_week) %>% as.character(),
sep = "/"
)
) |>
# muck the names
select(
`CHART WEEK` = chart_week,
`THIS WEEK` = current_week,
`TITLE` = title,
`PERFORMER` = performer,
`LAST WEEK` = last_week,
`PEAK POS.` = peak_pos,
`WKS ON CHART` = wks_on_chart
)
hot100_assignment %>% glimpse()Rows: 351,700
Columns: 7
$ `CHART WEEK` <chr> "1/1/2022", "1/1/2022", "1/1/2022", "1/1/2022", "1/1/20…
$ `THIS WEEK` <dbl> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, …
$ TITLE <chr> "All I Want For Christmas Is You", "Rockin' Around The …
$ PERFORMER <chr> "Mariah Carey", "Brenda Lee", "Bobby Helms", "Burl Ives…
$ `LAST WEEK` <dbl> 1, 2, 4, 5, 3, 7, 9, 11, 6, 13, 15, 17, 18, 0, 8, 25, 1…
$ `PEAK POS.` <dbl> 1, 2, 3, 4, 1, 5, 7, 6, 1, 10, 11, 8, 12, 14, 7, 16, 12…
$ `WKS ON CHART` <dbl> 50, 44, 41, 25, 11, 26, 24, 19, 24, 15, 31, 18, 14, 1, …
Export
Here we export the data for the assignment.
This is commented out so I don’t overwrite “current” data.
# hot100_assignment |> write_csv("data-out/hot100_assignment.csv")