My Blog: Reading and writting data

1.Load packages we will use

library(tidyverse)
library(here)
library(janitor)
library(skimr)

download co2 emissions per capita from Our world in Data into the directory for the post.
Assign the location of the file to file_csv. The data should be in the same directory as this file. -Read the data into R and assign emissions

file_csv <- here("_posts","2021-02-16-reading-and-writting-data","co-emissions-per-capita.csv")

emissions <- read_csv(file_csv)

show the first 10 rows (observation of) emissions

emissions

# A tibble: 22,383 x 4
   Entity      Code   Year `Per capita CO2 emissions`
   <chr>       <chr> <dbl>                      <dbl>
 1 Afghanistan AFG    1949                    0.00191
 2 Afghanistan AFG    1950                    0.0109 
 3 Afghanistan AFG    1951                    0.0117 
 4 Afghanistan AFG    1952                    0.0115 
 5 Afghanistan AFG    1953                    0.0132 
 6 Afghanistan AFG    1954                    0.0130 
 7 Afghanistan AFG    1955                    0.0186 
 8 Afghanistan AFG    1956                    0.0218 
 9 Afghanistan AFG    1957                    0.0343 
10 Afghanistan AFG    1958                    0.0380 
# … with 22,373 more rows

5.start with emissions data THEN - Use clean_names from the janitor package to make easier to work with - assign the output totidy_emissions - show the first 10 rows of tidy_emission

tidy_emissions  <- emissions %>% 
  clean_names()
tidy_emissions

# A tibble: 22,383 x 4
   entity      code   year per_capita_co2_emissions
   <chr>       <chr> <dbl>                    <dbl>
 1 Afghanistan AFG    1949                  0.00191
 2 Afghanistan AFG    1950                  0.0109 
 3 Afghanistan AFG    1951                  0.0117 
 4 Afghanistan AFG    1952                  0.0115 
 5 Afghanistan AFG    1953                  0.0132 
 6 Afghanistan AFG    1954                  0.0130 
 7 Afghanistan AFG    1955                  0.0186 
 8 Afghanistan AFG    1956                  0.0218 
 9 Afghanistan AFG    1957                  0.0343 
10 Afghanistan AFG    1958                  0.0380 
# … with 22,373 more rows

6 start with the tidy_emissions THEN -use filter to extract rows with years == 2000 -use skim to calculate the descriptive statistics

tidy_emissions %>% 
  filter(year==2000) %>% 
  skim()

Table 1: Data summary
Name	Piped data
Number of rows	219
Number of columns	4
_______________________
Column type frequency:
character	2
numeric	2
________________________
Group variables	None

Variable type: character

skim_variable	n_missing	complete_rate	min	max	empty	n_unique	whitespace
entity	0	1.00	4	32	0	219	0
code	12	0.95	3	8	0	207	0

Variable type: numeric

skim_variable	n_missing	complete_rate	mean	sd	p0	p25	p50	p75	p100	hist
year	0	1	2000.00	0.00	2e+03	2000.00	2000.00	2000.00	2000.00	▁▁▇▁▁
per_capita_co2_emissions	0	1	5.06	6.74	2e-02	0.71	2.82	7.97	58.39	▇▁▁▁▁

13 observations have a missing code. How are these observations different? -start with tidy_emissions then extract rows with year==2000 and are missing a code
```
tidy_emissions %>% 
  filter(year==2000,is.na(code))
```

# A tibble: 12 x 4
   entity                     code   year per_capita_co2_emissions
   <chr>                      <chr> <dbl>                    <dbl>
 1 Africa                     <NA>   2000                     1.11
 2 Asia                       <NA>   2000                     2.40
 3 Asia (excl. China & India) <NA>   2000                     3.35
 4 EU-27                      <NA>   2000                     8.46
 5 EU-28                      <NA>   2000                     8.61
 6 Europe                     <NA>   2000                     8.48
 7 Europe (excl. EU-27)       <NA>   2000                     8.47
 8 Europe (excl. EU-28)       <NA>   2000                     8.19
 9 North America              <NA>   2000                    14.6 
10 North America (excl. USA)  <NA>   2000                     5.39
11 Oceania                    <NA>   2000                    12.6 
12 South America              <NA>   2000                     2.32

11.Use bind_rowsto bind together the max_15_emitters and min_15_emitters - assign the output to max_min_15