Repository navigation
Expand file tree
/
Copy pathj2sr_02_quarto_python-wrangle.qmd
More file actions
123 lines (83 loc) · 3.61 KB
/
Copy pathj2sr_02_quarto_python-wrangle.qmd
File metadata and controls
123 lines (83 loc) · 3.61 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
---
title: "Wrangle rows and columns"
format: html
---
# Instructions
This week, you learned about several more pandas methods and attributes.
These functions help you work with the columns in your data set.
You'll use these skills to add a new column that represents percent change in
child health scores over time.
To learn more about the variables contained in the data set,
refer to the data dictionary [here](https://rstudio.github.io/academy-python-foreign-aid/assets/j2sr_dictionary.html).
# Milestone
```{python}
#| label: 'setup'
#| include: false
# Import your packages here
import pandas as pd
from matplotlib import rcParams
# Set some options
pd.set_option('display.max_columns', None)
rcParams.update({'savefig.bbox': 'tight'}) # Keeps plotnine legend from being cut off
```
## Wrangle rows and columns
In this milestone, you'll create a new variable to calculate annual percent change in
child health scores for a particular country.
## Recreation
### Part 1 - Import data
Before you begin, you will need to import your data set.
Use the code chunk below to read the data from the data file `j2sr.csv`,
which is stored in the `data/` folder in your working directory.
Be sure to save the data set to a variable named `j2sr`.
```{python}
#| label: 'recreation-import'
```
### Part 2 - Create and Summarize
Run the code below to see a table.
```{python}
#| label: 'recreate-this'
#| message: false
solution = pd.read_csv('data/milestone02.csv')
solution
```
Your task is to use what you've learned to transform `j2sr` into this table.
We want to look at how child health scores change across years for a particular country of interest.
We will use `Fiji` in this example.
You will need to:
1. Create a DataFrame called `country` to work with in the following steps.
`country` should contain the rows of `j2sr` where the `country` is `Fiji`.
2. Create a new column called `child_lag` in which each value represents the
value of the `child_health` column from the previous row.
*Hint*: You can use a method called `.shift()` to get the previous values.
3. Calculate percent change between the current year and the previous year.
The formula for percent change is `((current - previous) / previous) * 100`.
Name this new column `child_pct_change`.
```{python}
#| label: 'recreation-create'
```
#### Check your work
Run the following code chunk to test whether your table has the same dimensions as the solution:
```{python}
#| label: 'compare_shape'
# The result will be `True` if the dimensions are the same
solution.shape == country.shape
```
Run the following code chunk to test whether your table matches the solution:
```{python}
#| label: 'compare'
# If your answer is correct, the comparison should return an empty DataFrame.
country.reset_index(drop=True).round(10).compare(solution.round(10))
```
## Extension
Using the code chunk below, investigate a research question about this data, using the additional data wrangling skills you learned this week. Some ideas:
1. Which country had the highest annual percent change in child health? In which year did this occur?
2. How many countries had a negative percent change in child health in a given year?
3. Visualize child health over time for all countries -- does this match your calculations for annual percent change in child health?
4. [any other research question of interest]
Alternately, working with a data set of your own, complete the following:
1. Read in your data
2. Create at least one new variable in your data set using mathematical operations
3. Use your updated data set to create at least one graph and/or table
```{python}
#| label: 'extension'
```