| Type: | Package |
| Title: | Join World Bank Data, Country Codes and Maps on the ISO Spine |
| Version: | 3.0.0 |
| Description: | A complete toolkit for getting country data onto honest maps. Country names rarely line up across data sources ("US", "U.S.", "United States", "United States of America" are one country, but a naive join treats them as four), so 'countryatlas' makes ISO codes the universal join key. It generalises a one-call, map-ready table that stitches together 'ggplot2' map geometry, 'WDI' World Bank indicators and the 'countrycode' Rosetta stone; exposes the join machinery for the user's own data; ships curated reference data (metadata, group memberships, an indicator catalogue, flags and currencies); adds analysis helpers (per-capita, regional roll-ups, ranking, inequality and convergence statistics); and turns one hand-drawn choropleth into a full vocabulary of projected, area-honest maps (binned and quantile choropleths, proportional-symbol, spike, bivariate, value-by-alpha, cartogram, tile-grid, flow, small-multiple, animated, globe and interactive), and can hand its curated, ISO-reconciled tables to 'ggsql' for database-side spatial rendering. Honesty is treated as a feature rather than a slogan: classification methods can be compared side by side, missing data can be hatched rather than greyed, coverage and provenance travel with the plot, and the distortion each projection introduces can be measured and drawn. Heavy spatial dependencies stay optional, and a bundled offline snapshot lets every example, test and vignette run without the network. |
| License: | GPL (≥ 3) |
| URL: | https://pursuitofdatascience.github.io/countryatlas/, https://github.com/PursuitOfDataScience/countryatlas |
| BugReports: | https://github.com/PursuitOfDataScience/countryatlas/issues |
| Encoding: | UTF-8 |
| Language: | en-GB |
| LazyData: | true |
| Depends: | R (≥ 4.1.0) |
| Imports: | cachem, cli, countrycode, dplyr, ggplot2, grDevices, memoise, parallel, rlang, stats, tibble, tidyr, tools, utils, WDI |
| Suggests: | biscale, cartogram, cartogramR, classInt, comtradr, covr, cshapes, DBI, duckdb, eurostat, gganimate, ggiraph, ggpattern, ggrepel, ggsql, gifski, giscoR, gt, knitr, leaflet, magick, mapgl, mapproj, maps, nanoarrow, OECD, owidR, plotly, regions, rmapshaper, rmarkdown, rnaturalearth, rnaturalearthdata, scales, sf, spdep, stringdist, testthat (≥ 3.2.0), tmap, units, vdiffr, withr |
| VignetteBuilder: | knitr |
| Config/testthat/edition: | 3 |
| RoxygenNote: | 7.3.2 |
| NeedsCompilation: | no |
| Packaged: | 2026-10-01 15:03:42 UTC; youzhi |
| Author: | Youzhi Yu [aut, cre] |
| Maintainer: | Youzhi Yu <yuyouzhi666@icloud.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-01 15:50:18 UTC |
countryatlas: join World Bank data, country codes and maps on the ISO spine
Description
countryatlas exists to kill one recurring source of pain: country names
never line up across data sources. The package makes ISO codes the universal
join key and hands you a ready-to-map tibble that stitches together map
geometry (ggplot2::map_data() or Natural Earth sf), World Bank indicators
(WDI::WDI()) and the countrycode::countrycode() crosswalk.
Details
The happy path stays one call: world_data(). Everything else is opt-in.
Core data assembly
world_data(), country_data(), world_geometry(), locate_country(),
country_borders(), neighbors(), distance_between().
Data sources beyond the World Bank
fetch_indicator(), add_indicator(), compare_sources(),
country_sources(), register_country_source(),
remove_country_source(), and the adapters fetch_owid(),
fetch_eurostat(), fetch_oecd() and fetch_comtrade().
The join engine
standardize_country(), join_world(), attach_geometry(), country_join(),
country_join_all(), dissolve_country().
Diagnostics
check_country_match(), repair_country_names(), country_overrides(),
audit_coverage().
Reference data
convert_country(), country_codes(), country_groups(), in_group(),
wdi_search(), and the datasets country_meta, common_indicators,
country_groups_tbl, country_groups_history, world_snapshot,
world_tiles, historical_codes, disputed_territories.
Time
country_timeline(), audit_time_coverage(), historical_geometry().
Analysis helpers
per_capita(), aggregate_regions(), rank_countries(), complete_years(),
interpolate_missing(), growth_rate(), index_to(), share_of_world(),
lag_by_country(), diff_by_country(), deflate(), to_ppp(),
rate_check(), smooth_rates(), correlate_indicators(),
beta_convergence(), sigma_convergence(), convergence_club(), gini(),
theil().
Spatial statistics and networks
country_weights(), morans_i(), local_morans(), lisa_map(),
gearys_c(), getis_ord(), spatial_lag(); flow_matrix(),
country_network(), od_map().
Visualization
world_map(), globe_map(), spin_globe(), facet_map(), bubble_map(),
spike_map(), bivariate_map(), cartogram_map(), dorling_map(),
gridded_cartogram(), tile_map(), flow_map(), animate_world(),
interactive_map(), geom_country_labels(), theme_world_map(),
simplify_geometry().
Honest maps
classify_compare(), coverage_map(), value_by_alpha_map(),
projection_info(), projection_compare(), projection_distortion(),
tissot_map(), cartogram_diagnostics(), map_provenance(),
dispute_policy(), check_dispute_coverage().
Subnational
standardize_subnational(), nuts_geometry(), subnational_map().
Reporting
country_factsheet(), world_table().
Database rendering (ggsql)
as_ggsql_source(), world_query().
Performance & caching
clear_wdi_cache(), clear_country_cache().
Options
Six options change the package's behaviour. All are unset by default.
countryatlas.cache_dirWhere the persistent World Bank cache lives. Defaults to
tools::R_user_dir("countryatlas", "cache"); set it to""for session-only caching. Seeclear_wdi_cache().countryatlas.cache_max_ageHow long a persistent cache entry stays usable, in seconds. Defaults to 30 days. World Bank figures are revised, so an old entry is not merely stale on disk – it is a different answer from the one the API would give now. A single non-negative number;
Inffor no expiry.countryatlas.cache_max_sizeThe size cap on the persistent cache, in bytes. Defaults to 50 MB, past which the least-recently-used entries are dropped. CRAN policy allows a package cache under
tools::R_user_dir()only if its contents are actively managed. A single non-negative number;Inffor no cap.countryatlas.workersHow many processes fetch indicators in parallel (only when the cache is on disk – a memory-only memo cannot survive a fork). Defaults to one fewer than the available cores, and to 2 under
R CMD check, per CRAN policy. Must be a single finite number; values below one are clamped to one.countryatlas.gdp_compatSet to
TRUEto restore thegdp_per_capita_2015column thatworld_data()emitted in 1.0.0. A deprecation shim, off by default, and now warning when used.countryatlas.dispute_policyWhich map convention disputed territories are drawn under:
"none"(default),"de_facto","de_jure"or"neutral". Set it withdispute_policy()rather than directly, which also reports what the setting does and does not change.
Author(s)
Maintainer: Youzhi Yu yuyouzhi666@icloud.com
See Also
Useful links:
Report bugs at https://github.com/PursuitOfDataScience/countryatlas/issues
Fetch an indicator and join it to your data
Description
fetch_indicator() plus country_join() in one step: pull an indicator from
any registered source and attach it to a frame you already have, matched on
the ISO spine (and on year too, when both sides are panels).
Usage
add_indicator(data, source, indicator, countries = NULL, years = NULL, ...)
Arguments
data |
A frame with |
source, indicator, countries, years, ... |
Passed to |
Value
data with the indicator column(s) added.
How much countries actually saves
countries bounds the result, not necessarily the download. Only some
providers accept a country filter in the request:
| Source | countries reaches the provider? |
comtrade | yes -- sent as reporter |
wdi | partly -- the year range is sent, countries are filtered here |
owid, eurostat, oecd | no -- the full dataset is downloaded, then filtered |
So add_indicator(one_row, "owid", "life-expectancy") still transfers every
country and year that dataset holds in order to keep a single value. When
that matters, narrow with years (which the wdi and comtrade adapters do
push down), or fetch once into a variable and reuse it rather than calling
this per subset. The on-disk cache means a repeated wdi fetch is free;
the other sources are memoised for the session only.
See Also
fetch_indicator(), country_join()
Examples
## Not run:
world_snapshot$countries |>
add_indicator("wdi", c(unemployment = "SL.UEM.TOTL.ZS"), years = 2020)
## End(Not run)
Roll countries up to region / income / continent
Description
Aggregate a country-level value to a coarser grouping, optionally with population-weighted means.
Usage
aggregate_regions(data, value, by = "region", fun = "sum", weight = NULL)
Arguments
data |
A country-level data frame. |
value |
The value column to aggregate (unquoted). |
by |
Grouping column(s) (character), default |
fun |
Aggregation: |
weight |
Optional weight column (unquoted) for |
Value
A tibble of by plus the aggregated value.
Groups with no data
Missing values are dropped before aggregating, so a group is summarised from
whatever it does have. A group with no non-missing value returns NA
rather than a figure: sum() would otherwise report 0, mean() NaN and
min()/max() -Inf/Inf, each of which reads as a real total for a
region we simply have no data for. Use audit_coverage() to see where those
gaps are.
Examples
df <- data.frame(iso3c = c("USA", "CAN", "BRA"),
region = c("North America", "North America", "Latin America"),
gdp = c(21, 1.7, 1.4))
aggregate_regions(df, gdp, fun = "sum")
Animate a choropleth over time
Description
Given a panel from world_data(2000:2020, ...), animate the choropleth over
year via the optional gganimate package, or fall back to a faceted
small-multiple when it is not installed.
Usage
animate_world(data, fill, time = year, projection = "equal_earth", ...)
Arguments
data |
A panel map-ready frame (polygon or sf) with a |
fill |
The fill column (unquoted). |
time |
The time column (unquoted; default |
projection |
Projection for the sf backend. See |
... |
Passed to |
Value
A gganim object (if gganimate is available) or a faceted
ggplot.
Examples
## Not run:
world_data(2000:2005, c(gdp = "NY.GDP.PCAP.KD")) |>
animate_world(gdp)
## End(Not run)
Export a countryatlas table as a ggsql source
Description
Hand countryatlas's curated, ISO-reconciled, WDI-joined spatial table to
ggsql so it can be charted with DRAW spatial – the
bridge that lets ggsql draw maps of your override-corrected data instead of
its static bundled world. sf geometry is WKB-encoded so ggsql can decode it.
Usage
as_ggsql_source(
data,
name = "countryatlas_world",
format = c("duckdb", "parquet", "arrow"),
con = NULL,
path = NULL,
geometry_col = "geometry"
)
Arguments
data |
A map-ready frame (ideally |
name |
The table name to register/write (default |
format |
|
con |
An existing DuckDB |
path |
Output path for |
geometry_col |
Name for the WKB geometry column (default |
Value
Depending on format: a DuckDB connection (with the table written),
a Parquet file path, or a nanoarrow array stream.
You own the connection that format = "duckdb" returns, and duckdb
keeps its in-memory database alive until the handle is released, so close
it when you are done:
src <- as_ggsql_source(d, format = "duckdb") on.exit(DBI::dbDisconnect(src, shutdown = TRUE))
format = "parquet" needs no such care: it closes the connection it
opened before returning the path. Passing your own con leaves it open
in every case, since it was never ours to close.
Examples
## Not run:
# Curate in R, render in the database:
src <- world_data(2020, geometry = "sf") |> as_ggsql_source(format = "duckdb")
ggsql::ggsql_execute(src, world_query(gdp_per_capita))
## End(Not run)
Attach geometry to a country-level table
Description
The bridge between a one-row-per-country table (e.g. from country_data())
and plotting: bolts polygon or sf geometry onto your data, keyed on
iso3c.
Usage
attach_geometry(
data,
by = "iso3c",
geometry = c("polygon", "sf"),
scale = "small",
region = NULL,
projection = "equal_earth",
recenter = NULL,
overrides = country_overrides(),
year = NULL
)
Arguments
data |
A data frame with an |
by |
The join key (default |
geometry |
|
scale |
Natural Earth resolution for the |
region |
Optional region subset (see |
projection, recenter |
Projection, and optional central meridian, for
the |
overrides |
Name -> iso3c overrides applied when matching the geometry
backend's country names (default |
year |
Attach historical geometry for this year instead of present-day
borders, via |
Value
For "polygon", a tibble with long/lat/group plus your
columns, one row per polygon vertex. For "sf", an sf object, one row
per feature. Both carry every country the backend has – see How many
rows come back.
One row in, one row out
Geometry is attached once per row, not once per country. That is what a panel wants – one row per country-year, each carrying the shape – but it means a frame that repeats a country by accident draws that country more than once, and only the last one painted is visible. The package cannot tell the two apart (a panel's time column may be called anything), so reduce to one row per country yourself when that is what you meant.
Which countries have geometry
The join keeps only countries the chosen backend actually carries, so rows of
data with no matching geometry are dropped silently – worth checking first
when a country you expected is missing from the map. Coverage differs by
backend and, for "sf", by scale, which changes which countries are
present and not merely how detailed they look. Of the 215 countries in
world_snapshot, "polygon" carries 210, "sf" with scale = "small" (the
default, 110m) carries 169, and "sf" with scale = "medium" carries 214:
the 110m coastlines omit most small states, so scale = "medium" is the fix
when microstates matter – Hong Kong, Macao, Tuvalu and the British Virgin
Islands are each in no other backend. Gibraltar alone is in none of them.
How many rows come back
The result is the backend's whole map, not just your rows: every country the
backend carries is present, and the ones absent from data carry NA in
your columns. That is what makes them draw in na.value rather than vanish,
which is the point – a choropleth that quietly omits the countries you have
no data for reads as though they did not exist. It does mean the result is
much larger than data and is not something to summarise directly:
attach_geometry() on three countries returns 240 of them on the polygon
backend and 176 on "sf", whatever data held. The row count is larger
still: "polygon" gives one row per polygon vertex (about 99,000), and
"sf" one row per feature – usually one per country, but a divided
country appears more than once (Cyprus at scale = "small"; Cyprus and
India at "medium"), so an iso3c join against it can fan out. Summarise
data before attaching geometry, or use the verbs in this package, which
de-duplicate to one row per country first.
Examples
df <- data.frame(iso3c = c("USA", "CAN"), value = c(1, 2))
if (requireNamespace("maps", quietly = TRUE)) {
attach_geometry(df, geometry = "polygon")
}
Coverage / missingness audit
Description
What is missing, before you trust the map: which countries are unmatched,
the NA rate per indicator, and which World Bank regions / income groups are
under-covered – so a half-empty map is caught before it is published.
Usage
audit_coverage(data, indicator = NULL, by = c("region", "income", "continent"))
Arguments
data |
A country-level (or map-ready) data frame. |
indicator |
Optional character vector of value columns to report |
by |
Grouping for the coverage breakdown: |
Value
A list of class countryatlas_coverage, with elements unmatched,
na_rates and by_group. It has a print() method, so at the console you
see a formatted report rather than the raw list; reach into the elements by
name to use the numbers programmatically.
-
unmatched– one row per input value that did not resolve to an ISO code. Empty is the good case. -
na_rates– one row per indicator:nis the number of countries considered,n_missinghow many of them lack a value, andna_rateisn_missing / n. An infinite value counts as lacking one: no map can draw it, andcoverage_map()shows it as missing too. -
by_group– one row per group:n_countriesis how many countries are in that group, andna_rateis the share of those lacking a value. The counts sum ton.
Both n and by_group's n_countries are denominators – everything
counted, not everything present. That is the opposite orientation from
map_provenance(), whose n_countries is the numerator (countries drawn
with a value) and whose n_total is the denominator. The two verbs
report the same coverage from opposite ends, so on 215 countries with 24
missing this gives n = 215 where map_provenance() gives
n_countries = 191.
Examples
audit_coverage(countryatlas::world_snapshot$countries)
Does the data respect when countries existed?
Description
The time-aware counterpart to audit_coverage(). A join can succeed and
still be wrong about history: South Sudan with 1995 data, Czechoslovakia with
2001 data, the USSR with 2010 data. Those rows survive every check the package
had, because the country resolves and the year is a number.
Usage
audit_time_coverage(data, quiet = FALSE)
Arguments
data |
A panel with |
quiet |
Suppress the console summary and return the table silently.
(Unlike |
Value
A tibble of the offending rows: iso3c, country, year, issue
("before_existence" or "after_dissolution") and existed (a
human-readable span). Zero rows means the panel is clean.
What it can and cannot see
Dissolution dates come from historical_codes, which covers the entities the package curates (USSR, Yugoslavia, Czechoslovakia and the rest). Independence dates come from the same table read in reverse: a successor state is treated as not existing before its predecessor dissolved. Countries with no entry in the crosswalk – most of the world – are assumed to have existed throughout, so a clean result means "nothing the crosswalk knows about is wrong", not "every date is right".
See Also
audit_coverage(), dissolve_country(), country_timeline()
Examples
panel <- data.frame(
iso3c = c("SSD", "CZE", "FRA"),
year = c(1995L, 2001L, 2001L),
gdp = c(1, 2, 3)
)
audit_time_coverage(panel)
Beta convergence (growth regression)
Description
Do poor countries grow faster than rich ones? The classic unconditional
beta-convergence test: each country's average log growth rate is regressed
on its log initial level. A significantly negative beta is convergence;
the implied convergence speed and half_life (years to close half the
gap) are derived from it.
Usage
beta_convergence(data, value)
Arguments
data |
A panel with |
value |
The value column (unquoted); must be positive (log scale). |
Value
A one-row tibble: beta, se, t_value, p_value, r_squared,
n (countries), speed (annual convergence rate) and half_life
(years).
speed and half_life are NA in two cases: when beta >= 0, because
there is no convergence to put a rate on; and when the panel's per-country
spans are too heterogeneous for any single span to reconcile with the
fitted slope, which is warned about. beta and its inference are
unaffected in both – only the annualised figures need one common span, so
restrict the panel to a shared window if you need them. The fitted stats::lm() object is attached as the
"model" attribute.
See Also
sigma_convergence() for the dispersion-over-time counterpart.
Examples
set.seed(1)
start <- runif(20, 6, 11) # log initial level
growth <- 0.05 - 0.004 * start + rnorm(20, 0, 0.002) # poorer grow faster
panel <- data.frame(
iso3c = rep(sprintf("C%02d", 1:20), each = 2),
year = rep(c(2000L, 2020L), 20),
gdp = as.vector(rbind(exp(start), exp(start + growth * 20)))
)
beta_convergence(panel, gdp)
Two-variable bivariate choropleth
Description
A 2-D bivariate choropleth with a built-in 2-D legend (via the optional
biscale package), e.g. GDP per capita x life expectancy in one map.
Usage
bivariate_map(
data,
fill_x,
fill_y,
palette = "GrPink",
dim = 3,
projection = "equal_earth"
)
Arguments
data |
An |
fill_x, fill_y |
The two value columns (unquoted). |
palette |
A |
dim |
Bivariate dimension: classes per variable, 2, 3 (default) or 4.
A 4 x 4 map needs a palette that has one, such as |
projection |
Projection; see |
Value
A ggplot object (the map; combine with biscale::bi_legend() for a
standalone legend).
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE) &&
requireNamespace("biscale", quietly = TRUE)) {
attach_geometry(countryatlas::world_snapshot$countries, geometry = "sf") |>
bivariate_map(gdp_per_capita, life_expectancy)
}
Proportional-symbol (bubble) map
Description
Plots sized circles at country centroids – the right idiom for totals (population, total emissions, total GDP), which a choropleth misrepresents because big values hide in small countries and vice versa.
Usage
bubble_map(
data,
size,
color = NULL,
projection = "equal_earth",
backend = c("polygon", "sf"),
max_size = 18,
alpha = 0.7
)
Arguments
data |
A country-level frame with |
size |
The column controlling bubble size (unquoted). |
color |
Optional column controlling bubble colour (unquoted). |
projection |
Projection for the base map (sf path). See |
backend |
|
max_size |
Largest bubble size. |
alpha |
Bubble transparency. |
Value
A ggplot object.
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
bubble_map(snap, population)
}
Did the cartogram actually converge?
Description
Cartograms fail quietly. An under-converged one looks entirely plausible while still misrepresenting the areas it exists to make honest. This reports the residual error per country, so the failure is visible.
Usage
cartogram_diagnostics(x, weight = NULL)
Arguments
x |
A |
weight |
The weight column (unquoted). Required when |
Value
A tibble of iso3c, target_share (the country's share of the
weight), actual_share (its share of the cartogram's area) and
area_error (the relative difference). The summary – mean absolute error,
worst country – is attached as the "countryatlas_cartogram" attribute.
What counts as converged
A perfect cartogram has area_error of 0 everywhere. In practice a mean
absolute error under a few percent is good and under 10% is usually
acceptable; a systematically large error, or one concentrated in the small
countries, means the algorithm stopped early. Raise itermax and try again.
See Also
cartogram_map(), dorling_map(), gridded_cartogram()
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("cartogram", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
sfd <- attach_geometry(countryatlas::world_snapshot$countries,
geometry = "sf")
cg <- cartogram_map(sfd, population)
cartogram_diagnostics(cg)
}
Area-honest cartogram
Description
Resizes countries by weight (population, GDP, ...) via the optional
cartogram package, defeating the "big empty countries dominate the eye"
bias of world choropleths.
Usage
cartogram_map(
data,
weight,
type = c("contiguous", "dorling", "noncontiguous", "flow"),
fill = NULL,
projection = "equal_earth",
...
)
Arguments
data |
An |
weight |
The column to resize by (unquoted). |
type |
|
fill |
Optional fill column (unquoted); defaults to |
projection |
Projection; an equal-area CRS is recommended. See
|
... |
Passed to the underlying |
Value
A ggplot object.
Which algorithm
"contiguous" (Dougenik) and "dorling"/"noncontiguous" come from
cartogram. "flow" comes from cartogramR and implements the
Gastner-Seguy-More flow-based method, which is both the current state of the
art and far faster than diffusion-based approaches – prefer it for
contiguous cartograms when cartogramR is available.
Cartograms fail quietly: an under-converged one looks plausible while still
misrepresenting the areas it exists to make honest. Pass a larger itermax
if the result still looks close to the true map.
References
Gastner, M. T., Seguy, V. & More, P. (2018). Fast flow-based algorithm for creating density-equalizing map projections. Proceedings of the National Academy of Sciences 115(10), E2156-E2164. doi:10.1073/pnas.1712674115
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE) &&
requireNamespace("cartogram", quietly = TRUE)) {
attach_geometry(countryatlas::world_snapshot$countries, geometry = "sf") |>
cartogram_map(population, type = "dorling")
}
Pre-flight country-match report
Description
A report on what will and will not match before you trust the map: the
input, its iso3c, whether it matched, whether it is a historical
(dissolved) entity, and a suggestion (the closest known country name by
string distance) for misses. Surfaced automatically by join_world().
Usage
check_country_match(
x,
origin = "country.name",
custom_match = country_overrides(),
suggest = TRUE
)
Arguments
x |
A vector of country names or codes. |
origin |
How to read |
custom_match |
Overrides applied before matching (default
|
suggest |
Whether to compute closest-name suggestions for misses
(requires the optional |
Details
The historical flag matters even for rows that matched: countrycode
silently resolves "USSR" to Russia's RUS, so Soviet-era data becomes
Russian data without a warning. Rows flagged historical should usually be
routed through dissolve_country() instead.
Value
A tibble with columns input, iso3c, matched, historical,
suggestion.
See Also
dissolve_country() for resolving the entities this flags as
historical to their successor states, and repair_country_names() for
applying the suggestion column automatically.
Examples
check_country_match(c("USA", "Cote d'Ivoire", "Yugoslavia", "Wakanda"))
# "USSR" matches (to RUS!) but is flagged historical:
check_country_match("USSR")
Which disputed territories does your data touch?
Description
Cross-references your data against disputed_territories so a contested area does not pass unremarked. Reports both directions: the disputed territories your data covers, and those it is silent about.
Usage
check_dispute_coverage(data, quiet = FALSE)
Arguments
data |
A frame with |
quiet |
Suppress the console summary. |
Value
A tibble of every disputed territory the package knows about, with
in_data saying whether your data covers it. The scope caveat in
disputed_territories applies: this is a documented subset, not every
dispute in the world.
See Also
disputed_territories, dispute_policy(), audit_coverage()
Examples
check_dispute_coverage(countryatlas::world_snapshot$countries)
The same map under several classifications
Description
Small multiples of one choropleth, drawn once per classification method, plus the break table and the count of countries in each class. The point is that the choice is consequential and usually unexamined: Brewer & Pickle (2002) found quantiles among the best methods for general choropleth reading and natural breaks (Jenks) below 70% as accurate, which is the reverse of the common GIS default.
Usage
classify_compare(
data,
value,
methods = c("quantile", "jenks", "equal", "pretty"),
n = 5,
ncol = NULL,
...
)
Arguments
data |
A map-ready frame (polygon or |
value |
The value column (unquoted). |
methods |
Classification styles to compare. Any of |
n |
Number of classes (default |
ncol |
Number of facet columns. |
... |
Passed to |
Value
A faceted ggplot object, with the per-method break and class-count
table attached as the "countryatlas_classification" attribute (and
readable with map_provenance()).
References
Brewer, C. A. & Pickle, L. (2002). Evaluation of methods for classifying epidemiological data on choropleth maps in series. Annals of the Association of American Geographers 92(4), 662-681. doi:10.1111/1467-8306.00310
See Also
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
cmp <- attach_geometry(snap, geometry = "polygon") |>
classify_compare(gdp_per_capita)
attr(cmp, "countryatlas_classification")
}
Clear the cached downloads
Description
Empties the memoised in-session cache and, optionally, the on-disk one.
Usage
clear_country_cache(source = NULL, disk = FALSE)
Arguments
source |
Which source's cache to clear, or |
disk |
Also delete the on-disk cache (default |
Value
Invisibly TRUE.
What a global clear releases
Called with no source, this also drops the two cached geometry backends:
the Natural Earth sf layer held per scale (tens of megabytes at
scale = "medium") and the memoised map_data("world") tibble (about
99,000 rows per override set). Those are the largest things the package
keeps in memory, and in a long-lived process – a Shiny app or a plumber
API – this is the only way to release them. They rebuild on the next map.
Naming a source leaves geometry alone, since it is not a data source.
See Also
country_sources(), fetch_indicator()
Examples
clear_country_cache()
Clear the on-disk / in-memory WDI cache
Description
Forget memoised World Bank fetches, both in-session and (optionally) on disk.
Usage
clear_wdi_cache(disk = FALSE)
Arguments
disk |
Whether to also delete the persistent on-disk cache. |
Value
Invisibly TRUE.
Where the cache lives
The persistent cache goes in the standard per-user cache location,
tools::R_user_dir("countryatlas", "cache"). Point it elsewhere with
options(countryatlas.cache_dir = ), or skip the disk entirely by passing
cache = FALSE to world_data() / country_data(). The directory itself is
created the first time a cached fetch is attempted, whether or not the World
Bank answers; only a successful fetch leaves a response in it, and reading the
bundled world_snapshot never goes near it. Under R CMD check the whole
cache moves to the session temp directory, so a check never writes to the
user's file space.
The directory may hold other files too. The cache only ever writes, expires
and deletes its own entries (named by a hash, with the extension
.countryatlas), and disk = TRUE removes the directory itself only when
that leaves it empty.
How the cache is managed
The persistent cache expires its own contents, so it does not grow without
bound and does not serve stale figures indefinitely: an entry is dropped
once it is 30 days old, and if the directory exceeds 50 MB the
least-recently-used entries go first. Both limits are adjustable with
options(countryatlas.cache_max_age = ) (seconds) and
options(countryatlas.cache_max_size = ) (bytes). A dropped entry costs a
re-fetch, nothing more.
Expiry matters beyond disk space: World Bank observations are revised, so a figure cached long ago is not necessarily the figure the API would return today.
Examples
clear_wdi_cache() # forget the in-session memo
## Not run:
clear_wdi_cache(disk = TRUE) # also delete the persistent cache
## End(Not run)
Curated indicator catalogue
Description
A friendly-name to WDI-code lookup so indicator = common_indicators$population
beats memorising "SP.POP.TOTL".
Usage
common_indicators
Format
A tibble with columns name (friendly name), code (WDI indicator
code) and description.
Source
World Bank indicator catalogue.
Do two sources agree?
Description
Fetch the same indicator from several providers and put the answers side by side. Sources disagree more often than people expect – different vintages, PPP bases, territorial definitions, revision schedules – and on the ISO spine the comparison is one join. This is the verb that turns "I used OWID" into "I used OWID, and here is where it differs from the World Bank".
Usage
compare_sources(
indicator,
sources = c("wdi", "owid"),
year,
countries = NULL,
tolerance = 0.05
)
Arguments
indicator |
Either one indicator code used for every source, or a named
character vector giving each source its own code:
|
sources |
Source names to compare (default |
year |
A single year. |
countries |
Optional |
tolerance |
Relative difference above which a country counts as a
disagreement (default |
Value
A tibble with one row per country: the value from each source,
n_sources (how many reported it), rel_diff (max relative spread) and
disagrees. The correlation, coverage and disagreement summary is attached
as the "countryatlas_source_summary" attribute.
See Also
fetch_indicator(), country_sources()
Examples
## Not run:
cmp <- compare_sources(c(wdi = "NY.GDP.PCAP.KD", owid = "gdp_per_capita"),
sources = c("wdi", "owid"), year = 2020)
attr(cmp, "countryatlas_source_summary")
## End(Not run)
Fill or interpolate panel gaps
Description
Completes a panel so every country has every year, optionally filling missing
values by carry-forward ("locf") or linear interpolation ("linear") so
animations do not flicker on missing years.
Usage
complete_years(
data,
years = NULL,
value = NULL,
method = c("none", "locf", "linear")
)
Arguments
data |
A panel with |
years |
The full set of years to complete to. Defaults to the observed min:max. |
value |
Optional value column(s) (character) to fill. If |
method |
|
Value
A completed (and optionally filled) panel tibble – or an sf
frame, if data was one; each invented row carries its country's geometry.
Examples
df <- data.frame(iso3c = "USA", year = c(2000L, 2002L), gdp = c(1, 3))
complete_years(df, 2000:2002, method = "linear")
Convergence clubs
Description
Countries do not all converge to one steady state; they converge in groups. This implements the Phillips-Sul (2007) log-t procedure: a regression test for whether a set of countries is converging, applied iteratively to peel off clubs that converge internally even when the whole sample does not.
Usage
convergence_club(data, value, min_size = 2, alpha = 0.05)
Arguments
data |
A panel with |
value |
The value column (unquoted); usually income per head. |
min_size |
Smallest club to report (default |
alpha |
Significance level for the one-sided log-t test (default
|
Value
A tibble: iso3c, club (an integer, 1 = highest-level club, NA =
not classified), and the club's log_t statistic. The per-club test results
are attached as the "countryatlas_clubs" attribute. Every country in
data appears: one without a complete series (a missing or non-finite
value in any year) cannot be tested, so it comes back with club = NA and
a warning naming it.
The test
For each country form the relative transition path
h_{it} = y_{it} / \bar{y}_t, then regress
\log(H_1/H_t) - 2\log(\log t) on \log t over the last part of the
sample, where H_t is the cross-sectional mean of
(h_{it}-1)^2. The one-sided t statistic on \log t is the log-t
statistic: above -1.65 the group is converging. Clubs are then formed by
sorting countries on their final-period value and growing a core group while
the test still passes.
A panel needs a reasonable number of periods for this to mean anything – below roughly fifteen the test has very little power, and the function warns.
References
Phillips, P. C. B. & Sul, D. (2007). Transition modeling and econometric convergence tests. Econometrica 75(6), 1771-1855. doi:10.1111/j.1468-0262.2007.00811.x
See Also
beta_convergence(), sigma_convergence()
Examples
set.seed(1)
# two groups converging to different levels
panel <- expand.grid(iso3c = c(paste0("A", 1:5), paste0("B", 1:5)),
year = 2000:2024)
panel$y <- ifelse(startsWith(as.character(panel$iso3c), "A"), 100, 30) +
rnorm(nrow(panel), 0, 2)
convergence_club(panel, y)
Friendly country-code conversion
Description
A discoverable wrapper around countrycode::countrycode() exposing the full
set of schemes with first-class shortcuts for the high-value ones: flag
emoji, currency, top-level domain, continent/region and research codes
(Correlates of War, Polity, Gleditsch-Ward, V-Dem, IMF, FAO, FIPS, GAUL).
Usage
convert_country(
x,
to = "iso3c",
from = "country.name",
custom_match = country_overrides(),
warn = TRUE
)
Arguments
x |
A vector of country names or codes. |
to |
Destination scheme. A shortcut ( |
from |
Origin scheme (default |
custom_match |
Optional overrides (default |
warn |
Whether to warn about inputs that match no country (default
|
Value
A vector of converted codes.
Examples
convert_country(c("Japan", "Brazil"), to = "flag")
convert_country("Germany", to = "currency")
convert_country(c("USA", "France"), to = "continent")
convert_country(c("Germany", "United States"), to = "name_fr")
Pairwise correlation of indicators on the spine
Description
Which indicators move together across countries? Computes pairwise
correlations between indicator columns (pairwise-complete, so patchy
coverage doesn't shrink every pair to the common subset), with the per-pair
n reported so a headline r computed on 12 countries can't masquerade as
a world fact.
Usage
correlate_indicators(data, ..., method = c("pearson", "spearman"), min_n = 3)
Arguments
data |
A country-level (or map-ready) data frame; map-ready frames are
reduced to one row per country first, so the reported |
... |
< |
method |
|
min_n |
Minimum number of complete pairs for a correlation to be
reported (default |
Value
A tibble with one row per indicator pair: var_x, var_y, r,
n (complete pairs), sorted by |r| descending.
Examples
correlate_indicators(countryatlas::world_snapshot$countries)
Country adjacency (shared land borders)
Description
Which countries share a land border with which, as a tidy edge list –
built from polygon topology (sf::st_touches()), so it reflects the same
curated geometry as the rest of the package. Powers neighbors().
Usage
country_borders(scale = "small", region = NULL)
Arguments
scale |
Natural Earth resolution to compute adjacency from. This is not
a cosmetic choice – see Which countries the default leaves out below.
|
region |
Optional region subset (see |
Value
A tibble, one row per bordering pair: iso3c_a, country_a,
iso3c_b, country_b. Each unordered pair appears once, with
iso3c_a <= iso3c_b alphabetically.
Turning it into a graph
igraph::graph_from_data_frame() takes the first two columns as the edge
endpoints, so pass only the two code columns – handing it the whole tibble
would build edges from each country's code to its own name:
igraph::graph_from_data_frame(
country_borders()[, c("iso3c_a", "iso3c_b")], directed = FALSE)
Attaching igraph also masks neighbors(), which it exports too, so call
that one as countryatlas::neighbors() from then on.
Which countries the default leaves out
Adjacency is computed from Natural Earth polygons, and the default
scale = "small" (110m) has no polygon at all for the European microstates.
Andorra, Liechtenstein, Monaco, San Marino and the Vatican are therefore
absent from the table entirely – not merely missing a short border, but
contributing no rows, despite a land border being the whole of their
geography. The default reports 310 pairs over 153 countries; France comes
back with 8 neighbours rather than 10.
scale = "medium" (50m) has all five, giving 322 pairs over 162 countries
and France its full list. Use it whenever the microstates matter:
country_borders(scale = "medium")
The same 110m gap is why morans_i()'s contiguity weights exclude them –
see its n_excluded – and it is the land-border twin of the island problem
described in vignette("honest-maps"). Note the two Guiana borders are real,
not artefacts: French Guiana makes Brazil and Suriname neighbours of France.
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
head(country_borders(region = "Europe"))
# The whole world is only a little dearer: a fraction of a second.
nrow(country_borders())
}
The countrycode codelist as a tidy tibble
Description
The whole countrycode::codelist reshaped into a tidy, pipeable lookup you
can filter() / join() directly – one row per country.
Usage
country_codes(codes = NULL)
Arguments
codes |
Optional character vector of column names to keep (in addition
to |
Value
A tibble, one row per country.
Examples
country_codes()
country_codes(c("iso2c", "continent", "currency"))
Lightweight one-row-per-country table
Description
The analysis counterpart to world_data(): no polygons, one tidy row per
country (iso3c, iso2c, country, classifications and the requested
indicators). This is what you actually join() / mutate() / summarise()
/ rank() on; attach geometry only at draw time with attach_geometry().
Usage
country_data(
year,
indicator = NULL,
latest = FALSE,
panel = FALSE,
classify = c("income", "continent", "region"),
cache = TRUE,
language = "en",
parallel = TRUE
)
Arguments
year |
A single year or a range (with |
indicator |
A named character vector of WDI codes (or |
latest |
Use the most recent non- |
panel |
Return a panel keyed on |
classify |
Which classifications to add. |
cache |
Whether to use the WDI cache. |
language |
WDI language code. |
parallel |
Whether to fetch indicators in parallel. Ignored when the
cache is memory-only; see |
Value
A tibble, one row per country (or per country-year for a panel).
iso3c is the stable key; country is a label and its spelling depends on
where the row came from. A successful fetch carries the World Bank's names
("Korea, Rep.", "Congo, Dem. Rep."), while the country spine used when the
fetch returns nothing carries the countrycode names ("South Korea",
"Congo - Kinshasa") – as do convert_country(), standardize_country()
and the rest of the package. Match on iso3c, and relabel with
convert_country(iso3c, to = "country") if you need one consistent set.
Examples
country_data(2020, c(co2 = "EN.GHG.CO2.MT.CE.AR5"))
Everything the package knows about one country
Description
A single-country summary drawn from the bundled reference data and, if you ask, live indicators: codes, geography, groups, neighbours and any dissolution history.
Usage
country_factsheet(x, indicators = NULL, origin = "country.name")
Arguments
x |
One country name or code. |
indicators |
Optional indicator codes to fetch (named, as in
|
origin |
How to read |
Value
A countryatlas_factsheet object – a list of tibbles (identity,
geography, groups, neighbours, indicators) that prints as a
formatted block.
See Also
country_meta, country_groups(), neighbors(), world_table()
Examples
country_factsheet("Brazil")
Country-group membership
Description
Answers the constant question "is this country in the EU / OECD / G7 / G20 /
BRICS / ...?" from a curated membership table. By default that is the current
snapshot (country_groups_tbl); pass as_of to ask the question of a
particular date, which is what a panel needs.
Usage
country_groups(group = NULL, as_of = NULL)
Arguments
group |
One or more group names: any of |
as_of |
A date (or a year) at which to evaluate membership. |
Value
A tibble of group, iso3c, country; with as_of, also from
and to.
Membership changes, and which groups are dated
A snapshot silently misstates any panel that spans an accession. An EU panel over 2015-2020 either includes the United Kingdom throughout or excludes it throughout, and both are wrong:
"GBR" %in% country_groups("EU", as_of = 2016)$iso3c # TRUE
"GBR" %in% country_groups("EU", as_of = 2021)$iso3c # FALSE
country_groups_history carries dated membership for twelve groups: EU,
EuroZone, NATO, OECD, ASEAN, EFTA, GCC, Mercosur, Nordic, Visegrad, BRICS and
G7. Commonwealth, G20 and OPEC are not dated – their histories involve
suspensions, readmissions and contested dates that would have to be sourced
case by case, and a fabricated date is worse than an absent one. Asking for
as_of on those warns and falls back to the snapshot.
See Also
in_group(), country_groups_history, country_timeline()
Examples
country_groups("EU")
country_groups(c("G7", "BRICS"))
# the UK was a member in 2016 and not in 2021
nrow(country_groups("EU", as_of = 2016))
nrow(country_groups("EU", as_of = 2021))
Dated country-group membership
Description
When each country joined – and where applicable left – each of twelve international groups. The dated counterpart to country_groups_tbl, which is a single current snapshot.
Usage
country_groups_history
Format
A tibble with 176 rows:
- group
Group name.
- iso3c
ISO 3166-1 alpha-3 code.
- country
Country name.
- from
Date membership took effect.
- to
The first date on which the country was no longer a member (membership runs up to the day before), or
NAfor a current member. The United Kingdom's EUtois therefore 2020-02-01: it left at the end of 31 January 2020.
Details
A snapshot silently misstates any panel that spans an accession: an EU panel
over 2015-2020 either includes the United Kingdom throughout or excludes it
throughout, and both are wrong. country_groups() and in_group() read this
table when given as_of.
Scope, and what is deliberately absent
Twelve groups are dated: EU, EuroZone, NATO, OECD, ASEAN, EFTA, GCC,
Mercosur, Nordic, Visegrad, BRICS and G7. Commonwealth, G20 and OPEC are
not, and that is a decision rather than an omission – their histories
involve suspensions, readmissions and contested dates that would have to be
sourced case by case, and a fabricated date is worse than an absent one.
country_groups(as_of =) warns and falls back to the snapshot for those.
Dates are the treaty or accession date where one exists, otherwise 1 January of the accession year. The table is validated at build time against country_groups_tbl: the members current today must reproduce the snapshot exactly, for every group covered.
See Also
country_groups(), in_group(), country_timeline()
Examples
# EFTA is the instructive one: most of its founders left, for the EU
subset(country_groups_history, group == "EFTA")
Country-group membership (point-in-time)
Description
A curated, dated membership table for the common country groups.
Usage
country_groups_tbl
Format
A tibble with columns group, iso3c, country.
As of when
Membership is a snapshot taken on 2026-06-01, carried on the table itself:
attr(country_groups_tbl, "as_of") #> [1] "2026-06-01"
Read the attribute rather than this paragraph if you need the date in code;
it is set from a single constant in data-raw/build_datasets.R, so it
cannot drift from the data the way a hand-written date can. For membership
at any other date use country_groups_history and
country_groups(as_of = ), which is the table this
snapshot is a slice of.
Source
Curated from official membership lists, as of the date above. (This
used to point at the package NEWS for the reference date, where two
different dates were on record – an Rd should not delegate a fact to a
changelog.)
Reconcile and join two messy country tables
Description
The generic two-table version of the package's whole reason for being: join
any two data frames that each key on country names or codes, by reconciling
both sides to iso3c first. Tables keyed on "Czech Republic" vs
"Czechia", or "South Korea" vs "Korea, Rep.", just work.
Usage
country_join(
x,
y,
by_x,
by_y,
origin_x = "country.name",
origin_y = "country.name",
type = c("left", "inner", "full"),
suffix = c(".x", ".y"),
key = c("iso3c", "cowc", "cown", "gwn"),
warn = TRUE
)
Arguments
x, y |
Data frames to join. |
by_x, by_y |
The country columns in |
origin_x, origin_y |
How to read each key (countrycode origin schemes). |
type |
Join type: |
suffix |
Suffix for clashing non-key columns (default
|
key |
Which code system to join on. |
warn |
Whether to report values that resolve to no country (default
|
Value
A tibble joined on a reconciled iso3c key.
Joining historical data
the second spine:
ISO 3166 was first published in 1974 and never covered colonies, so iso3c
cannot key anything before about 1970. Correlates of War and Gleditsch-Ward
codes can, they run back to the nineteenth century, and
historical_geometry() is keyed on gwn. Setting key switches the join
onto one of those:
country_join(a, b, country, nation, key = "gwn")
The trade-off is real and worth stating: COW/GW codes cover states ISO never
did, but they omit the dependencies and non-sovereign territories ISO does
cover, so a modern dataset joined on gwn loses Hong Kong, Puerto Rico and
the rest – which the join warns about. Use iso3c unless you are working
before 1970.
Examples
a <- data.frame(country = c("Czechia", "South Korea"), gdp = c(1, 2))
b <- data.frame(nation = c("Czech Republic", "Korea, Rep."), pop = c(10, 51))
country_join(a, b, country, nation)
Join many messy country tables on the ISO spine
Description
The many-table generalisation of country_join(): reduce-join a list of data
frames that each key on country names or codes, reconciling every one to
iso3c first.
Usage
country_join_all(
tables,
by,
origin = "country.name",
type = c("full", "left", "inner"),
key = c("iso3c", "cowc", "cown", "gwn"),
warn = TRUE
)
Arguments
tables |
A list of data frames. |
by |
A single country-column name present in every table, or a character vector giving the column for each table. |
origin |
countrycode origin scheme(s) for the key column(s) (default
|
type |
Join type: |
key |
Which code system to join on, as in |
warn |
Whether to report values that resolve to no country (default
|
Value
A single tibble joined on key (clashing non-key columns get
dplyr's default .x/.y suffixes).
Examples
a <- data.frame(country = c("Czechia", "South Korea"), gdp = c(1, 2))
b <- data.frame(country = c("Czech Republic", "Korea, Rep."), pop = c(10, 51))
d <- data.frame(country = c("Czechia", "Korea"), area = c(79, 100))
country_join_all(list(a, b, d), by = "country")
Static per-country metadata
Description
One row per country with the facts people constantly need and currently scrape together by hand.
Usage
country_meta
Format
A tibble with one row per country and columns iso3c, iso2c,
country, continent, region, un_region, income, capital,
capital_lat, capital_lon, centroid_lat, centroid_lon, area_km2,
currency, tld, landlocked, flag.
Assembled from countrycode::codelist, so Kosovo (XKX) has no row –
countrycode has none either. The geometry backends and
convert_country() do handle it; distance_between(), which reads its
centroids from here, does not. Ten further territories have a row but no
centroid or area.
country therefore carries the English names from countrycode
("South Korea", "Congo - Kinshasa"), which differ from the World Bank's for
38 of the 215 countries in world_snapshot ("Korea, Rep.",
"Congo, Dem. Rep."). Each
table is faithful to its own source, so join on iso3c and keep whichever
label you want to display – reconciling the two is what country_join()
is for.
Source
Assembled from countrycode, WDI metadata and Natural Earth geometry.
Describe an origin-destination table as a network
Description
Node- and edge-level summaries of a bilateral flow table: who sends, who
receives, who is central. No igraph required – at country scale the dense
arithmetic is trivial.
Usage
country_network(
data,
from,
to,
weight = NULL,
origin = "country.name",
top_n = 20
)
Arguments
data |
An OD table. |
from, to |
The origin and destination country columns (unquoted). |
weight |
The flow column (unquoted). If omitted, every pair counts as 1. |
origin |
How to read |
top_n |
How many edges to return in the |
Value
A list of two tibbles:
-
nodes–iso3c,country,out_flow,in_flow,net_flow,out_degree,in_degreeandstrength_share(the country's share of all flow, in or out). -
edges–from,to,weight,shareandreciprocity(the opposite flow as a proportion of this one;NAwhere there is none).
See Also
flow_matrix(), od_map(), flow_map()
Examples
od <- data.frame(
from = c("China", "China", "Germany", "USA", "Mexico"),
to = c("USA", "Germany", "France", "Mexico", "USA"),
value = c(500, 100, 80, 300, 320)
)
country_network(od, from, to, value)
The registered data sources
Description
What fetch_indicator() can reach, whether each one's backing package is
actually installed, and how to cite it.
Usage
country_sources()
Value
A tibble: source, meta, key_col, key_type, cache,
available (is the
backing package installed?) and citation.
See Also
register_country_source(), fetch_indicator()
Examples
country_sources()
A country's existence span, predecessors and successors
Description
When did this country exist, and what came before and after it? Reads the bundled historical_codes crosswalk in both directions – so it answers both "what did the USSR become" and "what was Estonia part of".
Usage
country_timeline(x, origin = "country.name", warn = TRUE)
Arguments
x |
Country names or codes, current or historical ( |
origin |
How to read |
warn |
Whether to report inputs that match neither a historical entity
nor a modern country (default |
Value
A tibble, one row per input: input, iso3c, country,
dissolved (the year it ceased to exist, or NA if it still does),
predecessors and successors (list-columns of iso3c codes). An input
that resolves to neither spine keeps its input and gets NA elsewhere;
with warn = TRUE it is reported rather than left to be spotted.
Why the two directions are not mirror images
Both columns hold codes, so an entity ISO never coded cannot appear in
either. ISO 3166-3 only records changes from 1974 on, and three entities in
historical_codes predate it: the United Arab Republic (1961), Tanganyika
and Zanzibar (both 1964). Asking about one of those by name works –
country_timeline("Tanganyika") gives TZA as a successor, because the
table stores that side as a name – but the reverse does not:
country_timeline("TZA") reports no predecessors, since there is no code to
report. The other eleven entities carry codes and round-trip in both
directions. Naming them instead of coding them would break the column's
type; inventing codes for them would be a guess, which this package does
not make.
See Also
dissolve_country(), historical_codes, audit_time_coverage()
Examples
country_timeline(c("USSR", "Estonia", "France"))
Spatial weights on the country spine
Description
Build a reusable neighbour-weights object for morans_i(), local_morans(),
gearys_c(), getis_ord() and spatial_lag(). Four schemes, three of which
give every country at least one neighbour – which land-border contiguity, the
historical default, cannot do for an island.
Usage
country_weights(
type = c("contiguity", "knn", "distance", "custom"),
countries = NULL,
k = 5,
cutoff_km = NULL,
w = NULL,
style = c("W", "B"),
scale = "small"
)
Arguments
type |
|
countries |
Optional |
k |
Neighbours per country for |
cutoff_km |
Distance band for |
w |
For |
style |
|
scale |
Natural Earth resolution for |
Value
A countryatlas_weights object: the weights matrix plus the scheme
that built it. Inspect it by printing; as.matrix() gives the matrix.
Choosing a scheme
Contiguity encodes "shares a border", which is the right relation for
spillovers that cross borders by land. It is the wrong relation for a global
question, because it silently deletes the islands. "knn" is the safe
default for world-scale work: every country participates, and k controls how
local the neighbourhood is. "distance" is right when the process has a real
length scale. "custom" is right when geography is not the relevant space at
all.
See Also
morans_i(), local_morans(), lisa_map(), spatial_lag()
Examples
# k-nearest neighbours: no sf needed, and islands are included
w <- country_weights("knn", k = 4)
w
snap <- countryatlas::world_snapshot$countries
morans_i(snap, gdp_per_capita, weights = country_weights("knn", k = 5),
n_perm = 99)
Map the data availability itself
Description
A choropleth of whether a value is present, rather than what it is. The
companion to audit_coverage(), which reports the same thing as a table: a
world map with a large well-covered region and a systematically empty one is
telling you something about the indicator that the headline map hides behind
a uniform grey.
Usage
coverage_map(data, value, by = NULL, title = NULL, ...)
Arguments
data |
A map-ready frame (polygon or |
value |
The column whose availability to map (unquoted). |
by |
Optional grouping column for a panel: with a |
title |
Optional plot title (defaults to a generated one). |
... |
Passed to |
Value
A ggplot object.
See Also
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
attach_geometry(snap, geometry = "polygon") |>
coverage_map(gdp_per_capita)
}
Convert a money series to constant prices
Description
A nominal series is not comparable across years. deflate() divides by a
price index rebased to base_year, turning current-price values into constant
base_year prices.
Usage
deflate(data, value, base_year, deflator = NULL, suffix = "_real")
Arguments
data |
A panel with |
value |
The nominal value column (unquoted). |
base_year |
The year whose prices to express everything in. |
deflator |
Either a column in |
suffix |
Suffix for the new column (default |
Value
data with the constant-price column added.
Which deflator
The GDP deflator is the right default for aggregate output. For household
spending the CPI (FP.CPI.TOTL) is usually preferred, and for cross-country
level comparisons a deflator is not enough at all – you want
to_ppp() as well, because exchange rates do not equalise purchasing power.
Deflating and converting to PPP are different corrections for different
problems, and a cross-country panel over time generally needs both.
See Also
to_ppp(), per_capita(), index_to()
Examples
d <- data.frame(iso3c = "USA", year = 2000:2002,
gdp = c(100, 110, 120), defl = c(90, 100, 105))
deflate(d, gdp, base_year = 2001, deflator = defl)
State which map convention you are using
Description
Disputed territories are drawn differently by different conventions, and a
map that does not say which one it used is making a choice silently. This
sets the session's convention so world_map() can record it and
map_provenance() can report it.
Usage
dispute_policy(policy = NULL)
Arguments
policy |
One of:
Called with no argument, returns the current policy. |
Value
When called with no argument, the policy currently in effect. When
setting, the policy that was in effect before the call, invisibly – R's
convention for a setter, so
on.exit(dispute_policy(dispute_policy("neutral"))) restores it.
What this does and does not do
It records a choice and makes it visible. It does not redraw any boundary,
and selecting "de_jure" will not give you claimed-boundary geometry,
because the package does not have any – Natural Earth's auxiliary claim
lines are not bundled. Anyone publishing under an institutional convention
should verify the shapes against that institution's own basemap rather than
trusting a setting.
See Also
disputed_territories, check_dispute_coverage(), world_map()
Examples
old <- dispute_policy("neutral") # sets, and returns what it replaced
dispute_policy() # "neutral"
dispute_policy(old) # put it back
Disputed territories
Description
Territories whose status is contested, recorded so that a map can say so.
Usage
disputed_territories
Format
A tibble with 22 rows:
- territory
Common name of the territory.
- iso3c
ISO 3166-1 alpha-3 code where one exists, else
NA. Most disputed territories have none, which is why they cannot appear in aniso3c-keyed dataset at all.- administered_by
The party in de facto control, or
NAwhere none is. An ISO 3166-1 alpha-3 code where that party has one – see the note on codes below.- claimed_by
Semicolon-separated claimants, coded as for
administered_by.- status
One of
"un_member","un_observer","partially_recognised","administered"or"claimed".- note
One sentence of context, including why the row is here.
Details
This table records that a dispute exists and who the parties are. It does not adjudicate, rank claims, or imply that any claim is better founded than another. Where it says "administered by" it means de facto control as reported by the mapping sources the package already uses (Natural Earth), not recognition, legitimacy or endorsement.
What reads these columns
Nothing in the package does. The internal layer and note builders behind
world_map(disputes = ) and dispute_policy() key on iso3c alone, and no
verb filters on the parties. administered_by and
claimed_by are here for your own filtering – "show me everything Morocco
administers", "drop the territories with more than one claimant" – and are
documented precisely so that such a filter is writable. Their contents are
validated at build time against the placeholder list below.
The codes in administered_by and claimed_by
Mostly ISO 3166-1 alpha-3, so they join the iso3c spine directly – but not
entirely, and the exceptions are the point of the table. Six parties here are
entities ISO assigns no code to, and they are written with a mnemonic
placeholder instead: ABK (Abkhazia), CYP-N (Northern Cyprus), OST
(South Ossetia), PMR (Transnistria), SAH (the Sahrawi Arab Democratic
Republic) and SOL (Somaliland).
Five of them – ABK, CYP-N, OST, PMR and SOL – appear in
administered_by for the like-named territory, which has no ISO code of its
own: the entity administers itself and ISO codes neither. SAH is the
exception and worth knowing about: it appears only as a claimant, of
Western Sahara, which ISO does code (ESH) and which MAR administers.
None of the six are ISO codes, so none will resolve through
convert_country() or any other iso3c lookup. Filter them out with
%in% country_codes()$iso3c if you need a strictly ISO-keyed column.
Scope
A documented subset, not the roughly 188 disputed areas the EU's data-visualisation guidance counts. The selection criterion is mechanical and checkable: territories that appear as a distinct unit or a contested boundary in Natural Earth at 1:110m or 1:50m, and have an ISO 3166-1 code, a widely-used user-assigned code, or a standard "disputed" label in ISO, UN M49 or World Bank practice. That criterion requires nobody to judge the merits.
See Also
dispute_policy(), check_dispute_coverage(), world_map()
Examples
disputed_territories[, c("territory", "iso3c", "status")]
Resolve dissolved entities to their successor states
Description
Historical data is full of countries that no longer exist – the USSR,
Yugoslavia, Czechoslovakia – and they poison naive joins twice over: most
are silently unmatched, and some are silently mismatched
(countrycode resolves "USSR" to Russia alone, so Soviet-era totals
quietly become Russian totals). dissolve_country() resolves a mixed
vector of historical and modern names against the curated
historical_codes crosswalk: dissolved entities expand to one row per
successor state (one-to-many, dated), while modern names pass through as
single rows – so a whole messy column can be piped in unchanged.
Usage
dissolve_country(x, warn = TRUE)
Arguments
x |
A character vector of country names (historical and/or modern). |
warn |
Whether to warn about names that match neither a historical
entity nor a modern country (default |
Value
A tibble with one row per (input, successor) pair: input (as
given), historical (canonical dissolved-entity name, NA for modern
countries), dissolved (year the entity ceased to exist, NA for
modern), iso3c and country (the successor state). Unmatched inputs
yield one row with iso3c = NA.
See Also
historical_codes for the crosswalk itself and the successor
policy (e.g. Kosovo's inclusion in the Yugoslavia list);
check_country_match(), whose historical column flags these entities;
repair_country_names(), which deliberately leaves them alone.
Examples
dissolve_country(c("USSR", "Czechoslovakia", "France"))
# One-to-many: Yugoslavia expands to its successor territories
dissolve_country("Yugoslavia")
Great-circle distance between two countries
Description
Haversine distance (km) between two countries' centroids – the lightweight
companion to country_borders() for "how far apart" rather than "do they
touch". Works from the bundled country_meta centroids, so unlike most of
the spatial toolkit it needs neither sf nor the network.
Usage
distance_between(a, b, origin = "country.name")
Arguments
a, b |
Vectors of country names or codes. Either the same length, or one of them length 1 to compare one country against many. |
origin |
How to read |
Value
A numeric vector of great-circle distances in kilometres (NA for
any country that doesn't resolve to a known centroid).
Countries without a bundled centroid
country_meta carries no centroid for a handful of small or dependent
territories (Bouvet Island, the British Virgin Islands, Gibraltar, Hong Kong,
Macao, Svalbard and Jan Mayen, Tokelau, Tuvalu, the U.S. Minor Outlying
Islands and the Aland Islands), and no row at all for Kosovo, because
countrycode::codelist has none. Those inputs return NA here even though
the geometry backends do map them – so neighbors() and country_borders()
know about Kosovo while this function does not.
Examples
distance_between("France", "Germany")
distance_between("USA", c("Canada", "Mexico"))
Dorling cartogram (first-class verb)
Description
Non-overlapping proportional circles sized by weight, positioned to stay
as close as possible to each country's true location – arguably the most
legible cartogram variant, since a microstate's circle is exactly as
visible as a giant country's. A first-class verb for
cartogram_map()(type = "dorling") that surfaces the Dorling-specific
tuning knobs.
Usage
dorling_map(
data,
weight,
fill = NULL,
k = 5,
itermax = 1000,
projection = "equal_earth"
)
Arguments
data |
An |
weight |
The column controlling circle size (unquoted). |
fill |
Optional fill column (unquoted); defaults to |
k |
Share of the bounding box filled by the largest circle (default
|
itermax |
Maximum iterations of the circle-repulsion algorithm
(default |
projection |
Projection; an equal-area CRS is recommended. See
|
Value
A ggplot object.
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE) &&
requireNamespace("cartogram", quietly = TRUE)) {
attach_geometry(countryatlas::world_snapshot$countries, geometry = "sf") |>
dorling_map(population)
}
Small-multiple choropleths
Description
Facet a choropleth into small multiples (one panel per group or per year) –
the static counterpart to animate_world(), for print and side-by-side
comparison. Builds a world_map() and facets it on facet.
Usage
facet_map(data, fill, facet, ncol = NULL, ...)
Arguments
data |
A map-ready frame (polygon or sf) containing the |
fill |
The fill column (unquoted). |
facet |
The faceting column (unquoted; e.g. |
ncol |
Number of facet columns (passed to |
... |
Passed to |
Value
A faceted ggplot object.
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
mapdf <- attach_geometry(snap, geometry = "polygon")
facet_map(mapdf, gdp_per_capita, continent, style = "quantile")
}
Fetch an indicator from any registered source
Description
One verb, many providers. Whatever the source, the result comes back on the
ISO spine: iso3c, year where the data is a panel, and one column per
indicator.
Usage
fetch_indicator(source, indicator, countries = NULL, years = NULL, ...)
Arguments
source |
A registered source name (see |
indicator |
Indicator code(s). Name the vector to rename the columns,
as in |
countries |
Optional |
years |
Optional numeric year vector; |
... |
Passed to the source's own |
Value
A tibble keyed on iso3c (and year, for a panel).
See Also
add_indicator(), compare_sources(), country_sources()
Examples
## Not run:
fetch_indicator("wdi", c(gdp = "NY.GDP.PCAP.KD"), years = 2020)
fetch_indicator("owid", "life_expectancy", years = 2020)
## End(Not run)
Great-circle origin-destination flow map
Description
Draws great-circle arcs between country pairs from an origin-destination table (trade, migration, flights, remittances), resolving both endpoints to centroids automatically.
Usage
flow_map(data, from, to, weight = NULL, origin = "country.name", n = 50)
Arguments
data |
An OD table. |
from, to |
The origin and destination country columns (unquoted; names
or |
weight |
Optional column controlling arc width/alpha (unquoted). |
origin |
How to read |
n |
Points per arc (smoothness). |
Value
A ggplot object.
Examples
od <- data.frame(from = c("China", "Germany"),
to = c("United States", "France"),
value = c(500, 200))
if (requireNamespace("maps", quietly = TRUE)) {
flow_map(od, from, to, value)
}
An origin-destination table as a matrix
Description
Reshape a long bilateral table (trade, migration, flights, remittances) into
a square origin x destination matrix on the ISO spine, with both country
columns standardised. The natural input to a network analysis, and to
country_weights()(type = "custom") – which is how "countries near each
other in trade space" becomes a spatial weights object.
Usage
flow_matrix(
data,
from,
to,
weight = NULL,
origin = "country.name",
symmetric = FALSE,
fill = 0
)
Arguments
data |
An OD table. |
from, to |
The origin and destination country columns (unquoted). |
weight |
The flow column (unquoted). If omitted, every pair counts as 1. |
origin |
How to read |
symmetric |
If |
fill |
Value for pairs with no flow (default |
Value
A square numeric matrix with iso3c row and column names. Rows are
origins, columns destinations.
See Also
country_network(), flow_map(), country_weights()
Examples
od <- data.frame(
from = c("China", "China", "Germany", "USA"),
to = c("USA", "Germany", "France", "Mexico"),
value = c(500, 100, 80, 300)
)
flow_matrix(od, from, to, value)
Geary's C (spatial autocorrelation)
Description
The other classical global autocorrelation statistic. Where Moran's I is a
correlation-like measure centred on -1/(n-1), Geary's C is a
distance-like one centred on 1: below 1 means positive autocorrelation
(neighbours are similar), above 1 means negative. It is more sensitive than
Moran's I to local differences.
Usage
gearys_c(data, value, weights = NULL, n_perm = 999)
Arguments
data |
A country-level frame with |
value |
The value column (unquoted). |
weights |
A |
n_perm |
Permutations for the pseudo-p-value (default |
Value
A one-row tibble: c (observed), expected (always 1), n
(countries used), n_excluded (countries with data that the weights could
not reach), n_links (non-zero weights), p_value and an excluded
list-column of the excluded iso3c codes. n and n_excluded sum to the
countries supplied with a value, and mean the same here as in
morans_i().
p_value is one-sided on the lower tail:
(1 + \#\{C^{*} \le C_{obs}\}) / (n_{perm} + 1). The lower tail is
the clustered one, which is the opposite way round from Moran's I:
Geary's C runs from 0 (neighbours identical) through 1 (no
autocorrelation) upwards, so a small c is the evidence of positive
spatial association. Never exactly zero; the floor is
1/(n_{perm}+1). Set a seed beforehand for a reproducible p_value.
References
Geary, R. C. (1954). The contiguity ratio and statistical mapping. The Incorporated Statistician 5(3), 115-146. doi:10.2307/2986645
See Also
Examples
snap <- countryatlas::world_snapshot$countries
gearys_c(snap, gdp_per_capita, weights = country_weights("knn", k = 5),
n_perm = 99)
Centroid-anchored country labels
Description
A ggplot2 layer that places labels (names, ISO codes or flag emoji) at
country centroids, with optional ggrepel collision avoidance. Designed for
the polygon backend produced by world_data() / join_world(): it reads the
long, lat and group columns, so it errors on an sf frame and points at
ggplot2::geom_sf_text() instead. Placement is exact only while group is
present – that is what identifies each country's separate pieces, and the
label goes on the largest one.
Usage
geom_country_labels(
mapping = NULL,
data = NULL,
repel = TRUE,
flag = FALSE,
size = 3,
...
)
Arguments
mapping |
Aesthetic mapping; defaults to |
data |
Optional layer data, as for any |
repel |
Use |
flag |
If |
size |
Label text size. |
... |
Passed to the underlying text geom. |
Value
A ggplot2 layer.
Examples
library(ggplot2)
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
mapdf <- attach_geometry(snap, geometry = "polygon")
# Labelling all 188 countries at once is unreadable, and ggrepel responds
# by dropping nearly every label. Pass `data` to choose a subset ...
world_map(mapdf, gdp_per_capita) +
geom_country_labels(
data = ~ dplyr::filter(.x, iso3c %in% c("USA", "BRA", "CHN", "IND", "ZAF"))
)
# ... or zoom in, where there is room for every label.
europe <- attach_geometry(
dplyr::filter(snap, continent == "Europe"), geometry = "polygon")
world_map(europe, gdp_per_capita) +
geom_country_labels(size = 2.5) +
coord_quickmap(xlim = c(-25, 45), ylim = c(34, 72))
}
Getis-Ord G statistics (hot spots)
Description
Global G and local G_i^*: unlike Moran's I, these distinguish
clusters of high values from clusters of low ones, which is what
"hot spot" analysis usually wants.
Usage
getis_ord(data, value, weights = NULL, local = TRUE)
Arguments
data |
A country-level frame with |
value |
The value column (unquoted). |
weights |
A |
local |
If The global |
Value
With local = TRUE, a tibble of iso3c, gi_star, z_score and
p_value (two-sided, from the normal approximation), one row per country
used. With local = FALSE, a one-row tibble of g, expected, n
(countries used – the same count, so the local form returns n rows) and
n_links (non-zero weights).
References
Getis, A. & Ord, J. K. (1992). The analysis of spatial association by use of distance statistics. Geographical Analysis 24(3), 189-206. doi:10.1111/j.1538-4632.1992.tb00261.x
See Also
local_morans(), country_weights()
Examples
snap <- countryatlas::world_snapshot$countries
getis_ord(snap, gdp_per_capita, weights = country_weights("knn", k = 5))
Gini coefficient (population-weightable)
Description
The Gini index of inequality across countries, optionally weighted (weight by population and the statistic describes inequality between people assigned their country's mean, not between country units).
Usage
gini(x, weights = NULL, na.rm = TRUE)
Arguments
x |
A numeric vector (e.g. GDP per capita by country). |
weights |
Optional non-negative weights (e.g. population), either the
same length as |
na.rm |
Whether to drop |
Value
A single number in [0, 1]: 0 is perfect equality.
See Also
theil(), which adds a between/within-group decomposition.
Examples
snap <- countryatlas::world_snapshot$countries
gini(snap$gdp_per_capita) # between countries
gini(snap$gdp_per_capita, weights = snap$population) # between people
Orthographic globe choropleth
Description
The world as a globe (orthographic projection) centred on lon/lat – the
honest answer to "the whole world on a rectangle exaggerates the poles". Takes
the same fill / style options as world_map(). The default "sf" backend
gives the cleanest limb; the "polygon" backend draws the globe with
ggplot2::coord_map() and needs only maps + mapproj (no sf).
Usage
globe_map(
data,
fill,
lon = 0,
lat = 20,
backend = c("sf", "polygon"),
style = c("continuous", "binned", "quantile", "jenks", "categorical"),
palette = NULL,
n_bins = 5,
borders = TRUE,
title = NULL,
legend = NULL,
na_label = "No data",
interactive = FALSE
)
Arguments
data |
A map-ready frame: an |
fill |
The fill column (unquoted). |
lon, lat |
The longitude / latitude the globe is centred on (the face pointing at the viewer). |
backend |
|
style, palette, n_bins, borders, title, legend, na_label |
As in |
interactive |
If |
Value
A ggplot object.
Examples
# No sf required -- the polygon backend needs only maps + mapproj:
if (requireNamespace("maps", quietly = TRUE) &&
requireNamespace("mapproj", quietly = TRUE)) {
globe_map(countryatlas::world_snapshot$countries, continent,
backend = "polygon", style = "categorical")
}
## Not run:
# The sf backend gives the cleanest limb (needs a World Bank fetch):
world_data(2020, geometry = "sf") |>
globe_map(gdp_per_capita, lon = 10, lat = 30)
## End(Not run)
One square per N people
Description
A gridded (or "waffle") cartogram: the world redrawn as equal cells, each worth a fixed quantity, allocated to countries in proportion to their value and placed near where they belong. Where a Dorling cartogram preserves position and a contiguous one preserves adjacency, this preserves countability – the reader can literally count the cells.
Usage
gridded_cartogram(data, value, cells = 1000, fill = NULL, cell_size = 2.5)
Arguments
data |
A country-level or map-ready frame with |
value |
The column to allocate cells by (unquoted). |
cells |
Total number of cells to distribute (default |
fill |
Optional fill column (unquoted); defaults to |
cell_size |
Grid spacing in degrees (default |
Value
A ggplot object. The per-country cell allocation is attached as the
"countryatlas_cells" attribute – every placeable country, including the
ones that rounded to zero cells, so share sums to 1 and the rounding is
fully visible.
Rounding is the whole difficulty
Allocating a whole number of cells to each country cannot be exact, so the
remainder has to go somewhere. This uses the largest-remainder method, which
guarantees the cell total is exactly cells and that no country with a
positive value gets zero cells while a smaller one gets one. The attached
table reports each country's exact share alongside its integer allocation so
the rounding is inspectable rather than hidden.
Crowded neighbours overlap
Each country's block is centred on its own centroid, with no collision
avoidance between countries. That is deliberate – a global packing solve
would push countries away from where they belong – but it means blocks in
crowded regions are drawn on top of one another, and a partly hidden block
cannot be counted or compared. The effect is not marginal: at the defaults
(cells = 1000, cell_size = 2.5) about a third of the cells overlap a
cell of a different country, across some sixty countries, and it grows with
cells – at cells = 2500 it is roughly two thirds.
cell_size is the lever, because it scales the tiles without moving the
centroids: dropping it to 1.5 cuts the overlap at cells = 1000 to about
a tenth of the cells. Fewer cells also helps. Where exact areas matter
more than geographic position, dorling_map() resolves collisions by
displacing circles instead.
See Also
cartogram_map(), dorling_map(), tile_map()
Examples
snap <- countryatlas::world_snapshot$countries
gridded_cartogram(snap, population, cells = 400)
Year-on-year (or compound) growth rate
Description
Adds a growth-rate column to a panel: either the period-over-period change
("yoy") or the compound annual growth rate from the first observed year
("cagr"), computed per country.
Usage
growth_rate(data, value, type = c("yoy", "cagr"), suffix = "_growth")
Arguments
data |
A panel with |
value |
The value column (unquoted). |
type |
|
suffix |
Suffix for the new column (default |
Value
data with a growth-rate column added (a proportion, so 0.03 = 3%).
Rows come back sorted by iso3c then year: the calculation reads each
country's series in time order, so a row-aligned vector held alongside
data will no longer line up.
Examples
df <- data.frame(iso3c = "USA", year = 2000:2002, gdp = c(100, 110, 121))
growth_rate(df, gdp)
Historical / dissolved entities and their successor states
Description
A curated crosswalk from dissolved entities (Soviet Union, Yugoslavia,
Czechoslovakia, ...) to the modern states that succeeded them – one row per
(entity, successor) pair, dated, so historical panels can be brought onto
the modern ISO spine honestly instead of being silently dropped (or worse:
countrycode resolves "USSR" to Russia alone). Consumed by
dissolve_country() and flagged by check_country_match().
Usage
historical_codes
Format
A tibble with one row per (entity, successor):
- historical
Canonical name of the dissolved entity.
- iso3c_hist
The alpha-3 code the entity held at dissolution, where one existed (
SUN,YUG,CSK,DDR,ANT,SCG,YMD, ...); it may since have been inherited by a successor (e.g.YEM).- dissolved
Year the entity ceased to exist.
- iso3c, country
The successor state.
- relation
How the successor relates to the entity, per successor:
"succession"if it is a genuinely new state created atdissolved, or"continuation"if the same state carried on (possibly with less territory) or a state that already existed absorbed the entity. Sudan is both at once –SDNcontinued andSSDis new – which is why the relation is per successor rather than per entity. The distinction matters: a"continuation"was not created atdissolved, so testing its data against that year says nothing, and the code did not cease to exist either.audit_time_coverage()keys on this column for both reasons.
Details
Kosovo (XKX) is included among the Yugoslavia and Serbia-and-Montenegro
successors on a territory basis (its territory was part of both); filter
it out if your analysis follows strict UN-membership succession.
Source
Curated from ISO 3166-3 and the historical record.
Historical country boundaries
Description
Country polygons as they were, from CShapes 2.0 (Schvitz et al. 2022), which maps states and colonies and dependencies for 1886-2019 with per-polygon validity periods. A 1970 map with 2024 borders is a common and quiet error; this is the fix.
Usage
historical_geometry(year, dependencies = FALSE, projection = "equal_earth")
Arguments
year |
The year to draw, or a |
dependencies |
Include colonies and dependencies (default |
projection |
Projection to return the geometry in (see |
Value
An sf frame with gwcode, country, iso3c (where one can be
assigned – see below), status, from, to and geometry. Two further
CShapes columns are passed through when the installed version supplies
them, since they answer the questions this verb is usually asked:
-
owner– thegwcodeof the sovereign a dependency belonged to; a sovereign state carries its owngwcodehere. This is the column that makesdependencies = TRUElegible: without it a colony and its metropole are two unrelated rows.owner != gwcodepicks out the dependencies. -
capname– the capital's name at that date.
Both were returned but undocumented. Neither is guaranteed: cshapes
decides what its own table holds, so check with names() rather than
assuming.
The ISO spine does not reach back
ISO 3166 was first published in 1974 and never covered colonies, so a
historical map cannot be keyed on iso3c. CShapes uses Gleditsch-Ward
codes, which is why gwcode is the key here and iso3c is a best-effort
extra, read off the GW code: the modern code of the state that holds it.
So it is NA for an entity with no modern counterpart (the German
Democratic Republic, Czechoslovakia, Yugoslavia, the two Yemens), and also
where historical_codes says the modern state did not exist yet – GW give
the USSR and Russia one code, but the 1980 polygon is the Soviet Union, so
it carries NA rather than "RUS". A colony carries the code of the state
it became: CShapes names the 1950 Gold Coast "Ghana" and it comes back as
"GHA", with status saying it was a colony. Join historical data on
gwcode, not on iso3c, and use convert_country()(to = "gwn") to get
there from a modern code.
References
Schvitz, G., Girardin, L., Ruegger, S., Weidmann, N. B., Cederman, L.-E. & Gleditsch, K. S. (2022). Mapping the international system, 1886-2019: The CShapes 2.0 dataset. Journal of Conflict Resolution 66(1), 144-161. doi:10.1177/00220027211013563
See Also
world_geometry(), country_timeline(), world_map()
Examples
## Not run:
# Africa before decolonisation needs the dependencies
historical_geometry(1950, dependencies = TRUE)
## End(Not run)
Is a country in a group?
Description
A vectorised membership predicate built on country_groups().
Usage
in_group(x, group, origin = "country.name", as_of = NULL)
Arguments
x |
A vector of country names or codes. |
group |
A single group name (see |
origin |
How to read |
as_of |
A date or year at which to evaluate membership; |
Value
A logical vector the same length as x. A value origin cannot
resolve to an ISO code answers FALSE – the same as a country that is
genuinely outside the group – so run check_country_match() first if you
need to tell "not a member" from "not recognised".
See Also
country_groups(), country_groups_history
Examples
in_group(c("France", "United States", "Japan"), "EU")
# membership is a function of time
in_group("United Kingdom", "EU", as_of = 2016)
in_group("United Kingdom", "EU", as_of = 2021)
Rebase a series to an index (base year = 100)
Description
Rescales a value column so the chosen base year equals to (100 by default),
per country – the standard way to compare trajectories that start at very
different levels.
Usage
index_to(data, value, base_year, to = 100, suffix = "_index")
Arguments
data |
A panel with |
value |
The value column (unquoted). |
base_year |
The year set equal to |
to |
The index value the base year maps to (default |
suffix |
Suffix for the new column (default |
Value
data with an index column added. The column is NA for any
country whose series does not cover base_year (see the note there). A
negative base-year value is indexed as it is, so that country's index
runs opposite to its series.
Examples
df <- data.frame(iso3c = "USA", year = 2000:2002, gdp = c(50, 55, 60))
index_to(df, gdp, base_year = 2000)
# A base year the data does not have gives NA, not an error:
index_to(df, gdp, base_year = 1999)
Web-ready interactive choropleth
Description
An interactive choropleth with hover and zoom, for dashboards and
R Markdown / Quarto. Engines are all optional Suggests.
Usage
interactive_map(
data,
fill,
tooltip = NULL,
engine = c("plotly", "ggiraph", "leaflet", "ggsql", "mapgl"),
...
)
Arguments
data |
A map-ready frame (polygon or sf). The |
fill |
The fill column (unquoted). |
tooltip |
Optional tooltip column (unquoted). |
engine |
|
... |
Passed to |
Value
An interactive widget.
Examples
## Not run:
world_data(2020) |> interactive_map(gdp_per_capita)
world_data(2020, geometry = "sf") |>
interactive_map(gdp_per_capita, engine = "ggsql", transform = "log10")
## End(Not run)
Fill missing values, and say that you did
Description
Interpolate or carry forward missing observations in a panel. Every value this invents is flagged in a companion column, and that flag is not optional – an imputed value that travels through a pipeline looking like data is exactly the failure this package exists to prevent.
Usage
interpolate_missing(
data,
value = NULL,
method = c("linear", "locf", "none"),
max_gap = 3
)
Arguments
data |
A panel with |
value |
Column(s) to fill (character). |
method |
|
max_gap |
Longest run of consecutive missing years to fill. Gaps longer
than this are left alone, because interpolating across a decade is not
interpolation. Default |
Value
data with the gaps filled and, for each filled column, a logical
<column>_imputed companion. The map verbs count those columns when they
write provenance, so they keep working through any verb that preserves
columns. An "countryatlas_imputed" attribute lists the flag columns for
convenience, but nothing in the package reads it, and dplyr drops it as
it drops most attributes – rely on the columns, not the attribute. Rows
come back sorted by iso3c then year.
The hard rule
The flag cannot be turned off. world_map() reads it and refuses to draw
imputed values as though they were observed without at least noting it in the
caption. If you need values with no flag, compute them yourself – the
package will not hand you a frame where invented numbers are indistinguishable
from measured ones.
See Also
complete_years(), rate_check(), coverage_map()
Examples
p <- data.frame(iso3c = "USA", year = 2000:2005,
gdp = c(1, NA, NA, 4, NA, 6))
interpolate_missing(p, "gdp")
One call: your data, on a map
Description
Auto-detects the country column, standardises it to ISO codes (via
standardize_country()), attaches geometry and returns a plot-ready frame –
the function that fulfils the package's promise for your own data. Pipe the
result straight into world_map().
Usage
join_world(
data,
country_col = NULL,
origin = "country.name",
geometry = c("polygon", "sf", "none"),
scale = "small",
region = NULL,
projection = "equal_earth",
recenter = NULL,
warn = TRUE
)
Arguments
data |
A data frame keyed on country names or codes. |
country_col |
The country column (unquoted). If omitted, it is auto-detected. |
origin |
How to read |
geometry |
|
scale |
Natural Earth resolution for the |
region |
Optional region subset (see |
projection, recenter |
Projection, and optional central meridian, for
the |
warn |
Whether to report unmatched countries (default |
Value
A plot-ready frame: polygon tibble, sf object, or (for
geometry = "none") the standardised table.
Examples
rates <- data.frame(country = c("United States", "Brazil", "Kenya"),
vaccination_pct = c(0.7, 0.8, 0.6))
if (requireNamespace("maps", quietly = TRUE)) {
joined <- join_world(rates, country)
}
Panel lag / difference by country
Description
The two panel primitives everyone hand-rolls (and gets subtly wrong when the
frame isn't sorted): the value n years back, and the change since then –
grouped by iso3c, ordered by year, so country A's 1960 never leaks into
country B's first row.
Usage
lag_by_country(data, value, n = 1, suffix = NULL)
diff_by_country(data, value, n = 1, suffix = NULL)
Arguments
data |
A panel with |
value |
The value column (unquoted). |
n |
Number of periods to lag / difference over (default |
suffix |
Suffix for the new column. Defaults to |
Value
data with the lagged / differenced column added. Rows come back
sorted by iso3c then year: the calculation reads each country's series
in time order, so a row-aligned vector held alongside data will no
longer line up.
Examples
df <- data.frame(iso3c = "USA", year = 2000:2003, gdp = c(100, 110, 121, 133))
lag_by_country(df, gdp)
diff_by_country(df, gdp)
Map LISA clusters
Description
The map of local_morans(): countries coloured by cluster type, with
non-significant ones left neutral. Hot spots (High-High) and cold spots
(Low-Low) read immediately; the off-diagonal categories are the spatial
outliers.
Usage
lisa_map(data, value, weights = NULL, n_perm = 999, alpha = 0.05, ...)
Arguments
data |
A map-ready frame (polygon or |
value |
The value column (unquoted). |
weights |
A |
n_perm, alpha |
Passed to |
... |
Passed to |
Value
A ggplot object, with the local_morans() table attached as the
"countryatlas_lisa" attribute.
See Also
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
set.seed(1)
attach_geometry(snap, geometry = "polygon") |>
lisa_map(gdp_per_capita, weights = country_weights("knn", k = 5),
n_perm = 99)
}
Local Moran's I (LISA)
Description
Local Indicators of Spatial Association (Anselin 1995): one Moran statistic
per country, plus the cluster type it belongs to. Where morans_i() answers
"is there clustering anywhere", this answers "where, and of what kind".
Usage
local_morans(data, value, weights = NULL, n_perm = 999, alpha = 0.05)
Arguments
data |
A country-level frame with |
value |
The value column (unquoted). |
weights |
A |
n_perm |
Permutations for the pseudo-p-value (default |
alpha |
Significance threshold for the |
Value
A tibble, one row per country: iso3c, value, lag (the
neighbour average), ii (the local statistic), p_value and cluster
("High-High", "Low-Low", "High-Low", "Low-High" or "Not significant").
p_value is a two-sided pseudo-p from conditional permutation:
(1 + \#\{|I_i^{*}| \ge |I_i|\}) / (n_{perm} + 1), so it is never
exactly zero and its floor is 1/(n_{perm}+1) – with the default 999
permutations, 0.001. Two-sided because a local statistic is interesting at
both ends: a country surrounded by unlike neighbours is as much a finding
as one surrounded by like ones. cluster is "Not significant" wherever
p_value > alpha, and everywhere when n_perm = 0 leaves it NA. Set a
seed beforehand for a reproducible p_value.
References
Anselin, L. (1995). Local Indicators of Spatial Association – LISA. Geographical Analysis 27(2), 93-115. doi:10.1111/j.1538-4632.1995.tb00338.x
See Also
lisa_map(), morans_i(), country_weights()
Examples
snap <- countryatlas::world_snapshot$countries
set.seed(1)
local_morans(snap, gdp_per_capita, weights = country_weights("knn", k = 5),
n_perm = 99)
Tag coordinates with the country that contains them
Description
Point-in-polygon lookup: given longitude / latitude vectors (or an sf POINT
object), return the iso3c of the country each point falls in – the bridge
for getting point data (events, stations, observations) onto the country
spine so it can be joined, aggregated and mapped like everything else.
Usage
locate_country(
lon = NULL,
lat = NULL,
points = NULL,
scale = "small",
add = "country",
tolerance_km = 25
)
Arguments
lon, lat |
Equal-length numeric vectors of longitude / latitude, giving
one point per element (ignored if |
points |
Optional |
scale |
Natural Earth resolution for the lookup geometry. |
add |
Extra attributes to return alongside |
tolerance_km |
Snap an unmatched point to the nearest country when it
lies within this many kilometres of one (default |
Value
A tibble with one row per point: iso3c plus any add columns
(NA for points that fall in no country, e.g. open ocean).
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
locate_country(lon = c(2.35, -74.0), lat = c(48.85, 40.7)) # Paris, NYC
}
What went into this map
Description
Report the provenance of a world_map() (or any plot the package's map verbs
produced): the package version, the geometry backend and projection, the
classification method and its breaks, the fill column, and how many countries
are shown versus missing. These are the questions a reviewer asks first, and
the answers are already known at plot time – this just makes them readable.
Usage
map_provenance(x, value = NULL)
Arguments
x |
A plot returned by any of the package's map verbs –
|
value |
For a data frame, the column whose coverage to report (unquoted). Ignored for a plot, which already knows its own fill. |
Value
A one-row tibble of provenance fields, invisibly printed in a
human-readable block. Fields: countryatlas, fill, backend,
projection, style, n_bins, na_style, n_countries, n_missing,
n_total, uncertainty, disputes, dispute_policy, n_imputed,
breaks, missing_iso3c and snapshot_year.
The three counts are: n_countries, the countries actually drawn with a
value; n_missing, those drawn without one; and n_total, the two added
together – every country the map covers. n_countries is the numerator,
not the denominator, which its name does not say on its own.
Putting it on the plot
world_map()(footnote = "auto") prints the coverage line as a caption, and
classification_report = TRUE attaches the per-class counts. Together they
cover what a methods note needs:
p <- world_map(mapdf, gdp_per_capita, style = "quantile",
footnote = "auto", classification_report = TRUE)
map_provenance(p)
attr(p, "countryatlas_classification")
See Also
world_map(), audit_coverage(), coverage_map()
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
p <- attach_geometry(snap, geometry = "polygon") |>
world_map(gdp_per_capita, style = "quantile")
map_provenance(p)
}
Global Moran's I (spatial autocorrelation)
Description
Do neighbouring countries have similar values? Global Moran's I on the country
spine, with a permutation pseudo-p-value. No spdep required: at ~200
countries the dense arithmetic is trivial.
Usage
morans_i(data, value, scale = "small", n_perm = 999, weights = NULL)
Arguments
data |
A country-level data frame with |
value |
The value column (unquoted). |
scale |
Natural Earth resolution for the default contiguity adjacency
(see |
n_perm |
Number of permutations for the pseudo-p-value (default |
weights |
A |
Value
A one-row tibble: i (observed Moran's I), expected
(-1/(n-1) under no autocorrelation), n (countries used),
n_excluded (countries with data that the weights could not reach),
n_links, p_value (one-sided, P(I_{perm} \ge I_{obs}), computed as
(1 + \#\{I^{*} \ge I_{obs}\}) / (n_{perm} + 1), so never exactly
zero – the floor is 1/(n_{perm}+1)) and an excluded list-column of
the excluded iso3c codes. Set a seed beforehand for a reproducible
p_value.
Which countries are left out
The default weights are land-border contiguity, and an island has no land
border – so any country with no land neighbour present in data drops out
entirely. On the bundled world_snapshot that is around a quarter of the
countries with data: Japan, the United Kingdom, Australia, Indonesia,
Madagascar, New Zealand, the Philippines, Iceland, Cuba, Sri Lanka and every
small island state. The omission is systematic rather than random.
n_excluded and excluded report it, and country_weights() fixes it –
"knn" and "distance" give every country neighbours:
morans_i(snap, gdp_per_capita, weights = country_weights("knn", k = 5))
References
Moran, P. A. P. (1950). Notes on continuous stochastic phenomena. Biometrika 37(1/2), 17-23. doi:10.2307/2332142
See Also
country_weights(), local_morans(), gearys_c(), spatial_lag()
Examples
snap <- countryatlas::world_snapshot$countries
set.seed(42)
# every country included, no sf required
morans_i(snap, gdp_per_capita, n_perm = 99,
weights = country_weights("knn", k = 5))
A country's neighbours
Description
Which countries border a given country (or countries) – a vectorised
lookup built on country_borders().
Usage
neighbors(x, origin = "country.name", scale = "small", warn = TRUE)
Arguments
x |
A vector of country names or codes. |
origin |
How to read |
scale |
Natural Earth resolution to compute adjacency from. |
warn |
Whether to report values that do not resolve to a country
(default |
Value
A tibble with one row per (iso3c, neighbor) pair: the queried
country's iso3c, and each bordering country's iso3c and country
name (neighbor, neighbor_country). Countries with no land border
(islands, e.g. Japan, Madagascar) return zero rows.
Pass a vector, don't loop
Every call rebuilds the whole world's adjacency from polygon topology, so
asking about one country costs the same as asking about all of them. x is
vectorised, and adding countries to a single call only adds the filtering:
countryatlas::neighbors(c("FRA", "DEU", "ESP"), origin = "iso3c")
Looping instead pays that rebuild once per country – for every bordering
country in the world, roughly two orders of magnitude more work than one
vectorised call. The same applies to country_borders(), which does the work.
Name clash with igraph
igraph also exports a neighbors(), and it takes a graph and a vertex
rather than country names. Whichever package is attached later wins, so if
you use both – which country_borders() suggests, for building a graph of
the adjacency – qualify this one as countryatlas::neighbors().
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
neighbors("France")
neighbors(c("FRA", "JPN"), origin = "iso3c") # Japan has no land border
}
NUTS geometry for Europe
Description
European subnational boundaries from Eurostat's GISCO service via the
optional giscoR package. Nothing is bundled – the geometry is downloaded
and cached by giscoR itself.
Usage
nuts_geometry(
level = 2,
year = 2021,
countries = NULL,
resolution = "60",
projection = "equal_earth"
)
Arguments
level |
NUTS level: |
year |
NUTS vintage: |
countries |
Optional |
resolution |
GISCO resolution: |
projection |
Projection for the result (see |
Value
An sf frame with nuts_id, iso3c, name, level and geometry.
See Also
standardize_subnational(), subnational_map()
Examples
## Not run:
nuts_geometry(level = 2, countries = c("DEU", "FRA"))
## End(Not run)
Origin-destination small multiples
Description
One small map per origin, each showing where that origin's flow goes. The
answer to the arc map's central problem: past a few dozen flows,
flow_map() is a plate of spaghetti and an OD map is legible.
Usage
od_map(
data,
from,
to,
weight = NULL,
origin = "country.name",
origins = 6,
direction = c("out", "in"),
...
)
Arguments
data |
An OD table. |
from, to |
The origin and destination country columns (unquoted). |
weight |
The flow column (unquoted). If omitted, every pair counts as 1. |
origin |
How to read |
origins |
Which origins to draw. A character vector of names or codes,
or an integer giving how many of the largest to take (default Note that this is not |
direction |
|
... |
Passed to |
Value
A faceted ggplot object.
See Also
flow_map(), country_network(), facet_map()
Examples
od <- data.frame(
from = rep(c("China", "Germany", "USA"), each = 3),
to = c("USA", "Japan", "Brazil", "France", "Italy", "Poland",
"Mexico", "Canada", "Japan"),
value = c(500, 200, 90, 80, 70, 60, 300, 280, 120)
)
if (requireNamespace("maps", quietly = TRUE)) {
od_map(od, from, to, value, origins = 3)
}
Normalise an indicator by population
Description
Removes the "is this map just a population map?" footgun by dividing a value
column by population. If no population column is supplied, SP.POP.TOTL is
pulled automatically for the relevant countries and years.
Usage
per_capita(data, value, pop = NULL, suffix = "_per_capita", cache = TRUE)
Arguments
data |
A country-level (or panel) data frame with |
value |
The value column to normalise (unquoted). |
pop |
Optional population column (unquoted). If absent, population is fetched from WDI. |
suffix |
Suffix for the new column (default |
cache |
Whether to use the WDI cache when fetching population. |
Value
data with a new per-capita column.
Examples
df <- data.frame(iso3c = c("USA", "CHN"), year = 2020L,
co2 = c(5e6, 1e7), pop = c(331e6, 1402e6))
per_capita(df, co2, pop)
The same map under several projections
Description
Small multiples of one choropleth, drawn once per projection, so the cost of a projection choice is visible rather than asserted. The data and the classification are held fixed; only the CRS varies.
Usage
projection_compare(
data,
fill,
projections = c("equal_earth", "robinson", "winkel_tripel", "mercator"),
ncol = NULL,
labeller = c("name", "property"),
...,
projection = NULL
)
Arguments
data |
An |
fill |
The fill column (unquoted). |
projections |
Projections to compare (default: Equal Earth, Robinson,
Winkel tripel and Mercator – an equal-area, two compromises and a
conformal). See |
ncol |
Number of facet columns. |
labeller |
|
... |
Passed to |
projection |
Not an argument of this function. It exists only to catch
the singular spelling, which R's partial matching would otherwise bind to
|
Value
A faceted ggplot object.
See Also
projection_info(), tissot_map()
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
attach_geometry(countryatlas::world_snapshot$countries, geometry = "sf") |>
projection_compare(gdp_per_capita, style = "quantile")
}
Measure what a projection distorts
Description
The numeric companion to tissot_map(): distortion sampled on a grid, so it
can be summarised, compared or mapped rather than eyeballed. Computed by
projecting a small circle at each grid point and measuring what happens to
it, which is the definition rather than an approximation of it.
Usage
projection_distortion(
projection = "equal_earth",
measure = c("areal", "angular", "max_scale"),
spacing = 10,
max_lat = 85
)
Arguments
projection |
Projection to measure (see |
measure |
|
spacing |
Grid spacing in degrees (default |
max_lat |
Absolute latitude limit (default |
Value
A tibble of lon, lat and distortion, with the area-weighted
mean and the range attached as the "countryatlas_distortion" attribute.
Reading the numbers
An equal-area projection has "areal" distortion of 1 everywhere – that is
what equal-area means, and it is worth checking rather than trusting.
A conformal projection has "angular" distortion of 0 everywhere and
unbounded areal distortion. A compromise projection is bad at both by a
little, everywhere, which is the trade it makes.
Distortion is measured against the WGS84 ellipsoid, the datum every CRS the
package builds is defined on. Equal Earth, Gall-Peters and the three Lambert
azimuthal projections have ellipsoidal forms and read exactly 1. PROJ
implements Mollweide and Eckert IV with their spherical formulas, so on this
datum they are equal-area only to within about 0.7% ("areal" between
0.993 and 1.007): a property of the maps drawn here, and the kind of thing
this check exists to show.
See Also
tissot_map(), projection_info(), projection_compare()
Examples
if (requireNamespace("sf", quietly = TRUE)) {
d <- projection_distortion("mercator", measure = "areal")
attr(d, "countryatlas_distortion")
}
What a projection preserves
Description
Look up the properties of the projections world_map() understands: the
construction family, whether it is equal-area (a choropleth's ink is
proportional to ground area) or conformal (shapes are locally right), and the
PROJ string the package builds. Called with no arguments it returns the whole
table, which is the quickest way to see what is available.
Usage
projection_info(projection = NULL)
Arguments
projection |
A projection name, or |
Value
A tibble with one row per projection: projection, family,
property, equal_area, conformal, note and proj4.
Choosing one
For a world choropleth the honest choice is equal-area, because the eye
reads coloured area as quantity: a projection that inflates Greenland makes
Greenland's value look more important than it is. Equal Earth is the
recommended default (and the package's own) – it is equal-area, close to
Robinson in appearance, and cheap to evaluate (Savric, Patterson & Jenny
2019). "mercator" is conformal, not equal-area, and should not be used for
choropleths. Use projection_compare() to see the difference on your own
data, and tissot_map() to see it on the graticule.
References
Savric, B., Patterson, T. & Jenny, B. (2019). The Equal Earth map projection. International Journal of Geographical Information Science 33(3), 454-465. doi:10.1080/13658816.2018.1504949
See Also
projection_compare(), tissot_map(), world_map()
Examples
projection_info()
projection_info("mercator")
# every equal-area projection the package can build
subset(projection_info(), equal_area)$projection
Add rank, percentile and z-score
Description
Adds rank, percentile and z_score for a value column, optionally within
a group (region, year, ...), for "top 10" tables and labelling.
Usage
rank_countries(data, value, within = NULL, desc = TRUE)
Arguments
data |
A data frame. |
value |
The value column to rank (unquoted). |
within |
Optional grouping column(s) (unquoted or character) to rank within. |
desc |
Rank descending (largest = rank 1); default |
Value
data with rank, percentile and z_score columns added.
Examples
df <- data.frame(iso3c = c("USA", "CHN", "IND"), gdp = c(21, 17, 3))
rank_countries(df, gdp)
Flag rates computed over tiny denominators
Description
The small-number problem, named: a rate over a very small population is mostly noise, and on a map it shouts as loudly as a rate over a very large one. This reports which countries' rates you should not trust before you plot them.
Usage
rate_check(data, numerator, denominator, min_denominator = NULL, rate = NULL)
Arguments
data |
A country-level frame. |
numerator, denominator |
The count and the population it is over (unquoted). |
min_denominator |
Denominators below this are flagged. |
rate |
An existing rate column (unquoted), if you already computed it.
Otherwise the rate is |
Value
A tibble of iso3c, numerator, denominator, rate,
expected_se (the Poisson standard error of the rate, \sqrt{r/d};
NA, with a warning, for a negative rate) and flagged, sorted with the
least reliable first.
"Least reliable" is ordered on the standard error a single event would
imply, \sqrt{\max(y, 1)}/d, which is identical to expected_se for
every row with at least one event. Ordering on expected_se directly put
zero-count rows last, at the reliable end: it is \sqrt{y}/d, which
is exactly 0 when the count is 0, so one country with no events out of
251 people outranked another with one event out of the same 251. Observing
nothing is not evidence of precision. expected_se itself is still the
plain Poisson standard error, 0 and all.
flagged is TRUE for a
denominator below the threshold, FALSE above it or missing, and NA for
every row when no threshold could be computed at all – which is warned
about, and means sum(flagged) is NA rather than a misleading 0.
What to do about it
Three answers, in rough order of preference: smooth_rates() shrinks the
unreliable rates toward the global rate; value_by_alpha_map() leaves them
alone but fades them out; or drop them and say so in the caption. Plotting
them raw and unremarked is the one option that misleads.
References
Roth, R. E., Woodruff, A. W. & Johnson, Z. F. (2010). Value-by-alpha maps: an alternative technique to the cartogram. The Cartographic Journal 47(2), 130-140. doi:10.1179/000870409X12488753453372
See Also
smooth_rates(), per_capita(), value_by_alpha_map()
Examples
d <- data.frame(
iso3c = c("CHN", "IND", "TUV", "NRU"),
cases = c(50000, 42000, 3, 1),
pop = c(1.41e9, 1.39e9, 11000, 12000)
)
rate_check(d, cases, pop)
Register a data source on the country spine
Description
Teach countryatlas where to get indicators from. A source is a name plus a
fetch function; once registered it works with fetch_indicator(),
add_indicator() and compare_sources() exactly like the built-in ones.
Usage
register_country_source(
name,
fetch,
meta = NULL,
citation = NULL,
key_col = "iso3c",
key_type = "iso3c",
cache = TRUE
)
Arguments
name |
Short source name, used everywhere else (e.g. |
fetch |
A function |
meta |
Optional one-line description of the provider. |
citation |
Optional citation string; surfaced by |
key_col |
The country-key column |
key_type |
How to read that column: any |
cache |
Whether results should be memoised for the session
(default |
Details
This is deliberately a registry rather than more Suggests. It means a
provider with no CRAN package – V-Dem, the IMF, the World Inequality
Database – is reachable without this package depending on anything, and it
means an internal or proprietary feed is a first-class citizen.
Value
Invisibly, the source name.
The fetch contract
fetch(indicator, countries, years) must return a data frame with:
a country key column (named
key_col, holdingkey_typevalues – normally aniso3ccolumn ofiso3ccodes),a
yearcolumn (integer) if the data is a panel,one column per requested indicator, named after the indicator, and
one row per key –
iso3cfor a cross-section,iso3candyearfor a panel.
The key has to be unique because add_indicator() joins on it: a repeated
key multiplies the caller's rows, so it is reported rather than joined in
silence. Aggregate before returning if the provider serves several rows per
country-year.
countries is a character vector of iso3c codes or NULL for all;
years is a numeric vector or NULL for the provider's default. Returning
extra columns is fine – they travel through. Anything the provider cannot
serve should come back as NA, not as a missing row, wherever that is
cheap to arrange.
See Also
country_sources(), fetch_indicator(), compare_sources()
Examples
# A trivial in-memory source
register_country_source(
"demo",
fetch = function(indicator, countries = NULL, years = NULL) {
data.frame(iso3c = c("USA", "FRA"), year = 2020L, demo_value = c(1, 2))
},
meta = "A toy source for the examples"
)
country_sources()
fetch_indicator("demo", "demo_value")
# Registering is permanent for the session, so example() would leave "demo"
# and "demo_named" in the registry and country_sources() would report
# different rows afterwards. Clean up at the end -- see the last lines.
# A source keyed on country names rather than codes: `key_col` says which
# column holds the key, `key_type` says how to read it.
register_country_source(
"demo_named",
fetch = function(indicator, countries = NULL, years = NULL) {
data.frame(country = c("United States", "Japan"), year = 2020L,
demo_value = c(3, 4))
},
key_col = "country", key_type = "country.name",
meta = "A toy name-keyed source"
)
fetch_indicator("demo_named", "demo_value")
# Leave the registry as it was found.
remove_country_source(c("demo", "demo_named"))
Remove a registered data source
Description
The counterpart to register_country_source(). Registering is permanent for
the session, so without this there was no way to undo one – which made any
code that registers a source (an example, a test, an exploratory script)
leave the registry permanently changed, and country_sources() report
different rows afterwards.
Usage
remove_country_source(source)
Arguments
source |
One or more source names, as given to
|
Details
The five built-in sources cannot be removed: they are what the package
documents, and dropping one would make ?fetch_indicator wrong. Pass
cache = FALSE to a fetch instead if you want to bypass one.
Value
The names actually removed, invisibly.
See Also
register_country_source(), country_sources()
Examples
register_country_source("scratch", function(indicator, ...) NULL)
"scratch" %in% country_sources()$source
remove_country_source("scratch")
"scratch" %in% country_sources()$source
Auto-repair country names to their closest known match
Description
The "act on it" companion to check_country_match(): replaces unmatched
country names with their closest known country name (by string distance), but
only when the match is confident enough, and reports what it changed. Pipe the
result into standardize_country() / join_world().
Usage
repair_country_names(
x,
threshold = 0.2,
origin = "country.name",
verbose = TRUE
)
Arguments
x |
A vector of country names. |
threshold |
Maximum string distance to accept a repair (0 = identical,
1 = unrelated). Lower is stricter; default Install |
origin |
countrycode origin scheme (default |
verbose |
Whether to message the substitutions made (default |
Value
A character vector the same length as x, with confident misses
replaced by the closest known country name (others left unchanged). The
applied substitutions are attached as the attribute "repairs".
See Also
check_country_match() for the report this acts on, and
dissolve_country() for dissolved entities, which are deliberately not
repaired.
Examples
repair_country_names(c("United States", "Brzil", "Germny"))
Each country's share of the world total
Description
Adds a column giving each country's value as a share of the (year's) world
total – e.g. share of global emissions or GDP. Operates within year when a
panel is supplied. A dplyr grouping on data is ignored – the denominator
is always the world (or the year's) total, never the group's.
Usage
share_of_world(data, value, suffix = "_share")
Arguments
data |
A country-level (or panel) data frame. |
value |
The value column (unquoted). |
suffix |
Suffix for the new column (default |
Value
data with a share column added: a proportion in [0, 1] when no
value is negative. Negative values are the caller's business and are
summed as they are, so their shares fall outside that range.
Examples
df <- data.frame(iso3c = c("USA", "CHN"), co2 = c(5, 10))
share_of_world(df, co2)
Sigma convergence (dispersion over time)
Description
Is the cross-country distribution actually narrowing? Reports the dispersion
of a (positive) indicator across countries for every year of a panel –
falling dispersion is sigma convergence. The natural companion to
beta_convergence(): beta convergence is necessary but not sufficient for
sigma convergence.
Usage
sigma_convergence(data, value, measure = c("sd_log", "cv"))
Arguments
data |
A panel with |
value |
The value column (unquoted). |
measure |
|
Value
A tibble with one row per year: year, n (countries with
positive values) and sigma.
See Also
beta_convergence() for the growth-regression counterpart.
Examples
df <- data.frame(
iso3c = rep(c("A", "B", "C"), 2),
year = rep(c(2000L, 2010L), each = 3),
gdp = c(1, 10, 100, 2, 11, 60) # dispersion falls
)
sigma_convergence(df, gdp)
Simplify (thin) geometry for faster plotting
Description
Reduce the vertex count of an sf object via the optional rmapshaper
package (falling back to sf::st_simplify()), for fast web/plotting.
Usage
simplify_geometry(x, keep = 0.05, ...)
Arguments
x |
An |
keep |
Proportion of vertices to keep: greater than 0 and at most 1
( |
... |
Passed to the underlying simplifier. |
Value
A simplified sf object.
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
world_geometry(geometry = "sf") |> simplify_geometry(keep = 0.1)
}
Shrink unreliable rates toward the global rate
Description
Empirical-Bayes smoothing: a rate computed over a small denominator is pulled
toward the overall rate in proportion to how little information stands behind
it, while a rate over a large denominator is left essentially alone. Standard
practice in disease mapping, and the statistical counterpart to
value_by_alpha_map()'s visual answer.
Usage
smooth_rates(
data,
numerator,
denominator,
method = c("eb", "none"),
suffix = "_smoothed"
)
Arguments
data |
A country-level frame. |
numerator, denominator |
The count and its denominator (unquoted). |
method |
|
suffix |
Suffix for the new columns (default |
Value
data with <numerator>_rate and <numerator>_smoothed columns
added, plus <numerator>_shrinkage – the weight given to the country's own
rate, between 0 (fully shrunk to the global rate) and 1 (untouched). A row
with no finite, positive denominator or with a negative count has no rate:
all three are NA there, with a warning.
The model
A Poisson-gamma model: counts y_i \sim \mathrm{Poisson}(d_i \theta_i)
with \theta_i \sim \mathrm{Gamma}, whose mean and variance are estimated
from the data by the method of moments. The posterior mean is
w_i r_i + (1 - w_i)\bar{r} with w_i = d_i / (d_i + \alpha), so the
shrinkage weight is exactly the "how much do we believe this country"
quantity that rate_check() flags. Where the between-country variance is
estimated as non-positive (rates no more dispersed than Poisson noise alone),
every rate shrinks fully to the global mean, which is the right answer:
the data contain no evidence of real between-country variation.
On a panel the prior is estimated separately for each year, so every
rate is shrunk toward its own year's global rate and every row is kept.
See Also
rate_check(), per_capita(), value_by_alpha_map()
Examples
d <- data.frame(
iso3c = c("CHN", "IND", "TUV", "NRU"),
cases = c(50000, 42000, 3, 1),
pop = c(1.41e9, 1.39e9, 11000, 12000)
)
smooth_rates(d, cases, pop)
Built-in source adapters
Description
Thin wrappers that put a provider's data on the ISO spine. Each is
Suggests-gated on that provider's own client package – countryatlas does
not reimplement any of them. All four are registered as sources, so the usual
route is fetch_indicator()("owid", ...) rather than calling these
directly; they are exported because calling them directly is sometimes what
you want.
Usage
fetch_owid(indicator, countries = NULL, years = NULL, ...)
fetch_eurostat(indicator, countries = NULL, years = NULL, ...)
fetch_oecd(indicator, countries = NULL, years = NULL, ...)
fetch_comtrade(indicator, countries = NULL, years = NULL, ...)
Arguments
indicator |
Indicator code(s), optionally named to rename the output columns. |
countries |
Optional |
years |
Optional numeric year vector. |
... |
Passed to the underlying client. |
Value
A tibble on the ISO spine: iso3c, year and one column per
indicator.
Which provider needs what
| Adapter | Needs | Notes |
fetch_owid() | owidR | Our World in Data; indicator is an OWID chart slug |
fetch_eurostat() | eurostat | European coverage only; geo codes are harmonised to iso3c |
fetch_oecd() | OECD | indicator is a dataset id; OECD's own filters go through ... |
fetch_comtrade() | comtradr | UN trade flows; needs an API token (see comtradr::set_primary_comtrade_key())
|
See Also
fetch_indicator(), register_country_source(), compare_sources()
Examples
## Not run:
fetch_owid("life-expectancy", years = 2020)
fetch_eurostat("demo_pjan", years = 2020)
## End(Not run)
The neighbour average, as a column
Description
The spatially lagged value: for each country, the (weighted) mean of its neighbours. The building block behind every statistic here, and useful on its own – "what is happening around this country" as a regressor, a map layer or a scatter-plot axis against the country's own value (the Moran scatterplot).
Usage
spatial_lag(data, value, weights = NULL, suffix = "_lag")
Arguments
data |
A country-level frame with |
value |
The value column (unquoted). |
weights |
A |
suffix |
Suffix for the new column (default |
Value
data with the lagged column added. Countries the weights cannot
reach get NA – and since that NA is indistinguishable from one caused
by a missing input value, the codes themselves are attached as the
"countryatlas_excluded" attribute, the frame-shaped counterpart to the
excluded column morans_i() returns.
See Also
country_weights(), local_morans()
Examples
snap <- countryatlas::world_snapshot$countries
spatial_lag(snap, gdp_per_capita, weights = country_weights("knn", k = 5))
Spike map (heights at country centroids)
Description
The classic "population spikes" display: a triangular spike at each country
centroid whose height encodes the value. Like bubble_map() it is the
honest idiom for totals, with a different visual trade-off: spikes
overplot less in dense regions (Europe, the Caribbean) because they only
grow upward. Uses the polygon backend, so it needs only maps.
Usage
spike_map(
data,
height,
max_height = 20,
width = 1.6,
color = "#B2182B",
alpha = 0.65
)
Arguments
data |
A country-level frame with |
height |
The column controlling spike height (unquoted). |
max_height |
Height of the tallest spike, in degrees of latitude
(default |
width |
Base width of each spike, in degrees of longitude (default
|
color |
Spike colour (default a warm red). |
alpha |
Spike fill transparency. |
Value
A ggplot object.
Examples
if (requireNamespace("maps", quietly = TRUE)) {
spike_map(countryatlas::world_snapshot$countries, population)
}
Spin the globe
Description
An animated GIF of the world rotating on its axis: a sequence of orthographic
globe_map() frames at evenly spaced central longitudes, assembled into a
looping animation with the optional gifski (preferred) or magick package.
Embeds directly in R Markdown / Quarto / a README.
Usage
spin_globe(
data,
fill,
lat = 20,
n_frames = 60,
fps = 15,
backend = c("polygon", "sf"),
width = 480,
height = 480,
file = NULL,
...
)
Arguments
data |
A map-ready frame (see |
fill |
The fill column (unquoted). |
lat |
The latitude the globe is tilted toward (the viewer's eye line). |
n_frames |
Number of frames in one full 360 degrees rotation. |
fps |
Frames per second of the output animation. |
backend |
|
width, height |
Pixel dimensions of the animation. |
file |
Optional output path ( |
... |
Passed to |
Value
The path to the written GIF, invisibly.
Examples
# Six frames rather than the default 60, so this stays quick enough to be
# checked: \dontrun{} meant the example was never executed by anything, and
# an example nothing runs is an example free to rot.
if (requireNamespace("maps", quietly = TRUE) &&
requireNamespace("mapproj", quietly = TRUE) &&
(requireNamespace("gifski", quietly = TRUE) ||
requireNamespace("magick", quietly = TRUE))) {
# No sf required on the polygon backend.
gif <- spin_globe(world_snapshot$countries, continent,
backend = "polygon", style = "categorical",
n_frames = 6, width = 200, height = 200)
file.exists(gif) # written to a temporary file
}
Add ISO codes and classifications to any data frame
Description
The package's mission, exposed for your data: take a data frame keyed on
messy country names (or codes) and attach standardised ISO codes plus useful
classifications, reconciling spellings via countrycode::countrycode() and
the curated country_overrides() table. The result joins cleanly to anything
else keyed on iso3c.
Usage
standardize_country(
data,
country_col,
origin = "country.name",
add = c("iso3c", "iso2c", "continent", "region"),
custom_match = country_overrides(),
warn = TRUE
)
Arguments
data |
A data frame / tibble. |
country_col |
The column holding country names or codes (unquoted, tidy-eval). |
origin |
How to read |
add |
Character vector of attributes to add. Defaults to
|
custom_match |
A named character vector of name -> iso3c overrides;
defaults to |
warn |
Whether to warn about unmatched countries (default |
Value
data with the requested columns added (and existing same-named
columns overwritten).
Examples
df <- data.frame(nation = c("U.S.", "S. Korea", "Czechia"), value = 1:3)
standardize_country(df, nation)
Standardise subnational region names to ISO 3166-2
Description
The subnational counterpart to standardize_country(): resolve messy region
names within a country to ISO 3166-2 codes, so subnational data can be joined
on a real key instead of on spelling.
Usage
standardize_subnational(
data,
region,
country,
origin = "country.name",
warn = TRUE
)
Arguments
data |
A data frame with a region column. |
region |
The region-name column (unquoted). |
country |
The country column (unquoted), or a single country name/code applying to every row. ISO 3166-2 codes are only unique within a country, so this is required. |
origin |
How to read |
warn |
Warn about regions that do not resolve (default |
Value
data with iso3c and iso_3166_2 columns added. Unresolved
regions get NA, never a guess.
Coverage, stated plainly
Two things resolve, and it is worth being blunt about how little that is.
A region value that is already an ISO 3166-2 code ("DE-BY", "US-CA")
passes through, upper-cased and trimmed, provided its country prefix
matches the country the row gives – these codes are unique only within a
country, which is why country is required. A mismatch is reported and
left as NA.
A region name resolves only through the optional regions package's
crosswalk, and only when the installed version exposes a name-to-code pair
this function recognises. As of regions 0.1.8 none of them do:
nuts_lau_2019 offers lau_name_national / lau_name_latin and
all_valid_nuts_codes has no name column, so no region name resolves at
all, anywhere – the function says so once per session. The package
carries no ISO 3166-2 name table of its own, deliberately: the datasets that
do pair names with codes key on NUTS codes (DE2) rather than ISO 3166-2
(DE-BY), and filling iso_3166_2 from those would put a different code
system in the column.
So: pass codes if you have them. If you have names, expect NA until
regions ships a usable crosswalk. This function returns NA rather than a
plausible-looking wrong code, and audit_coverage() on the result is the
right next step.
See Also
subnational_map(), nuts_geometry(), standardize_country()
Examples
d <- data.frame(region = c("Bavaria", "Hesse", "Nowhere"), value = 1:3)
if (requireNamespace("regions", quietly = TRUE)) {
standardize_subnational(d, region, country = "Germany")
}
Map subnational data
Description
A choropleth below the country level, joining your data to NUTS geometry on
the region code. The subnational counterpart to world_map(), scoped to
where a maintained code system and free geometry actually exist.
Usage
subnational_map(
data,
fill,
by = "nuts_id",
level = 2,
year = 2021,
countries = NULL,
resolution = "60",
...
)
Arguments
data |
A frame with a NUTS/ISO 3166-2 code column. |
fill |
The fill column (unquoted). |
by |
The code column in |
level, year, countries, resolution |
Passed to |
... |
Passed to |
Value
A ggplot object.
See Also
nuts_geometry(), standardize_subnational(), world_map()
Examples
## Not run:
d <- data.frame(nuts_id = c("DE21", "DE22"), value = c(1, 2))
subnational_map(d, value, level = 2, countries = "DEU")
## End(Not run)
Theil index, with between/within decomposition
Description
The Theil T inequality index – less famous than Gini, but it decomposes exactly into a between-group and a within-group component, answering "how much of world inequality is between continents vs within them?" in one call. Weight by population to describe inequality between people rather than between country units.
Usage
theil(x, weights = NULL, groups = NULL, na.rm = TRUE)
Arguments
x |
A positive numeric vector (log scale; zero/negative values are dropped with a warning). |
weights |
Optional non-negative weights (e.g. population), either the
same length as |
groups |
Optional grouping vector (e.g. continent), the same length as
|
na.rm |
Whether to drop |
Value
Without groups: a single non-negative number (0 = perfect
equality). With groups: a tibble with components "total",
"between" and "within" (total = between + within) and each
component's share of the total (NA when the total is 0, i.e.
perfect equality, and the shares are undefined).
When there is nothing to compute (no values left after na.rm, a zero
total weight, an infinity in x or weights, or, with na.rm = FALSE, a
missing value or group), the result is a single NA whatever groups
says, so reach for the components only after checking is.data.frame().
See Also
gini() for the more familiar single-number summary, which does not
decompose.
Examples
snap <- countryatlas::world_snapshot$countries
theil(snap$gdp_per_capita, weights = snap$population)
theil(snap$gdp_per_capita, weights = snap$population, groups = snap$continent)
A clean theme for world maps
Description
Strips axes, panel grid and background so the map is the focus. Applied by
every plotting function in the package except bivariate_map(), which uses
biscale::bi_theme() so the map matches its own legend, and exported here for
reuse on plots you build yourself.
Usage
theme_world_map(base_size = 12, base_family = "")
Arguments
base_size |
Base font size. |
base_family |
Base font family. |
Value
A ggplot2 theme object.
Examples
library(ggplot2)
ggplot() + theme_world_map()
Equal-area world tile grid
Description
A statebins-style equal-area tile grid of the world (one square per country)
so tiny states are actually visible. Uses the bundled world_tiles layout.
For small multiples of a tile grid, facet the result as you would any other
ggplot (or see facet_map() for the choropleth equivalent).
Usage
tile_map(data, fill, label = TRUE)
Arguments
data |
A country-level frame with |
fill |
The fill column (unquoted). |
label |
Whether to draw ISO codes on the tiles (default |
Details
Every tile in the layout is drawn, taking the scale's na.value fill where
data has no row for it. The converse also holds and is quieter: data rows
keyed on one of the 10 countries with no tile are dropped without a warning
(see world_tiles for which).
Value
A ggplot object.
Examples
tile_map(countryatlas::world_snapshot$countries, gdp_per_capita)
Tissot's indicatrix: what a projection does to the ground
Description
Draws Tissot indicatrices – small circles of equal ground radius, placed on a graticule and projected with everything else. On an equal-area projection every ellipse encloses the same area (though shapes shear); on a conformal projection every ellipse stays circular but sizes explode. It is the standard cartographic device for showing what a projection costs, and it makes the package's "honest maps" claim visible instead of asserted.
Usage
tissot_map(
projection = "equal_earth",
spacing = 30,
radius_km = 500,
max_lat = 75,
fill = "#B2182B",
color = "grey20"
)
Arguments
projection |
Projection to illustrate (see |
spacing |
Degrees between indicatrices (default |
radius_km |
Ground radius of each circle in kilometres (default |
max_lat |
Absolute latitude limit for the circle centres (default |
fill, color |
Fill and outline colour for the ellipses. |
Value
A ggplot object: the world outline with indicatrices on top.
See Also
projection_info(), projection_compare()
Examples
if (requireNamespace("sf", quietly = TRUE) &&
requireNamespace("rnaturalearth", quietly = TRUE)) {
tissot_map("mercator") # circles stay round, and grow enormously
tissot_map("equal_earth") # equal areas, sheared shapes
}
Convert to purchasing-power-parity terms
Description
Market exchange rates do not equalise what money buys. to_ppp() divides a
local-currency series by the PPP conversion factor, putting every country on
comparable international dollars – the correction that makes a cross-country
level comparison meaningful.
Usage
to_ppp(data, value, factor = NULL, suffix = "_ppp")
Arguments
data |
A panel with |
value |
The local-currency value column (unquoted). |
factor |
Either a column holding the PPP conversion factor (unquoted),
or |
suffix |
Suffix for the new column (default |
Value
data with the PPP-converted column added.
See Also
Examples
d <- data.frame(iso3c = c("IND", "USA"), year = 2020L,
gdp_lcu = c(1e5, 1e4), ppp = c(21.9, 1))
to_ppp(d, gdp_lcu, factor = ppp)
Value-by-alpha: equalise a rate by its denominator
Description
A choropleth where colour carries the value and opacity carries an equalising variable (usually population), over a neutral background. It is the answer to the small-number problem – a rate computed over eleven thousand people shouts as loudly as one computed over a billion – and unlike a cartogram it solves it without distorting geometry, which is the main objection to cartograms. Roth, Woodruff & Johnson (2010) introduced it for exactly this purpose.
Usage
value_by_alpha_map(
data,
value,
equalize,
style = c("quantile", "continuous", "binned", "jenks"),
palette = NULL,
n_bins = 5,
alpha_range = c(0.15, 1),
transform = c("rank", "log10", "identity"),
background = "grey20",
title = NULL,
legend = NULL,
projection = "equal_earth"
)
Arguments
data |
A map-ready frame (polygon or |
value |
The value column, carried by colour (unquoted). |
equalize |
The equalising column, carried by opacity (unquoted) – population, total counts, or whatever denominator the rate was built on. |
style |
Classification for the colour channel (as |
palette |
Optional viridis palette name. |
n_bins |
Number of colour bins for the binned styles. |
alpha_range |
Minimum and maximum opacity (default |
transform |
Transform applied to |
background |
Colour behind the countries, which shows through where opacity is low (default a dark neutral). |
title, legend |
Optional plot title and legend title. |
projection |
Projection for the |
Value
A ggplot object.
References
Roth, R. E., Woodruff, A. W. & Johnson, Z. F. (2010). Value-by-alpha maps: an alternative technique to the cartogram. The Cartographic Journal 47(2), 130-140. doi:10.1179/000870409X12488753453372
See Also
cartogram_map() and dorling_map() (the geometry-distorting
answers to the same problem), world_map()
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
attach_geometry(snap, geometry = "polygon") |>
value_by_alpha_map(gdp_per_capita, population)
}
Search World Bank indicators
Description
A tidy, pipeable wrapper on WDI::WDIsearch() for discovering indicator
codes.
Usage
wdi_search(pattern, field = c("name", "indicator"), cache = NULL)
Arguments
pattern |
A regular expression to search indicator names/codes for. |
field |
Which field to search: |
cache |
Optional cached |
Value
A tibble of matching indicator codes and names.
Examples
# Searches WDI's bundled indicator list, so this needs no connection.
wdi_search("CO2 emissions")
Curated country-name overrides (replaces the silent drop-list)
Description
A documented custom_match table for entities that map backends
(ggplot2::map_data() and Natural Earth) get wrong or leave without an ISO
code. Earlier versions of the package deleted these regions; now they are
matched instead, so they stop silently disappearing from maps.
country_overrides() is the current name, as of the package's rename to
countryatlas. wdj_overrides() is deprecated and warns once per session;
it returns the same table and will be removed. The help page kept describing
it as "a backward-compatible alias" after the code had started warning.
Usage
wdj_overrides(extra = NULL)
country_overrides(extra = NULL)
Arguments
extra |
An optional named character vector of additional overrides
(names are country/region names, values are |
Details
The table maps a country/region name (as spelled by the geometry backends) to
an ISO 3166-1 alpha-3 code. Pass the result as the custom_match argument to
standardize_country(), world_data() and friends. Every downstream code
(iso2c, continent, region, flag, ...) is derived from this iso3c, so a
single override is enough.
Value
A named character vector suitable for countrycode(custom_match=).
Accented names and locales
Every name in this table is plain ASCII, and that is deliberate: ASCII
spellings match in any locale. Accented spellings ("Curacao" with a
cedilla, "Saint Barthelemy" with an acute) are matched natively by
countrycode::countrycode() in a UTF-8 locale, which is why they are not
listed here – but in a non-UTF-8 locale (LC_CTYPE=C) they cannot be
compared reliably and resolve to NA.
Accented spellings also come in two Unicode forms that look identical: the
accent can be one precomposed code point (NFC) or a base letter followed by
a combining mark (NFD, which macOS returns for filenames). Only NFC matches
countrycode::countrycode()'s tables, so a name that resolves to nothing is
retried with its combining marks stripped, which turns an NFD spelling into the
ASCII spelling that resolves anywhere. Only unresolved names are retried, so
this never changes a name that already matched.
If your input may contain accented country names, run in a UTF-8 locale.
De-accenting with iconv(x, to = "ASCII//TRANSLIT") gives ASCII spellings
that resolve everywhere, but it is not an escape from the locale problem:
//TRANSLIT is itself locale-dependent, so under LC_CTYPE=C it returns
NA (or, given an explicit from = "UTF-8", replaces each accent with ?)
and nothing resolves. De-accent while still in a UTF-8 locale, or supply the
ASCII spellings directly.
Examples
# `country_overrides()` is the current name; `wdj_overrides()` warns.
country_overrides()
country_overrides(c(Somaliland = "SOM"))
Map-ready, enriched country tibble
Description
The package's headline function, generalised but backward-compatible. Returns
a tibble that already stitches together map geometry, World Bank indicators
and the countrycode crosswalk, keyed on the ISO spine – ready to pipe into
world_map() or ggplot2.
Usage
world_data(
year,
indicator = c(gdp_per_capita = "NY.GDP.PCAP.KD"),
geometry = c("polygon", "sf", "none"),
scale = c("small", "medium", "large"),
region = NULL,
classify = c("income", "continent", "region"),
projection = "equal_earth",
recenter = NULL,
latest = FALSE,
cache = TRUE,
language = "en",
parallel = TRUE,
overrides = country_overrides()
)
Arguments
year |
A single year or a range (e.g. |
indicator |
A named character vector of WDI codes. Names drive column
names, e.g. |
geometry |
|
scale |
Natural Earth resolution for the |
region |
Optional subset: a continent, group name, |
classify |
Which classifications to add (any of |
projection, recenter |
Projection, and optional central meridian, for
the |
latest |
If |
cache |
Whether to use the memoised / on-disk WDI cache. |
language |
WDI language code (default |
parallel |
Whether to fetch multiple indicators in parallel. Ignored
when the cache is memory-only (an unwritable |
overrides |
Name -> iso3c overrides for geometry matching (default
|
Details
world_data(2020) keeps its original behaviour (polygon backend, GDP per
capita). Everything else is opt-in: any indicator(s), a span of years (a
panel), an sf backend with real projections, and region subsetting.
Value
A tibble (polygon backend), sf object (sf backend) or country-level
tibble (geometry = "none").
iso3c is the stable key; country is a label and its spelling depends on
where the row came from. A successful fetch carries the World Bank's names
("Korea, Rep.", "Congo, Dem. Rep."), while the country spine used when the
fetch returns nothing carries the countrycode names ("South Korea",
"Congo - Kinshasa") – as do convert_country(), standardize_country()
and the rest of the package. Match on iso3c, and relabel with
convert_country(iso3c, to = "country") if you need one consistent set.
Examples
# geometry = "polygon", the default, comes from the suggested `maps`
# package, so guard the call: an example may not assume a Suggests is
# installed (R CMD check runs \donttest{} blocks, and CRAN has a
# check flavour with no suggested packages at all).
if (requireNamespace("maps", quietly = TRUE)) {
world_data(2020)
}
# geometry = "none" needs nothing beyond the hard dependencies.
world_data(2020, indicator = c(life_exp = "SP.DYN.LE00.IN"),
geometry = "none")
Geometry without the data
Description
Sometimes you just want the canvas: country polygons, label-ready centroids, coastlines, internal borders, a graticule or an ocean rectangle – already projected, region-subset and antimeridian-safe. This is the building block the plotting functions sit on, exposed for power users.
Usage
world_geometry(
what = c("countries", "centroids", "coastline", "borders", "graticule", "ocean"),
geometry = c("polygon", "sf"),
scale = "small",
region = NULL,
projection = "equal_earth",
recenter = NULL,
year = NULL
)
Arguments
what |
What to return: |
geometry |
|
scale |
Natural Earth resolution for the |
region |
Optional subset: a continent, a group name, a vector of |
projection |
Projection for the |
recenter |
Optional central meridian (e.g. |
year |
Draw the world as it was in this year, via |
Value
A tibble (polygon backend) or sf object (sf backend), with columns
depending on what:
"countries"polygon:
long,lat,group,order,region,subregion,iso3c,iso2c. sf:iso3c,iso2c,name_long."centroids"the same identifier columns plus
centroid_lonandcentroid_lat."coastline","borders","ocean","graticule"sf only.
The centroid columns are in the coordinate system of the object
returned, so on the sf backend they are projected metres, not degrees –
centroid_lon for France is 174097, not 2.1. For centroids in degrees
use country_meta$centroid_lon / $centroid_lat, which is also what the
polygon backend returns.
A few Natural Earth features have no ISO code and so come back with iso3c
NA – Somaliland at every scale, plus the Indian Ocean Territories and
Ashmore and Cartier Islands from "medium" on. They are kept so the land is
still drawn; drop or country_overrides() them if you group by iso3c.
"orthographic" is the one genuinely hemispheric projection: the countries
on the far side have no image and come back as empty geometries, and the
ones on the horizon are cut there (correctly, but sf::st_coordinates()
cannot read a column that mixes empty and non-empty – drop them first).
The other three azimuthal projections
("azimuthal_equal_area", "north_polar", "south_polar") are Lambert
equal-area and draw the whole globe, the far side stretched around the
rim rather than dropped, so pass region if you want a polar view of the
northern countries alone.
"ocean" is a whole-globe background rectangle. It is unavailable in all
four azimuthal projections – "orthographic" has no image for it, and the
Lambert three cut the globe at the antipode, which collapses the rectangle's
outline – and it cannot be recentred; both cases error rather than
returning an invisible layer.
Examples
if (requireNamespace("maps", quietly = TRUE)) {
head(world_geometry("countries", geometry = "polygon"))
}
One-line choropleth, several honest styles
Description
Encapsulates the choropleth boilerplate and goes beyond a single style.
Auto-detects the polygon vs sf backend, applies theme_world_map(), and –
for sf – a real projection via ggplot2::coord_sf(). Binned / quantile /
jenks styles are offered because a continuous fill on a skewed indicator
hides almost all the variation; binning is the honest default for
choropleths.
Usage
world_map(
data,
fill,
style = c("continuous", "binned", "quantile", "jenks", "categorical"),
projection = "equal_earth",
palette = NULL,
n_bins = 5,
borders = TRUE,
title = NULL,
legend = NULL,
na_label = "No data",
recenter = NULL,
na_style = c("grey", "hatched", "outline", "omit"),
footnote = NULL,
classification_report = FALSE,
uncertainty = NULL,
n_uncertainty = 3,
disputes = c("ignore", "mark"),
engine = c("ggplot2", "tmap")
)
Arguments
data |
A map-ready frame from |
fill |
The fill column (unquoted). |
style |
|
projection |
For the |
palette |
Optional palette name passed to the relevant |
n_bins |
Number of bins for binned/quantile/jenks styles. |
borders |
Draw country borders (default |
title, legend |
Optional plot title and legend title. |
na_label |
Legend key label for missing data, used by the styles with
a discrete legend ( |
recenter |
Optional central meridian for the |
na_style |
How to draw countries with no data: |
footnote |
Optional caption. |
classification_report |
If |
uncertainty |
Optional uncertainty column (unquoted) – a standard error, a confidence half-width, anything where larger means less certain. Supplying it switches the fill to a value-suppressing uncertainty palette (Correll, Moritz & Heer 2018): the value range contracts as uncertainty rises, so an uncertain estimate cannot claim an extreme colour, and the legend becomes the value x uncertainty grid. |
n_uncertainty |
Number of uncertainty levels for the VSUP (default |
disputes |
|
engine |
|
Value
A ggplot object.
Missing data is not zero
The default grey reads as "low" to many people, which is exactly wrong for
"unknown". na_style = "hatched" draws diagonal hatching instead –
unambiguous, and it survives greyscale printing. "omit" leaves a hole,
which is honest but can be mistaken for ocean. Whichever you pick,
footnote = "auto" states the count in words:
world_map(mapdf, gdp_per_capita, na_style = "hatched", footnote = "auto")
coverage_map() goes further and maps availability itself.
Choosing a classification
The classification changes what readers conclude, and not by a little.
Brewer & Pickle's 56-subject study over nine map series found quantiles
among the best methods for general choropleth reading, and natural breaks
(Jenks) below 70% as accurate – the opposite of the common GIS default.
style = "quantile" is therefore the safe choice for a general audience.
Jenks earns its place on strongly clustered distributions, where quantiles
would split a natural group across two colours. Use classify_compare() to
see the difference on your own data before committing.
References
Brewer, C. A. & Pickle, L. (2002). Evaluation of methods for classifying epidemiological data on choropleth maps in series. Annals of the Association of American Geographers 92(4), 662-681. doi:10.1111/1467-8306.00310
Correll, M., Moritz, D. & Heer, J. (2018). Value-suppressing uncertainty palettes. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 1-11. doi:10.1145/3173574.3174216
See Also
classify_compare(), coverage_map(), projection_compare(),
map_provenance(), dispute_policy()
Examples
snap <- countryatlas::world_snapshot$countries
if (requireNamespace("maps", quietly = TRUE)) {
mapdf <- attach_geometry(snap, geometry = "polygon")
world_map(mapdf, gdp_per_capita, style = "quantile")
}
Emit a ggsql spatial query for a country map
Description
Build a ggsql query string that draws a choropleth from
a registered countryatlas source – the same idea as world_map(), but the
map is rendered in the database (DuckDB) and returned as a web-ready
Vega-Lite widget, so the geometry never has to come back into R. Pure string
builder with no dependencies; pair it with as_ggsql_source() +
ggsql::ggsql_execute(), or drop the string into a {ggsql} chunk.
Usage
world_query(
fill,
source = "countryatlas_world",
projection = "equal_earth",
palette = "viridis",
transform = NULL,
title = NULL,
draw = "spatial",
layer = c("choropleth", "bubble", "binned"),
facet = NULL,
size = NULL,
n_bins = NULL
)
Arguments
fill |
The fill column (unquoted or a string). |
source |
The table/source name registered with ggsql (default
|
projection |
A projection ggsql's |
palette |
A scale ggsql's |
transform |
Optional scale transform for |
title |
Optional plot title ( |
draw |
The spatial layer (default |
layer |
|
facet |
Optional column to facet the query by, e.g. |
size |
Column driving symbol size for |
n_bins |
Number of classes for |
Value
A ggsql_query string (prints as the formatted query).
Executing the query
Building the string needs nothing installed. Running it needs
ggsql >= 0.4.1, the version that added the DRAW spatial clause; older
ggsql releases parse the query and reject that clause. As of August 2026
that clause has shipped in the ggsql engine but not yet in the ggsql R
package (still 0.3.3), so interactive_map()(engine = "ggsql") will refuse
until the bindings catch up. PROJECT TO additionally needs a spatial
backend – for DuckDB, its spatial extension.
Examples
world_query(gdp_per_capita, projection = "equal_earth",
palette = "magma", transform = "log10",
title = "GDP per capita")
Offline snapshot of world data
Description
A small, lazy-loaded, one-row-per-country snapshot of a curated indicator set for one recent year. It lets every example, test and vignette run offline and deterministically, without the World Bank API.
Usage
world_snapshot
Format
A list with three elements:
- countries
A tibble, one row per country, with
iso3c,iso2c,country, the classificationscontinent,regionandincome, and the curated indicatorsgdp_per_capita,population,life_expectancyandco2_per_capita.- sf
NULLin the released package – geometry is not bundled twice. Attach it on demand withattach_geometry():attach_geometry(world_snapshot$countries, geometry = "sf")pulls the same Natural Earth 110m polygons fromrnaturalearth.- year
The reference year.
country carries the World Bank's own names, which differ from the
countrycode names used by country_meta for 38 countries.
Source
World Bank via WDI; geometry from Natural Earth via rnaturalearth. Snapshot year: 2024.
A publication-ready table from a map-ready frame
Description
The tabular counterpart to world_map(): take the same curated frame and
produce a ranked, formatted table instead of a picture. Uses gt when it is
installed and a plain tibble otherwise, so it never becomes a hard dependency.
Usage
world_table(
data,
value = NULL,
top_n = 20,
desc = TRUE,
columns = NULL,
engine = c("gt", "tibble"),
title = NULL,
subtitle = NULL
)
Arguments
data |
A country-level or map-ready frame. |
value |
The column to rank on (unquoted). Tied values share a rank, as
in |
top_n |
How many rows (default |
desc |
Sort descending (default |
columns |
Extra columns to keep, beyond |
engine |
|
title, subtitle |
Optional table title and subtitle ( |
Value
A gt table, or a tibble when gt is unavailable or
engine = "tibble".
See Also
rank_countries(), country_factsheet(), world_map()
Examples
world_table(countryatlas::world_snapshot$countries, gdp_per_capita,
top_n = 5, engine = "tibble")
Equal-area world tile-grid layout
Description
A statebins-style equal-area tile layout: one square per country, positioned
on a row/col grid derived from country centroids. Used by tile_map().
Usage
world_tiles
Format
A tibble with columns iso3c, country, row, col; one row per
country, with row/col unique across the grid.
Details
The grid holds one row for each of the 239 countries in country_meta that
has a bundled centroid; the 10 without one (ALA, BVT, GIB, HKG,
MAC, SJM, TKL, TUV, UMI, VGB – see country_meta) have no tile
and so cannot be drawn by tile_map().
Source
Derived from Natural Earth country centroids.