Fixing a Rate Map Dominated by Tiny, Nearly Empty Areas

Problem statement

You map a rate by county and the darkest colour lands on remote, sparsely populated places: West Texas, the Great Plains, Nevada. The pattern looks striking and it is mostly the denominator.

Measured on 2023 crude death rates for the 3,144 US counties, from the Census Bureau's county population estimates:

  • Loving County, Texas, has the highest rate: 65.2 per 1,000, from 3 deaths among 46 residents. The national rate is 9.4.
  • The 20 counties with the highest rates have a median population of 2,104.
  • Counties under 10,000 people are 24% of counties and 32.3% of the land, but hold 1.20% of the population. On a choropleth they take a third of the ink.

Nothing in the data is wrong. The map is answering "which small counties had an unusual year?" when the reader thinks it answers "where is mortality high?".

Quick answer

Diagnose first, then choose a fix that matches what the map must say:

flag = unreliable(counties, "DEATHS2023", "POPESTIMATE2023")        # Example 3
top = top_class_profile(counties, "raw", "POPESTIMATE2023", "km2")  # Example 2
306 areas with a 95% half-width above 25% of the rate; median population 2,582
   raw: top class median population 12,388; 129 of 629 under 5,000; 18.9% of the land; ...

Then pick one:

  • Keep the observed rates, and mark the unreliable ones with hatching or transparency.
  • Pool several years to grow the denominators before computing the rate.
  • Smooth with Empirical Bayes, and check against held-out years that it helped.
  • Map the counts instead, with proportional symbols, if the question is where events happen.
Table of the share of US counties, land area and population in counties under 5,000, 10,000 and 25,000 people.
Half the counties, and more than half the map, hold one person in twenty.

Step-by-step solution

1. Profile the darkest class

Look at who is in it before looking at where it is. On a five-class quantile map of raw 2023 death rates, the top class held 629 counties with a median population of 12,388; 129 of them had fewer than 5,000 people. The top 20 counties had a median of 2,104.

A top class whose median population is far below the median of all counties is the first sign that the map is ranking denominators.

2. Compare land with people

counties under     share of counties   share of land   share of people
   5,000                 10%               13.2%            0.26%
  10,000                 24%               32.3%            1.20%
  25,000                 48%               54.6%            5.08%

A choropleth weighs each county by its area. Counties under 25,000 people cover more than half of the map and hold about one person in twenty, so their rates set the visual impression regardless of what the rates are.

3. Measure which rates are unreliable

Deaths in a year behave roughly like a Poisson count, so the standard error of a rate is the square root of the count divided by the population. Flag rates whose 95% interval is wider than a quarter of the rate itself:

306 areas with a 95% half-width above 25% of the rate; median population 2,582

Loving County's rate of 65.2 has a standard error of 37.7. Those 306 counties cover 13.8% of the land and 0.27% of the people, and 79 of them sit in the darkest class of the raw map.

4. Show the uncertainty on the map

If the map must show observed rates โ€” a statutory report, a dashboard of recorded values โ€” keep them and draw the flagged counties differently:

ax = counties[~flag].plot(column="raw", scheme="quantiles", k=5, cmap="Reds", legend=True)
counties[flag].plot(ax=ax, column="raw", scheme="quantiles", k=5, cmap="Reds",
                    alpha=0.35, hatch="///", edgecolor="grey", linewidth=0.2)

Faded, hatched polygons tell the reader that the colour is there but not to be trusted. Say so in the legend.

5. Pool years to grow the denominator

Four years of deaths over four years of population is still a crude rate, with four times the evidence:

method          top class median pop   under 5,000 in top class   top class land   top-20 median pop   range
raw 2023              12,388                    129                    18.9%              2,104         0.0โ€“65.2
2021โ€“2024 pooled      12,962                    127                    14.8%              2,844         0.0โ€“55.8

Pooling cut the land in the darkest class from 18.9% to 14.8% and the maximum from 65.2 to 55.8, but the top 20 were still small counties. For a county of 46 people, four years is not enough.

6. Smooth, and check that it helped

Empirical Bayes pulls each rate towards a prior in proportion to how little evidence it rests on:

method          top class median pop   under 5,000 in top class   top class land   top-20 median pop   range
global EB             15,970                     77                    14.2%             15,210         2.7โ€“30.0
spatial EB            15,152                     90                    14.0%             10,826         2.8โ€“30.3

The top 20 became counties with a median of 15,210 people, and the map stopped being led by Loving County. But smoothing can also remove real differences: against deaths in other years, global Empirical Bayes made small-county rates less accurate on this data, because small rural counties really are older. Run the held-out check in the smoothing guide before publishing a smoothed map.

7. Map counts when the question is about where

If the question is where deaths occur โ€” for planning services, not comparing risk โ€” map the number of deaths with circles sized by count. Large counties then dominate, correctly, and area plays no part.

Bar chart of the median population of the 20 counties with the highest death rates under raw, pooled, spatial Empirical Bayes and global Empirical Bayes rates.
The top of a raw rate map is a list of tiny counties; smoothing replaces it with counties of ordinary size.

Code examples

Example 1 โ€” who owns the map

import numpy as np
import pandas as pd
import mapclassify


def who_owns_the_map(frame, population, area, thresholds=(5000, 10000, 25000)):
    """Share of areas, land and people below each population threshold."""
    total_area, total_pop = frame[area].sum(), frame[population].sum()
    for t in thresholds:
        s = frame[frame[population] < t]
        print(f"under {t:,}: {len(s) / len(frame):.0%} of areas, {s[area].sum() / total_area:.1%} of land, "
              f"{s[population].sum() / total_pop:.2%} of people")
under 5,000: 10% of areas, 13.2% of land, 0.26% of people
under 10,000: 24% of areas, 32.3% of land, 1.20% of people
under 25,000: 48% of areas, 54.6% of land, 5.08% of people

area is land area in kmยฒ, here ALAND / 1e6 from the boundary file. Run it on any set of areas before choosing a choropleth.

Example 2 โ€” profile the darkest class

def top_class_profile(frame, column, population, area, k=5, small=5000):
    """Who sits in the darkest class of a quantile map of `column`."""
    classes = mapclassify.Quantiles(frame[column], k=k)
    top = frame[classes.yb == k - 1]
    print(f"{column:>6}: top class median population {top[population].median():,.0f}; "
          f"{int((top[population] < small).sum())} of {len(top)} under {small:,}; "
          f"{top[area].sum() / frame[area].sum():.1%} of the land; "
          f"top-20 median population {frame.nlargest(20, column)[population].median():,.0f}; "
          f"range {frame[column].min():.1f}-{frame[column].max():.1f}")
    return top
   raw: top class median population 12,388; 129 of 629 under 5,000; 18.9% of the land; top-20 median population 2,104; range 0.0-65.2
pooled: top class median population 12,962; 127 of 629 under 5,000; 14.8% of the land; top-20 median population 2,844; range 0.0-55.8
    eb: top class median population 15,970; 77 of 629 under 5,000; 14.2% of the land; top-20 median population 15,210; range 2.7-30.0
   seb: top class median population 15,152; 90 of 629 under 5,000; 14.0% of the land; top-20 median population 10,826; range 2.8-30.3

Running it on each candidate column turns "which map looks better" into a comparison of who ends up in the darkest colour.

Example 3 โ€” flag the rates the data cannot support

def unreliable(frame, events, population, per=1000, share=0.25):
    """Areas whose 95% Poisson interval is wider than `share` of their own rate."""
    rate = frame[events] / frame[population] * per
    se = np.sqrt(frame[events].clip(lower=1)) / frame[population] * per
    flag = 1.96 * se > share * rate
    print(f"{int(flag.sum())} areas with a 95% half-width above {share:.0%} of the rate; "
          f"median population {frame.loc[flag, population].median():,.0f}")
    return flag
306 areas with a 95% half-width above 25% of the rate; median population 2,582

Clipping the count at 1 gives areas with zero events a non-zero standard error, so a rate of 0 from a tiny population is flagged rather than presented as certain. For survey estimates, use the published margin of error instead of the Poisson approximation.

Explanation

Why the extremes come from small places

With few people, a single event moves a rate a long way. One death more or fewer in Loving County changes its rate by 21.7 per 1,000. Large counties average out the year-to-year chance; small ones display it. Across counties, the raw rate's correlation with the logarithm of population was โˆ’0.438, and the standard deviation of rates was 6.88 in counties under 2,500 people against 2.34 in counties of 100,000 or more.

Why the map amplifies it

Choropleth colour is spread over land, and small populations live on large areas. The least populated tenth of counties covers 13.2% of the country, so whatever their rates show fills an eighth of the image while describing a quarter of one percent of the people. Classification makes it worse: quantile breaks put the most extreme values in the darkest class, and extremes are exactly what small denominators produce.

Why pooling helps less than expected

Pooling multiplies the evidence by the number of years, which halves the standard error with four years. That is a large improvement for a county of 5,000 people and a small one for a county of 46, whose four-year denominator is still under 200 person-years. The maximum pooled rate was still 55.8.

Why smoothing needs a check

Smoothing assumes that most of the spread in small areas is noise around a shared mean. When small areas genuinely differ โ€” here, older populations in rural counties โ€” shrinking them towards the national mean removes signal along with noise. The smoothed map looks calmer and can be less accurate. A held-out comparison against other years is what distinguishes the two cases.

Decision diagram choosing between flagging unreliable rates, pooling or smoothing, and mapping counts, depending on what the map must show.
There is no single fix: the right one depends on whether the map reports, compares or locates.

Edge cases or notes

  • Zero rates are unreliable too. Kalawao County, Hawaii, had 0 deaths among 81 people; a rate of 0 there is not evidence of immortality.
  • Quantile maps are the most affected classification because the extremes set the darkest class.
  • Crude rates confound age. Rural counties are older; age-standardised rates answer a different and usually better question.
  • Hatching needs a legend entry. A reader cannot infer what faded polygons mean.
  • Suppressing small areas hides a third of the map. Flagging keeps the geography visible.
  • Pooling changes the period. Label a 2021โ€“2024 rate as such; it is not a 2023 rate.
  • Survey rates have margins of error. Use them instead of Poisson standard errors for ACS data.
  • A cartogram or a dot map sidesteps area. Both size the display by people rather than land, at the cost of a less familiar shape.

FAQ

Why do the darkest areas on my rate map all have tiny populations?

Because small denominators produce extreme rates by chance. The 20 US counties with the highest 2023 death rates had a median population of 2,104, led by Loving County with 3 deaths among 46 people.

How do I know which rates are unreliable?

Estimate a standard error. For counts of events, the square root of the count divided by the population works; 306 counties had a 95% interval wider than a quarter of their own death rate.

Should I just remove the small areas?

Usually not. Counties under 10,000 people cover a third of US land; removing them leaves a map with holes. Flag them with hatching or transparency instead.

Does pooling several years fix it?

It helps: the darkest class covered 14.8% of the land instead of 18.9%. But the top 20 were still small counties, because even four years of a tiny population is a small denominator.

Is Empirical Bayes smoothing the answer?

It removed the tiny counties from the top of the ranking, raising the top 20's median population to 15,210. Check it against other years first; on this data it made small-county rates less accurate overall.

When should I map counts instead of rates?

When the question is where events happen rather than where risk is high. Proportional symbols sized by count show that without letting land area drive the picture.