← Back to Blog
DataAnalysisDeveloperNeighborhood SafetyReal Estate

Three Burglaries Is Not a Trend

📅 August 24, 2026·⏱ 11 min read·By SpotCrime

In 1999 Andrew Gelman and Phillip Price published two maps of United States counties. One shaded the counties in the highest tenth of kidney-cancer death rates. The other shaded the counties in the lowest tenth. The maps look alike. Both light up the rural Midwest and the Great Plains, the same thinly settled places, again and again. Their paper is called All Maps of Parameter Estimates Are Misleading, and the explanation has nothing to do with kidney cancer.

A county of a thousand people can hold the worst rate in the country on four deaths and the best rate in the country on zero. Both are easy. Neither means anything. Small denominators produce extreme rates in both directions, and a map of raw rates is mostly a map of which places are small.

Crime data has this disease worse. The unit everyone wants is not a county of a thousand people. It is a block.

The county that is worst and best at once

Gelman and Price were writing for epidemiologists, in Statistics in Medicine, about the practice of shading a map by observed rate. Their argument is arithmetic. When sample sizes vary across areas, the highest and lowest observed rates land disproportionately in the poorly sampled ones, because those are where a single extra event moves the rate the most. Nothing about the underlying risk needs to differ at all for the map to look like a story.

The part people forget is the second half. Gelman and Price also showed that adjusting for small-sample noise flips the bias rather than removing it. Shrink each county toward the national average in proportion to how noisy it is, and now the highest adjusted rates cluster in the well-sampled places, because those are the only ones allowed to stay extreme. The title is literal. There is no map here that is not misleading in some direction. The job is to pick which distortion you can defend.

The finding in one sentence

Raw rates make small places look extreme. Adjusted rates make big places look extreme. Both maps are wrong, and neither is wrong at random.

What a block actually holds

Most US incident feeds publish location at the hundred-block level, a convention we walked through in the geocoding problem. A hundred-block segment is one side of one street between two cross streets. Call it twenty to forty addresses. In a dense city, a few hundred residents.

Now count burglaries there. In most of the country the answer for any given year is zero. Sometimes one. Occasionally two or three. A hundred-block segment that reliably produces double-digit burglaries every year exists, and there are not many of them, and everyone in the precinct already knows where they are.

So every block-level statistic anyone builds is computed on counts in the single digits. Percent changes, year-over-year comparisons, rankings, safety scores, the little red arrow on a listing page. All of it rests on numbers small enough that chance is the loudest thing in them.

Andy Wheeler did this arithmetic in 2014

A former crime analyst who now runs CRIME De-Coder, Wheeler wrote a post called Understanding Uncertainty that should be required reading for anyone shipping a crime product. It opens with the request every analyst gets: something bad happened, the media is calling it a wave, please provide some related analysis.

His examples are the whole argument.

A jurisdiction averages two homicides a year and has five so far. That sounds like a crisis. Under a Poisson model the probability of seeing five or more is 5.7 percent. Uncommon, not remotely impossible. An agency averaging fifteen robberies a month sees twenty-one. The probability is about 8 percent, which means you should expect a month like that roughly once a year, forever, with nothing changing underneath.

The third example is the one that should worry anyone running alerts. A jurisdiction averages three domestic assaults a week and one week posts nine. The probability of nine or more in a given week is 0.003. Genuinely rare. But over fifty-two weeks the probability of it happening at least once is about 18 percent, and over a decade you would expect it about twice. The event is rare. Seeing the event somewhere in a long series is not.

Wheeler's phrasing for what this is for: the analyst's job is to “avoid chasing noise.” He also flags the caveat honestly. Crime counts are frequently over-dispersed rather than Poisson, with more zeros and bigger bursts than the model expects, in which case the probabilities above are, if anything, too alarmed. Check whether the variance tracks the mean before you trust the p-value.

How much noise you are carrying

For a Poisson count, the standard deviation is the square root of the mean. That single fact governs everything downstream. Relative noise falls as 1 over the square root of the count, which means it falls slowly, which means small counts stay hopeless for a long time. Below are approximate central 95 percent ranges for the observed count when the true underlying rate is fixed and nothing at all has changed.

0–7
Observed range when the true rate is 3
1–10
True rate 5
4–17
True rate 10
16–35
True rate 25
81–120
True rate 100
58% → 10%
Relative noise, from a true rate of 3 down to a true rate of 100

Read the first tile again. A block whose true burglary rate never budges will show you zero one year and seven the next, and both are ordinary. Everything a user might infer from that pair is wrong.

Percent change is the wrong statistic here

One incident to three is a 200 percent increase. Three to one is a 67 percent decrease. Both sentences are true, both are publishable, and neither carries information. Percent change divides by the smallest, noisiest quantity in the calculation, which is exactly backward from what you want.

It is also the format that travels. Nobody screenshots a confidence interval. They screenshot the arrow. We wrote about a related failure in crime seasonality, where a month-over-month comparison reports the calendar rather than crime. Small counts and seasonality compound. A single block, compared across two months, is telling you about July and about luck, in some unknown mixture, and about the neighborhood barely at all.

Why the worst block always improves

Pick the top fifty blocks in a city by last year's incident count. Do nothing. Come back in a year. They will be down.

This is regression to the mean, and it is not a subtlety. Selecting on a noisy observation selects partly for high true risk and partly for a lucky year, and the lucky year does not repeat. The steeper the noise, the more of the apparent improvement is mechanical. Any evaluation that measures a hot-spot intervention by comparing the targeted blocks before and after is measuring its own selection rule.

The same trap sits inside product surfaces that look nothing like a policing evaluation. A neighborhood page that highlights the three worst blocks this quarter will show improvement next quarter and a different three worst blocks, indefinitely, on a stationary city. A real-estate valuation model fed block-level deltas will learn that high-crime blocks improve, which is a fact about the selection, not about the blocks. We covered the geographic half of this in hotspot mapping: bin choice and bandwidth already move the picture before any of the statistics start.

The test

Before shipping any block ranking, simulate it. Take the city's own annual totals, redistribute them randomly across blocks in proportion to a fixed rate, and run your ranking on the fake data. If it produces a confident-looking list of worst blocks and a satisfying year-over-year improvement, the ranking is measuring the random number generator.

Why the national numbers survive all this

None of the above is an argument against crime statistics. It is an argument about denominators, and the national series have enormous ones.

The Real-Time Crime Index reports murder down 17.3 percent for January through June 2026 against the same period in 2025, drawn from 590 agencies covering 119.4 million people. Motor vehicle theft is down 18.7 percent, burglary 15.1 percent, robbery 14.3 percent, and property crime overall 10.1 percent. Violent crime is down 5.9 percent, with aggravated assault the laggard at 3.1 percent. Those figures rest on tens of thousands of events. Relative Poisson noise on a count in the thousands is a couple of percent, which is small next to a 17 percent move.

The same square root that ruins a block rescues a country. This is why murder is the anchor series in nearly every credible analysis, a point we made in the dark figure of crime, and why the state-level spread in the geography of the crime decline is interpretable while the block-level spread mostly is not. The caveat runs in the obvious direction: the least populous states sit closer to the edge of this than the maps suggest, and a single bad year in a state with a few dozen murders is not the same evidence as a single bad year in California.

The trend and the block are different products with different error budgets. Trouble starts when a pipeline treats them as the same query at different zoom levels.

Seven things to do about it

  1. Set a display floor. Below some count, refuse to show a percent change or a rank at all. Thirty events is a defensible floor for a two-period comparison. Under that, show the raw count and say what window it covers. A number with no ranking attached is honest. A rank built on four events is not.
  2. Buy counts with time or space. The only two ways to get above the floor are a longer window or a bigger area. A 36-month window at block level, or a twelve-month window over a few adjoining blocks, both work. Say which one you chose, in the response payload, per record.
  3. Return an interval, not a point. If the endpoint emits a rate or a score, emit its uncertainty in the same object. Consumers who ignore it are no worse off. Consumers who use it can stop a bad inference before it reaches a user.
  4. Use a real test for two-period comparisons. Wheeler's Poisson e-test is designed for exactly this and takes about four lines of code. It answers whether the change is distinguishable from chance, which percent change never attempts.
  5. Use funnel charts for cross-sectional comparison. Plot rate against population with control limits that widen as the denominator shrinks. The shape of the funnel teaches the reader the small-numbers problem without a word of explanation, which a league table actively hides.
  6. Shrink, and disclose which way it bends. Empirical Bayes pulls each block toward the local mean in proportion to its noise, and it is the right default for a score. It also does the thing Gelman and Price described: it moves the extremes toward well-measured areas. Document the direction of the bias you chose. Our own weighting and aggregation choices are written up in the SpotScore methodology.
  7. Count your comparisons. Scanning twenty thousand blocks at a 5 percent threshold yields about a thousand flagged blocks in a city where nothing happened. Alerting systems scan continuously, which multiplies the exposure again. Adjust the threshold, or accept that most of what you fire is noise.

The model will not save you

One more, because it is now the common failure. A language model asked to summarize a block's crime history will produce fluent narration of a random walk. It has no access to the variance, and it will not volunteer that four incidents cannot support the sentence it just wrote.

Gio Circo's work on classifier calibration using NEISS injury data found that model confidence scores and token probabilities run overconfident, which is the same problem one layer up. The division we argued for in AI in law enforcement holds here exactly: deterministic code owns the arithmetic and the thresholds, the model owns the language, and a human owns the decision. Compute the interval in code. Hand the model the interval. Let it write the sentence.

The related trap is what happens when the alert reaches a person. On April 13, 2026, CrimeRadar pushed an active-shooter alert to parents in Mount Vernon, Missouri, based on a misheard radio transmission. No shooting occurred. We wrote that up in speed versus accuracy. A false positive from a bad transcription and a false positive from an unremarkable Poisson fluctuation arrive on the same phone, looking the same.

What honest looks like

An honest block-level product says less than users expect and says it more precisely. Three burglaries on this block in the last three years. Comparable to the surrounding twenty blocks. Not enough events to say whether it is getting better or worse.

That last sentence is the one nobody wants to ship. It is also the only one on the page that is definitely true.

Somewhere in your database is a block that had three burglaries last year and one this year. The chart renders a steep green line. Nothing happened to that block. The line is the finding.

Access Address-Level Crime Data

Real-time incidents · SpotScore™ safety ratings · 36-month trends · 22,000+ US cities. Normalized and verified — because raw data isn't enough.