← Back to Blog
DataAnalysisDeveloperCrime TrendsAPI

The Enforcement Confound: Why Part of Every Crime Feed Measures Policing, Not Crime

πŸ“… August 24, 2026·⏱ 12 min readΒ·By SpotCrime

Between 2019 and 2024 the adult drug arrest rate fell by roughly half. Adult violent crime arrests over the same window fell about 16%. Nobody thinks drug activity in the United States dropped by 50% in five years. The gap between those two numbers is not a crime trend β€” it is a measurement of how much discretionary enforcement police departments were able and willing to perform. Every crime data feed in the country carries some of that signal, mixed in with the thing you actually wanted to measure.

This is the enforcement confound, and it is the most under-discussed structural problem in applied crime data. It is not a data quality issue in the usual sense β€” the records are accurate, the agencies are reporting in good faith, the pipeline is clean. The problem is that a subset of offense categories has no independent existence in the record at all. They appear when an officer looks, and they do not appear when no officer looks. Feed those categories into a block-level safety score, a trend chart, or a retrieval tool for a language model, and you will confidently report a change in neighborhood conditions when what changed was the size of an academy class.

Two framings up front, because they govern everything below. First, this is a claim about measurement, not about whether proactive policing works; those are different arguments and we will separate them explicitly near the end. Second, the confound is not uniform β€” it is close to zero for some offenses and close to total for others, and the practical response is to know which is which rather than to distrust the feed globally.

Two channels into the record

Every incident in a police records management system got there through one of three doors. A victim report: someone was harmed and called. A third-party report: a witness, a neighbor, a business, an alarm company, a hospital called. Or officer observation: an officer initiated a stop, a search, a patrol check, or a proactive operation and generated the record themselves.

The first two channels are demand-side. Their volume is governed by how much crime occurs and by what share of it victims choose to report β€” the reporting rate, which is itself unstable and which we covered in detail in the dark figure of crime. The third channel is supply-side. Its volume is governed by how many officers are on the street, how they are deployed, what the department's enforcement priorities are this quarter, and what the state legislature has decided is illegal.

Nothing in a standard incident feed labels which door a record came through. The schema gives you an offense code, a timestamp, a location, and β€” if you are lucky β€” a report type. The initiation channel has to be inferred from the offense category, and for most categories the inference is straightforward.

The distinction in one sentence

A homicide enters the record whether or not anyone is patrolling. A drug possession offense enters the record only because someone was patrolling.

What the arrest data shows

The clearest evidence for the confound comes from arrest statistics, because arrests are the purest expression of enforcement activity. The Council on Criminal Justice's four-decade arrest analysis (1980–2024) puts total US arrests at roughly 10.1 million in 2019 and roughly 7.5 million in 2024 β€” a decline of about 25% in five years, and roughly 50% below the 1997 peak of about 15 million. On a population-adjusted basis, the 2024 arrest rate sits about 30% below 2019 and about 71% below its 1994 peak.

The interesting part is the composition of that decline, not its size.

βˆ’50%
Adult drug arrest rate, 2019 β†’ 2024
βˆ’16%
Adult violent crime arrests vs. 2019
βˆ’14%
Adult property crime arrests vs. 2019
591
Adult drug arrests per 100k, 2024 (62% below 2006 peak)
~377k
Adult violent arrests, 2024
~825k
Adult property arrests, 2024

Drug offenses fell three times faster than violent offenses. That ratio is the confound made visible. Violent arrests are anchored to violent incidents, which are anchored to victims who call. Drug arrests are anchored to nothing except the decision to conduct a stop. When enforcement capacity contracts, the anchored series moves a little and the unanchored series moves a lot.

A methodological note that matters if you go looking at the underlying files: the FBI reported 831,446 drug-related arrests nationally in 2024, of which about 187,792 were marijuana possession, per NORML's reading of the 2024 release. That count is not directly comparable to a 591-per-100,000 adult rate: one is a count of arrests reported by participating agencies, the other is a modeled rate normalized to the adult population. Differencing them against each other produces a number that means nothing. This is the same participation problem that runs through everything post-2021, and it is worth re-reading our NIBRS explainer before treating any national arrest series as a continuous time series.

An elasticity ordering

It is useful to think of each offense category as having an enforcement elasticity: the share of its recorded volume that would change if patrol activity changed while underlying behavior held constant. Nobody has published a rigorous per-category estimate of this, and the numbers below are a qualitative ordering rather than measured coefficients. Treat them as a design heuristic.

  • Near zero. Murder and nonnegligent manslaughter. A body is discovered and a case is opened regardless of staffing. This is why murder is the anchor series in nearly every serious crime analysis.
  • Low. Motor vehicle theft, residential burglary, robbery. Insurance claims and direct victimization force a report. Reporting rates for these are high and comparatively stable.
  • Moderate. Aggravated assault, simple assault, larceny-theft, vandalism. Real victimization drives the volume, but police discretion drives the classification β€” the aggravated / simple boundary in particular is a judgment call that varies by agency and by era. See the taxonomy normalization problem.
  • High to near-total. Drug possession, drug sale, weapons offenses, DUI, prostitution, loitering, trespass, disorderly conduct, liquor law violations, resisting arrest. These are frequently called victimless offenses; the more precise term for our purposes is officer-initiated. Their recorded volume is close to a direct readout of patrol intensity.

The practical rule follows directly: if you are building anything that is supposed to describe conditions at a place, the bottom category of that list is measuring your data supplier, not your subject.

The supply side, quantified

The confound only matters if enforcement supply actually moves, and it has moved substantially. The Council on Criminal Justice's policing dashboard puts the national sworn officer count at roughly 800,000 at its 2017 peak and roughly 770,000 in 2024. Survey work by the Police Executive Research Forum found staffing among responding agencies down 5.5% between 2020 and 2024, with 2025 recovering slightly but still below 2019–2020 levels.

The direction has now reversed. PERF's 2026 survey reports that participating agencies hired about 17% more officers in 2025 than in 2024, and roughly 40% more than in 2021, with resignations below their 2021 peak. Aggregate staffing among responders has climbed back to a bit above 2022 levels. Recent coverage on the SpotCrime blog has tracked the accompanying shift toward faster hiring pipelines and revised academy training.

Now put the two halves together. Officer-initiated offense counts fell hard between 2019 and 2024 while staffing contracted. Staffing is now expanding. Absent any change in underlying behavior, the arithmetic expectation is that officer-initiated offense counts rise over the next several reporting years β€” in some cities substantially β€” while victim-reported offenses continue whatever trajectory they were already on.

And that trajectory is downward. The Real-Time Crime Index, drawing on 590 agencies covering 119.4 million people, reports the first half of 2026 down 5.9% in violent crime and 10.1% in property crime year over year, with murder down 17.3%, robbery down 14.3%, burglary down 15.1%, and motor vehicle theft down 18.7%. That is consistent with the longer FBI-based picture at USAFacts, which put the 2024 violent crime rate at 359 per 100,000 and the property crime rate at 1,760 per 100,000, down 5.4% and 9% respectively year over year and down 49.1% overall since 2001.

The scenario to design for

A city fills three academy classes. Total recorded incidents on a given block rise 20% year over year, driven entirely by drug, weapons, and disorderly conduct entries. Victim-reported categories on that block are flat or falling. An unweighted safety score reads that block as materially more dangerous. It is not. It is more patrolled.

Legalization is the clean experiment

The cleanest demonstration that these categories are constructed rather than observed is marijuana. Adult use is now legal in 24 states plus the District of Columbia. In each of those jurisdictions, on a specific date, a high-volume offense category dropped to near zero β€” not because behavior changed, but because a legislature changed a definition. National marijuana possession arrests fell from about 200,306 in 2023 to about 187,792 in 2024, and that gradual national figure conceals much sharper discontinuities inside individual states.

For a time series, a statutory change is a structural break: the series before the break and the series after it are measuring different things, and any model fit across the boundary will produce a spurious trend. This is the same class of defect we documented in the LAPD handoff gap and in the drone-as-first-responder disposition shift. The mechanisms differ; the failure mode in your pipeline is identical.

The NIBRS wrinkle

One technical detail catches people who work across the 2021 transition. Under the old Summary Reporting System, drug offenses entered the national data essentially only as arrests β€” there was no reported-offense channel for them. Under NIBRS, drug and narcotic violations are Group A offenses, reported as incidents in their own right with associated properties and arrestee segments.

The consequence is that an apparent increase in drug offenses across the SRS-to-NIBRS boundary can be entirely an artifact of the reporting standard adding a channel that did not previously exist. If you are joining pre-2021 and post-2021 data on drug categories without accounting for this, the join is invalid. The same applies to any local feed that migrated its records system in that window.

What to do about it

Six practices, in rough order of how much they buy you.

1. Tag every normalized category with an initiation channel. Add a field β€” initiation with values like victim_reported, third_party, and officer_initiated β€” at the taxonomy layer, not at the query layer. It costs a lookup table and it makes every downstream consumer capable of asking the right question.

2. Split the series before you chart it.Publish victim-reported and officer-initiated counts as separate lines. A combined total is not wrong, but it is uninterpretable when the two components can move in opposite directions for unrelated reasons. Practical tooling for this kind of splitting and reporting β€” pandas, SQL, automated CompStat-style outputs β€” is well covered in Andrew Wheeler's Python data science guide for crime analysts.

3. Weight officer-initiated categories near zero in any place-descriptive score. This is the position we take in the SpotScore methodology. A drug possession arrest tells you something real about a location, but what it tells you is entangled with deployment decisions in a way that a burglary report is not. If you cannot separate the two contributions, the honest weight is small.

4. Anchor on the low-elasticity series. When a block-level or city-level count moves, check whether murder, motor vehicle theft, and residential burglary moved with it. If the anchors are flat and the total is up, you are almost certainly looking at an enforcement shift. This is the same anchoring logic that makes murder the reference series in national trend work.

5. Monitor a discretion ratio per agency. Compute the officer-initiated share of monthly incidents for each source feed and alert on step changes. A ratio that jumps from 0.18 to 0.27 in one quarter is a policy event β€” a new chief, a reconstituted street crimes unit, a consent decree, a legalization date, an academy graduation β€” and it should annotate every chart drawn from that feed afterward. Change-point detection on the ratio is more informative than change-point detection on the raw count.

6. Keep the model away from the arithmetic.If a language model is summarizing your data for a user, the enforcement confound is exactly the sort of nuance it will smooth over β€” and its stated confidence will not reflect the uncertainty. Gio Circo's analysis of LLM classifier calibration using NEISS injury data found that model confidence scores and token probabilities are systematically overconfident relative to actual accuracy. Deterministic code should own the counts and the channel split; the model should own the sentence.

The causal caveat, stated plainly

Everything above is an argument about measurement. It is not an argument that proactive enforcement is ineffective, and the distinction is important enough to be explicit.

There is a substantial body of research suggesting that concentrated, place-based enforcement reduces crime in treated areas, with contested findings on displacement and diffusion. If that literature is broadly right, then a contraction in officer-initiated activity is not purely a recording artifact β€” it may also be a partial cause of subsequent changes in victim-reported crime, in whichever direction the causal effect actually runs. The measurement problem and the causal question are entangled in the real world even though they are separable in a schema.

But notice that the entanglement does not rescue the naive reading. If officer-initiated counts fall because patrols contracted, and violent crime subsequently rose because patrols contracted, then the officer-initiated series moved opposite to the risk it is being used as a proxy for. A safety score that reads fewer drug arrests as a safer block would be wrong in the most consequential direction available. Both stories β€” pure artifact and partial cause β€” argue for taking these categories out of the score, not for reading them naively.

The honest position is that we do not have per-category enforcement elasticities, we do not have a clean national panel of patrol intensity, and the 2020–2024 period bundled staffing changes, legalization, pandemic disruption, the NIBRS transition, and a historic crime decline into the same five years. Anyone offering a confident decomposition of that is overselling. What we can say with confidence is narrower and still useful: the categories that require an officer to exist behave differently from the categories that do not, that difference is large, and it is visible in the arrest data.

The short version

A crime feed is not a single instrument. It is two instruments taped together β€” one pointed at victimization, one pointed at enforcement β€” reporting into the same table with no column telling you which is which. Most of the time the mixture is stable enough that nobody notices. It is not stable now: staffing fell 5.5% and is rebounding, arrests are 25% below 2019, two dozen states have deleted a high-volume offense category by statute, and the reporting standard changed underneath all of it.

Add the column. Split the series. Anchor on murder. Keep the officer-initiated categories out of anything that claims to describe a place. None of this is expensive, and it is the difference between a product that tracks conditions on the ground and one that tracks the size of the department's next academy class.

Access Address-Level Crime Data

Real-time incidents Β· SpotScoreβ„’ safety ratings Β· 36-month trends Β· 22,000+ US cities. Normalized and verified β€” because raw data isn't enough.