🚨 CityCrimeMap

Predictive layer

Where is crime likely to happen next?

CityCrimeMap publishes a machine-learning-based predicted risk overlay on every city map. It highlights the small share of city blocks where reported incidents are most likely to concentrate in the near future.

On San Francisco the model's top 10% of city cells captures roughly 59% of reported incidents — a PAI of 5.92 against a random baseline of 1.0.

How the model works

  1. Grid the city into small square cells. For dense cities we use 300 m cells; for smaller cities 200–250 m so the map has visible contrast. A typical US city ends up with somewhere between 5,000 and 15,000 grid cells.
  2. Compute a smoothed density of past incidents per cell. This is a kernel density estimate (KDE) — a standard statistical technique for turning a scatter of points into a continuous surface. We rescale it (log-transform, normalize) so the machine-learning model can actually learn from it.
  3. Train a random forest model to predict the incident count per cell in a held-out future window, using the KDE feature plus each cell's coordinates and lagged counts. Random forests are ensemble models built from hundreds of decision trees; they're a well-established baseline for spatial prediction because they handle non-linearities and sparse features gracefully.
  4. Show only the top 10% of cells by predicted risk. Cells with zero predicted risk are dropped entirely so the GeoJSON payload stays small.

The methodology is a machine-learning adaptation of Risk Terrain Modeling, published in Wheeler & Steenbeek (2021), Journal of Quantitative Criminology. Our full implementation is open source at tidycop-hotspots (MIT).

What the numbers mean

We publish the Predictive Accuracy Index (PAI) for every city we cover. PAI = (share of incidents captured) ÷ (share of area used). A PAI of 1.0 means the model is no better than picking cells at random; 2.0 means it's twice as good; 3.0+ is a strong lift. It's the standard metric in the predictive-policing academic literature.

Concretely, if the top 10% of San Francisco cells has a PAI of 4.6, that means those cells together contain about 46% of the reported incidents in the held-out test window — almost 5× what you'd get by guessing.

CityPAI @ 10%VerdictTrain setGrid
San Francisco, CA5.92excellent lift1001 rows32 cells · 300 m
Boston, MA4.88excellent lift833 rows43 cells · 200 m
Seattle, WA4.87excellent lift965 rows45 cells · 300 m
Providence, RI4.74excellent lift450 rows19 cells · 200 m
Gainesville, FL4.46excellent lift349 rows22 cells · 250 m
Washington, DC, DC4.45excellent lift667 rows38 cells · 300 m
Minneapolis, MN4.38excellent lift883 rows44 cells · 300 m
Hartford, CT4.32excellent lift859 rows28 cells · 250 m
Pittsburgh, PA4.00very strong lift1094 rows39 cells · 300 m
Denver, CO3.66very strong lift1002 rows57 cells · 300 m
Cleveland, OH3.44very strong lift1006 rows51 cells · 300 m
Baltimore, MD3.20very strong lift1002 rows57 cells · 300 m
Los Angeles, CA3.06very strong lift1002 rows64 cells · 500 m
Indianapolis, IN2.52very strong lift951 rows74 cells · 300 m
Detroit, MI2.51very strong lift1001 rows67 cells · 300 m
Kansas City, MO2.10strong lift1001 rows40 cells · 300 m
Dallas, TX1.92strong lift1214 rows63 cells · 400 m
Houston, TX1.79strong lift1017 rows78 cells · 300 m
Rochester, NY1.68strong lift191 rows15 cells · 250 m
Chicago, IL1.51strong lift1002 rows79 cells · 300 m
Cincinnati, OH1.09modest lift over random207 rows16 cells · 250 m
New Orleans, LAPredictive layer not enabled for this city yet.

Sorted by PAI, best model first. Training sets are short right now (a few hundred rows per city) because we're only ingesting the last ~45 days. Longer training histories are on the roadmap and will improve every one of these numbers.

What the model does not do

We're deliberately narrow about what the predictive layer is for. It's a research-backed tool for public awareness, not a policing product.

Compared to other crime maps

Most public crime maps show only historical dots. That's useful for a look-back but doesn't tell you where risk concentrates going forward. To our knowledge, CityCrimeMap is the only free, open-source, city-agnostic crime map that ships a peer-reviewed predictive layer.

See predictions for your city

Have questions about the model? Read the full methodology page, the FAQ, or file an issue on GitHub.

Free email alerts

Get your city crime alerts in your inbox

SpotCrime sends a daily email with crimes reported near any address you choose. Free, no app required, unsubscribe anytime.

Sign up for your city alerts on SpotCrime →

Opens spotcrime.com in a new tab.