Chicago · 2015–2025 · July 2026

A Tale of Two AI Minds:
Heat, Crime, and Chicago

A day 10°F hotter than others in the same month shows about +5.6% more reported violent crime. Old Freakonomics-style question. Modern open data. Max model starts the build. Smart cheaper model finishes the audit.

Portfolio analysis · public data · no API keys · offline after first download

+5.6%violent crime per +10°F
~1.5×violent vs property response
+6.1%cumulative hot-spell effect
2.76Mcrimes · 4,018 days
01

TL;DR

The short version

Take a day in Chicago that is 10°F hotter than other days in the same month of the same year. Hold day-of-week, holidays, rain, and snow fixed. Reported violent crime is about 5.6% higher on that hotter day. Property crime moves too, but less (~3.7%). The gap is real.

This is not “crime is high in July.” July is already stripped out. The model uses leftover day-to-day weather noise inside each month.

Battery drives the violent result. Hot days in a row compound rather than steal crime from tomorrow. Future temperature does not predict today’s crime. Dropping 2020, the unrest week, or any single year barely moves the number.

What this is not: proof of pure aggression, a neighborhood map, or a claim about unreported crime. It is a clean reduced-form fact about reported crime and heat at the city-day level.

02

The question that data can answer

Most things that move crime (poverty, policing, demographics) change slowly. Weather jumps overnight. That makes heat one of the cleaner daily shocks a city gets.

Why this topic

Heat and crime is an old Freakonomics-style question: take a behavioral claim people argue about, find a daily shock that is nearly as good as random, and measure it. The topic is not new. The fun is updating it with free city crime APIs, free reanalysis weather, a full 2015–2025 panel, month×year fixed effects, and a written stress-test suite. Not nostalgia. A clean same-month identification story on modern data.

Three bets, written before the models:

H1 · Level

Hotter days have more reported crime, all else equal.

CONFIRMED

H2 · Composition

Violent crime responds more, in percent terms, than property crime.

CONFIRMED · ~1.5×

H3 · Shape

The response is not a straight line forever. Extreme heat might flatten or invert.

FLATTENS ~80°F · NO INVERTED-U
Plain English

Imagine two Tuesdays in the same July. One is 78°F. One is 88°F. Same month, same year, same day-of-week structure, same rain flag. The second Tuesday is the experiment. Weather chose it. We count what the crime desk logged.

Why this design

Routine-activity theory says heat puts more people outdoors and in contact. Heat-aggression research says heat shortens fuses. Both predict more interpersonal crime on hot days. Property crime should respond less once people (and their stuff) are already outside. City-day data cannot fully separate those channels. The gap between crime types is the tell that something behavioral is happening, not just “more tickets on nice weather.”

03

The data, and why you can trust it

2,757,885 reported crimes across 4,018 days (2015–2025), joined to ERA5 daily weather at the Loop. No API keys. Offline after first download.

Pull counts, not 2.8M rowsChicago portal SoQL aggregates server-side: day × crime type × count(*).
Verify the aggregationIndependent count(*) checks for spot months, annual totals, and THEFT probes.
Lock the cacheSHA-256 manifest so silent file drift fails loud.
Own the weather objectOne Loop grid point for a whole city. Measurement error attenuates. Estimates are conservative on that margin.
Technical

Crime: City of Chicago Data Portal ijzp-q8t2 (CPD CLEAR). Weather: Open-Meteo ERA5, America/Chicago, °F. Primary predictor: daily maximum temperature. Violent = battery, assault, robbery, homicide, criminal sexual assault (both portal spellings). Property = theft, burglary, motor vehicle theft, criminal damage. Window ends 2025-12-31 to avoid recent reporting lag.

/* The decision that makes the project runnable on a laptop */
$select = date_trunc_ymd(date) AS day, primary_type, count(*) AS n
$group  = day, primary_type
/* ~125k rows instead of ~2.8M incident rows */
Data detail that mattered

Criminal sexual assault appears under two spellings for years (rename transition). Both labels merge into violent. Each source incident has one primary type, so this is not double counting. “Total” crime is ~26% “other” (narcotics, deceptive practice, etc.). Headline modeling focuses on violent and property.

04

What the raw data shows (and what it does not prove)

Every summer, crime and temperature rise together. That chart is not the result. It is why you need fixed effects.

Two-panel chart: monthly mean violent and property crime 2015-2025 on top, monthly mean daily high temperature below. Both series crest every summer and trough every winter. Early 2020 is shaded.
Raw co-movement. Every summer both crime series crest with temperature. COVID and the late-May 2020 unrest week leave scars. Property crime also climbs after 2022 with the motor-vehicle-theft wave. None of that is identified as a heat effect yet.
Plain English

If you only look at “hot months have more crime,” you mixed heat with daylight, school calendars, tourism, and summer life. The model throws those away and keeps day-to-day weather surprises inside each month. Within a typical month, daily highs still swing about 9°F on average. Plenty of identifying variation.

Four scatter plots of residual log crime against residual temperature for battery, robbery, theft, and assault after removing month-year means. Battery and assault show the strongest upward residual slopes.
After stripping the calendar. Within-month residual heat still tracks residual crime, strongest for battery and assault, weaker for theft. This is the variation the regressions use.
05

The model, stated cleanly

One workhorse specification. Several stress tests. No black-box “the computer found heat” story.

Technical

For daily count Ct and daily high Tt:

log C_t = β · (T_t / 10) + α_month×year + δ_day-of-week + γ'W_t + ε_t

W holds holiday, rain (≥1 mm), and snow-day flags. Report 100×(eβ − 1) as percent change per +10°F. Errors: Newey-West HAC with data-selected lag 9 (lag 14 as sensitivity). Count-model check: Poisson QMLE and Negative Binomial. Violent-vs-property gap: stacked regression with day-clustered errors.

M1 → M2

Calendar controls, then full month×year fixed effects. Tightening barely moves β. A pure seasonal artifact would collapse here.

STABLE

M3 counts

Poisson and NegBin match log-OLS. The result is not an artifact of logging counts.

MATCHES

M4 shape

Temperature bins plus AIC-selected spline (df=6). Formal kink search for violent crime at 80°F.

PLATEAU

M5 – M6

Stacked gap test. Hottest-decile dummy and summer-only ≥90°F vs 70s. Not “just summer.”

HOLDS
Plain English

We ask whether unusually hot days, compared with other days in the same month, show more crime reports. Then whether violence moves more than property, whether the curve flattens at the top, and whether the result dies when we change the estimator or drop weird years.

06

Results

Same-day heat moves reported crime. Violence moves more. Hot spells add memory. Extreme heat plateaus rather than reverses.

Horizontal bar chart of percent effects per plus 10 Fahrenheit for total, violent, property, battery, and theft with 95 percent confidence intervals.
Same-day effects per +10°F daily max under month×year fixed effects. Battery is the engine of the violent result. Theft looks like the property story.
Outcome% per +10°F95% CINote
Total crime+4.0%[+3.5, +4.4]Mixture of motives
Violent+5.6%[+5.1, +6.0]Public headline
Property+3.7%[+3.2, +4.2]~1.5× smaller than violent
Battery only+6.0%[+5.5, +6.5]59% of violent volume
Theft only+3.4%[+2.8, +4.0]Property core
Hot spell (cumulative, L*=1)+6.1%[+5.6, +6.6]Same-day + lag memory
Bar chart of distributed lag coefficients for violent crime at lag 0 and lag 1 selected by AIC.
Hot spells compound. AIC selects one lag for violence. Yesterday’s heat still adds. Cumulative effect about +6.1% per +10°F spell. Not pure displacement from tomorrow.
Two panels of AIC versus break temperature for violent and property crime piecewise linear models.
Shape, chosen by AIC. Violent crime: break near 80°F. Slope about +6.0% per +10°F below the break, near flat above. Property crime: break near 50°F (cold suppression, then shallow). No inverted-U in Chicago’s observed range.
Summer check

Within June–August only, days ≥90°F still show higher violent crime than same-summer 70s days (~+2.5%). Property shows nothing in that cut. Heat is not a costume for “it was July.”

07

Stress tests

If a finding only lives in one specification, it is not a finding.

TestWhat we askedWhat happened
F1 Leads Does tomorrow’s heat predict today’s crime? Same-day +5.1%. Lead-1 +0.6%. Lead-2 +0.3%. PASS
F2 Components Is “violent” battery-heavy behavior? Battery +6.0%. Assault +5.7%. Robbery +3.5%. Theft +3.4%. PASS
F3 Heat index Does “feels like” overturn dry-bulb? Per residual SD, tmax / apparent / heat-index sit near 4.6–5.1%. Tmax stays public. PASS
F4 Inference Do HAC-9, HAC-14, week clusters, HC1 disagree? All print +5.6% for violence. PASS
F5 Permutation Shuffle temperature inside each month 300 times. Observed residual slope is extreme. p ≈ 0. PASS
F6 Influence Drop unrest, New Year’s, top 1% days, all of 2020. Every shift under 0.1 percentage points. PASS
F7 FDR Many outcomes, honest multiple testing. Benjamini-Hochberg still clears total, violent, property, battery. PASS
Histogram of residual slope estimates under within-month temperature shuffles, with observed slope marked far in the right tail.
Permutation null. Within each month×year cell, shuffle temperature and refit. The real residual association sits outside the null cloud.
Bar chart of residual correlations between log violent crime and temperature for same day and three future leads. Same day is largest.
Leads are small. Residual correlation peaks on the same day and falls for future heat. Reverse causality would have looked different.
08

What we tuned (and what we refused to fake)

Bandwidth, lag order, and shape flexibility. Chosen with Newey-West theory and AIC/BIC. Not a random forest sold as causality.

HAC lag = 9

Newey-West rule of thumb for T≈4,018, confirmed on an SE grid. Original lag 14 was slightly conservative. Point estimate unchanged.

Distributed lag L* = 1

AIC and BIC agree for violence. Same-day + lag memory → cumulative +6.1%.

Kink at 80°F

Piecewise linear search by AIC. Rises, then flattens. Not an inverted-U.

Per-1SD treatment race

Mean temperature looks “bigger” per +10°F because means move less. On a within-month SD scale, tmax and tmean are close. Public treatment stays daily max.

Leave-one-year-out forest plot and rolling three-year window path for the violent heat coefficient.
Stability. Leave-one-year-out violent effects sit between about +5.4% and +5.7%. No year is load-bearing.
# Workhorse, once the panel exists
log(violent) ~ tmax10 + C(ym) + C(dow) + holiday + rain + snow_day
# cov: HAC(maxlags=9)
# report: 100 * expm1(coef['tmax10'])
09

A tale of two AI minds

July 2026. One human directing. Two models. Not a contest. A handoff with a real constraint.

Two July projects

This heat study is one of two AI-built Freakonomics-style portfolio pieces worked in July 2026. The other is Do UFOs Follow the News Cycle? (618k reports, media and misidentification, Claude-led). Neither write-up was invented on publish day. Both were planned as a pair: public data, behavioral framing, full attack surface, hand-built HTML articles.

The combo, stated flat

Max model starts. Smart cheaper model finishes. Claude Fable 5 on Claude Max opens a blank folder and builds the first complete analysis. Grok 4.5 on SuperGrok takes that draft and does the second-pass work: validation depth, residual EDA, bandwidth and lag selection, falsification, shared modules, and this article. Neither is “the better model.” Different tools, different tasks.

Why the handoff is practical, not ideology

Claude Max is excellent for greenfield velocity, but it rate-limits on a roughly five-hour cycle. Mid-project, that clock stops a long session cold. SuperGrok / Grok 4.5 is available on the X side (poster access, paid). When Max is limited, or when the job is audit-and-attack rather than blank-page build, the work continues on Grok instead of waiting. Capacity and task fit, not a scoreboard of which model is smarter.

ToolTask fitWhat that pass shipped
Claude Fable 5
Max · 2 heavy prompts
Greenfield velocity. Super-prompt to working portfolio notebook. Until the 5-hour Max limit hits. SoQL aggregation, weather join, full M1–M6 ladder, first figures, sensitivity table, first findings. Scaffold that made the project real in one sitting.
Grok 4.5
SuperGrok · rest of the work
Keep building when Max is limited. Audit, selection, stress tests, packaging. Data QA and manifest, residual EDA, Model Lab (HAC, DLAG, kink, LOYO), F1–F7 falsification, panel.py / models.py, integration, this page.
Process

Plan first. Run second. Human decides what ships. Notebook runs offline top to bottom. Model lab is one command. This page is hand-written HTML, not a notebook dump. For the code path a strong DS can follow in the repo, see the project README.

10

Limits, stated plainly

11

Reproduce

Story vs code

This page is the write-up. The Python path (panel, FE fit, lab, falsification) lives in the GitHub folder with a full follow-along README: heat-and-crime/README.md. Fastest check: examples/headline_fit.py.

cd heat-and-crime
pip install -r requirements.txt
python examples/headline_fit.py      # +5.6% path
python examples/leads_placebo.py     # F1 leads
python run_model_lab.py              # full G0-G7 + F1-F7
python run_eda_diagnostics.py

With data/ present, runs are offline. Public APIs only. Manifest hashes lock the cache.

12

Scoreboard

What to remember

  • +10°F daily max → +5.6% violent crime the same day, inside month×year cells. Total +4.0%. Property +3.7%.
  • Violence responds ~1.5× property. Battery +6.0% is the core.
  • Hot spells compound to about +6.1% cumulative (AIC lag order 1).
  • Shape: rises then flattens near 80°F. No inverted-U in this range.
  • Not just summer: within Jun–Aug, 90°F+ days still elevate violence vs 70s.
  • Stress tests: leads small, permutation extreme, influence cuts tiny, FDR clear.
  • Build: Claude Fable scaffolded the notebook in two prompts (July 2026). Grok 4.5 did the audit, lab, stress tests, and this article. Max starts. Smart cheap finishes. Different tools, different tasks.

Back to the top

Home