The point: nobody's ever checked the boring stuff at this scale
This project never asks what people saw. It asks a question data can actually answer: why do people report?
A UFO report is a human decision. Someone saw something, sat with it (sometimes for 30 years), and one day typed it into a form. The news, the movies, the neighbors, the embarrassment: all of it shapes that decision. So a UFO database is secretly one of the best behavioral datasets on Earth. 618,316 rows of attention, fear, and nerve.
Picture a sky that never changes. Reports would still spike on July 4, jump after a hearing, and surge when SpaceX launches. Before anyone can claim something weird is up there, you have to measure all that human noise first. We did. What's left after the noise is the good part, for both camps.
One promise, kept the whole way: nothing here calls any witness a liar, and nothing here confirms a single craft. Skeptics get receipts. Believers get an honest count of what the boring explanations can and can't absorb. Everyone gets the same numbers.
The four bets we made before running anything
Bet 1: Reports follow the news
Big UAP moments should spike reports within days.
Result: real objects in the sky spike reports 2.2x. Hearings don't create new sightings at all. They make people finally report old ones, at 2x. Half right, in the most interesting way possible.
SPLITBet 2: People, not portals
State report counts should track population, not famous "hotspots."
Result: population explains 88% of which states report. The leftovers are dark skies and Roswell tourism, not mystery.
WONBet 3: Military misidentification
Reports should cluster near military airfields.
Result: they do, 1.39x more than population predicts. But civilian airports score 1.57x. It's an aviation story, not a secret base story.
SPLITBet 4: The trend is plumbing
Long-run rises and falls should track reporting technology, not the sky.
Result: web forms up, old archives dead, Reddit up 1000x. The famous "2024 decline" was two databases going stale.
WONThe data, and why you can trust it
618,316 deduplicated reports from six databases (NUFORC, MUFON, UFOCAT, UPDB, UFO-search, r/UFOs), merged by the open source UFOSINT project. We audited it before drawing a single conclusion.
The audit caught the traps early: 1.85% duplicate overlap between NUFORC and MUFON (so we never pool them for stats), backfilled old MUFON records (so its panel starts in 2006), and a pileup of reports dated "the 1st of the month" from people typing June 1997 as 06/01/1997. Remember that last one. It comes back later disguised as a mystery.
And because "some database from GitHub" deserves suspicion, we checked our copy against a completely independent public scrape of NUFORC:
Two centuries of reports, six charts, zero mysteries needed
Every chart below changed how we built the analysis. None of them needed aliens.
The model: a weather forecast for UFO reports
One honest statistical model, stress-tested against a fancy machine-learning challenger. The boring one won, and we can prove it.
The model learns what a normal Tuesday in March should produce, given the season, the weekday, and the holidays. After a big event we ask: did the next days beat the forecast? Then the acid test: would 2,000 random fake event dates have "beaten" it too? If fake dates work as well as real ones, there's no effect.
Counts get a count model
Daily report counts are jumpy integers. Variance runs 9x the mean. The negative binomial model prices that wildness in. Simpler models pretend it away.
CHOSEN BY THE DATAAll events at once
All 18 events get before and after windows, fitted together, so overlapping events compete instead of double-counting. The "before" windows catch hype that starts early.
NO DOUBLE COUNTINGFake dates as the referee
Every effect gets ranked against thousands of placebo dates, with a correction for testing many events. Cheap, brutal, hard to fool.
PLACEBO-FIRSTTwo databases must agree
NUFORC and MUFON are run by different people with different systems. A result that shows in only one isn't a result.
REPLICATION RULEWe refit the headline model with plain Poisson, the lazy default. Its confidence interval: [2.02, 2.28]. The honest one: [1.39, 3.55]. The lazy model is 7.7x too confident. Run this data through Poisson and you'd "discover" significant effects everywhere. Skeptics, aim your fire at anyone's UFO stats that don't handle this.
We gave machine learning its shot too
A tuned gradient-boosted model got new features (moon phase, eight meteor showers, days since the last Starlink launch) and a training set through 2022. Then it faced 2023 to 2026, unseen. It lost to the plain seasonal model. A hybrid tied it. Here's the punchline: random day-to-day noise caps any model's daily accuracy at a correlation of 0.43, and the calendar alone already gets you 41% of the way there. There's almost nothing left for the machine to find.
The full specification, for the stats people
Negative binomial GLM, log link. Dispersion by Cameron-Trivedi auxiliary regression (alpha 0.25 NUFORC, 0.13 MUFON). Day-of-week dummies, three Fourier harmonics of day of year, year fixed effects, holiday and meteor-shower dummies. Per event: pre [-14,-1], [0,7], [8,30] indicators fitted jointly. Newey-West HAC(14) errors, insensitive to lag 7 or 28. Primary inference: 2,000-draw placebo ranking at 7 and 30 day horizons, Benjamini-Hochberg FDR. Era-matched placebo pools as robustness. Category pooling for power. Leave-one-out stability. Minimum detectable effects reported for null results. One notebook, one command.
Congress can't make you see a UFO. It can make you finally admit you saw one.
Pool the events by type and ask the real question: what actually moves reports?
When something real flies over, reports jump in both databases, every way we slice it. When Congress holds a UAP hearing, new sightings don't budge. And it's not a power problem: anything bigger than a 1.48x hearing effect is ruled out by the data.
But remember the three date fields. The data tracks when sightings happened, when they were filed, and which filings describe old events. Switch from the "happened" clock to the "filed" clock and the hearings light up:
Hearings don't change the sky. They change the embarrassment math. Attention works like an amnesty: it makes reporting feel safe, and people come forward with what they've been sitting on for decades. That's a finding about stigma. It doesn't debunk a single witness, and it doesn't confirm a single craft. Both camps get to keep it.
Three deadpan footnotes. The hyped 2021 ODNI report is the only event that lowered reports (0.6x, the week after its official shrug). Modern movies do nothing (Independence Day: 0.70x). And Storm Area 51, two million RSVPs and all, produced statistical silence.
SpaceX accidentally ran the biggest UFO experiment in history
384 launches on a public schedule. Each one is a test nobody signed up for.
A fresh batch of Starlink satellites crosses the night sky as a bright string of pearls for a few nights after launch. Launch dates are public and set by orbital mechanics, not UFO culture. So here's the test: if misidentification drives reports, "Formation" shaped reports should jump for 0 to 3 nights after each launch. Shapes a satellite train can't imitate (a triangle, a disc) shouldn't move at all.
Formation reports: 1.70x [1.26, 2.30] in launch windows. Triangle control: 1.00. Pre-Starlink placebo years: 0.96. The formation-vs-triangle difference-in-differences, the estimate weather or hype can't fake: 1.45x [1.04, 2.02]. About 1 in 10 of every formation-shaped UFO report in the era sits inside SpaceX's published schedule.
The sky learned a new trick in 2019. Exactly one column of the database responded, exactly on schedule, and never before the launch. It's the cleanest result in the whole project because the experiment repeats 66 times with a control group.
The 20 strangest days in 30 years, named one by one
Flip the whole project around. Let the model pick the days it can't explain. Then hunt down every single one.
| Of the top 20 | Count | What they were |
|---|---|---|
| named causes | 9 | Phoenix Lights. The 2015 SoCal missile test. The 1999 Great Lakes fireball parade. A NASA rocket that painted a glowing cloud over the East Coast (2009). A Chinese booster burning up over three states (2016). A Starlink window. The tail of the 2017 NYT story. The Tinley Park Lights, a famous case the model rediscovered on its own. And, newest addition: the Nov 14, 1997 night, solved during fact-checking of this very page. Story at the bottom. |
| data-entry artifacts | 9 | The "1st of the month" pileups the audit predicted. June 1997 typed as 06/01/1997, dressed up as statistical significance. |
| mixed | 1 | A spy-balloon aftershock landing on a month-rounded date. |
| still unexplained | 1 | Sep 23, 1998: a 5-second fireball at 9:10 pm reported from Seattle to Victoria BC, plus a Texas cluster the same evening. One bright object, logged before the internet kept fireball records. Nobody has named it. |
#1 on the board: the Phoenix Lights
March 13, 1997. Two events in one Arizona evening. Around 8 pm, a silent V of lights crossed a 300-mile corridor from Nevada through Phoenix to Tucson. Around 10 pm, a row of bright lights hung over the city. Thousands watched. Wikipedia calls it perhaps the most widely witnessed UFO event in history.
The official accounting: the 10 pm lights were LUU-2 illumination flares dropped by Maryland Air National Guard A-10s over the Barry M. Goldwater Range. The 8 pm V matched five A-10s flying in formation with steady lights. Plenty of witnesses never accepted the aircraft answer, and the case has the best character arc in UFO history: Governor Fife Symington mocked it in 1997 with an aide in an alien costume, then said in 2007 that he saw the V himself and called it "bigger than anything that I've ever seen."
What our data adds: that one night, NUFORC logged 84 reports against 3 expected, past 1-in-a-trillion under the model's null, and the week after ran 5.4x (placebo q = 0.009). The strongest single event in the entire study. It sits on the "named" list because real things were verifiably in the sky that night: flares and aircraft. Whether the 8 pm V was only aircraft is an argument about testimony, and counts can't settle testimony. Which is exactly where this project draws its line.
The strangest days on record turn out to be rockets, fireworks, and paperwork. Plus one orphan night. Skeptics get names and dates. Believers get something rarer: a certified, model-audited shortlist. One night in 30 years beats arguing about 618,316 rows.
Where UFOs live: dark skies, Roswell, and airports
Population explains 88% of which states report. The leftovers have addresses.
And yes, cow country shows up
Ask the data about rural America and it answers loudly. Denser states report fewer UFOs per person: each doubling of population density trims about 6% off the per-capita rate. City light eats the sky. The per-capita champions are places with cows, mountains, and darkness: Montana, Vermont, Maine, West Virginia, Alaska. The lore says strange things happen near remote ranches. The data says remote ranches are where the sky is actually visible. The cows are innocent bystanders with great seats.
Now the loaded question. Do reports cluster near military bases?
Under the strictest symmetric treatment (drop the biggest metro pileups from both lists), it's 1.57x airports vs 1.39x bases. Same order, narrower gap, non-overlapping intervals. If proximity to civilian aviation predicts reports at least as well as proximity to Area 51's mythology, the simple reading is misidentified aircraft, not leaked secrets.
The "New Jersey" drone wave was 83% not from New Jersey
Winter 2024. The biggest UFO story in a decade. The reports know something the headlines didn't.
Add up the national surge and 83% of it was filed outside NJ, NY, PA, and CT. California, Florida, and Texas led the pack, with zero confirmed drone activity. The rest of the country peaked at 2.5x its own baseline on nothing but coverage. A local stimulus became a national reporting event in eight weeks. Same amnesty mechanism as the hearings, playing out on a map in fast forward.
Drones got better. Sightings got weirder.
Rewind ten years and a drone was an expensive toy. Now the FAA has hundreds of thousands of them registered, the true fleet is larger, and every one is a silent cluster of lights that can hover, dart sideways, and park in the sky. A machine purpose-built to generate UFO reports. Our shape data catches the shift: the modern era belongs to lights, orbs, and formations, the exact profile of a quadcopter at night.
The 2024 wave shows the full loop. Some real drones over New Jersey. Then wall-to-wall coverage. Then a country full of people stepping outside, seeing ordinary aircraft and hobby drones, and filing. Better drones didn't just fill the sky. They gave every light up there a plausible new name.
Then we attacked everything we just told you
Fourteen objections, the strongest we could write, from three angles: hostile statistician, professional debunker, committed believer. Every one tested with computation. Here is the actual prompt that ordered the attack:
"Your database is an unaudited hobby project."
Monthly counts match an independently scraped mirror at r = 0.999. The waves are real waves.
DEFEATED"Your December hearing effect is just a meteor shower."
A hole we found in our own model: the Geminids peak sits inside one key window. Refit with all eight showers controlled. Result barely moved (1.43 to 1.45).
TESTED, SURVIVES"Airports-beat-bases is a geocoding artifact."
Partly fair. City-centroid coordinates flatter airports. We reran both lists under identical rules and rewrote the claim with the narrower numbers.
CONCEDED & FIXED"Your hearing null just means weak stats."
No. The data rules out any hearing effect bigger than 1.48x. An interval, not a shrug.
QUANTIFIED"Your placebo dates come from quieter years."
Redrew them era-matched, within 3 years of each event. All five significant events survive.
DEFEATED"You overstated your own ML result."
True. We'd written "nearly saturated" when the math says 41% of a 0.43 ceiling. The sentence got replaced. The red team applies to us too.
SELF-CORRECTEDThe document polices itself going forward. A claims ledger at the end of the notebook recomputes all 21 headline numbers on every run and fails loudly if the text and the data drift apart. Continuous integration, but for conclusions.
Seven prompts made everything here, including this page
One person with ideas and standards. One AI doing the labor. Here's the receipt, starting with the prompt that started it all:
| # | The prompt, in one line | What came out |
|---|---|---|
| 1 | Build the whole project | Pinned data, cleaning pipeline, event study, geography, first findings |
| 2 | Review your own work like a senior engineer, then keep going | The audit that created the two-database replication rule, plus all the models |
| 3 | Max genius mode: find the wow | Starlink experiment, strangest-days leaderboard, the 120-year-old memory, the drone map |
| 4 | Now attack everything you found | 14 attacks, 3 amendments, the claims ledger |
| 5 | Prepare the final defense | The full Q&A dossier inside the notebook |
| 6 | Build the page that wows readers | The first version of this page |
| 7 | Now make it human | This rewrite: plain words, the TL;DR, and these receipts |
The workflow each time: plan mode first (the AI proposes, the human approves or redirects), then execute. The human's fingerprints are on every fork: which findings deserved the spotlight, when the writing sounded like a robot, and the entire idea of turning the analysis on itself.
Honest footnote: those are the seven big ones. Getting to done took a stack of small follow-up prompts on top. Fix this button. Cut that box. Double-check that 1998 date. The big prompts built the machine; the little ones are where the taste lives. And one of those little fact-checks ended up solving a 28-year-old mystery. See the end of the page.
What did it cost?
What actually got paid: nothing extra. This ran on a Claude Max subscription in summer 2026, which still carried good Fable 5 rates, so the real price was the flat plan plus some patience with rate limits (we hit a few mid-build). The number below answers a different question: what would a report like this cost on the pure pay-per-token API?
We can't see the meter from inside the session, so here's the arithmetic at list prices for the model used (Claude Fable 5: $10 per million input tokens, $50 per million output, cached re-reads about $1 per million). Assume roughly 150 model calls across the 7 prompts, each re-reading about 120K tokens of context with 95% served from cache, plus 0.5 to 1 million tokens of output (text, code, and the model's always-on reasoning). That lands around $70 to $105 total, call it $10 to $15 per prompt. For scale: a fully audited data project, by API meter, for about the price of one statistics textbook.
The verdict, for both ends of the table
If you came in skeptical
- Report volume obeys population (0.83 elasticity), weekends, fireworks (5x), the Moon (4% fewer at full), and SpaceX's launch calendar (1.70x). Human behavior, measured.
- The 20 strangest days in 30 years reduce to rockets, fireworks, and typos, with names and dates attached.
- Hearings and headlines move filing, not seeing. The "2024 decline" and the "national drone invasion" both dissolve under honest accounting.
- A tuned ML model can't beat a plain seasonal forecast. At the daily level there's nothing hidden to find.
If you came in believing
- Nothing here debunks a witness. We tested counts. We never touched testimony.
- Two nights in 30 years have a tight one-place, one-hour signature and no identified cause. They're named, dated, and waiting for you.
- The backfill result is about stigma: people sit on sightings for decades until attention makes reporting feel safe. That's an argument for better reporting channels.
- Everything below daily counts, the level where interesting things would hide, is outside this analysis. We say so plainly.
Limits, stated plainly
Reports aren't objects, and reporters aren't a random sample of witnesses. Coordinates are city-level. The event calendar holds 18 modeling events, so power is honest but finite. Hearings get scheduled because attention is already rising, so post-event bumps bound the effect rather than isolate pure cause. And the "1-in-a-trillion" day scores rank days; they aren't literal betting odds.
Run it yourself
One notebook, one command. Pinned data with checksum verification, five documented public datasets, and the claims ledger standing guard. Every number and figure on this page comes from that single reproducible run.
TL;DR
We took 618,316 UFO reports and checked what really makes people report UFOs. Answer: fireworks, SpaceX, the news, and finally feeling safe enough to talk. One night in 30 years still has no explanation.
The big picture: yes, UFO sightings are correlated with real events in the sky. When something actually flies over (a missile test, a satellite train, a spy balloon), reports jump within days, in both databases, right on schedule. The news moves reports too, but mostly by unlocking old sightings people were sitting on. And after every known cause is subtracted, 1 day out of 11,354 still spikes with nothing to pin it on: Sep 23, 1998. It had company until we fact-checked this page. Correlated where the world supplies a cause. Unexplained where it doesn't.
- July 4 runs 5x a normal day. Fireworks work.
- Starlink launches spike "formation" UFO reports 1.7x for three nights. On SpaceX's published schedule. 1 in 10 formation reports of the era sits inside it.
- Hearings create zero new sightings but double the reports of decades-old ones. Attention is an amnesty. One filed sighting was 120 years old.
- Population explains 88% of which states report. Dark skies explain most of the rest.
- New Mexico is #1 per capita, and its excess is 2.08x concentrated around Roswell. Tourism, in the data.
- Airports beat military bases, 1.57x vs 1.39x. It's an aviation story.
- The "New Jersey" drone wave was 83% filed from other states. Panic travels faster than drones.
- The 20 strangest days in 30 years: 9 named causes, 9 typos, 1 mix, 1 genuine unknown (Sep 23, 1998). The other unknown fell during fact-checking: a Russian rocket stage.
- A tuned ML model lost to a plain seasonal forecast on unseen years. The full moon cuts reports 4%.
- Seven big prompts to one AI built all of it, including this page, plus a pile of small fix-it prompts to finish, for roughly the price of a textbook. Every number is machine-verified on every run.
Two mysteries entered this article. One survived.
Earlier drafts ended with two unexplained nights. Then a fact-check did its job. November 14, 1997 looked like this in the raw records: 38 reports in one evening, 30 from Washington state, 27 at 9 pm, shape "Formation," durations of 30 seconds to 3 minutes, seen from Oregon to British Columbia. Too long for a meteor. Too widespread for flares. That's the signature of a rocket stage breaking up on re-entry. Our launch catalog showed a Russian Proton launch two days earlier, and the satellite observers' decay archive confirms it: the Kupon launch platform came down over the Pacific Northwest that night, and the archive even notes it "resulted in a number of UFO reports." Mystery number one: solved, 28 years later, by a follow-up prompt.
September 23, 1998 survived the same treatment. A 2-to-8-second fireball at 9:10 pm, reported independently from Seattle, Tacoma, Everett, Bothell, and Victoria BC inside the same 20 minutes, with a second cluster over Texas that evening. 43 raw reports, original timestamps intact, filed within a day or two. No holiday, no meteor-shower peak, no launch, no re-entry we can find, no news event, no data-entry artifact. The likely truth is boring: a bright bolide that fell before the internet kept records of falling things. But likely isn't named. Out of 11,354 days, this one is still open. If you want a mystery, start there.