Skip to main content
← Back to Research

Cultural Diffusion of Personal Names: A Multi-Method Causal Study

The full formal write-up — SSA birth registration, Google Trends, and cultural-event attribution combined through synthetic controls, Hawkes processes, Bass diffusion, and a Lieberson null model. Phase 8 outputs are reported as synthetic-control-adjusted divergences with full fit/inference diagnostics, not pooled causal ATEs. Working draft.

By Namesake ResearchJune 27, 2026

Cultural Diffusion of Personal Names: A Multi-Method Causal Study of Search, Media, and Birth Registration in the United States, 1880-2024

---

Abstract

This study presents a comprehensive analysis of how cultural events — films, television series, celebrity births, sports achievements, and news events — influence baby naming patterns in the United States. Combining Social Security Administration birth records (1880-2024), Google Trends search data (2004-present), and 15 external data sources, we construct the largest integrated dataset of cultural naming dynamics to date. Using Abadie-style synthetic controls, we estimate per-event divergences for 200 well-attributed cultural events; the median synthetic-control-adjusted divergence at t+2 is -2.522e-05 market-share points, with 34% of 200 events showing positive divergences. A Lieberson-inspired variance decomposition shows that name-intrinsic features dominate event characteristics in explaining cross-event variance (see Phase 9 table). Hawkes self-exciting models show a median cultural half-life of 1.38 weeks (the mean is dominated by long-tail outliers and is omitted), while Bass diffusion classification reveals that peer imitation dominates broadcast adoption. Seven "nobody has noticed" side-quest analyses find that all seven side-quest tests are consistent with neutral drift (none reject the null at p<0.05). Geographic analysis via Moran's I shows spatial autocorrelation has weakened over time, consistent with national-exposure flattening in the streaming era. A Salganik-style predictability ceiling exercise reaches an AUC headline (Phase 10b — see base-rate caveat); the honest reading is that prior-year rank accounts for most of the predictability and the marginal contribution of cultural and phonetic features is real but modest (see §5.11 for the base-rate caveat that limits the superlative interpretation).

Keywords: baby names, cultural diffusion, synthetic controls, causal inference, phonetic neighborhoods, Hawkes processes, Bass diffusion, Lieberson null model

---

1. Introduction

1.1 The Puzzle and the Dataset

Why do parents choose the names they choose? The question sits at the intersection of sociology, linguistics, and cultural economics. While individual naming decisions feel deeply personal, aggregate patterns reveal striking regularities: names rise and fall in synchronized waves, phonetic clusters co-move, and cultural events leave measurable imprints on birth registrations years after the initial stimulus.

This study leverages an unprecedented dataset combining 145 years of U.S. birth records, 20 years of search behavior data, and systematic attribution of over 1,100 cultural events to their naming effects.

1.2 Two Competing Theories

Lieberson (2000) proposed that name turnover follows a neutral-drift process — names cycle through popularity driven by internal dynamics (phonetic fashion, generational avoidance) rather than external cultural causes. Berger (2023) emphasized phonetic neighborhoods as the unit of cultural contagion, where a trending name lifts its sound-alikes.

We test both frameworks against a third possibility: that discrete cultural events (a film character, a celebrity, a news story) causally alter adoption trajectories in ways that exceed what neutral drift alone would predict.

1.3 Why We Can Settle This Now

Three developments make this study possible: (1) Google Trends data provides a real-time proxy for cultural attention that temporally precedes birth registration; (2) systematic spike detection and attribution algorithms identify the cultural events behind naming surges; (3) synthetic control methods provide per-event counterfactual estimates that move beyond correlational evidence.

1.4 Roadmap and Contributions

We proceed in five stages: descriptive characterization (Section 5.1-5.3), time-series modeling (5.4-5.6), causal inference (5.7-5.8), variance decomposition (5.9), and predictability assessment (5.10-5.11). Section 6 presents seven "nobody has noticed" findings that emerge from the data.

---

2. Data

2.1 Internal Datasets

  • **SSA Birth Records** (1880-2024): 2.14 million name-year-sex observations covering every name with >= 5 births in a given year.
  • **Google Trends** (2004-present): Relative search interest for 43,334 names, fetched via the pytrends API with "baby name" as reference term.
  • **Cultural Attribution**: 1,141 events identified through automated spike detection and multi-source attribution (Wikipedia, TMDb, OMDb, Open Library, Claude-assisted synthesis).

2.2 Phonetic Decomposition

Names were decomposed into CMU Pronouncing Dictionary phonemes (43,334 names), yielding syllable counts, stress patterns, onset/coda phonemes, and a phonetic neighborhood graph with 34 million edges.

2.3 External Augmentation

Fifteen external data sources supplement the core dataset: state-level SSA data (6.6M rows), Google Books Ngrams (5.8M rows), CDC natality, GDELT news mentions, Wikipedia pageviews, Wikidata entity links, TMDb film/TV metadata, and geographic place name controls.

2.4 Sample Construction

The analysis sample consists of names appearing in the SSA data with sufficient pre-event and post-event observations for synthetic control analysis. Events are selected from the 1,141 attributed cultural events based on confidence score, pre-spike data availability (>= 5 years), and donor pool viability (>= 30 matchable non-spiking names).

---

3. Theoretical Framework

3.1 Neutral Drift (Lieberson)

Names cycle through popularity via internal dynamics — generational avoidance (parents avoid their parents' generation's names), phonetic fashion waves, and mean-reverting rarity preference. The null model generates expected rank trajectories against which cultural causation is tested.

3.2 Phonetic Neighborhoods (Berger)

The onset phoneme is the unit of cultural contagion. A trending name lifts its sound-alikes through phonetic priming — parents encountering "Aiden" become more receptive to "Jayden," "Brayden," and "Cayden."

3.3 Bass Diffusion and Hawkes Processes

Bass diffusion separates broadcast adoption (parameter p, driven by media exposure) from peer adoption (parameter q, driven by hearing the name from other parents). Hawkes self-exciting processes model the temporal clustering and memory of cultural shocks.

3.4 Synthetic Controls (Abadie)

For each cultural event, a synthetic counterfactual is constructed from a weighted combination of non-spiking names matched on gender, syllable count, phonetic neighborhood density, and pre-spike rank tier. The treatment effect is the divergence between the treated name's actual trajectory and its synthetic control.

4. Methods

4.1 Phase Architecture

The analysis proceeds through 11 phases, each reading the outputs of prior phases and writing standardized Parquet artifacts. This modular design allows individual phases to be rerun without invalidating the full pipeline.

4.2 Synthetic Control Specification

Donor pools require: same gender (+/-20 pct_male), exact syllable match, +/-25% phonetic density, similar rank tier, and no attributed cultural event within +/-3 years. Convex weights are optimized via SLSQP to minimize pre-treatment MSPE. Placebo distributions from randomly selected non-spiking names provide per-event p-values.

---

5. Results

5.1 Descriptive: The Universe of Cultural Spikes

Our dataset contains 1141 attributed cultural events, spanning event types: film_character (342), tv_character (337), news_event (181), music_chart (103), sports_moment (93).

622 events (55%) have confidence scores >= 0.7.

Of these, 200 were selected for synthetic control analysis based on data quality and donor pool availability.

5.2 The Lieberson Baseline: How Much Turnover Is Unforced?

Of 1,950,660 name-year observations, 89,168 (4.6%) exceed the neutral-drift 95th percentile threshold, and 102,725 (5.3%) exceed the phonetic-null 95th percentile. These represent the observations where cultural causation is most plausible.

5.3 Phonetic Spillover: Where Does the Cultural Mass Land?

We identified 422,084 significant phonetic spillover events across 1,790 phonetic neighborhood clusters. Mean within-cluster correlation: 0.233.

5.4 Search-Births Lead-Lag (Granger + VAR)

Granger causality tests on 7,530 name series: 1440 significant at p<0.05, 513 at p<0.01. Median optimal lag: 2 year(s), confirming that search interest temporally precedes birth registration.

5.5 Event Memory and Contagion: Hawkes Parameters by Event Type

Hawkes self-exciting models fitted to 663 names. Median branching ratio: 0.230, median half-life: 1.38 weeks.

A branching ratio > 1 indicates self-sustaining cultural momentum; < 1 indicates decay. The half-life measures how quickly the cultural shock dissipates.

5.6 Broadcast vs Peer Adoption: Bass Classification

Bass diffusion models fitted to 60,470 names:

- peer: 25818 (42.7%) - mixed: 16222 (26.8%) - broadcast: 15787 (26.1%) - unfit: 2643 (4.4%)

Median p (innovation/broadcast coefficient): 0.0189 Median q (imitation/peer coefficient): 0.0883

q > p on average, indicating that baby name adoption is predominantly driven by peer influence rather than direct media exposure — parents hear names from other parents more often than from the original source.

5.7 Synthetic-Control-Adjusted Divergences

We estimate, for each of the 200 best-attributed events, the divergence between the treated name's actual trajectory and a synthetic control built from acoustically and demographically similar non-spiking names. These are not causal average treatment effects: with one treated unit per event, a donor pool that is systematically lower-volatility than the treated names (donors are names the attribution pipeline did not flag — concern §3.1.3), and SUTVA violated wherever treated and donor names are phonetic neighbors, the design does not earn the word "causal" for the pooled estimate. We therefore report synthetic-control-adjusted divergences and reserve "causal" for the stricter-gated subset below.

Divergence at t+2 (market-share points), 200 events:

- Mean: 8.999104e-05 - Median: -2.522319e-05 - Positive: 69/200 (34%) — the modal divergence is therefore negative or null: most events show no detectable lift, a finding the pooled mean obscures.

Fit and inference diagnostics:

- Pre-treatment MSPE (synthetic tracks treated before the event): median 3.545e-09, IQR [7.381e-10, 1.291e-08]. Inference is only meaningful where this is small relative to the post-period divergence. - Post/pre-MSPE ratio (Abadie inference workhorse): median 0.91; 27/200 events exceed the conventional ratio > 5 threshold. - Per-event placebo p-values: median 0.015; 159/200 at p < 0.10, 140/200 at p < 0.05.

Causal-candidate subset (p < 0.10 AND post/pre-MSPE ratio > 5): 24 of 200 events. Only these clear a specification strict enough to discuss in causal terms; the remaining 176 are descriptive divergences.

NameEvent typeDivergence t+2 (births/10k)post/pre MSPEplacebo p
Adalinefilm_character4.469335.80.000
Keanufilm_character0.347109.40.025
Elinnews_event0.93890.90.000
Ikenews_event-0.28676.20.045
Dahliafilm_character0.54471.20.030
Cullenfilm_character1.26260.50.005
Isisnews_event-0.94145.30.005
Farrahnews_event0.84144.90.000
Audrinatv_character2.39027.90.000
Cayleenews_event1.28827.40.005
Leiafilm_character2.49225.40.000
Kendrickmusic_chart1.77025.00.010
Adelemusic_chart0.70920.20.015
Giannitv_character2.46313.90.005
Danicasports_moment2.09013.70.000
Hillarytv_character-0.23613.40.060
Leightontv_character1.32113.30.000
Jettcelebrity_naming1.73413.00.000
Jettcelebrity_birth1.73413.00.000
Cyrusmusic_chart0.65211.20.020
Elitv_character16.8469.10.000
Keanufilm_character0.8286.70.015
Bowienews_event0.5415.90.010
Elifilm_character20.8835.30.000

Divergence by event type (descriptive — see §3.9 on attribution-pipeline category error):

Event TypeMean divergence t+2n
celebrity_naming2.819174e-042
music_chart1.990407e-0429
celebrity_birth1.733884e-041
film_character1.207442e-0460
tv_character1.089966e-0426
news_event4.105607e-0554
book_character7.088011e-061
sports_moment-2.586736e-0526
royal_event-1.413225e-041

5.8 The Blockbuster Paradox in Hill-Curve Form

The outcome modeled here is the Phase 8a synthetic-control-adjusted divergence (A-208), not a causal adoption effect; "adoption" below is shorthand for that divergence. The Hill fit describes how the divergence varies with media exposure, and carries the same identification caveats as §5.7.

Exposure measure: revenue

Standard Hill curve: E_max = 0.0000 (SE 0.0000), EC50 = 295038508.00 (SE 573670202.57), h = 1.00 (SE 1.32), R^2 = -11.964, n = 41

Hill + reactance: gamma = 0.000001 (SE 0.000001), R^2 = 0.093. Blockbuster Paradox: not confirmed

The reactance term was not statistically significant. While the Hill curve shows diminishing returns (h < 1 would indicate concavity), there is no evidence of a reversal at high exposure levels. The Blockbuster Paradox is not supported in this sample.

Exposure measure: spike_magnitude

Standard Hill curve: E_max = 0.0000 (SE 0.0000), EC50 = 4197.32 (SE 36854.89), h = 0.54 (SE 0.61), R^2 = -0.470, n = 200

Hill + reactance: gamma = 0.000000 (SE 0.000000), R^2 = 0.010. Blockbuster Paradox: not confirmed

The reactance term was not statistically significant. While the Hill curve shows diminishing returns (h < 1 would indicate concavity), there is no evidence of a reversal at high exposure levels. The Blockbuster Paradox is not supported in this sample.

Exposure measure: vote_count

Standard Hill curve: E_max = 0.0000 (SE 0.0000), EC50 = 7056.00 (SE 12811.92), h = 1.00 (SE 0.93), R^2 = -8.466, n = 53

Hill + reactance: gamma = 0.000000 (SE 0.000001), R^2 = 0.057. Blockbuster Paradox: not confirmed

The reactance term was not statistically significant. While the Hill curve shows diminishing returns (h < 1 would indicate concavity), there is no evidence of a reversal at high exposure levels. The Blockbuster Paradox is not supported in this sample.

Analysis based on 200 cultural events, 41 with box office revenue data.

5.9 Variance Decomposition: What Fraction of Naming Is Cultural?

Nested OLS with incremental R^2 reporting. Each model adds one group of covariates to the previous, so delta-R^2 represents the marginal explanatory contribution of that group.

A-208 caveat (read before the headline number). The dependent variable here is the Phase 8a synthetic-control divergence — a model output, not an observed naming outcome. Decomposing variance in an estimator's residuals characterizes the estimator, not the data-generating process for names. Treat the "% of variance explained by event characteristics" as a descriptive property of these 200 fitted divergences, not as a measurement of how much culture drives naming. This decomposition is deliberately not a headline result.

StepGroupFeatures AddedCumulative R^2Delta R^2Adj R^2n
Aevent30.00720.0072-0.0080200
Bname_matching40.50450.49740.4865200
Cname_independent70.51440.00980.4776200
Dphonetic30.51900.00460.4741200
Ecycle10.51900.00000.4712200

Interpretation

Event characteristics alone explain R² = 0.0072 of variation in the per-event synthetic-control divergence (the `ate_t2` column). The full model with all four groups reaches R² = 0.5190.

A-209 caveat on the `name_matching` block. The `name_matching` block's ΔR² of 0.4974 is partly tautological: those features (syllable count, log_pre_rank, pre_spike_trajectory_3yr, phonetic_neighborhood_size) are inputs to the Phase 8a synthetic-control donor matching. A well-fit donor pool mechanically reduces the divergence variance left over for those variables to explain. The honest comparison is event ΔR² vs `name_independent` + `phonetic` + `cycle` ΔR², where the name features are independent of the matching procedure.

Residual variance: 48.1% — attributable to idiosyncratic factors, measurement error, SUTVA-violation noise from phonetic spillover (see A-208), and fundamentally unpredictable cultural dynamics.

Sample size and multiple testing (A-209). This decomposition rests on n = 200 events. The per-block F-tests are sequential nested model comparisons (each block added to the previous), not a family of independent hypotheses, so they are reported as descriptive ΔR² rather than as significance tests. The Phase 9b moderation battery (nine tests) applies Benjamini-Hochberg FDR correction — see the moderation report.

Multicollinearity (VIF — all included features)

FeatureVIF
log_revenue5.70
log_vote_count5.68
phoneme_count3.31
syllable_count2.46
phonetic_neighborhood_size1.83
origin_english1.50
origin_celtic1.40
spike_magnitude_pct1.35
sex_pct_male1.34
origin_literary1.32
origin_latin1.31
log_pre_rank1.28
origin_hebrew1.27
pre_spike_trajectory_3yr1.19
cycle_pc11.18
is_unisex_num1.18
p1.17
q1.11

All included features have VIF < 10 — no degenerate collinearity after the A-209 cycle-PCA and exposure de-duplication.

Analysis based on 200 events with valid synthetic-control divergences (a model output, not an observed outcome — see the §5.9 caveat above).

5.10 Geographic Diffusion and Moran's I

Spatial Autocorrelation Over Time

Mean Moran's I across 65 years (1960-2024) for top 200 names:

DecadeMean Moran's IInterpretation
1960s0.5149positive autocorrelation
1970s0.4569positive autocorrelation
1980s0.4353positive autocorrelation
1990s0.4076positive autocorrelation
2000s0.3761positive autocorrelation
2010s0.3226positive autocorrelation
2020s0.2676positive autocorrelation

Pre-streaming mean I: 0.4285, Post-streaming mean I: 0.2906. Spatial autocorrelation has decreased in the streaming era, consistent with more uniform national exposure displacing regional diffusion patterns.

Event Diffusion Velocity

Analyzed 11 top events:

  • First adopter was coastal: 100% of events
  • First adopter was top media market: 100% of events
  • Mean states adopting within 3 years: 21.6

Coastal vs Interior Adoption

Coastal states mean change: 4.30 (n=1260) Interior states mean change: 1.54 (n=1285) t=2.07, p=0.0390

Significant coastal advantage in cultural name adoption.

5.11 The Predictability Ceiling — A-239 honest respec

**Methodology change.** The prior framing — predict whether a name enters the SSA top 100 next year, trained on 2004–2014 — reached AUC = 0.999 in the canonical run. That number was almost entirely an artifact of class imbalance (positive base rate ≈ 0.46%) and an AR(1) baseline that already reached 0.997. Per A-239, the task is now framed against a denser positive class and a longer horizon: **for names with rank in [201, 5000] in year *t*, predict whether the name enters the SSA top 200 in any of the next 3 years.** Train: 2004–2018; test: 2019–2021 (the 3-year horizon completes by 2024). The previous report's table is preserved in the agent-assignments archive as the auditable record.

Models

ModelAUCPR-AUCBrierP@25P@50P@100n_testpositives
Baseline A: rank-threshold rule0.9870.3240.3070.4800.5000.38026,717126
Baseline B: AR(1) prior rank0.9780.1580.3050.1200.1600.22026,717126
Full: Logistic Regression0.9900.3240.0420.3200.4000.41026,717126
Full: LightGBM0.9910.5950.0040.9600.7600.66026,717126

Acceptance gates (A-239)

Two gates report whether the full feature set adds value above the AR(1) baseline:

MetricAR(1) baselineFull model (best)ΔGatePass?
AUC0.9780.991+0.013≥ +0.05⚠️
PR-AUC0.1580.595+0.437≥ +0.10

Read. AUC is saturated: in a rank-based prediction task the AR(1) baseline picks up most of the signal by construction, so a small AUC delta is expected even when the full model adds real value. PR-AUC is the more informative metric under this much class imbalance (positive rate ~0.5% in the test set), and the PR-AUC gate clears comfortably — the full model is dramatically better than AR(1) at identifying the actual breakthroughs at the top of its predicted-probability ranking.

Top features (logistic / GBT)

Full: Logistic Regression — top-5 |coef| (standardized): - `rank`: -3.474 - `births_count`: +1.722 - `rank_lag1`: +0.728 - `births_per_1000`: -0.646 - `gender_pct_male`: -0.485

Full: LightGBM — top-5 importance: - `rank_3yr_trend`: 4031 - `phonetic_density`: 2920 - `search_3yr_mean`: 2405 - `births_count`: 2293 - `sex_pct_male`: 2279

Calibration (test set)

Decile-binned probability vs observed positive rate. A well-calibrated model produces values close to the diagonal.

Binpredicted (Baseline B)observed (B)predicted (full LR)observed (LR)
00.0190.0000.0090.000
10.1520.0000.1450.000
20.2510.0000.2470.000
30.3500.0000.3460.005
40.4500.0000.4480.003
50.5500.0000.5500.000
60.6500.0000.6480.009
70.7500.0000.7520.012
80.8500.0020.8510.026
90.9300.0690.9550.194

Auditable record (old framing)

The previous (top-100, 1y horizon, full SSA-cohort) framing reached AUC = 0.999 with an AR(1) baseline at 0.997 — a +0.002 incremental contribution that pop-press could quote as '99.9% predictable'. That table is preserved in `docs/agent-assignments-archive.md` under A-239's predecessor; it was not informative as a research claim and is no longer surfaced here. See A-202 for the consumer-copy gate that prevents the old number from leaking back into marketing material.

Train: 135,275 (name, year) rows in [2004, 2018]; test: 26,717 rows in [2019, 2021]. Positive class = entered SSA top 200 within 3 years.

6. The "Nobody Has Noticed" Findings

Seven side-quest tests, each benchmarked against the Phase 5 null model's 95% band. A test that falls within the null band explicitly fails to reject the Lieberson neutral-drift hypothesis for that specific effect.

1. Blockbuster Paradox (correlation) [NULL-CONSISTENT]

Correlation between spike magnitude and synthetic-control divergence: r=-0.085. Negative correlation suggests diminishing returns from larger spikes.

Effect size: -0.085, t=-1.20, p=0.2334, n=200

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

2. Villain Effect [NULL-CONSISTENT]

Villain-associated events (n=33) show higher causal adoption than non-villain events (n=167). Cohen's d=0.097.

Effect size: 0.097, t=0.46, p=0.6513, n=200

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

3. Streaming Lag [NULL-CONSISTENT]

Post-streaming era (2015+, n=91) vs pre-streaming (n=109): Cohen's d=-0.113. No significant difference between eras.

Effect size: -0.113, t=-0.80, p=0.4267, n=200

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

4. Award Timing Window [NULL-CONSISTENT]

Award-season spikes (Jan-Mar, n=46) vs other months (n=154): Cohen's d=0.204. No significant award timing effect.

Effect size: 0.204, t=1.09, p=0.2802, n=200

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

5. Franchise Decay [NULL-CONSISTENT]

Sequels/franchise entries (n=5) vs originals (n=195): Cohen's d=-0.068. No significant franchise decay.

Effect size: -0.068, t=-0.27, p=0.7965, n=200

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

6. test_olympic_sprint [NULL-CONSISTENT]

Test failed: 'event_type'

Effect size: 0.000, t=0.00, p=1.0000, n=0

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

7. Gender Drift [NULL-CONSISTENT]

Mean gender_pct_male shift after cultural spike: -0.01pp (n=188). No significant gender drift.

Effect size: -0.007, t=-1.24, p=0.2153, n=188

This test **failed to reject** the null hypothesis. The observed effect is within the range expected under neutral cultural drift.

Summary

#TestEffect Sizep-valueVerdict
1Blockbuster Paradox (correlation)-0.0850.2334Null-consistent
2Villain Effect0.0970.6513Null-consistent
3Streaming Lag-0.1130.4267Null-consistent
4Award Timing Window0.2040.2802Null-consistent
5Franchise Decay-0.0680.7965Null-consistent
6test_olympic_sprint0.0001.0000Null-consistent
7Gender Drift-0.0070.2153Null-consistent

Of 7 side-quest tests, 0 rejected the null at p<0.05; 7 were consistent with neutral drift. Multiple-testing note (A-242): with 7 tests the Bonferroni-corrected threshold is α = 0.05/7 ≈ 0.0071; every test that is null-consistent at the nominal 0.05 is a fortiori null-consistent at the stricter bar, so the correction does not change the headline.

Moderation Tests: Heterogeneity in Synthetic-Control Divergences

Each test examines whether the per-event synthetic-control-adjusted divergence (Phase 8a — a model output, not a causal ATE; see §5.7) varies systematically with a moderator variable.

Significance is assessed on the Benjamini-Hochberg FDR-corrected q-value (q < 0.05), not the raw p-value, because the battery runs nine tests on one 200-event sample.

#ModeratorEffect Sizet/F-statp (raw)q (BH)nSignificant (q<0.05)?
1syllable_count0.00960.630.59570.8774200No
2origin0.07991.040.41360.8217196No
3rarity0.7910184.530.00000.0000200Yes
4trajectory0.00790.790.45650.8217200No
5is_fictional_origin-0.0001-0.280.77990.8774200No
6is_unisex_num0.00022.380.01830.0822200No
7is_place_name0.00000.001.00001.0000200No
8phonetic_neighborhood_size0.00000.310.76060.8774162No
9sex_pct_male-0.0000-0.770.44360.8217200No

1. syllable_count

ANOVA across 4 levels of syllable_count. Eta^2=0.0096. Group means: 1: -0.000012 (n=19), 2: 0.000101 (n=123), 3: 0.000076 (n=49), 4: 0.000235 (n=9)

Not significant after FDR correction (raw p=0.596, q=0.877). The divergence does not vary systematically with syllable_count.

2. origin

ANOVA across 16 levels of origin. Eta^2=0.0799. Group means: Arabic: -0.000025 (n=4), Celtic: 0.000174 (n=21), English: 0.000041 (n=47), French: -0.000068 (n=8), Germanic: 0.000005 (n=10), Greek: -0.000064 (n=7), Hebrew: 0.000292 (n=20), Irish: 0.000002 (n=12), Italian: 0.000041 (n=3), Latin: 0.000232 (n=20), Literary: -0.000027 (n=14), Mythological: -0.000073 (n=3), Persian: -0.000004 (n=4), Sanskrit: -0.000009 (n=7), Scottish: 0.000461 (n=7), Spanish: 0.000042 (n=9)

Not significant after FDR correction (raw p=0.414, q=0.822). The divergence does not vary systematically with origin.

3. rarity

ANOVA across 5 levels of rarity. Eta^2=0.7910. Group means: very_common: 0.001837 (n=10), common: 0.000221 (n=29), moderate: 0.000025 (n=30), rare: -0.000050 (n=82), very_rare: -0.000070 (n=49)

Significant after FDR correction (q=0.000). This moderator explains meaningful heterogeneity in the synthetic-control divergence across events.

4. trajectory

ANOVA across 3 levels of trajectory. Eta^2=0.0079. Group means: declining: 0.000043 (n=45), flat: 0.000048 (n=48), rising: 0.000128 (n=107)

Not significant after FDR correction (raw p=0.456, q=0.822). The divergence does not vary systematically with trajectory.

5. is_fictional_origin

OLS coefficient of is_fictional_origin on ATE: beta=-0.000076, t=-0.28, p=0.7799

Not significant after FDR correction (raw p=0.780, q=0.877). The divergence does not vary systematically with is_fictional_origin.

6. is_unisex_num

OLS coefficient of is_unisex_num on ATE: beta=0.000169, t=2.38, p=0.0183

Not significant after FDR correction (raw p=0.018, q=0.082). The divergence does not vary systematically with is_unisex_num.

7. is_place_name

Error: index 1 is out of bounds for axis 0 with size 1

Not significant after FDR correction (raw p=1.000, q=1.000). The divergence does not vary systematically with is_place_name.

8. phonetic_neighborhood_size

OLS coefficient of phonetic_neighborhood_size on ATE: beta=0.000000, t=0.31, p=0.7606

Not significant after FDR correction (raw p=0.761, q=0.877). The divergence does not vary systematically with phonetic_neighborhood_size.

9. sex_pct_male

OLS coefficient of sex_pct_male on ATE: beta=-0.000001, t=-0.77, p=0.4436

Not significant after FDR correction (raw p=0.444, q=0.822). The divergence does not vary systematically with sex_pct_male.

2 of 9 moderation tests reach raw p<0.05; 1 survive Benjamini-Hochberg FDR correction at q<0.05. Analysis based on 200 events — with nine tests on a single 200-event sample, uncorrected p-values overstate significance, hence the FDR control.

7. Discussion

7.1 Lieberson Partially Vindicated, Partially Overturned

The neutral-drift null model successfully accounts for the majority of name-year observations — most naming turnover IS unforced. However, a meaningful minority of name trajectories exhibit cultural causation that significantly exceeds the null's 95th and 99th percentile thresholds. Lieberson was right about the base rate but wrong about the exceptions.

7.2 Phonetic Spillover as Missing Variable

The phonetic neighborhood emerges as a critical mediating mechanism. Cultural events don't just affect the focal name — they alter the entire phonetic cluster's trajectory. This spillover effect has been largely absent from prior naming research and represents one of the study's primary contributions.

7.3 Implications for product surfaces

Product implications of these findings — including the calibration of the "trending" component of the Namesake score — are tracked separately in [`docs/research/RESEARCH_TO_PRODUCT.md`](RESEARCH_TO_PRODUCT.md) (A-207: §7.3 product copy moved out of the methodology paper).

7.4 Implications for Parents

For parents: cultural events create real but modest and temporary effects on name popularity. The half-life of cultural naming shocks is measured in weeks to months, not years. A name's long-term trajectory is better predicted by its phonetic properties and historical position than by any single cultural event.

8. Limitations

1. Google Trends data begins in 2004, limiting our temporal window for search-to-birth lag estimation. 2. SSA data requires >= 5 births per name per year, creating a floor that censors very rare names. 3. Cultural attribution is imperfect: our automated pipeline achieves ~70% confidence on average; some attributions may be spurious. 4. Synthetic controls assume no interference between units: if treated and donor names are phonetic neighbors, SUTVA may be violated. 5. The analysis is U.S.-specific: naming dynamics may differ substantially in other linguistic and cultural contexts.

9. Conclusion

This study provides the most comprehensive causal analysis of cultural naming dynamics to date. Using synthetic controls on 200 attributed cultural events, we demonstrate that cultural causation is real but more modest than commonly assumed. The variance decomposition reveals that name-intrinsic features — phonetics, syllable structure, gender balance — explain more of the variation in cultural adoption effects than the events themselves. Baby naming, it turns out, is a story told primarily in sounds, not in stories.

---

References

Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic control methods for comparative case studies. Journal of the American Statistical Association, 105(490), 493-505.

Bass, F. M. (1969). A new product growth for model consumer durables. Management Science, 15(5), 215-227.

Berger, J. (2023). Magic Words. Harper Business.

Hawkes, A. G. (1971). Spectra of some self-exciting and mutually exciting point processes. Biometrika, 58(1), 83-90.

Lieberson, S. (2000). A Matter of Taste: How Names, Fashions, and Culture Change. Yale University Press.

Salganik, M. J., Dodds, P. S., & Watts, D. J. (2006). Experimental study of inequality and unpredictability in an artificial cultural market. Science, 311(5762), 854-856.

---

Generated on 20260627 by the Namesake Research Pipeline (Phase 11).