Leah Gerber

Resources

What I’m reading.

A working library, with a note on what each one actually shows and where it stops. These come out of five verified research passes run in August 2026. Every pass also failed to cover things, and those gaps are recorded alongside the findings rather than smoothed over. Sources marked free can be read without a university login.

Within-person variation

  • Fleeson (2001), Journal of Personality and Social Psychology author copy free Shows: sampling people five times a day for weeks, one person’s variation in personality states rivals the differences between people’s averages. Stops: small student samples, self-report, and the result flips when measured against standard questionnaires. His own abstract also shows person averages are almost perfectly stable, which most citations leave out.
  • Fisher, Medaglia & Jeronimus (2018), PNAS free · PMC6142277 Shows: group and individual estimates diverge in spread and in correlations. Stops: the famous 7.85-to-1 ratio it is cited for is contested — the between-person figures appear to be spreads of averages, not of people, and I do not cite that number.
  • Adolf & Fried (2019) and the authors’ reply, PNAS free · PMC6452692, PMC6452707 Shows: the argument narrowing in real time. The reply concedes group-to-individual generalizability lies on a continuum. Stops: it is an exchange of letters, not a settled result.
  • Anvari et al. (2025), Personality and Social Psychology Bulletin free · PMC12044207 Shows: what people believe about their own variability and what repeated measurement finds correlate at only .14 to .30. Stops: seven days of sampling, so nothing about longer horizons.
  • Molenaar (2004), Measurement paywalled Shows: the mathematical argument that group statistics do not apply to individuals by default. Stops: it is an entailment about inference, not a measurement of how much anyone varies. I have only reached the abstract so far.

Sleep and insight

  • Wagner et al. (2004), Nature author copy free · counts at PMC3902672 Shows: 59% of a sleep group found a hidden rule against about 23% awake. Stops: 22 people per cell, one omnibus test at p = .03, no effect size or confidence interval anywhere in the paper.
  • Schonauer et al. (2018), Frontiers in Human Neuroscience free · PMC5834438 Shows: no sleep effect on classic insight puzzles, effect sizes near zero. Stops: a three-hour daytime sleep window rather than a full night.
  • Brodt et al. (2018), Sleep paywalled Shows: time away from a problem helped, and spending it asleep added nothing. The single most relevant paper here. Stops: I have read it through a secondary source and say so until I get the full text.
  • Lacaux et al. (2021), Science Advances free · PMC8654287 Shows: people who drifted into early-stage sleep found the rule far more often. Stops: nobody was assigned to a sleep stage, so it cannot separate the stage causing insight from insight-prone people drifting there.
  • Lowe et al. (2025), PLoS Biology free · PMC12200826 Shows: a preregistered failure to replicate the sleep-onset claim, with an unplanned deeper-sleep effect instead. Stops: the new effect was exploratory and carries the same sorting problem.
  • Cordi & Rasch (2021), Current Opinion in Neurobiology partly free Shows: the sleep-and-memory literature revised downward by one of its own principal authors. Stops: a review, not new data.

Person-environment fit

  • Kristof-Brown, Zimmerman & Johnson (2005), Personnel Psychology free copy online Shows: fit predicts how people feel about work, .44 to .56, and barely predicts what they produce, .07. Stops: correlational throughout, and the effects shrink hard when different people rate each side.
  • Kristof-Brown, Schneider & Su (2023), Personnel Psychology free preprint Shows: the field’s own stocktaking, eighteen years on, revising none of the numbers and calling causal direction unresolved. Stops: only a handful of within-person studies existed to review.
  • Bloom, Moen and colleagues, the STAR trial free · PMC6719311 Shows: a randomized workplace trial that gave employees schedule control. Performance barely moved. Stops: one intervention in one firm, and the outcomes were self-reported.
  • Edwards (1994), Organizational Behavior and Human Decision Processes free green copy Shows: the field’s own methods critique — difference scores confound what they claim to measure. Stops: methodological, so it bounds other findings rather than making its own.

Judging people, official versions

  • National Research Council (2003), The Polygraph and Lie Detection free at nationalacademies.org Shows: the polygraph beats chance on specific incidents and fails for security screening, because base rates defeat it. Also that the states it measures arise without deception. Stops: 2003, and lab-heavy.
  • Bond & DePaulo (2006), Personality and Social Psychology Review paywalled · abstract free Shows: people average 54% at spotting lies in real time, professionals included, and do better with ears than eyes. Stops: lab studies at even base rates, so not an operational estimate.
  • GAO reviews of TSA behavior detection (2010–2017) free at gao.gov Shows: 98% of the sources the agency cited for its behavioral indicators did not provide valid evidence. Stops: reviews of the evidence base, not of whether the program ever caught anyone.
  • Lenzenweger (2015), Journal of Personality Assessment free author copy Shows: the famous 1948 OSS assessment structure does not reproduce with modern methods. Stops: a reanalysis of the same published matrix, and it says nothing either way about whether the assessments predicted field performance. Nothing verified does.

Judgment and sequence

  • Danziger, Levav & Avnaim-Pesso (2011), PNAS free · PMC3084045 Shows: the famous hungry-judges pattern, favorable parole rulings falling across a session. Stops: archival, cases were not randomly ordered, and the reported effect is about fifty times larger than the mechanism proposed to explain it. I report it as contested and do not cite it as causal.
  • Glöckner (2016), Judgment and Decision Making free Shows: the effect-size arithmetic, d of roughly 1.96 against a proposed mechanism whose registered replication sits at 0.04, plus simulations where rational scheduling produces part of the curve. Stops: simulation, without access to the raw data.
  • Chen, Moskowitz & Shue (2016), Quarterly Journal of Economics free NBER working paper Shows: position in a sequence moves real decisions even when order is randomized and the merits are held fixed — loan officers, umpires, asylum judges. Stops: the mechanism is the gambler’s fallacy, not fatigue, and popular coverage merges the two.

Behind each entry sits a claim-by-claim audit trail with verbatim quotes, so anything I say from these can be traced to a page. If a note above overstates what a source shows, tell me and I will correct it in public.