Leah Gerber

Reading research without the vocabulary problem.

Every term here has come up in the actual papers behind this project, and each one is anchored to a real number from them. Learn it against something concrete and it stays. The green boxes say what that term did in your own reading.

The numbers you will see

These five carry most of the weight in a results section.

Correlation r

How much two things move together, on a scale from minus one to one. Zero means no relationship at all. One means perfectly locked together.

To feel how big it is, square it. That gives you roughly the share of the variation it accounts for. An r of .56 is about 31%. An r of .07 is about 0.5%.

In your reading. Fit and job satisfaction is .56. Fit and job performance is .07. Squaring them is the clearest way to see why that gap matters. One explains a third of what is going on. The other explains almost nothing.

Standard deviation SD

The typical distance from the average. If average sleep is seven hours with a standard deviation of one hour, most people land between six and eight.

Small standard deviation means everyone is bunched together. Large means they are spread out.

Standard error, and the square root of n

This is the one behind the argument about the 7.85 ratio, so it is worth a minute.

Individual scores bounce around a lot. Averages of many scores bounce around much less, because averaging cancels out the extremes. How much less? Divide by the square root of how many things you averaged together.

Average 4 things, the spread halves, because the square root of 4 is 2. Average 100 things, the spread shrinks tenfold. Average 78 things and it shrinks by about 8.8.

In your reading. The famous 7.85 to 1 ratio compared one person’s raw scores against the spread of a set of averages. Those averages each pooled 78 observations, and the square root of 78 is 8.83. That may be most of what the ratio is measuring. Contested, and not settled.

p-value p

The probability of seeing a result at least this big if nothing real were going on. A p of .03 means 3%. By convention anything under .05 gets called significant, which is a bad word for it.

Two things it does not tell you, and the second catches nearly everybody. It says nothing about how big the effect is. And it is not the probability that there is nothing there. It only says how surprising your data would look if there were nothing there. A tiny meaningless effect measured in enough people produces a beautiful p-value.

Red flag. A paper that reports p-values and never reports an effect size is telling you it found something without telling you how much. Wagner 2004 did exactly that.

Effect size

How big, as opposed to whether. The p-value answers is there anything here. The effect size answers is it worth caring about.

They come in several currencies. Correlation is one. So is Cohen’s d, and partial eta-squared, which is the share of variation one factor accounts for.

In your reading. The study that found no sleep effect on insight reported a partial eta-squared of .005. That is 0.5%. It is the numerical way of saying nothing happened.

Confidence interval, and credibility interval CI

A confidence interval is a range for how uncertain one estimate is. If it crosses zero, then no effect at all is still on the table.

A credibility interval is a different animal that gets confused with it. When many studies are pooled, it describes how much the true effect varies from setting to setting rather than how uncertain the average is. A wide one means the effect really does differ by context.

In your reading. Fit and performance had a credibility interval running from -.04 to .44. Because it includes zero, that means in some workplaces the effect is nothing.

N and k

N is how many people. Small n sometimes means how many in one group.

k is how many studies, and you only see it when someone is pooling studies together.

In your reading. Fit and satisfaction rested on k of 65 studies and N of nearly 43,000 people. Fit and performance rested on k of 22 and N of under 6,000. The weak result is also the thinly evidenced one.

How a study is built

The design decides what the study is allowed to conclude. This matters more than the numbers.

Cross-sectional

One snapshot. Many people, measured once. Cheap, common, and it cannot tell you what caused what or how anything changed.

Longitudinal

The same people followed over time. Better for direction, because you can at least see which thing moved first.

Between-person, and within-person

Between-person compares different people to each other. Within-person compares one person to themselves at another moment.

They answer different questions and are constantly confused for one another. This confusion is the subject of your whole project.

In your reading. The fit literature’s own 2023 review said the research is mostly cross-sectional and between-person, and that only a handful of studies were within-person. That single sentence is why nobody can say fit causes anything.

Randomized

The researcher decides by chance who gets what. This is the thing that licenses saying one caused the other. Without it you have a pattern, not a cause.

Post-hoc sorting

Dividing people into groups afterward, based on what happened to them rather than what you assigned. It produces tables that look identical to a real experiment and proves far less.

In your reading. Both sleep-stage studies sorted people by whichever stage they happened to drift into. So they cannot separate that stage causing insight from insight-prone people being the ones who drift there. The giveaway is uneven group sizes, 49, 24 and 14. Nobody designs that.

Preregistration

Writing down what you predict, and how you will test it, before you collect the data. It stops a researcher quietly changing the question once they see the results.

A preregistered result carries more weight than one that was not. And a finding a paper stumbles into afterward is exploratory, which means suggestive rather than established, even inside a preregistered study.

In your reading. Lowe 2025 preregistered a prediction about N1 sleep and it failed. They then found something at N2 that they had not predicted. The first result is strong evidence. The second is a lead.

Replication

Someone running it again. Direct means the same design as closely as possible. Conceptual means the same idea tested a different way.

A finding that has never been directly replicated is a finding that has been agreed with, not confirmed.

Null result

The study found nothing. This is real information and it is much harder to publish, which means the published record leans toward things having worked.

Meta-analysis

Pooling many studies into one estimate. Stronger than any single study, and it inherits every flaw in the studies it pools. If they all measured the thing badly, the meta-analysis measures it badly with more decimal places.

Corrected correlation

A correlation adjusted upward to account for imperfect measurement. Always larger than the raw number.

Not cheating, but it means you cannot compare a corrected number from one paper to an uncorrected one from another.

The traps

These are the ways a result can look real and not be.

Confound

Something else that could explain the result. The two things you measured both move, and a third thing you did not measure is moving them.

In your reading. Your own example. Sleep drops and ideas arrive during mania. Sleep and ideas are correlated, and the state is moving both. Neither one is causing the other.

Common-method bias, or same-source bias

When one person supplies both measurements, their answers agree with each other more than reality warrants. People are consistent about their own opinions.

In your reading. The clearest example you have. Fit and performance is .22 when the same person rates both, and -.02 when different people do. Nearly all of the apparent effect tracked with who was doing the rating.

Moderator

Something that changes the size of an effect. Not a cause of the outcome. A dial on how strongly the cause works.

In your reading. Age moderates the sleep and insight effect. It showed up in young adults and was much weaker in older ones. The effect did not vanish, it depends.

Underpowered

Too few people to detect the effect reliably. If it finds nothing, that is not evidence the effect is absent, only that this study could not see it. If it finds something, the size is probably overstated.

In your reading. 22 people per group, in the study behind every sleep-on-it article you have read.

Who paid for it

Near the end of most papers there is a funding statement and a conflict-of-interest declaration. Read them. They tell you who sponsored the work and whether the authors have a stake in the answer.

Funding does not make a finding false, and independent work can still be wrong. What a sponsor buys is usually subtler than faked results. It shapes which questions get asked, which comparisons get run, and which findings get submitted at all. So the check is not who paid, it is whether the design quietly favors the payer — the comparison chosen, the outcome measured, the follow-up length.

In your reading. TSA's evidence for behavior detection was reviewed by GAO, an auditor with no stake in the program surviving. The agency’s own report to Congress was written to release withheld funding — and still conceded the evidence was missing, which is exactly why it is the most quotable document in the file. An admission against interest is the strongest kind.

Regression to the mean

Extremes drift back toward average on their own. Measure people on a terrible day and they will look better next time whatever you do to them.

It is the reason a self-experiment started because things were going badly will tend to show improvement.

The words specific to your project

Nomothetic, and idiographic

Nomothetic means studying what is true of people in general. Idiographic means studying one person in depth.

Almost all of psychology is the first. Your project is arguing for more of the second.

Ergodicity

A borrowed term from physics. A process is ergodic if what is true across a group at one moment is also true of each individual over time.

The conditions it requires are strict and human behaviour often fails them. The argument is not that group findings are always wrong about you. It is that you cannot assume they hold for you without checking, and almost nobody checks.

Careful. This is a mathematical point about what you are permitted to infer. It is not a measurement of how much people vary. Those get mixed up constantly and the distinction is what keeps your version honest.

Intraclass correlation ICC

How much of the total variation sits between people rather than within them. An ICC of .7 means 70% of the variation is differences between people and 30% is people differing from themselves.

Experience sampling, and EMA

Pinging people several times a day during ordinary life and asking what is happening right now. It beats asking someone to remember last month.

EMA stands for ecological momentary assessment, which is the same thing in a lab coat.

N-of-1

An experiment with one participant, run properly. Not a diary. It means repeating the switch between conditions several times, ideally without knowing which one you are in, so the pattern cannot be wishful thinking.

Heritability

The share of the differences among people, in one population at one time, that tracks with genetic differences among them.

It is not the share of you that came from your genes. There is no such quantity. And the figure moves when the environment changes, without any genes changing.

Getting hold of the papers

Open access, paywalled, preprint, green copy

Open access means free and legal. Paywalled means the journal wants money, usually $30 to $50 for one article.

A preprint is the version before peer review, posted free. Useful, and it may differ from the final paper, so check.

A green copy is the author’s own accepted version, posted on a university site. Legal, free, same content, different formatting.

DOI, PMC, and PMID

A DOI is a permanent address for a paper, starting with 10-point-something. It is the thing to search when a link dies.

PMC and PMID are identifiers in the free US National Library of Medicine database. A PMC number means a free full text exists.

Where to look, in order

Unpaywall first, which you already have. Then PubMed Central, then Semantic Scholar, then the author’s own university page. Then a public library card, which gets you into databases from home.

Then email the author. They own the right to send you a copy and most are pleased to be asked.