Publication bias is the tendency for a study to get published because of how its results turned out rather than how well it was done. Big, clean, surprising effects get into journals. Null results, messy results and failed replications stay on someone's hard drive. So the published literature is not a sample of the research that was carried out. It is a sample of the research that looked good.

That matters to you the moment you put "studies show" in a deck. On September 2, 2026, Psychological Science retracted one of the most quoted findings in the productivity canon: Dan Ariely and Klaus Wertenbroch's 2002 paper "Procrastination, Deadlines, and Performance: Self-Control by Precommitment." Retraction Watch reported that the data had been tampered with or fabricated. The paper had been cited close to 1,000 times in Web of Science, and by a broader count reported by Ynetnews, more than 2,100 times. It sat in the literature for 24 years. Publication bias is a large part of why nobody looked.

Why the literature tilts toward exciting results

The mechanism runs in one direction at every step.

A researcher finishes a study. If the result is null, submitting it is a poor use of a year, because journals take it less often: papers with statistically significant results are roughly three times more likely to be published than papers with null results, per the summary of the evidence on Wikipedia. So the null goes in a drawer. The psychologist Robert Rosenthal named this the file drawer problem in 1979; the statistician Theodore Sterling had described the pattern as early as 1959.

The significant result goes out, and reviewers weigh it partly on how novel and how clean it is. FORRT's glossary puts it plainly: publishability gets judged on the outcome rather than on methodological quality. Dickersin and Min's 1993 definition is "the failure to publish results based on the direction or strength of the study findings."

Then the citation machinery amplifies. A striking finding gets quoted in textbooks, talks and business books. Each citation makes the next one easier, because citing a famous study feels safe. And the thing that would correct the record, a replication that finds nothing, is itself the kind of paper that is hard to publish.

The important consequence is the one people skip: a high citation count measures how quotable and publishable a result was, not how true it is. Those two things come apart, and when they do, nothing in the system automatically notices.

The deadlines study, and 24 years of nobody checking

Picture the version of this you have probably sat through. A marketer is building a project-planning deck and adds a slide: "Research shows evenly spaced deadlines nearly double output, so we're setting weekly checkpoints rather than one delivery date." The citation is the 2002 paper. She has not read it. She has read three books that mention it, which felt like enough.

The finding she is repeating came from two small experiments: 99 MIT professionals writing papers, and 60 students proofreading texts, 20 per condition. In the proofreading study, the evenly spaced group averaged 136.1 corrections against 71.1 for the group with a single final deadline. Nearly double, as advertised.

A replication by Kyle Hyndman and Alberto Bisin, published in the same journal on July 15, 2026, found the deadline structure had negligible effects on performance. That failure prompted the research-fraud blog Data Colada, run by Uri Simonsohn, Joe Simmons and Leif Nelson, to examine the original files. Their August 31 post lists what they found.

d = 2.5
Effect size in the proofreading study
Larger than the average height difference between men and women (about 1.8)
18 of 20
Final-deadline participants with a 'twin'
Identical error counts across all three tasks, with ID numbers exactly 10 apart. No such pattern in the other conditions.
11.7% vs 85%
Rounded self-reported times
Original data versus the replication. Real people round; random number generators do not.
13 of 49
Grades altered in the other study
Changed after the 2001 submission, 12 of 13 in the direction favouring the hypothesis, while preserving the section mean of 85.67

Data Colada's verdict on the proofreading experiment was that the data "were severely tampered with or fabricated," and that they were "unable to generate a benign explanation for all anomalies presented here." Wertenbroch, who says he never personally accessed the raw data, requested the retraction on July 23, 2026. Ariely, in a statement posted on August 7, acknowledged "serious anomalies" and said the documentary record and his memory after two decades could not resolve them. In a video statement he said "the deeper analysis of our data show that these data cannot be trusted." The retraction notice gives the journal's reason as being that the authenticity and completeness of the dataset cannot be confirmed.

Note what publication bias did and did not do here. It did not fabricate anything. It made the fabricated version publishable, cite-worthy and famous, and it made the boring correction unrewarding for 24 years. An effect of d = 2.5 from 20 people per cell should have read as implausible on sight. Instead it read as a great result.

One nuance worth keeping: Wertenbroch distinguishes between people's demand for precommitment deadlines, which he says does replicate, and claims about how well those deadlines work, which do not.

The same tilt in medicine

The antidepressant reboxetine looked effective in the published trials. A 2010 meta-analysis that included unpublished negative trials sponsored by its manufacturer, Pfizer, found it ineffective. The drawer contained the answer the whole time.

A review of the Cochrane Library found positive findings were 27% more likely to be included in efficacy studies, and that in safety studies, results showing no adverse effects were 78% more likely to appear than findings of actual harm. And as of 1998, no acupuncture trial conducted in China or the former USSR had ever reported a negative result, which tells you about the publication process rather than about acupuncture.

This is why major journals began requiring clinical trials to be registered in advance in 2004. If the study is on the record before the results exist, a disappointing result cannot quietly vanish.

Not the same thing as p-hacking, and not the same thing as fraud

Who does it What it distorts
Publication bias The system: authors, reviewers, editors, citers Which true studies you get to see. Each paper may be fine; the set is skewed.
P-hacking One researcher, inside one dataset A single result, by trying analyses until one clears significance
Fabrication One person, deliberately The data itself. No statistical correction fixes invented numbers.

The 2002 case involved the third and was sustained by the first. Keeping them separate matters, because the fixes differ: preregistration and registries address publication bias, while fabrication is only caught by people going back to the raw files, which is what happened here.

Three questions before you cite a famous finding

Is the effect too clean for the sample size? A large, tidy difference from 60 people is not stronger evidence than a modest one from 6,000. It is usually a reason to look harder. This is the check that would have flagged d = 2.5 in 2002.

Has anyone repeated it, and what happened? Search the finding's name plus "replication" and "failed to replicate" before you search for supporting quotes. If the only evidence is the original paper plus people quoting the original paper, you have one study, not a literature.

Is the citation count doing the arguing? "Cited 2,000 times" is a claim about popularity. If you would not accept it as evidence from a colleague, do not accept it from a footnote.

And when a finding you have used gets retracted, the useful move is not embarrassment. It is going back to the decisions you made on its strength and asking whether they still make sense without it. Weekly checkpoints may still suit your team. They just no longer have a study behind them.