Yes, on average, and by considerably less than the internet implies. The largest meta-analysis pooled 146 randomized trials and found an effect of r = .075. The number has been shrinking for thirty years, not because journaling stopped working, but because the studies got bigger and more honest.
We sell a writing practice with a subscription attached, so read the number above as the number we would most like to have been larger. Everything below is sourced, and every source is linked so you can check that we reported it against ourselves rather than for ourselves.
The effect is real, small, free, and mostly not about the writing you think it is about.
In 2006 Joanne Frattaroli collected every randomized study of experimental disclosure she could find, published and unpublished, and pooled 146 of them covering 10,994 participants. The average effect was r = .075, which is a Cohen's d of about .151, with a 95% confidence interval of .052 to .098 (Frattaroli, Psychological Bulletin, 2006).
Translated out of statistics, using the display Frattaroli herself uses: in a study of 100 writers and 100 controls, about 54 of the writers improve and about 46 of the controls improve. That is the honest picture. Not nothing, and not a life rebuilt.
Two things are worth saying immediately, because the number is smaller than people expect and both directions of misreading are common.
Small is not the same as unimportant. Frattaroli makes the comparison herself: taking a daily aspirin after a heart attack to prevent a second one has an effect size of about r = .034, less than half of this, and no one calls that treatment pointless. When an intervention is free, private, takes an hour total and has no side effect worse than a rough evening, a small effect is worth having.
And the average hides a wide range. The 146 study effects ran from −.291 to .592. Some of those trials found the control group did better. Frattaroli notes that the average includes studies run under poor conditions, including conditions in which the writing appears to have been harmful. Which conditions those are turns out to be the most useful part of the whole literature.
This is the part that explains the gap between what the research says and what you have read on the internet, and it is worth understanding, because the same pattern is playing out right now in the research on AI and inner work.
| Year | Analysis | Studies | Effect |
|---|---|---|---|
| 1986 | Pennebaker and Beall, the original experiment | 1 | Writers visited the student health center at about half the rate of controls |
| 1998 | Smyth, first meta-analysis | 13 | d = .47 (r = .23) |
| 2006 | Frattaroli, largest meta-analysis | 146 | r = .075 (d = .151) |
| 2018 | Pennebaker's own retrospective | 100 plus | d of about .16 |
Nothing dishonest happened here. It is what a maturing literature looks like. The first study, Pennebaker and Beall in 1986, asked students to write for 15 minutes on four consecutive days about the most traumatic experience of their lives, and found they went to the health center about half as often over the following six months (Journal of Abnormal Psychology, 95(3), 274 to 281). It was a striking result and it launched a field.
Writing about it thirty-two years later, Pennebaker described that first study as "horribly underpowered," and said that what followed over the next several years was replications and failures to replicate, with criticism of both the methodology and the theory. He also reports the current figure plainly: across more than 100 studies, an average effect of about .16 in Cohen's d (Pennebaker, Perspectives on Psychological Science, 2018). In the same piece he notes that the inhibition theory that motivated the whole program is one he never found evidence for, and abandoned.
Frattaroli gives a specific reason her number came out below Smyth's: she went looking for unpublished work and found a lot of it. Forty-eight percent of her studies were unpublished, against 23 percent in the 1998 analysis, and unpublished studies had smaller effects. In her own data, published studies averaged r = .095 and unpublished studies averaged r = .054. She also says outright that because underpowered null studies stay in file drawers, the true effect on psychological health may be smaller still than what she estimated.
A founder of the field reporting a shrinking effect in print, and a meta-analyst arguing her own number is probably too high, is what functioning science looks like. Treat any wellness claim that has only ever gotten larger with the opposite of that respect.
This is the actionable part, and it is more specific than the advice normally attached to journaling. From Frattaroli's moderator analysis, effects were larger when:
Stack those and the picture changes materially. Frattaroli identifies eight studies that ran under a majority of the optimal conditions at once. Their average effect was r = .200, close to triple the overall figure. The honest summary is not that journaling barely works. It is that most of the trials, and almost certainly most of the people, are doing a diluted version of it.
Two findings that get left out of every listicle.
It does not change health behaviors. In Frattaroli's data this is the only outcome category that failed to reach significance at all: r = .007 across 10 studies. Smyth found the same thing in 1998 and called it surprising. Writing about your drinking does not reduce your drinking. Insight and behavior change are different machines, and this literature is clear about which one it powers.
And it does not reduce health care use in people who are already ill. A separate 2006 meta-analysis by Alex Harris pooled 30 randomized samples covering 2,294 participants and found expressive writing reduced health care visits in healthy samples, at g = 0.16, but not in samples defined by a medical diagnosis and not in samples screened for psychological criteria (Harris, Journal of Consulting and Clinical Psychology, 2006). The people most often sold journaling as a coping tool are the group where this particular benefit has not been demonstrated.
Writing about an upsetting event reliably makes you feel worse in the moment, and the size of that effect dwarfs the benefits. Smyth's meta-analysis measured the change in distress from just before writing to just after: d = .84, r = .39, significantly larger than any of the health outcomes it found (Smyth, 1998).
Now the part that matters. Smyth also tested whether feeling worse predicted doing better, since that is the folk theory of catharsis and it was the field's own hypothesis. It did not. The amount of short-term distress was unrelated to every long-term outcome measured. Some distress may be required to engage the material at all, but more of it does not buy you more.
So the session that left you shaking was not the productive one. It was just the loud one. If you have been using intensity as your measure of whether the writing is working, this is the finding to keep, and it is the same conclusion the stop rules arrive at from the safety side.
Here is the thing we would rather not print, and the reason it belongs on a page written by a company that sells a journaling practice.
Every study above tests one specific thing: a person writing privately, for 15 to 20 minutes, on three or four days, about a real upsetting event, with nobody reading it and nobody answering it. That is the intervention. That is what has an evidence base.
Frattaroli's inclusion criteria are explicit on this point. Studies in which experimenters gave the writer any oral or written feedback about what they had disclosed were excluded from the meta-analysis, on the stated grounds that such an intervention closely resembles psychotherapy and falls outside the scope. Fourteen studies were dropped on that criterion alone. Her own data adds a second, quieter version of the same point: studies that did not collect the writing afterward had larger effects than studies that did.
Read that against the industry. Journaling apps, guided prompt subscriptions, AI journals and LUX all do the one thing the evidence base defined itself by excluding: they answer you. Somebody, or something, reads what you wrote and responds to it.
The research everyone in this category cites as validation specifically removed the studies that resemble the product being sold.
That does not mean responding is worse. Psychotherapy also involves being answered, and it has a far stronger evidence base than expressive writing does. It means the transfer has not been demonstrated, and anyone citing Pennebaker at you while charging you a subscription, us included, is citing a body of work about unanswered writing to sell you the answered kind. You should know that before you pay anyone, and the free version of the tested protocol is below.
Four days in a row. Twenty minutes each. Alone, at home, in a room by yourself. Paper or a plain text file. Nothing that responds to you.
Write about one specific thing that is genuinely bothering you and that you have not already talked through with several people. Recent beats ancient. Ongoing beats resolved.
Write about the event and about how you feel about it, not just the facts. Do not stop to fix the grammar. Nobody is going to read it, including you, unless you decide otherwise later.
Expect to feel worse for the rest of the evening, and to notice nothing for weeks. The outcomes in these trials were measured a month to several months out, not the next morning.
If it is making you significantly worse for days rather than hours, stop. That is the same line the stop rules draw, and it applies here.
Reading the outcome categories together, a shape emerges that is more useful than a verdict. Effects showed up in psychological health, in reported physical symptoms, in some immune and physiological measures, in work and school outcomes, and in how people felt about the intervention. They did not show up in health behaviors, and they were weakest in the categories closest to fixed beliefs. Frattaroli's own summary is that disclosure helps psychological outcomes tied to emotion more than those tied to cognition.
Which is to say: writing is good at metabolizing something that is sitting on you. It is not a decision procedure, it is not a behavior change program, and it does not reliably revise what you believe about yourself. Those are different jobs and they need different tools, some of which are people.
If you have never done the four-day protocol above, do that before you pay anyone for journaling, ours included. It costs nothing, it takes about eighty minutes total, and it is the only version of this with 146 randomized trials underneath it.
What we do is a different job: naming the pattern first, so the writing has a subject. The free reading is six written questions, about eight minutes, and returns one word for the pattern running underneath how you answered. No card, and no requirement to ever pay us.
Take the free readingNot therapy, not a clinician, and no crisis handling. Nothing on this page is a clinical claim, and the trials above are about expressive writing, not about us.
Joanne Frattaroli, "Experimental disclosure and its moderators: A meta-analysis", Psychological Bulletin, 132(6), 2006, pp. 823 to 865. Free full text: PDF. Used for the 146 studies, r = .075, the confidence interval, the range, the 54 against 46 display, the aspirin comparison, all moderator figures, the r = .200 optimal-conditions subset, the published against unpublished gap, and the feedback exclusion criterion.
Joshua M. Smyth, "Written emotional expression: Effect sizes, outcome types, and moderating variables", Journal of Consulting and Clinical Psychology, 66(1), 1998, pp. 174 to 184. Used for d = .47 across 13 studies, the null result on health behaviors, and the short-term distress finding of d = .84 being unrelated to long-term outcomes.
Alex H. S. Harris, "Does expressive writing reduce health care utilization? A meta-analysis of randomized trials", Journal of Consulting and Clinical Psychology, 74(2), 2006, pp. 243 to 252. Used for g = 0.16 in healthy samples and the absence of the effect in medical and psychologically screened samples.
James W. Pennebaker, "Expressive Writing in Psychological Science", Perspectives on Psychological Science, 13(2), 2018, pp. 226 to 229. Used for his description of the first study, the replication history, the abandoned inhibition theory, and the figure of about .16 in Cohen's d across more than 100 studies.
James W. Pennebaker and Sandra K. Beall, "Confronting a traumatic event: Toward an understanding of inhibition and disease", Journal of Abnormal Psychology, 95(3), 1986, pp. 274 to 281. The original experiment.
Last verified 28 August 2026. Every figure on this page was read out of the source paper on that date.