Measured, published, and lower than the marketing. Language models reading your writing land somewhere between r = 0.29 and about r = 0.48 against your own questionnaire answers. That is a real signal, it is not identification, and the yardstick itself has a hole in it that nobody selling this points at.
We sell a reading produced by exactly this kind of inference, so we have a direct interest in you believing these numbers are high. Here they are with sources, along with the three things about them that are bad for us.
Every accuracy figure in this field measures how well a machine predicts what you would say about yourself. That is a strange yardstick for a tool whose entire purpose is the part you would not say.
| Study | Input | Sample | Accuracy |
|---|---|---|---|
| Youyou, Kosinski and Stillwell, 2015 | Facebook Likes | 86,220 | r = 0.56, against r = 0.49 for the person's friends |
| Peters and Matz, 2024 | Facebook status updates | 1,000 | r = 0.29 average, range 0.22 to 0.33 |
| Wright et al., 2026 | Open-ended spoken narratives | 60 and 108 | Average across models of about 0.48 and 0.44 |
In 2015 Wu Youyou, Michal Kosinski and David Stillwell used 86,220 volunteers who had completed a 100-item personality questionnaire. A model working from Facebook Likes predicted those answers at r = 0.56, while the participants' own Facebook friends, filling in the same questionnaire about them, reached r = 0.49. The paper reports the tipping points precisely: the model needed roughly 10, 70, 150 and 300 Likes to outperform an average work colleague, an average cohabitant or friend, an average family member, and an average spouse (PNAS, 112(4), 2015, pp. 1036 to 1040; free full text).
The authors state a limitation that is usually dropped when this result gets repeated: their human judges could only describe the person using a 10-item questionnaire, and in reality those people might know more than that instrument captured. The headline is that a computer beat your friends. The paper says a computer beat a 10-item form filled in by your friends.
Heinrich Peters and Sandra Matz tested GPT-3.5 and GPT-4 on the status updates of 1,000 Facebook users, with no training, and got an average correlation with self-reported Big Five scores of r = 0.29, ranging from 0.22 to 0.33. That is roughly the level reached by supervised models built specifically for the job, which is the interesting part: general models arrive at the same ceiling with no special training (PNAS Nexus, 3(6), 2024, article pgae231; free full text).
They also found the accuracy is not evenly distributed. Predictions were more accurate for women and for younger participants on several traits, which the authors attribute to training data or to differences in how people express themselves online. If you are older or male, the number that applies to you is below the average.
The strongest result comes from unstructured personal narrative rather than social media. Aidan Wright and colleagues had seven commercial language models score Big Five traits from spontaneous streams of thought and from nightly video diaries. Averaging the models produced convergent correlations with self-report of roughly 0.48 in one sample and 0.44 in the other, which the authors describe as comparable to or better than established benchmarks including agreement between a person and their own friends and family (Wright et al., Nature Human Behaviour, 2026).
The samples were 60 and 108 people, with mean ages of 24.5 and 28.0. The authors say plainly that the samples were not representative, that they were too small to compare performance across demographic groups, and that whether these models score people from different groups equivalently is an unstudied question.
This is where almost every article on this topic stops being useful, so here it is concretely.
Square the correlation and you get the share of variation the two measures hold in common. At r = 0.29 that is about 8 percent. At r = 0.44, about 19 percent. At r = 0.48, about 23 percent. At the celebrated 0.56, about 31 percent. So even the best published system leaves roughly seven tenths of the variation in your questionnaire scores unexplained.
And there is a second, larger caveat that survives any improvement in the models. These are population statistics. A correlation of 0.48 describes how a method tracks a sample of a hundred people. It says nothing about whether the specific reading you received this morning is right. There is no published number, for any of these systems, that tells an individual how much to trust their own result. Anyone implying otherwise, in either direction, is going beyond the evidence.
What the correlations do support is a modest and useful claim: your writing carries real signal about you, at a strength somewhere between a stranger's impression and a friend's. That is worth something. It is not a scan.
In nearly every study above, accuracy means agreement with how you answered a questionnaire about yourself. Self-report is the criterion. A model earns a high score by predicting what you would say about you.
Sit with what that implies for anything sold as self-discovery. If a system told you something true that you do not believe about yourself, that would count against its accuracy, not for it. The measurement rewards agreement with your self-image. The entire premise of shadow work is that your self-image has gaps in it on purpose.
The researchers know this. Wright and colleagues note that self and other ratings are blends of true variance, self-concept, impression management and observability, and that perfect agreement between a self-rating and another method is neither expected nor desirable. They also state that their study was not designed to maximize agreement with self-report, and that much higher agreement is probably achievable with prompts engineered to elicit a person's self-concept.
Read that last sentence as a product warning. Any company in this category can raise its apparent accuracy by getting better at telling you what you already think. That is a real and available option, it improves the metric, and it makes the tool worse at the only job that would justify it.
Every product in this space, ours included, shows you a rationale. Here is the line, here is why it mattered, here is what it suggests. That presentation implies the reasoning came first and the conclusion followed.
Wright and colleagues address this directly, because their own method uses model-generated rationales and quotes. Their assessment is that such explanations are post hoc, that they do not provide a direct summary of the model's internal reasoning, and that they should not be interpreted as reflecting the actual processes the model used to produce its result. They compare them to post hoc explainability techniques, and say process-level understanding requires mechanistic interpretability methods that have not been applied here.
The quote a system shows you is real. The story about why that quote produced this conclusion is generated afterward, and it is a plausible account rather than a record.
This is not an accusation against any particular product. It is how these systems work, it applies to ours, and the practical consequence is worth carrying: judge a reading by the evidence it points at in your own words, not by how convincing its reasoning sounds. The reasoning is the part that is easiest to make sound convincing and hardest to verify.
Having spent the page lowering expectations, here is what survives, because a page that only debunks is as useless as one that only sells.
There is good evidence that outside observers know things about you that you do not. Simine Vazire's self-other knowledge asymmetry work had 165 people rated by four friends and by strangers in a round-robin design, and found the pattern is not uniform: you are the best judge of traits that are hard to observe and unflattering to admit, such as anxiety, while other people are better judges of traits that are highly evaluative, such as intellect (Vazire, Journal of Personality and Social Psychology, 98(2), 2010, pp. 281 to 300).
That is the honest shape of the argument for reading from the outside. Not that an outside view is more accurate in general, but that self-knowledge and outside knowledge fail in different places, and the failures do not overlap. Where your view of yourself is most managed is where another view is worth most.
A model reading your text is a third thing again: not you, not someone who knows you, and not accountable to either. On the traits where the literature says the self is the better judge, it has no advantage over you at all.
None of the figures on this page are ours. There is no published study of LUX, no external validation, no peer review, and no accuracy figure we can honestly quote. Every correlation above belongs to somebody else's system measured on somebody else's data, and we have cited them because they are the best available evidence about the category, not because they are evidence about us.
We could commission something that produced a number. We have not, and until we do the honest position is the one stated here. If you see any company in this space quoting an accuracy figure, ask whose study it was, what the criterion was, and whether it was their product that was tested.
LUX asks six written questions, about eight minutes, and returns one word for the pattern running underneath how you answered. It reads the rhythm of the answering, not only the content of the answers. It is a starting point for looking, and it is not a diagnosis, a score, or a validated instrument.
The test we would rather you apply is the one this page argues for: keep the word, date it, and see whether it keeps surfacing in your own writing over the next few months. That check costs you nothing and it is worth more than any correlation we could quote at you.
Take the free readingFree, no card. Not therapy, not a clinician, and no crisis handling.
Wu Youyou, Michal Kosinski and David Stillwell, "Computer-based personality judgments are more accurate than those made by humans", PNAS, 112(4), 2015, pp. 1036 to 1040 (free full text). Used for the sample of 86,220, r = 0.56 against r = 0.49, the 10, 70, 150 and 300 Likes thresholds, and the authors' caveat about the 10-item judge questionnaire.
Heinrich Peters and Sandra C. Matz, "Large language models can infer psychological dispositions of social media users", PNAS Nexus, 3(6), 2024, article pgae231 (free full text). Used for r = 0.29 with a range of 0.22 to 0.33 across 1,000 users, the comparison with supervised models, and the higher accuracy for women and younger participants.
Aidan G. C. Wright, Whitney R. Ringwald, Colin E. Vize, Johannes C. Eichstaedt, Mike Angstadt, Aman Taxali and Chandra Sripada, "Assessing personality using zero-shot generative AI scoring of brief open-ended text", Nature Human Behaviour, 10(3), 2026, pp. 541 to 555. Used for the seven models, the samples of 60 and 108, the average convergent correlations of about 0.48 and 0.44, the stated limitations on representativeness and measurement invariance, the note that higher agreement with self-report is achievable with engineered prompts, and the statement that model-generated explanations are post hoc.
Simine Vazire, "Who knows what about a person? The self-other knowledge asymmetry (SOKA) model", Journal of Personality and Social Psychology, 98(2), 2010, pp. 281 to 300. Used for the finding that the self is the better judge of low-observability traits such as neuroticism while friends are better judges of highly evaluative traits such as intellect.
Last verified 28 August 2026. Every figure on this page was read out of the source paper on that date.