The Portrait of You That Lives in a Database
What it means when an algorithm predicts you better than your friends.

There is a test that psychologists sometimes run that I find quietly disturbing. They take a person's social media activity, feed it into a machine learning model, and ask the model to predict the person's responses to a standard personality questionnaire. Then they compare the model's predictions to what the person actually scores. In study after study, the model's predictions are more accurate than the predictions made by the person's close friends and family, and in some cases more accurate than the person's own self-assessment. The model does not know the person. It has never met them. It cannot sense the particular quality of their presence in a room or understand the private reasons behind the choices they make. And yet it knows, in the narrow but important sense of accurate prediction, things about them that the people who love them cannot see.
The most accurate portrait of who you are may live inside a server you have never visited, owned by a company you signed up to years ago without reading the terms. The AI systems that build these models are not doing something mysterious. They are doing something that falls out naturally from the combination of scale, continuity, and pattern recognition. Human beings are remarkably consistent. Our language, our choices, our timing, our associations, our responses to novelty and stress vary within ranges that are characteristic of who we are, and those ranges leave signatures in digital behaviour that add up over time. A system that can read those signatures across millions of data points, and compare them to the signatures of millions of other people whose outcomes are known, can make predictions no human observer could make, because no human observer has access to that much signal.

The specific things AI systems can infer are more extensive and more accurate than most people appreciate. From social media posts and timing alone, researchers have shown accurate prediction of personality traits on the standard Big Five dimensions. From purchasing behaviour, it is possible to infer relationship status, pregnancy, health conditions, and financial stress with enough reliability to be commercially useful. From location data, sleep patterns and health conditions can be inferred. From the specific way you interact with a phone, including typing speed and time between responses, emotional state can be estimated in real time. From the combination of all of these streams, something that functions like a reasonably complete psychological profile can be assembled and continuously updated.
The accuracy is not uniform, and the field is careful about overclaiming even where industry applications are less restrained. Personality prediction from social media is better for some traits than others, and the confidence intervals around individual predictions are wider than the aggregate accuracy numbers suggest. Inferences about sensitive attributes tend to be more accurate in aggregate than for any specific person, which means the system may be reliably wrong about you while being right about the category it placed you in. The error profile is not random either. Systems trained on majority populations tend to be less accurate for minority groups, and the commercial applications built on these inferences apply them to everyone regardless of whether the inference is reliable for the specific person in front of it.

The psychological research on self-knowledge is more humbling than most of us are comfortable with. Human beings are not particularly accurate observers of their own behaviour, preferences, motivations, or even factual history. We misremember our past selves. We overestimate our consistency. We confabulate explanations for choices we made for reasons we were not aware of. In the specific sense of predicting future behaviour, AI models built on behavioural data may genuinely outperform human self-knowledge. Not because they understand you, but because they have seen the pattern of what people like you do next.
This creates a strange epistemic situation. The most accurate model of your likely future behaviour may exist outside you, in a system that has no concept of what it is like to be you but has an excellent model of what people with your pattern of behaviour tend to do. If you want to know whether you will stick to an exercise commitment, or whether a relationship pattern you are repeating will resolve the same way it did before, or whether your stated preference for a certain kind of work actually predicts your job satisfaction, the behavioural model may be a better predictor than your own assessment. Whether you should want to consult it, what it would mean to calibrate your self-understanding against it, and what would happen to the experience of living your own life if you did, are questions with no easy answers.

What concerns me most is not the existence of these models but the asymmetry of access to them. The companies that build and maintain AI models of individuals have detailed, continuously updated, operationally useful knowledge of those individuals. The individuals themselves have access to a much less comprehensive and much less accurate understanding of what those models contain, what they infer, and how they are used. That asymmetry is not accidental. It is built into the business model. The value of behavioural models comes partly from the information advantage they provide to their owners, and the mechanisms that would give people meaningful access to and control over models built on their data would undermine that advantage.
The terrifying reading of all this is the one that most privacy advocates reach for, and it is not wrong. A system that can predict your behaviour, infer your vulnerabilities, and model your psychological weak points, in the hands of commercial entities whose interests are not aligned with yours, creates conditions for manipulation that are qualitatively different from anything that existed before. The advertising system that knows you are in a period of emotional vulnerability before you fully know it. The insurance company that knows your health trajectory better than your doctor. The employer that knows your likelihood of quitting before you have consciously considered leaving.
There is also a genuinely useful version of the same technology. Personalised medicine that can predict your specific response to a treatment before you take it saves lives. A learning system that identifies exactly where your understanding of a subject breaks down and addresses that gap is more effective than instruction that treats you as the average student. A mental health tool that detects early warning signs of depression from your behaviour gives you the chance to seek help earlier. The knowing itself is not the problem. The conditions under which it is deployed, who controls it, who benefits from it, and what accountability exists when it goes wrong, are where the ethical questions actually live. Terrifying and useful are not really alternatives. They are simultaneous realities of the same technology, and which one dominates depends entirely on choices we have not yet made well.
You might also like
View all
Are We Giving AI Too Much Control?
Control does not transfer in one moment. It seeps out one small decision at a time.

The Internet Is Being Flooded With AI Slop
Near-zero production costs, and what happens to the signal beneath the flood.