Spotify might know whether you smoke, drink, or exercise. It might also be able to guess your age group, gender, financial situation, marital status, and even something about your personality. Or rather, the public playlists you create on Spotify can reveal all of this, with a level of accuracy that is far from negligible. And whatever Spotify can know, anyone else can potentially know too, simply by drawing on the data contained in your playlists and letting artificial intelligence do the rest.
This is what emerges from From Beats to Breaches: How Offensive AI Infers Sensitive User Information from Playlists, a study by Stefano Cecconello, Mauro Conti, Luca Pajola, Luca Pasa, and Pier Paolo Tricomi, from the University of Padua and SpritzMatter, a university spin-off, and part of a broader line of research the group has been pursuing in recent years. In this study, the authors developed musicPIIrate, an Offensive AI system capable of using public music playlists to infer information that is anything but public: demographic data, personal habits, and some of the so-called Big Five, the five major personality traits used in psychology.
To be clear: we are not talking about what Spotify already knows about its users through the data collected directly by the service. In the model envisioned by the researchers, the attacker simply knows the public profile of the person being targeted and can access their shared playlists, the tracks they contain, and the related metadata. So there is no vulnerability to exploit, nor any stolen password. As Conti and his colleagues make clear, the attack relies exclusively on data that is public and voluntarily shared.
“In recent years, one of the research areas we have focused on most is precisely this: understanding how, starting from public information that seems to reveal nothing, it is actually possible to extract private information,” says Mauro Conti, Professor of Cybersecurity at the University of Padua and coordinator of the SPRITZ Security and Privacy Research Group. The group had already worked on cases that may seem even more science-fictional: from the sound produced by a keyboard during a conference call, for example, it is possible to try to reconstruct which key is being pressed; from the audio of a conversation, meanwhile, information can be inferred about the physical environment in which the speaker is located.
So it was only a matter of time before the same question was applied to music: if we leave dozens of playlists built up over the years publicly available, what do they say about us? What potentially sensitive data are we sharing with the world?
The idea that musical tastes say something about personality is not, in itself, a recent discovery. The paper draws on a long tradition of studies that have observed statistical correlations between musical preferences and psychological characteristics: as early as 2003, Rentfrow and Gosling associated greater openness to experience, for example, with a preference for complex genres such as jazz and classical music, while extroversion appeared more strongly linked to pop, energetic, and mainstream music. The leap lies in the fact that correlations like these, when fed into systems capable of analysing enormous quantities of information, can become a tool for automated profiling.
“Artificial intelligence allows us to analyse large quantities of data and learn from the data,” Conti explains. The process begins with people whose characteristics are already known, such as age, occupation, financial situation, and habits, and lets the model identify recurring patterns that a human being might not even be capable of spotting. Once learned, those characteristics can then be searched for in the data of other users who have never explicitly provided that information.
To test musicPIIrate, the researchers used a dataset created for a previous study by the same group: more than 10,000 playlists belonging to 739 users, containing over 200,000 songs and 55,000 artists in total. Each user has an average of almost fourteen playlists. Using that material, the system attempts to reconstruct fifteen different characteristics: age group, country, financial status, gender, marital status, and occupation; alcohol consumption, smoking habits, physical activity, and Premium account ownership; and finally openness, conscientiousness, extroversion, agreeableness, and neuroticism.
And this is where the playlist we use for running, the one we play in the car, or the one perhaps simply called “summer 2022” suddenly stop looking quite so harmless.
For smoking, for example, some of the models tested achieve an F1-score of around 0.60–0.61, compared with 0.44 for a random classifier. For alcohol consumption, the score reaches around 0.59 compared with 0.31; for exercise, about 0.35 compared with 0.16. Some personality traits also show substantial differences: the best-performing model reaches 0.43 for conscientiousness, 0.42 for extroversion, and 0.44 for neuroticism.
There is one misunderstanding that should be avoided, however: an F1-score of 0.60 does not mean that the program looks at John Smith’s profile and determines with 60 percent certainty that John Smith is a smoker. F1 is a statistical measure of the classifier’s overall performance, combining precision with its ability to correctly identify positive cases. MusicPIIrate, then, is not a lie detector, but in truth, it does not need to become one.
“The performance is not one hundred percent precise,” Conti stresses, but compared with a random guess, that is, simply guessing, there is still “a significant advantage.” And it is scale that changes the meaning of that advantage. If the goal were to prove in court that a single person smokes, it obviously would not be enough. But if the aim is to divide hundreds of thousands of people into groups that are statistically more likely to buy a certain product, click on a certain message, or respond to a certain stimulus, even an imperfect prediction can become valuable. Marketing, yes, but the greater risk is phishing.
An attacker could use the inferred information to create more personalised communications: knowing something about the victim’s musical tastes and personality, Conti notes, means being able to craft a more convincing email, perhaps using concert tickets or other elements consistent with their interests as bait.
But Spotify is not the only possible source.
A would-be profiler can take what they manage to infer from playlists, combine it with what they find on Instagram, Facebook, LinkedIn, or any other public source, and gradually build an increasingly detailed portrait. Whoever carries out the attack, Conti explains, can extract one piece of information from Spotify and another from Instagram, combining different sources of information to create increasingly refined profiles.
The truly sensitive piece of data, in short, may not be any single piece of information we chose to make public, but rather what emerges when our public traces are combined. And this is precisely where the second part of the research comes in. If a collection of playlists works like a kind of involuntary statistical fingerprint, why not try deliberately smudging that fingerprint?
The countermeasure developed by the researchers is called JamShield, and it is based on a simple principle: if a model learns to recognise us by observing recurring patterns in our data, introducing plausible but false patterns may be enough to confuse it, making its predictions less reliable.
JamShield is designed to add specially selected decoy playlists to a user’s profile in order to confuse the model. To simplify considerably, if certain playlists are statistically associated more frequently with a particular category of users, adding them to the profile of someone who belongs to the opposite category can shift the prediction in the wrong direction. The same principle can be applied to age, habits, and personality traits.
To make the whole system even more sophisticated, the authors decided not to create artificial playlists perfectly designed to fool the system. They would certainly have been more effective, but also easier to identify as decoys. JamShield instead uses playlists that actually exist and are taken from the dataset, not something that is merely realistic, as Dr Stefano Cecconello, a postdoctoral researcher at the University of Padua, points out, but something real.
Cecconello imagines the application almost like a privacy slider built into Spotify: the more protection the user chooses, the more decoy playlists the system adds to the profile displayed externally, without necessarily cluttering the personal experience of the person using the service.
In the tests, adding just one playlist already reduced the effectiveness of the attack by an average of around ten F1-score points. As the number of decoy playlists increases, the effect generally becomes stronger, although in some cases it quickly reaches a plateau; the researchers stopped at four precisely to avoid designing a defence that might be extremely effective in theory but completely unrealistic in everyday use.
For now, however, JamShield is not a feature we will find in the Spotify app. It is an experimental countermeasure proposed for implementation by the platform, although in principle individual users could also protect themselves by adding appropriately selected playlists, or in the future by relying on an application that does so automatically.
But all of this comes at a cost. A platform that introduced false data into public profiles would make the information that anyone outside the platform can read from those profiles less reliable. At the same time, a user who independently adopted a similar strategy would end up presenting others with a profile full of playlists that do not really belong to them. Privacy, after all, often alters some aspect of the very mechanism it is trying to protect.
For years, we learned to think of personal data as something fairly intuitive: name, address, phone number, location, credit card. Then social media arrived, and we had to get used to the idea that even a like, a search, or a photograph could reveal far more than it seemed. The advent of the latest AI systems has changed the rules once again, because information can now become sensitive not because it is inherently sensitive, but because there is a machine capable enough to connect it with other information.
A playlist will, naturally, continue to be a compilation for a party, a declaration of love, a folder of depressing songs for November, or the messy archive of whatever we were listening to in 2017. But viewed alongside thousands of others, it can suddenly become a kind of questionnaire we do not remember filling out.
With one difference: when asked, “Do you smoke?”, we never actually answered.
Pierluigi Fantozzi