Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Psychology

Ethical Challenges of Digital Trace Data in Social Research

Quick fact

In 2006, AOL released search logs for research, but after a journalist re-identified users from the data, the project was quickly pulled—highlighting the real risk of re-identification even when data is supposedly anonymized.

Why this is interesting

Your every click, search, and post leaves a trace that researchers can analyze. But should they be allowed to use it without asking you?

Read the full explanation

Understanding Ethical Challenges of Digital Trace Data in Social Research

Digital trace data refers to the records of human activity left behind on platforms, such as browsing history, social media posts, likes, and location pings. Unlike traditional research data, which is often collected through surveys or interviews with explicit consent, these traces are byproducts of people's everyday online behavior. Researchers can access this data in various ways, including using public APIs, scraping publicly available profiles, or partnering with companies. The ethical challenges emerge because this data is sensitive: it can reveal intimate details about relationships, opinions, health, and even personal habits. The key issue is that the people who generate the data rarely expect it to be used for research, and they are seldom asked for permission. A common sense analogy is this: if you walk down a street, you might be captured by a CCTV camera, but you likely don't expect a sociologist to analyze your every movement without your consent. Just because a platform's terms of service allow data collection for operational purposes, it doesn't automatically make it ethical to use that data for unrelated research.

A deeper explanation

The core ethical challenge is that the very nature of digital trace data conflicts with traditional research ethics principles. Informed consent, which is normally a cornerstone, is often impracticable because the data is generated passively and at scale. For example, scraping tweets for a study of political discourse would require obtaining permission from millions of users, which is impossible. Privacy is another issue: even when data is stripped of direct identifiers like usernames, the combination of attributes—such as location, timestamps, and content—can make individuals uniquely identifiable, a process known as re-identification. Harm can arise not only from the exposure of the data but also from the way findings are disseminated, potentially stigmatizing groups or communities. Furthermore, the data itself may encode biases (e.g., not everyone uses social media), and algorithms used to analyze the data can amplify these biases, leading to skewed conclusions that affect people's lives. The underlying principle here is that the digital world has eroded the boundaries of public and private spaces. What is considered 'public' on the internet is non-trivial to define, and users may have a different expectation of privacy than what the platform's terms imply. Researchers also face the issue of data ownership: who owns the digital traces—the users, the platforms, or the public? This ambiguity complicates the ethical decision. Addressing these challenges requires researchers to adopt a pragmatic approach: they must weigh the potential benefits of the research against the risks, seek alternative consent mechanisms (like opt-out), and be vigilant in de-identifying data and mitigating biases. It is an ongoing process of balancing curiosity with respect for human dignity.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.