ReputatioLab study · Society 30 January 2025

Business leaders: the personal data freely available that exposes you

Leaked addresses, official registers, metadata left behind without a second thought: a good part of an executive's life is now freely accessible online. Cross-referenced, these voluntary, involuntary or inherited traces become a physical risk factor.

Illustration from the ReputatioLab study on the abduction of the co-founder of Ledger and the personal data of wealthy executives
The Ledger case: the facts
21 January 2025 Day of the abduction David Balland, co-founder of Ledger, and his wife, abducted from their home by an armed gang.
$10m Ransom demanded Ten million dollars in cryptocurrency.
1,001 residents Population of the executive's municipality In 2021, according to the study: a good chance of running into him outside his seminars.
Case study

The case that revealed the problem

Ledger logoLedger

On 21 January 2025, David Balland, co-founder of Ledger, and his wife were abducted from their home by an armed gang. Ransom demanded: $10 million in cryptocurrency. A scene from a film. But behind this drama, one cannot help wondering: how did the criminals know where to go?

Ledger logo taken from Wikimedia Commons, Creative Commons BY-SA licence, unaltered.

Lessons

Key takeaways

  1. Many companies and individuals wrongly believe that information is protected simply because it is drowned in the mass of public data.
  2. Protecting your data demands constant vigilance and substantial resources, whereas someone looking for information needs only a single flaw to exploit a target.
  3. The transparency of company registers is important, and the right to information also exists, but there is no framework to arbitrate between them.

On 21 January 2025, David Balland, co-founder of Ledger, and his wife were abducted from their home by an armed gang. Ransom demanded: $10 million in cryptocurrency. A scene from a film. But behind this drama, one cannot help wondering: how did the criminals know where to go?

How did the criminals know where to go?

ReputatioLab, Abduction of the co-founder of Ledger

David Balland was not a flamboyant crypto entrepreneur flaunting his success on social media. Not the type to feed an Instagram account somewhere between Lamborghinis and seminars in Dubai. He led a quiet life, far from the spotlight. Yet his kidnappers seemed to know exactly where to strike.

Addresses retrieved from leaked databases, files accessible on official registers, metadata left on websites without a second thought. Personal data has become a physical risk factor.

Personal data has become a physical risk factor.

ReputatioLab, Abduction of the co-founder of Ledger

What types of data?

Manon El Assaidi, our director of operations, has handled many cases of this kind using what is known as “OSINT” (Open Source Intelligence).

Online identity

Voluntary digital traces

Posts, interactions, personal information

Involuntary digital traces

Browsing, location, metadata

Inherited digital traces

By others, archives

Analysis and profilingDirect or indirect

Offline identification

Illustration of the OSINT (Open Source Intelligence) work carried out by Manon El Assaidi, director of operations at Saper Vedere.
Figure from the ReputatioLab study, redrawn for the website

There are three types of digital traces:

Voluntary Voluntary digital traces Information we deliberately share online, such as social media posts.
Involuntary Involuntary digital traces Data collected without our explicit intent, such as the metadata attached to files or browsing information.
Inherited Inherited digital traces Information about us published by others, such as photos or mentions.

These traces can indeed reveal sensitive information. For example, a photo of a house shared on social media (a voluntary trace) may contain geolocation metadata (an involuntary trace), potentially exploitable by malicious individuals.

It is therefore essential to manage these traces carefully to protect our privacy and our online security.

We have established a typology of digital traces

Voluntary digital traces

Posts and shares

  • Photos, videos, status updates, comments, geolocation on social media
  • Blog posts, forums, comments on websites
  • Reviews of products or services

Online interactions

  • Likes, comments, shares, mentions
  • Private messages, emails, online discussions

Personal information

  • Social media profile, online CV, registration forms, addresses and other private information

Involuntary digital traces

Browsing data

  • Search history, websites visited, pages viewed
  • IP address, device type, operating system
  • Cookies, unique identifiers, fingerprints

Location data

  • GPS geolocation, geotagged photos
  • Social media check-ins, navigation apps

Metadata

  • Creation date, modification date, author of a document
  • Photo Exif data, GPS data in audio/video files

Inherited digital traces

Posts by other people

  • Photos, videos, mentions in posts
  • Comments, tags, content shares, geolocation

Archived information

  • Old websites, forums, discussion groups
  • Press articles, mentions in public documents

Personal information

  • Social media profile, online CV, registration forms, addresses and other private information

Identification from digital traces

Direct

Direct identification stems mainly from voluntary digital traces. This data, often shared openly by the individual, may contain explicit information allowing immediate and unambiguous identification of the person in the real world.

Indirect

Indirect identification is generally the result of analysing and correlating involuntary and inherited digital traces. Although this information, taken in isolation, does not always allow direct identification, aggregating and analysing it can reveal patterns, behaviours and social ties that lead to the individual being identified.

The typology of digital traces established by the study: voluntary, involuntary, inherited, and the two routes of identification.
Figure from the ReputatioLab study, redrawn for the website

How can data be used to move from one piece of metadata to another?

Specific data is searched for through various elements, then common variables between two metadata networks are used to obtain additional information about someone.

  1. In company data, I identify the name of an executive.
  2. This executive has a presence on LinkedIn, where I retrieve their previous positions and the people they interact with.
  3. Thanks to this, I identify siblings.
  4. These siblings have photographs on Facebook showing a family dinner, with data that make it possible to geolocate the executive's house.

Examples that can be found in company databases:

ID

NameSirenSirenName

Location

Website front endDocument metadata linksAPI

Output

Personal information

Each has its own specific features

Examples of data associated with an executive, found in company databases.
Figure from the ReputatioLab study, redrawn for the website

Of course, some data is of low interest or sensitivity, and some is highly sensitive:

Type of information Low Medium High
Direct categorisation
Posts and shares on social media, blogs, etc. Non-compromising status updates, non-sensitive photos Political opinions, religious affiliations Medical information, financial details, postal addresses
Online interactions
(comments, private messages)
Conversations with no personal disclosures Sensitive discussions (religion, politics) Personal disclosures, confessions
Personal information
(profile, online CV)
Basic information (name, age) Detailed professional history Full contact details, social security number, signature, etc.
Indirect categorisation
Browsing data
(search history, IP address)
General searches and browsing paths Medical history, legal consultations Sensitive searches (trauma, crimes)
Location data
(GPS geolocation)
Public places Travel habits Places frequented in confidence
Metadata
(creation date, photo Exif data)
Publication dates Places frequented, regular contacts Places and people frequented in confidence
Ranking of data by sensitivity, from low to high, as presented in the study.
Figure from the ReputatioLab study, redrawn for the website

Artificial intelligence will make this even more accessible

Alert

Artificial intelligence can process this previously inaccessible data on a massive, automatic and exhaustive scale, by sending hundreds of pre-set prompts to dedicated sources. Even in a very basic way, though, it can already yield information:

Capture Screenshot of a conversation with a consumer AI, reproduced in the study
Information obtained automatically through a basic artificial intelligence prompt.

And let us apply it to the case in hand:

Capture Screenshot of a conversation with a consumer AI, reproduced in the study
Application of this artificial intelligence method to the case studied by ReputatioLab.

If I want to work with visuals, videos or other material, I already have the list of things to look at:

Capture Screenshot of a conversation with a consumer AI, reproduced in the study
List of visual and video elements to examine, generated by artificial intelligence.

In a town of 1,001 inhabitants (in 2021), that leaves a good chance of running into him outside the seminars he gives in his field!

What issues does it raise?

Once the theoretical elements are clearly explained, it has to be said that there is a whole series of problems:

  • Differences in framing:How can harm be seen in posting family photos? How can an administrative document on page 25 of Google be seen as a digital reputation problem? Many companies and individuals wrongly believe that information is protected simply because it is drowned in the mass of public data. Yet OSINT tools make it possible to automate searches and make the invisible visible.
  • Asymmetry between offence and defence: protecting your data demands constant vigilance and substantial resources, whereas someone looking for information needs only a single flaw to exploit a target. Moreover, there are very many entry points: in the Ledger case, the kidnappers did not go after Eric Larchevêque. And there were other possibilities.
  • Difficulty of taking action: while an investigation will always uncover sensitive data, it is impossible to predict whether a study will yield interesting results. And even where highly sensitive data comes to light, the ability to act on it is rather limited.
  • No regulation on the subject: the transparency of company registers is important, and the right to information also exists, but there is no framework to arbitrate between them. And launching legal proceedings would directly produce a Streisand effect. (In trying to prevent the disclosure of information that some would like to hide, the opposite result occurs: the hidden fact becomes widely known.)