Evidence: well sourced. Imported from the supplied 65-Case Master Edition, dated September 19, 2026. Source links and classifications are retained as an attributed case account; import is not an independent source review.

Case at a glance

Case number
023
Date / range
2013
Sector
Re-identification and inference
Genetic asset
Y-chromosome markers and public demographic information
Security principle
Quasi-Identifier Fusion

Event summary

Researchers demonstrated that Y-chromosome markers could sometimes be linked to surnames through genealogy databases, then combined with age, geography, and family information to identify supposedly anonymous genome donors.

Source: pubmed.ncbi.nlm.nih.gov — Gymrek Surname Inference source 1.

Source: doi.org — Gymrek Surname Inference source 2.

The case in context

The surname-inference research connected genetic markers with genealogy records and demographic clues. Its significance was the joining of information across datasets: a name and a genomic record did not have to appear together in the same release to become linkable.

The method's reach depends on the population, the available genealogy information, and the clues accompanying the record. A surname candidate is not a guaranteed identity. The broader security question is which external datasets can supply the missing connection after obvious identifiers have been removed.

Acquisition and processing

anonymous male genome → Y-STR profile → genealogy database → surname candidate → demographics/public records → identity hypothesis

The sequence of events

  1. anonymous male genome
  2. Y-STR profile
  3. genealogy database
  4. surname candidate
  5. demographics/public records
  6. identity hypothesis

What became inferable or exposed

Y-chromosome markers and public demographic information

Researchers demonstrated that Y-chromosome markers could sometimes be linked to surnames through genealogy databases, then combined with age, geography, and family information to identify supposedly anonymous genome donors.

Security dimensions

Confidentiality

The confidentiality question concerns y-chromosome markers and public demographic information. Exposure and further inference must be distinguished from the fact of collection or availability.

Integrity

The integrity question is whether the described material, permissions, processing, or interpretation can be relied upon. Quasi-Identifier Fusion identifies the particular boundary examined here.

Availability

Access and continuity are assessed for the described event; potential effects are not presented as confirmed outages or losses.

Provenance

The relevant chain follows y-chromosome markers and public demographic information through the stages shown below. Missing public detail is not proof that internal records did not exist.

GeneticSecurity.org analysis

Genetic Exposure Radius

Not assessed

No single level is assigned where the supplied dossier gives a range, conditional outcome, or broad institutional consequence. The affected parties and proposed assessment are shown separately.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Genetic Persistence Risk

Not assessed

Persistence depends on the specific biological material or information retained. A potential effect is not treated as an observed genomic disclosure.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Genetic Provenance Integrity

Not assessed

A numeric provenance level is not inferred from the existence of a source or court record. It requires evidence of the relevant custody and processing controls.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Proposed classification and its limits

Suggested GER: GER-2/3. Suggested GPR: GPR-5. GPI: not central.

These are proposed classifications from the supplied case dossier. Conditional scores describe an assumed exposure; they are not evidence that it occurred. A single numeric value is left unassigned when the asset or outcome is not sufficiently bounded.

What this case does not prove

Surname inference does not work for everyone, is culturally and demographically uneven, and produces leads rather than guaranteed identities.

Mitigations and lessons

  • Remove unnecessary quasi-identifiers
  • Model auxiliary-data attacks
  • Controlled access
  • Family-risk disclosure
  • Resistant query design
  • Continuous re-identification testing

Primary sources

Secondary sources

No additional source listed. See the evidence notes for limitations.

Policy and standards

Genetic Security Policy and Standards

Review and correction history

Source edition: September 19, 2026. Imported case account; no substantive corrections recorded.

Correction policy and log

Cite this case

GS-CASE-023. Surname Inference: How an Anonymous Genome Found a Family Name. GeneticSecurity.org. https://geneticsecurity.org/cases/023-surname-inference-anonymous-genomes/