Evidence: well sourced. Imported from the supplied 65-Case Master Edition, dated September 19, 2026. Source links and classifications are retained as an attributed case account; import is not an independent source review.
Case at a glance
- Case number
- 023
- Date / range
- 2013
- Sector
- Re-identification and inference
- Genetic asset
- Y-chromosome markers and public demographic information
- Security principle
- Quasi-Identifier Fusion
Event summary
Researchers demonstrated that Y-chromosome markers could sometimes be linked to surnames through genealogy databases, then combined with age, geography, and family information to identify supposedly anonymous genome donors.
Source: pubmed.ncbi.nlm.nih.gov — Gymrek Surname Inference source 1.
The case in context
The surname-inference research connected genetic markers with genealogy records and demographic clues. Its significance was the joining of information across datasets: a name and a genomic record did not have to appear together in the same release to become linkable.
The method's reach depends on the population, the available genealogy information, and the clues accompanying the record. A surname candidate is not a guaranteed identity. The broader security question is which external datasets can supply the missing connection after obvious identifiers have been removed.
Acquisition and processing
anonymous male genome → Y-STR profile → genealogy database → surname candidate → demographics/public records → identity hypothesis
The sequence of events
- anonymous male genome
- Y-STR profile
- genealogy database
- surname candidate
- demographics/public records
- identity hypothesis
What became inferable or exposed
Y-chromosome markers and public demographic information
Researchers demonstrated that Y-chromosome markers could sometimes be linked to surnames through genealogy databases, then combined with age, geography, and family information to identify supposedly anonymous genome donors.
Affected parties and consent
- Direct parties
- Genome donors involved in the research demonstration
- Indirect parties
- Relatives and connected participants may be relevant where the asset contains relationship information.
- Direct count
- Unknown / not assigned
- Indirect count
- Unknown / not assigned
- Consent status
- Permission to contribute a record does not settle what may be inferred about a nonparticipant or relative. Public availability and authorization for a particular use should be distinguished.
Security dimensions
Confidentiality
The confidentiality question concerns y-chromosome markers and public demographic information. Exposure and further inference must be distinguished from the fact of collection or availability.
Integrity
The integrity question is whether the described material, permissions, processing, or interpretation can be relied upon. Quasi-Identifier Fusion identifies the particular boundary examined here.
Availability
Access and continuity are assessed for the described event; potential effects are not presented as confirmed outages or losses.
Provenance
The relevant chain follows y-chromosome markers and public demographic information through the stages shown below. Missing public detail is not proof that internal records did not exist.
Consent, persistence, and relational exposure
Consent
Permission to contribute a record does not settle what may be inferred about a nonparticipant or relative. Public availability and authorization for a particular use should be distinguished.
Persistence
Later reuse depends on the actual asset and links to other records; no future misuse is asserted.
Relational exposure
Relatives and connected participants may be relevant where the asset contains relationship information.
Case-specific assessment
confidentiality critical; integrity low; consent high; persistence high; relational exposure extended-family scale.
GeneticSecurity.org analysis
Genetic Exposure Radius
No single level is assigned where the supplied dossier gives a range, conditional outcome, or broad institutional consequence. The affected parties and proposed assessment are shown separately.
Confidence: not assigned. Classification: GeneticSecurity.org analysis.
Genetic Persistence Risk
Persistence depends on the specific biological material or information retained. A potential effect is not treated as an observed genomic disclosure.
Confidence: not assigned. Classification: GeneticSecurity.org analysis.
Genetic Provenance Integrity
A numeric provenance level is not inferred from the existence of a source or court record. It requires evidence of the relevant custody and processing controls.
Confidence: not assigned. Classification: GeneticSecurity.org analysis.
Proposed classification and its limits
Suggested GER: GER-2/3. Suggested GPR: GPR-5. GPI: not central.
These are proposed classifications from the supplied case dossier. Conditional scores describe an assumed exposure; they are not evidence that it occurred. A single numeric value is left unassigned when the asset or outcome is not sufficiently bounded.
What this case does not prove
Surname inference does not work for everyone, is culturally and demographically uneven, and produces leads rather than guaranteed identities.
Mitigations and lessons
- Remove unnecessary quasi-identifiers
- Model auxiliary-data attacks
- Controlled access
- Family-risk disclosure
- Resistant query design
- Continuous re-identification testing
Primary sources
- PRIMARY SOURCE pubmed.ncbi.nlm.nih.gov — Gymrek Surname Inference source 1
- PRIMARY SOURCE doi.org — Gymrek Surname Inference source 2
Secondary sources
No additional source listed. See the evidence notes for limitations.
Policy and standards
Genetic Security Policy and StandardsReview and correction history
Source edition: September 19, 2026. Imported case account; no substantive corrections recorded.
Correction policy and logCite this case
GS-CASE-023. Surname Inference: How an Anonymous Genome Found a Family Name. GeneticSecurity.org. https://geneticsecurity.org/cases/023-surname-inference-anonymous-genomes/