Evidence: well sourced. Imported from the supplied 65-Case Master Edition, dated September 19, 2026. Source links and classifications are retained as an attributed case account; import is not an independent source review.

Case at a glance

Case number
022
Date / range
2008
Sector
Re-identification and inference
Genetic asset
Dense SNP data and aggregate statistics
Security principle
Membership Inference from Summary Data

Event summary

Homer and colleagues showed that, under certain conditions, an individual's contribution to a complex DNA mixture could be detected using dense SNP data. The work triggered restrictions on some public aggregate genomic datasets and changed research-data risk assumptions.

Source: doi.org — Homer Et Al. And Aggregate GWAS Data source 1.

Source: grants.nih.gov — Homer Et Al. And Aggregate GWAS Data source 2.

The case in context

The Homer study challenged the assumption that summary genomic information was necessarily harmless to release. Under the study's conditions, a target's known genetic information could be compared with aggregate data to assess whether that person contributed to the group.

The sensitive conclusion may be membership itself. Association with a particular cohort can reveal something that an individual-level file was intended to protect. The lesson is not that all statistics identify people, but that a release assessment needs to consider outside knowledge, the structure of the dataset, and what can be inferred from their combination.

Acquisition and processing

target genotype + aggregate case/control statistics → statistical comparison → membership inference → sensitive-study association

The sequence of events

  1. target genotype + aggregate case/control statistics
  2. statistical comparison
  3. membership inference
  4. sensitive-study association

What became inferable or exposed

Dense SNP data and aggregate statistics

Homer and colleagues showed that, under certain conditions, an individual's contribution to a complex DNA mixture could be detected using dense SNP data. The work triggered restrictions on some public aggregate genomic datasets and changed research-data risk assumptions.

Security dimensions

Confidentiality

The confidentiality question concerns dense snp data and aggregate statistics. Exposure and further inference must be distinguished from the fact of collection or availability.

Integrity

The integrity question is whether the described material, permissions, processing, or interpretation can be relied upon. Membership Inference from Summary Data identifies the particular boundary examined here.

Availability

Access and continuity are assessed for the described event; potential effects are not presented as confirmed outages or losses.

Provenance

The relevant chain follows dense snp data and aggregate statistics through the stages shown below. Missing public detail is not proof that internal records did not exist.

GeneticSecurity.org analysis

Genetic Exposure Radius

Not assessed

No single level is assigned where the supplied dossier gives a range, conditional outcome, or broad institutional consequence. The affected parties and proposed assessment are shown separately.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Genetic Persistence Risk

Not assessed

Persistence depends on the specific biological material or information retained. A potential effect is not treated as an observed genomic disclosure.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Genetic Provenance Integrity

Not assessed

A numeric provenance level is not inferred from the existence of a source or court record. It requires evidence of the relevant custody and processing controls.

Confidence: not assigned. Classification: GeneticSecurity.org analysis.

Proposed classification and its limits

Suggested GER: GER-0/1. Suggested GPR: GPR-4. GPI: not central.

These are proposed classifications from the supplied case dossier. Conditional scores describe an assumed exposure; they are not evidence that it occurred. A single numeric value is left unassigned when the asset or outcome is not sufficiently bounded.

What this case does not prove

It does not make every aggregate statistic re-identifying. Risk depends on sample size, marker density, reference knowledge, query design, and defensive controls.

Mitigations and lessons

  • Formal privacy review
  • Minimum cohort sizes
  • Query budgets
  • Controlled access
  • Differential-privacy research
  • Audit logging
  • Attack-aware release testing

Primary sources

Secondary sources

No additional source listed. See the evidence notes for limitations.

Policy and standards

Genetic Security Policy and Standards

Review and correction history

Source edition: September 19, 2026. Imported case account; no substantive corrections recorded.

Correction policy and log

Cite this case

GS-CASE-022. Homer 2008: When Aggregate Genomic Data Stopped Looking Anonymous. GeneticSecurity.org. https://geneticsecurity.org/cases/022-homer-aggregate-gwas-membership-inference/