atlas-data_uniqueness

Data Uniqueness

Measures the degree to which each sample in a dataset is unique: sample identifiers are present and unique, and no two records share the same content across the non-ID columns.
Tags:
Data Quality

Overview

The Data Uniqueness evaluation measures how unique the samples in a dataset are. It runs two checks over the whole dataset at once. The first checks that each sample carries a present, unique identifier. The second checks that no two samples share the same content across the columns that define them. Each check produces a metric, so identifier problems and content duplication stay separate.

The evaluation works with no configuration: the sample ID is read from the id column, and every remaining column except the system-generated sample_id feeds the content hash. The hash is computed by the evaluation, so the dataset does not need a precomputed hash column.

Metrics

Sample ID Uniqueness

The fraction of records whose sample ID is present and appears exactly once (range: 0.0 to 1.0).

Sample ID Uniqueness
0.01.0
0.0
0.8
0.95
1.0
0.0No record has a unique sample ID, or the dataset has no ID column. Every record would collapse to the same identifier.
0.8Significant ID duplication - 20% or more of records share an identifier. Records cannot be addressed or deduplicated reliably.
0.95Minor ID duplication - 5% of records share an identifier. Impact depends on whether the affected records are critical to downstream processing.
1.0Every record has a present, unique sample ID.

Data Uniqueness

The fraction of records whose content hash appears exactly once in the dataset (range: 0.0 to 1.0).

Data Uniqueness
0.01.0
0.0
0.8
0.95
1.0
0.0Every record is a content duplicate of another record - the dataset contains no unique entries.
0.8Significant duplication - 20% or more of records are content duplicates. Deduplication is required before training or evaluation.
0.95Minor duplication - 5% of records are duplicates. Impact depends on dataset size and whether duplicates cluster around specific categories.
1.0No content duplicates detected - every record is unique across its non-ID columns.

Motivation

Duplicate records cause AI models to overfit to repeated examples, distort class distributions, and produce inflated evaluation metrics that do not reflect true generalisation performance. A record that appears ten times in training receives ten times the gradient signal of a unique record, silently skewing the learned decision boundaries. Because the content hash covers every non-ID column, exact duplicates are caught whether or not their identifiers differ.

Identifiers matter for the same reason. A dataset whose sample IDs repeat, or that has no identifier at all, cannot be deduplicated, joined, or traced reliably. Those two checks answer different questions, so the evaluation reports them as separate metrics rather than blending them into one score.

Methodology

  1. Samples: Each record in the dataset is scored against all other records in the dataset.
  2. Scoring: The Uniqueness Scorer reads the sample ID from the configured id_column_data_uniqueness (default id), then checks ID uniqueness and content uniqueness independently. The content hash is computed over a canonical JSON serialisation of the record's columns, excluding the ID column, any configured columns_to_exclude_data_uniqueness_hash, and sample_id. sample_id is always excluded because the platform adds it to every sample when the dataset does not provide one; to check a dataset's own sample_id column, set id_column_data_uniqueness to sample_id.

Each metric is the fraction of records that pass the corresponding check, averaged across the full dataset.

Scoring

Sample ID Uniqueness Scorer

Sample ID Uniqueness
Score valueExplanation
1.0The record's sample ID is present and no other record shares it.
0.0The record's sample ID is missing, or another record in the dataset uses the same ID.

Data Uniqueness Scorer

Data Uniqueness
Score valueExplanation
1.0No other record shares this record's content across its non-ID, non-excluded columns.
0.5The record's content matches one or more other records after the ID column is ignored. The records may be distinct entries that describe the same underlying object.
0.0The record is an exact content duplicate of another record and should be removed before training or evaluation.

Examples

Unique record - unique ID and unique content (passing)

Sample
idV-0041
makeBMW
modelX3
year2019
km_driven54200
dealer_idD-07
Sample ID Uniqueness Scorer
1.0

id "V-0041" does not appear on any other record in the dataset.

Data Uniqueness Scorer
1.0

No other record has the same make, model, year, mileage, and dealer combination. This record is unique by content.

Exact content duplicate with a different ID (failing)

Sample
idV-0099
makeBMW
modelX3
year2019
km_driven54200
dealer_idD-07
Sample ID Uniqueness Scorer
1.0

id "V-0099" is not used by any other record, even though the content matches record V-0041.

Data Uniqueness Scorer
0.0

Ignoring the id column, this record is identical to V-0041. The content hash collides, so the record is an exact duplicate.

Reused sample ID over different content (failing)

Sample
idV-0041
makeAudi
modelA4
year2021
km_driven31800
dealer_idD-02
Sample ID Uniqueness Scorer
0.0

id "V-0041" is already used by another record in the dataset. Identifiers must be unique even when the content differs.

Data Uniqueness Scorer
1.0

The content does not match any other record, so the record is unique by content. The failure here is the reused ID.

Run Evaluation in LatticeFlow AI Platform

Use the following CLI command to initialize and run the evaluation in LatticeFlow AI Platform.
Requires LatticeFlow AI Platform CLI
lf init eval --key atlas-data_uniqueness

Metrics

Sample ID UniquenessData Uniqueness

Don't have the LatticeFlow AI Platform?

Contact us to see this evaluation in action:
Contact Us