Data Uniqueness
Overview
The Data Uniqueness evaluation measures how unique the samples in a dataset are. It runs two checks over the whole dataset at once. The first checks that each sample carries a present, unique identifier. The second checks that no two samples share the same content across the columns that define them. Each check produces a metric, so identifier problems and content duplication stay separate.
The evaluation works with no configuration: the sample ID is read from the id
column, and every remaining column except the system-generated sample_id feeds
the content hash. The hash is computed by the evaluation, so the dataset does not
need a precomputed hash column.
Metrics
Sample ID Uniqueness
The fraction of records whose sample ID is present and appears exactly once (range: 0.0 to 1.0).
Data Uniqueness
The fraction of records whose content hash appears exactly once in the dataset (range: 0.0 to 1.0).
Motivation
Duplicate records cause AI models to overfit to repeated examples, distort class distributions, and produce inflated evaluation metrics that do not reflect true generalisation performance. A record that appears ten times in training receives ten times the gradient signal of a unique record, silently skewing the learned decision boundaries. Because the content hash covers every non-ID column, exact duplicates are caught whether or not their identifiers differ.
Identifiers matter for the same reason. A dataset whose sample IDs repeat, or that has no identifier at all, cannot be deduplicated, joined, or traced reliably. Those two checks answer different questions, so the evaluation reports them as separate metrics rather than blending them into one score.
Methodology
- Samples: Each record in the dataset is scored against all other records in the dataset.
- Scoring: The Uniqueness Scorer reads the sample ID from the configured
id_column_data_uniqueness(defaultid), then checks ID uniqueness and content uniqueness independently. The content hash is computed over a canonical JSON serialisation of the record's columns, excluding the ID column, any configuredcolumns_to_exclude_data_uniqueness_hash, andsample_id.sample_idis always excluded because the platform adds it to every sample when the dataset does not provide one; to check a dataset's ownsample_idcolumn, setid_column_data_uniquenesstosample_id.
Each metric is the fraction of records that pass the corresponding check, averaged across the full dataset.
Scoring
Sample ID Uniqueness Scorer
Data Uniqueness Scorer
Examples
Unique record - unique ID and unique content (passing)
id "V-0041" does not appear on any other record in the dataset.
No other record has the same make, model, year, mileage, and dealer combination. This record is unique by content.
Exact content duplicate with a different ID (failing)
id "V-0099" is not used by any other record, even though the content matches record V-0041.
Ignoring the id column, this record is identical to V-0041. The content hash collides, so the record is an exact duplicate.
Reused sample ID over different content (failing)
id "V-0041" is already used by another record in the dataset. Identifiers must be unique even when the content differs.
The content does not match any other record, so the record is unique by content. The failure here is the reused ID.