qualityMetricsFingerprint — canonicalization spec
Status. Pre-Phase-1a planning artifact. Drafted 2026-05-28. Will be folded into openspec/changes/batch-level-inventory-and-transforms/design.md when Phase 1a kicks off.
Purpose. The merge match rule (Decision 14 in end-to-end-trading-flow.md) keys on qualityMetricsFingerprint. Without rigorous canonicalization, two clearly-same products will fail to merge because of trivial encoding differences (9 vs 9.0, key reordering, whitespace, Prisma client serialization). This spec defines exactly how the fingerprint is computed so the match rule is deterministic across:
- different deploys (Node version changes, Prisma upgrades)
- different write paths (user input vs admin tool vs system-generated)
- different times (re-fingerprinting a 6-month-old batch produces the same value)
Audience. Anyone implementing or modifying the fingerprint function in Phase 1a or beyond.
1. Why this matters
Postgres treats '{"moisture": 9}'::jsonb = '{"moisture": 9.0}'::jsonb as TRUE, but JSON.stringify produces '{"moisture":9}' vs '{"moisture":9.0}'. SHA-256 of those bytes gives two different fingerprints. The matching graph fragments silently. The user files a support ticket. The fix is a one-line data migration to re-fingerprint everything — but only if every batch's metrics survive the round-trip correctly.
Worse: rules that work by accident (string trimming via a library that happens to handle null defensively) break the day someone replaces the library. Spec it now, write it once, test it exhaustively.
2. Algorithm — overview
qualityMetrics: object → drop null + missing fields
→ for each field, normalize per its
MetricField schema type
→ sort keys alphabetically
→ serialize to compact JSON
(no whitespace)
→ SHA-256(UTF-8 bytes)
→ first 16 hex chars
────────────────────────────
qualityMetricsFingerprint: stringResult: a 16-char hex string. 64 bits of entropy. At 100M products, collision probability ~10⁻⁵ — acceptable for a match-key (collisions cause merges, not data loss; humans can detect + split).
3. Versioning
const FINGERPRINT_ALGO_VERSION = 'v1';Stored alongside the fingerprint on Product:
model Product {
// ...
qualityMetricsFingerprint String
qualityMetricsFingerprintVersion String // 'v1' today
}The match rule includes the version: two products with v1 fingerprints match if their fingerprints are equal AND both are v1. A future v2 (e.g., if we change normalization rules) doesn't accidentally merge into v1 products. Migration tool re-fingerprints all v1 → v2 in a batch job.
4. Per-type normalization rules
Quality metrics are typed per CHG-015's MetricField system (ENUM | NUMBER | YEAR | TEXT | DATE). The fingerprint algorithm uses the leaf's metric schema to know how to normalize each field.
4.1 NUMBER
raw → parsed JS Number → String(n) (shortest decimal repr)| Raw input | After parse | Canonical string |
|---|---|---|
9 | 9 | "9" |
9.0 | 9 | "9" |
9.00 | 9 | "9" |
9.10 | 9.1 | "9.1" |
9.1234567890123 | 9.123456789012 (lost precision) | "9.123456789012" |
"9" (string with number-y value) | rejected — must be Number | — |
null | dropped (see §4.7) | — |
Why String(Number(n)) over n.toFixed(2):
toFixed(2)would canonicalize9→"9.00", which forces a decimal even where one wasn't intended. Bad for grades like"4 suta".String(Number(n))gives the shortest representation that round-trips to the same Number. Industry standard (RFC 8259 / 8785).
Floating-point precision warning. 0.1 + 0.2 = 0.30000000000000004. In practice, quality metrics are user-entered (whole or one-decimal numbers like moisture %), so this isn't a concern. If a future write path computes a value from arithmetic, round-before-fingerprint at the call site.
4.2 ENUM
raw → lowercase → trim → validate against leaf's enum valuesENUM values come from the leaf's metric schema (e.g., pack_type: 'CARTON' | 'BAG' | 'JUTE_SACK' | 'PP_BAG'). User input may be display-cased; the fingerprint normalizes to lowercase.
| Raw input | Canonical |
|---|---|
"CARTON" | "carton" |
"Carton" | "carton" |
" carton " | "carton" |
"carton2" (invalid) | reject during validation — fingerprint never runs on invalid data |
Why lowercase + trim: consistent regardless of which UI surface wrote the value. Display layer can re-case for presentation.
4.3 TEXT
raw → trim → collapse internal whitespace runs → preserve caseFree-form text. Case matters here (a "Mithilanchal Special" might be different from "mithilanchal special" if it's a proper noun). Internal whitespace runs collapse to single spaces to handle copy-paste noise.
| Raw input | Canonical |
|---|---|
"Handpicked premium" | "Handpicked premium" |
" Handpicked premium " | "Handpicked premium" |
"" (empty after trim) | dropped (see §4.7) |
4.4 YEAR
raw → parse to integer → String(n)Integer year (e.g., 2025). Treated like NUMBER but always integer.
4.5 DATE
raw → ISO 8601 date-only normalized to UTC midnight → YYYY-MM-DD| Raw input | Canonical |
|---|---|
"2025-10-15" | "2025-10-15" |
"2025-10-15T14:30:00+05:30" | "2025-10-15" (timezone stripped — date-only) |
"15-Oct-2025" | reject (must be ISO) |
"2025-10-15T20:00:00Z" (which is 2025-10-16 IST) | "2025-10-15" — UTC date preserved, NOT timezone-converted |
Date semantics. Quality metrics like harvest_date represent a calendar date, not an instant. Two batches harvested "the same day" by users in IST and UTC should fingerprint identically. We pick UTC date as the canonical form for consistency.
4.6 BOOLEAN
raw → true | false → "true" | "false"Straightforward. JSON booleans serialize as true / false.
4.7 Null + missing + undefined — dropped
These three are treated identically: the field is excluded from the canonical form. Two products that have {moisture: 9} and {moisture: 9, grade: null} MUST fingerprint identically — null means "this field has no value," same as not having the field at all.
function shouldInclude(value: unknown): boolean {
if (value === null || value === undefined) return false;
if (typeof value === 'string' && value.trim() === '') return false;
return true;
}Why: in real-world data entry, "did not fill in" vs "filled in null" vs "field added in later schema version with no default" are all the same human concept. Treating them differently in the match key produces split rows for no good reason.
4.8 Arrays — preserve order, recurse
Arrays appear in quality metrics for fields like flavors: ['salt', 'pepper'] where order matters (label-list order). Two arrays with the same elements in different order are DIFFERENT for fingerprint purposes.
raw → for each element, apply per-type normalization → join in original orderIf a future metric needs order-independent set semantics, the schema declares it as a SET type (not yet in MetricFieldType but reserved). Fingerprint sorts SET elements before serializing.
4.9 Nested objects — recurse
Rare in qualityMetrics today, but the spec handles them: apply the same per-type normalization to each leaf, sort keys alphabetically, drop nulls.
5. Serialization
After per-field normalization, build the canonical JSON:
function canonicalize(normalized: Record<string, string>): string {
const sortedKeys = Object.keys(normalized).sort();
// Compact serialization — no spaces, no newlines.
return JSON.stringify(
Object.fromEntries(sortedKeys.map((k) => [k, normalized[k]]))
);
}No whitespace. {"grade":"7-suta+","moisture":"9"} not { "grade": "7-suta+", "moisture": "9" }. The serializer's default behaviour (JSON.stringify(obj) without indent argument) gives compact form.
Key ordering is alphabetic UTF-16 code-unit order — the default Array.prototype.sort() semantics. Stable across Node versions.
6. Hashing
import { createHash } from 'node:crypto';
function fingerprintBytes(canonical: string): string {
return createHash('sha256').update(canonical, 'utf8').digest('hex').slice(0, 16);
}update(canonical, 'utf8')— explicit encoding. UTF-8 is the only correct choice (matches Postgres column encoding, matches Buffer.from default)..digest('hex').slice(0, 16)— 16 hex chars = 64 bits. Sufficient for a match-key.
7. Reference implementation (Phase 1a will inline this)
// apps/api/src/lib/qualityMetricsFingerprint.ts
import { createHash } from 'node:crypto';
import type { MetricField } from '@prisma/client'; // CHG-015 schema
export const FINGERPRINT_ALGO_VERSION = 'v1' as const;
export interface FingerprintResult {
fingerprint: string; // 16-char hex
version: typeof FINGERPRINT_ALGO_VERSION;
canonical: string; // for debugging / migrations only
}
export function fingerprintMetrics(
metrics: Record<string, unknown>,
schema: MetricField[],
): FingerprintResult {
const schemaByKey = new Map(schema.map((f) => [f.key, f]));
const normalized: Record<string, string> = {};
for (const [key, value] of Object.entries(metrics)) {
if (!shouldInclude(value)) continue;
const field = schemaByKey.get(key);
if (!field) continue; // unknown field — defensive, drop
const canonical = normalizeByType(value, field);
if (canonical !== null) normalized[key] = canonical;
}
const sortedKeys = Object.keys(normalized).sort();
const canonical = JSON.stringify(
Object.fromEntries(sortedKeys.map((k) => [k, normalized[k]]))
);
const fingerprint = createHash('sha256')
.update(canonical, 'utf8')
.digest('hex')
.slice(0, 16);
return { fingerprint, version: FINGERPRINT_ALGO_VERSION, canonical };
}
function shouldInclude(v: unknown): boolean {
if (v === null || v === undefined) return false;
if (typeof v === 'string' && v.trim() === '') return false;
return true;
}
function normalizeByType(value: unknown, field: MetricField): string | null {
switch (field.type) {
case 'NUMBER':
return String(Number(value));
case 'YEAR':
return String(Math.trunc(Number(value)));
case 'ENUM':
return String(value).toLowerCase().trim();
case 'TEXT':
return String(value).trim().replace(/\s+/g, ' ');
case 'DATE':
return new Date(String(value)).toISOString().slice(0, 10);
case 'BOOLEAN':
return value ? 'true' : 'false';
default:
return null;
}
}8. Test fixtures (Phase 1a's unit tests should include these)
const FIXTURES = [
// Same metrics, different number form
{ in: { moisture: 9 }, out: 'expected_fp_A' },
{ in: { moisture: 9.0 }, out: 'expected_fp_A' }, // same
{ in: { moisture: 9.00 }, out: 'expected_fp_A' }, // same
// Key ordering doesn't matter
{ in: { moisture: 9, grade: '7-suta+' }, out: 'expected_fp_B' },
{ in: { grade: '7-suta+', moisture: 9 }, out: 'expected_fp_B' }, // same
// Null + missing collapse
{ in: { moisture: 9, grade: null }, out: 'expected_fp_A' },
{ in: { moisture: 9, grade: undefined }, out: 'expected_fp_A' }, // same
// Case normalization for ENUM
{ in: { pack_type: 'CARTON' }, out: 'expected_fp_C' },
{ in: { pack_type: 'carton' }, out: 'expected_fp_C' }, // same
{ in: { pack_type: 'Carton' }, out: 'expected_fp_C' }, // same
// TEXT preserves case but trims + collapses whitespace
{ in: { notes: 'Handpicked premium' }, out: 'expected_fp_D' },
{ in: { notes: ' Handpicked premium ' }, out: 'expected_fp_D' }, // same
{ in: { notes: 'handpicked premium' }, out: 'expected_fp_NOT_D' }, // different
// DATE — timezone stripped
{ in: { harvest_date: '2025-10-15' }, out: 'expected_fp_E' },
{ in: { harvest_date: '2025-10-15T14:30:00+05:30' }, out: 'expected_fp_E' }, // same
// Different leaf metrics → not the same
{ in: { moisture: 9 }, out: 'expected_fp_A' },
{ in: { moisture: 9, broken_pct: 2 }, out: 'expected_fp_F' }, // different
// BOOLEAN
{ in: { handpicked: true }, out: 'expected_fp_G' },
{ in: { handpicked: false }, out: 'expected_fp_NOT_G' }, // different
// Empty after normalization → same fingerprint as truly empty
{ in: {}, out: 'expected_fp_EMPTY' },
{ in: { grade: null }, out: 'expected_fp_EMPTY' }, // same
{ in: { grade: ' ' }, out: 'expected_fp_EMPTY' }, // same
];The expected values are populated by running the reference implementation once and recording the outputs. After that, the fixtures are the canonical record — if the algorithm ever produces a different value for the same input, the test fails loudly.
9. Migration story when rules change (v1 → v2)
If a future bug requires a fix to the algorithm:
- Bump
FINGERPRINT_ALGO_VERSIONto'v2'. - Add new normalization logic alongside (don't delete v1).
- Background migration: walk all Products, recompute fingerprint under v2, write to
qualityMetricsFingerprint+ bumpqualityMetricsFingerprintVersion = 'v2'. - During migration window, match rule reads BOTH versions for any given Product (transitional).
- When 100% are v2, drop the v1 read path.
This is exactly the pattern Postgres uses for its pg_stat_statements_info etc. Versioned, migratable, never silently changes.
10. What this spec deliberately does NOT cover
- Quality-metric VALIDATION — that's the job of
validateMetricsAgainstSchema(already exists in CHG-015). Fingerprint runs AFTER validation; bad data never reaches fingerprint. metricsCompleteflag computation — separate concern, lives on Product, computed bycomputeMetricsComplete.- Display formatting — fingerprint canonicalization is for matching, not display. UI uses its own formatting per leaf.
- Cross-leaf merging — leaf change always means new Product row (Decision 14). Fingerprint comparison only happens within a single leaf.
Drafted 2026-05-28 during Phase 0 backlog work. Will move into Phase 1a's openspec/changes/batch-level-inventory-and-transforms/design.md when that change is scaffolded.
