Skip to content

qualityMetricsFingerprint — canonicalization spec ​

Status. Pre-Phase-1a planning artifact. Drafted 2026-05-28. Will be folded into openspec/changes/batch-level-inventory-and-transforms/design.md when Phase 1a kicks off.

Purpose. The merge match rule (Decision 14 in end-to-end-trading-flow.md) keys on qualityMetricsFingerprint. Without rigorous canonicalization, two clearly-same products will fail to merge because of trivial encoding differences (9 vs 9.0, key reordering, whitespace, Prisma client serialization). This spec defines exactly how the fingerprint is computed so the match rule is deterministic across:

  • different deploys (Node version changes, Prisma upgrades)
  • different write paths (user input vs admin tool vs system-generated)
  • different times (re-fingerprinting a 6-month-old batch produces the same value)

Audience. Anyone implementing or modifying the fingerprint function in Phase 1a or beyond.


1. Why this matters ​

Postgres treats '{"moisture": 9}'::jsonb = '{"moisture": 9.0}'::jsonb as TRUE, but JSON.stringify produces '{"moisture":9}' vs '{"moisture":9.0}'. SHA-256 of those bytes gives two different fingerprints. The matching graph fragments silently. The user files a support ticket. The fix is a one-line data migration to re-fingerprint everything — but only if every batch's metrics survive the round-trip correctly.

Worse: rules that work by accident (string trimming via a library that happens to handle null defensively) break the day someone replaces the library. Spec it now, write it once, test it exhaustively.

2. Algorithm — overview ​

   qualityMetrics: object           → drop null + missing fields
                                    → for each field, normalize per its
                                      MetricField schema type
                                    → sort keys alphabetically
                                    → serialize to compact JSON
                                      (no whitespace)
                                    → SHA-256(UTF-8 bytes)
                                    → first 16 hex chars
   ────────────────────────────
        qualityMetricsFingerprint: string

Result: a 16-char hex string. 64 bits of entropy. At 100M products, collision probability ~10⁻⁵ — acceptable for a match-key (collisions cause merges, not data loss; humans can detect + split).

3. Versioning ​

   const FINGERPRINT_ALGO_VERSION = 'v1';

Stored alongside the fingerprint on Product:

prisma
model Product {
  // ...
  qualityMetricsFingerprint        String
  qualityMetricsFingerprintVersion String  // 'v1' today
}

The match rule includes the version: two products with v1 fingerprints match if their fingerprints are equal AND both are v1. A future v2 (e.g., if we change normalization rules) doesn't accidentally merge into v1 products. Migration tool re-fingerprints all v1 → v2 in a batch job.

4. Per-type normalization rules ​

Quality metrics are typed per CHG-015's MetricField system (ENUM | NUMBER | YEAR | TEXT | DATE). The fingerprint algorithm uses the leaf's metric schema to know how to normalize each field.

4.1 NUMBER ​

   raw  →  parsed JS Number  →  String(n)  (shortest decimal repr)
Raw inputAfter parseCanonical string
99"9"
9.09"9"
9.009"9"
9.109.1"9.1"
9.12345678901239.123456789012 (lost precision)"9.123456789012"
"9" (string with number-y value)rejected — must be Number—
nulldropped (see §4.7)—

Why String(Number(n)) over n.toFixed(2):

  • toFixed(2) would canonicalize 9 → "9.00", which forces a decimal even where one wasn't intended. Bad for grades like "4 suta".
  • String(Number(n)) gives the shortest representation that round-trips to the same Number. Industry standard (RFC 8259 / 8785).

Floating-point precision warning. 0.1 + 0.2 = 0.30000000000000004. In practice, quality metrics are user-entered (whole or one-decimal numbers like moisture %), so this isn't a concern. If a future write path computes a value from arithmetic, round-before-fingerprint at the call site.

4.2 ENUM ​

   raw  →  lowercase  →  trim  →  validate against leaf's enum values

ENUM values come from the leaf's metric schema (e.g., pack_type: 'CARTON' | 'BAG' | 'JUTE_SACK' | 'PP_BAG'). User input may be display-cased; the fingerprint normalizes to lowercase.

Raw inputCanonical
"CARTON""carton"
"Carton""carton"
" carton ""carton"
"carton2" (invalid)reject during validation — fingerprint never runs on invalid data

Why lowercase + trim: consistent regardless of which UI surface wrote the value. Display layer can re-case for presentation.

4.3 TEXT ​

   raw  →  trim  →  collapse internal whitespace runs  →  preserve case

Free-form text. Case matters here (a "Mithilanchal Special" might be different from "mithilanchal special" if it's a proper noun). Internal whitespace runs collapse to single spaces to handle copy-paste noise.

Raw inputCanonical
"Handpicked premium""Handpicked premium"
" Handpicked premium ""Handpicked premium"
"" (empty after trim)dropped (see §4.7)

4.4 YEAR ​

   raw  →  parse to integer  →  String(n)

Integer year (e.g., 2025). Treated like NUMBER but always integer.

4.5 DATE ​

   raw  →  ISO 8601 date-only normalized to UTC midnight  →  YYYY-MM-DD
Raw inputCanonical
"2025-10-15""2025-10-15"
"2025-10-15T14:30:00+05:30""2025-10-15" (timezone stripped — date-only)
"15-Oct-2025"reject (must be ISO)
"2025-10-15T20:00:00Z" (which is 2025-10-16 IST)"2025-10-15" — UTC date preserved, NOT timezone-converted

Date semantics. Quality metrics like harvest_date represent a calendar date, not an instant. Two batches harvested "the same day" by users in IST and UTC should fingerprint identically. We pick UTC date as the canonical form for consistency.

4.6 BOOLEAN ​

   raw  →  true | false  →  "true" | "false"

Straightforward. JSON booleans serialize as true / false.

4.7 Null + missing + undefined — dropped ​

These three are treated identically: the field is excluded from the canonical form. Two products that have {moisture: 9} and {moisture: 9, grade: null} MUST fingerprint identically — null means "this field has no value," same as not having the field at all.

ts
function shouldInclude(value: unknown): boolean {
  if (value === null || value === undefined) return false;
  if (typeof value === 'string' && value.trim() === '') return false;
  return true;
}

Why: in real-world data entry, "did not fill in" vs "filled in null" vs "field added in later schema version with no default" are all the same human concept. Treating them differently in the match key produces split rows for no good reason.

4.8 Arrays — preserve order, recurse ​

Arrays appear in quality metrics for fields like flavors: ['salt', 'pepper'] where order matters (label-list order). Two arrays with the same elements in different order are DIFFERENT for fingerprint purposes.

   raw  →  for each element, apply per-type normalization → join in original order

If a future metric needs order-independent set semantics, the schema declares it as a SET type (not yet in MetricFieldType but reserved). Fingerprint sorts SET elements before serializing.

4.9 Nested objects — recurse ​

Rare in qualityMetrics today, but the spec handles them: apply the same per-type normalization to each leaf, sort keys alphabetically, drop nulls.

5. Serialization ​

After per-field normalization, build the canonical JSON:

ts
function canonicalize(normalized: Record<string, string>): string {
  const sortedKeys = Object.keys(normalized).sort();
  // Compact serialization — no spaces, no newlines.
  return JSON.stringify(
    Object.fromEntries(sortedKeys.map((k) => [k, normalized[k]]))
  );
}

No whitespace. {"grade":"7-suta+","moisture":"9"} not { "grade": "7-suta+", "moisture": "9" }. The serializer's default behaviour (JSON.stringify(obj) without indent argument) gives compact form.

Key ordering is alphabetic UTF-16 code-unit order — the default Array.prototype.sort() semantics. Stable across Node versions.

6. Hashing ​

ts
import { createHash } from 'node:crypto';
function fingerprintBytes(canonical: string): string {
  return createHash('sha256').update(canonical, 'utf8').digest('hex').slice(0, 16);
}
  • update(canonical, 'utf8') — explicit encoding. UTF-8 is the only correct choice (matches Postgres column encoding, matches Buffer.from default).
  • .digest('hex').slice(0, 16) — 16 hex chars = 64 bits. Sufficient for a match-key.

7. Reference implementation (Phase 1a will inline this) ​

ts
// apps/api/src/lib/qualityMetricsFingerprint.ts

import { createHash } from 'node:crypto';
import type { MetricField } from '@prisma/client';  // CHG-015 schema

export const FINGERPRINT_ALGO_VERSION = 'v1' as const;

export interface FingerprintResult {
  fingerprint: string;          // 16-char hex
  version: typeof FINGERPRINT_ALGO_VERSION;
  canonical: string;            // for debugging / migrations only
}

export function fingerprintMetrics(
  metrics: Record<string, unknown>,
  schema: MetricField[],
): FingerprintResult {
  const schemaByKey = new Map(schema.map((f) => [f.key, f]));
  const normalized: Record<string, string> = {};

  for (const [key, value] of Object.entries(metrics)) {
    if (!shouldInclude(value)) continue;
    const field = schemaByKey.get(key);
    if (!field) continue;  // unknown field — defensive, drop
    const canonical = normalizeByType(value, field);
    if (canonical !== null) normalized[key] = canonical;
  }

  const sortedKeys = Object.keys(normalized).sort();
  const canonical = JSON.stringify(
    Object.fromEntries(sortedKeys.map((k) => [k, normalized[k]]))
  );
  const fingerprint = createHash('sha256')
    .update(canonical, 'utf8')
    .digest('hex')
    .slice(0, 16);

  return { fingerprint, version: FINGERPRINT_ALGO_VERSION, canonical };
}

function shouldInclude(v: unknown): boolean {
  if (v === null || v === undefined) return false;
  if (typeof v === 'string' && v.trim() === '') return false;
  return true;
}

function normalizeByType(value: unknown, field: MetricField): string | null {
  switch (field.type) {
    case 'NUMBER':
      return String(Number(value));
    case 'YEAR':
      return String(Math.trunc(Number(value)));
    case 'ENUM':
      return String(value).toLowerCase().trim();
    case 'TEXT':
      return String(value).trim().replace(/\s+/g, ' ');
    case 'DATE':
      return new Date(String(value)).toISOString().slice(0, 10);
    case 'BOOLEAN':
      return value ? 'true' : 'false';
    default:
      return null;
  }
}

8. Test fixtures (Phase 1a's unit tests should include these) ​

ts
const FIXTURES = [
  // Same metrics, different number form
  { in: { moisture: 9 },       out: 'expected_fp_A' },
  { in: { moisture: 9.0 },     out: 'expected_fp_A' },  // same
  { in: { moisture: 9.00 },    out: 'expected_fp_A' },  // same

  // Key ordering doesn't matter
  { in: { moisture: 9, grade: '7-suta+' }, out: 'expected_fp_B' },
  { in: { grade: '7-suta+', moisture: 9 }, out: 'expected_fp_B' },  // same

  // Null + missing collapse
  { in: { moisture: 9, grade: null },     out: 'expected_fp_A' },
  { in: { moisture: 9, grade: undefined }, out: 'expected_fp_A' },  // same

  // Case normalization for ENUM
  { in: { pack_type: 'CARTON' }, out: 'expected_fp_C' },
  { in: { pack_type: 'carton' }, out: 'expected_fp_C' },  // same
  { in: { pack_type: 'Carton' }, out: 'expected_fp_C' },  // same

  // TEXT preserves case but trims + collapses whitespace
  { in: { notes: 'Handpicked premium' },     out: 'expected_fp_D' },
  { in: { notes: ' Handpicked  premium ' }, out: 'expected_fp_D' },  // same
  { in: { notes: 'handpicked premium' },    out: 'expected_fp_NOT_D' },  // different

  // DATE — timezone stripped
  { in: { harvest_date: '2025-10-15' }, out: 'expected_fp_E' },
  { in: { harvest_date: '2025-10-15T14:30:00+05:30' }, out: 'expected_fp_E' },  // same

  // Different leaf metrics → not the same
  { in: { moisture: 9 },                    out: 'expected_fp_A' },
  { in: { moisture: 9, broken_pct: 2 },     out: 'expected_fp_F' },  // different

  // BOOLEAN
  { in: { handpicked: true },  out: 'expected_fp_G' },
  { in: { handpicked: false }, out: 'expected_fp_NOT_G' },  // different

  // Empty after normalization → same fingerprint as truly empty
  { in: {},                  out: 'expected_fp_EMPTY' },
  { in: { grade: null },     out: 'expected_fp_EMPTY' },  // same
  { in: { grade: '   ' },    out: 'expected_fp_EMPTY' },  // same
];

The expected values are populated by running the reference implementation once and recording the outputs. After that, the fixtures are the canonical record — if the algorithm ever produces a different value for the same input, the test fails loudly.

9. Migration story when rules change (v1 → v2) ​

If a future bug requires a fix to the algorithm:

  1. Bump FINGERPRINT_ALGO_VERSION to 'v2'.
  2. Add new normalization logic alongside (don't delete v1).
  3. Background migration: walk all Products, recompute fingerprint under v2, write to qualityMetricsFingerprint + bump qualityMetricsFingerprintVersion = 'v2'.
  4. During migration window, match rule reads BOTH versions for any given Product (transitional).
  5. When 100% are v2, drop the v1 read path.

This is exactly the pattern Postgres uses for its pg_stat_statements_info etc. Versioned, migratable, never silently changes.

10. What this spec deliberately does NOT cover ​

  • Quality-metric VALIDATION — that's the job of validateMetricsAgainstSchema (already exists in CHG-015). Fingerprint runs AFTER validation; bad data never reaches fingerprint.
  • metricsComplete flag computation — separate concern, lives on Product, computed by computeMetricsComplete.
  • Display formatting — fingerprint canonicalization is for matching, not display. UI uses its own formatting per leaf.
  • Cross-leaf merging — leaf change always means new Product row (Decision 14). Fingerprint comparison only happens within a single leaf.

Drafted 2026-05-28 during Phase 0 backlog work. Will move into Phase 1a's openspec/changes/batch-level-inventory-and-transforms/design.md when that change is scaffolded.

Last updated:

Internal technical documentation — Cropto