Spots

Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's…

TL;DR We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 signature filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.

We'll build a dedup worker that catches re-encoded

We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 signature filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.

A dedup.py module with one entry point, check(path)

A dedup.py module with one entry point, check(path) -> Verdict, that returns exact, near, contains, or new, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't pip install or apt install.

The reason for three stages instead of one

The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced. If that line is missing, your FFmpeg build was configured without it; grab a static build. 2. Make some test inputs 🎬 We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free. sha256sum fixtures/*.mp4 gives you four different hashes for what a human would call two videos. That's the problem. 3. Stage 1: the byte hash (keep it, it's free)

This catches the accidental double upload and nothing

This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives. 4. Stage 2: perceptual hash + Hamming gate 🔍

videohash2 samples one frame per second, shrinks each

videohash2 samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.

You'll see the re-encode and the watermark land

You'll see the re-encode and the watermark land close to zero and unrelated land far away. The interesting one is excerpt: it will not be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.

⚠️ Note: pick the Hamming threshold conservatively and

⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.

⚠️ Note: pick the Hamming threshold conservatively and

⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence. 5. Stage 3: MPEG-7 signature for containment

FFmpeg's signature filter fingerprints every frame and, given

FFmpeg's signature filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and store a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate. Then compare two inputs in one run: Expected shape of the output (offsets will differ):

News

Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's MPEG-7 signature

TL;DR We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have.

@spots #dev
Source: Dev.to
See more like this