From DedupTool to hdw-dedup-engine: Turning a Duplicate Detector into a Reusable Python Package

DedupTool started as an application. Its duplicate-detection engine has now become infrastructure

Over time, the same image-deduplication logic began serving more than one of my projects. DedupTool obviously needed it, but ChronoName also uses duplicate detection as part of its photo-archive workflow, while the Person Recognition desktop app uses it to protect its recognition knowledge base from redundant image copies.

Keeping that engine inside an application repository increasingly created the wrong dependency:

wrong dependency

ChronoName did not really depend on DedupTool,  it just depended on its duplicate-detection engine. So, the solution was to extract that engine from the tool and into its own Python package:

right dependecy

The result is hdw-dedup-engine 0.1.0: an independently installable Python package for image duplicate detection, with its own versioning, regression suite, build artifacts and GitHub repository. As a regular Python package, it’s hosted on PyPI.  

Why extract the engine?

The first versions of this code were never designed as a library. Duplicate detection was simply functionality inside DedupTool. That was perfectly reasonable while DedupTool was its only consumer.

The architecture became less convincing when ChronoName started using the same functionality.

ChronoName is a photo-archive utility that renames photos and videos according to their capture timestamp so that files remain chronologically sortable without depending on a particular photo-management database. Duplicate detection is useful there too, particularly when cleaning up archives containing camera originals, exported images, messaging-app copies and compressed derivatives.

Copying the duplicate logic into ChronoName would have created two implementations.

Importing it directly from the DedupTool source tree avoided that duplication, but introduced another problem: an application was now being used as a library.

Something like D:\Coding\ChronoName imports D:\Coding\DedupTool\deduptool works on one development machine. However, it is not a clean package boundary. So, the engine needed its own identity.

The new package boundary

The repository is now structured as a conventional Python package:

hdw-dedup-engine package

The distribution and import names follow the usual Python convention:

PyPI distribution: hdw-dedup-engine

Python package:  hdw_dedup_engine

Installation is therefore simply: pip install hdw-dedup-engine

Rather than reaching into another application’s source directory, application code can now work against an explicit package API:

from hdw_dedup_engine import (
    DedupConfig,
    DedupRunOptions,
    plan_duplicates,
)
Python

What is actually inside the engine?

The package contains the reusable parts of the duplicate-detection system, not the DedupTool-specific workflows or its GUI.

The package includes the core pipeline for:

  • file discovery and feature extraction;
  • perceptual hashing;
  • exact hashing;
  • duplicate matching;
  • clustering;
  • keeper selection;
  • SQLite feature caching;
  • reporting;
  • quarantine/move actions;
  • configuration and capability detection.

All application-specific functionality remains outside the package.

For example, DedupTool still owns its PyQt interface, Qt worker integration, application paths, logging and executable packaging, besides using the packages for its core business: finding duplicates.

ChronoName similarly remains responsible for its own photo-renaming workflow, journaling, undo logic, filing audit, ExifTool integration and user interface, using the engine package for cleaning purposes.

The Person Recognition App has its own functions like searching for people or finding look-alikes in photo collections while using the dedup-engine to prevent low-res duplicates from degrading its knowledge base.  

Thus, hdw-dedup-engine is not DedupTool without a GUI, it is the common engine underneath different applications that need duplicate-image analysis each for their own purposes.

The matching policy came with it

Extracting the engine also meant preserving some lessons learned while developing DedupTool. One of the more interesting was a real false-positive produced by dHash.

Two completely unrelated images — a landscape viewed through a car window and a photograph of a document — ended up only four bits apart in dHash space.

An earlier implementation effectively treated a sufficiently close dHash as proof of duplication. That pair therefore matched. The current engine instead treats dHash as evidence that requires corroboration. Conceptually:

checks & controls

The package retains the hardened policy in which dHash is no longer sufficient evidence by itself: candidate matches must also satisfy the pHash and wHash corroboration thresholds, with SSIM available as an optional additional check. This was particularly important for the package extraction because a false positive is more dangerous once pairwise matches feed clustering. One incorrect edge can potentially connect otherwise separate groups.

Choosing the keeper is part of the problem too

Finding a group of duplicate images is only half of a practical deduplication system.

The next question is: Which copy should be kept? That also evolved during DedupTool development.

An earlier ranking gave image sharpness too much influence. In practice, Laplacian-variance sharpness can reward compression artifacts and other characteristics that do not necessarily indicate the best source image. Therefore, the current keeper ranking instead prioritizes:

keeper policy

Sharpness remains useful, but as a late tie-breaker rather than the dominant signal.

That policy now lives in one package and is therefore shared consistently by every application using the engine.

Caching matters once the archive gets large

Perceptual hashing thousands of images repeatedly is wasteful when most of the archive has not changed. The package therefore includes its SQLite-backed feature index.

On a test archive of roughly 4,700 files, an uncached scan previously took about 17.1 seconds. After the feature cache had been populated repeating the scan took roughly 1.7 seconds. Adding several new files meant the next run could reuse the existing 4,712 cached feature records and hash only a few new files.

The important part is that the cache is a reusable engine capability, while each consuming application remains responsible for where its database and other application data live.

Maintenance at package level

Improvements to matching, keeper selection, caching or other shared functionality now happen in one place. Implement and test the change once in hdw-dedup-engine, and every consuming application can benefit from it without maintaining its own copy of the algorithm.

DedupTool became the first real consumer

After establishing the standalone package and its regression baseline, DedupTool itself was migrated. That may sound slightly backwards: the engine originated in DedupTool, now it is imported as an external dependency. But that is precisely the architecture I wanted.

I removed the old embedded engine after the application passed its integration tests.

The Windows PyInstaller application was also rebuilt with hdw_dedup_engine bundled into the executable. So, users of the desktop application do not need to install Python or manually install the engine package. The package separation is an architectural and developer-facing improvement; the packaged application remains self-contained.

ChronoName and Person Recognition also become consumers

The next migration was ChronoName. Instead of depending on DedupTool’s source tree it imported the hdw_dedup_engine package containing the core functions needed. ChronoName can use duplicate analysis and quarantine functionality without knowing anything about DedupTool’s GUI, CLI or application structure. Likewise, future changes to DedupTool itself no longer touch ChronoName.

For the Managing Persons in Photo Collections desktop app the same holds true: just import from the package the needed functions. No need to know anything about DedupTool.

Optional capabilities stay optional

I also wanted to avoid turning a relatively focused image-processing library into a huge dependency installation. The core engine primarily depends on NumPy and Pillow. Additional functionality can use optional packages such as: pillow-heif, scikit-image, send2trash.

SSIM, for example, remains an optional additional check rather than a mandatory part of the core matching policy. Applications that need it can install the scikit-image dependency, while applications that do not need SSIM avoid inheriting that heavier scientific-computing dependency stack.

There is now one implementation of duplicate detection, one regression suite protecting its behaviour while allowing for optional parts, and multiple applications that can evolve independently around it.

Validation & Testing

We validated the extraction at two levels. The engine itself has a 79-test regression suite protecting matching, clustering, keeper selection, caching, reporting and actions. Each consuming application then tests its own integration boundary. DedupTool added 17 application-level integration tests, while the Person Recognition migration added nine consumer tests plus an end-to-end GUI smoke test covering report-only detection and real quarantine on disposable data.

hdw-dedup-engine 0.1.0

The first package release deliberately starts at 0.1.0. The underlying duplicate-detection code is older and I have already been exercising it extensively, but the package itself — its public API, independent versioning and distribution contract — is new. Starting at 0.1.0 makes that distinction explicit.

The next steps are less about adding algorithms and more about proving the package boundary through real consumers:

✓ Extract engine;

✓ Build regression suite;

✓ Package wheel & source distribution;

✓ Migrate DedupTool;

✓ Migrate ChronoName;

✓ Migrate Person Recognition;

→ Let future consumers validate the package boundary;

→ Evolve the public API from real usage rather than speculation.

That is also why I resisted trying to design a large abstract library API upfront. The package API should evolve from actual consumers.

A summary of this article has also been publish on LinkedIn.

Related Stories