- SHA-256 hashing with multiprocessing (configurable workers) - Resumable via checkpoints (saves every N files, auto-resumes) - Reports exact matches, hash-only matches, same-name/different-content, unique - --move (quarantine to __duplicates__/) or --delete - Python 3.8+, stdlib only
78 lines
2.5 KiB
Markdown
78 lines
2.5 KiB
Markdown
# dedup
|
|
|
|
Find and remove duplicate files between two directories using SHA-256 hashing with parallel workers.
|
|
|
|
## Requirements
|
|
|
|
- Python 3.8+
|
|
- No external dependencies (standard library only)
|
|
|
|
## Usage
|
|
|
|
```bash
|
|
# Compare two directories (read-only, generates a report)
|
|
python3 dedup.py /path/to/backup /path/to/original
|
|
|
|
# Move duplicates to a __duplicates__/ quarantine folder
|
|
python3 dedup.py --move /path/to/backup /path/to/original
|
|
|
|
# Permanently delete duplicates
|
|
python3 dedup.py --delete /path/to/backup /path/to/original
|
|
|
|
# Control parallelism and checkpointing
|
|
python3 dedup.py --workers 8 --checkpoint-every 500 /large/backup /original
|
|
|
|
# Restrict to specific extensions
|
|
python3 dedup.py --ext .jpg,.png,.cr2 /path/to/backup /original
|
|
|
|
# Top-level files only (no recursion)
|
|
python3 dedup.py --no-subdirs /path/to/backup /original
|
|
|
|
# Remove checkpoint data and start fresh
|
|
python3 dedup.py --clean /path/to/backup /original
|
|
```
|
|
|
|
## Options
|
|
|
|
| Flag | Description |
|
|
|------|-------------|
|
|
| `--move` | Move duplicates to `__duplicates__/` folder (safe, reversible) |
|
|
| `--delete` | Permanently delete duplicates (irreversible) |
|
|
| `--no-subdirs` | Only compare top-level files |
|
|
| `--ext .jpg,.png` | Restrict to specific extensions (default: all media) |
|
|
| `--report PATH` | Write report to PATH (default: `dirA/duplicates_report.txt`) |
|
|
| `--workers N` | Number of parallel hash workers (default: CPU count) |
|
|
| `--checkpoint-every N` | Files between checkpoints, 0 to disable (default: 1000) |
|
|
| `--clean` | Remove checkpoint data and exit |
|
|
| `--version` | Show version |
|
|
|
|
## How It Works
|
|
|
|
1. **Phase 1** — Hash all files in the reference directory (dir B)
|
|
2. **Phase 2** — Hash all files in the source directory (dir A)
|
|
3. **Phase 3** — Compare hashes and generate a report
|
|
|
|
Files are classified as:
|
|
- **Exact matches** — same filename and same content hash
|
|
- **Hash matches** — same content, different filename
|
|
- **Same name, different content** — same filename, different content (not duplicates)
|
|
- **Unique to dirA** — no matching file in dirB
|
|
|
|
## Checkpointing
|
|
|
|
For large directories, progress is saved every `--checkpoint-every` files
|
|
(default: 1000) to `dirA/.dedup_checkpoint/`. If the process is interrupted,
|
|
re-run the same command to resume from where it left off.
|
|
|
|
Checkpoints are automatically cleaned after a successful run. Use `--clean`
|
|
to manually discard them.
|
|
|
|
## Output
|
|
|
|
A text report is written with sections for each match category, showing
|
|
file paths and SHA-256 hashes.
|
|
|
|
## License
|
|
|
|
MIT
|