This is a work in progress. Support for public domain books downloaded from:
- HathiTrust
- Archive.org/Microsoft
This removes the info page and watermarks from every page. Starting parameter is a directory path name, enclosed in quotes if there are any spaces.
All PDFs in that path, and any subpaths, will be processed. This will create a [filename]_clean.pdf for each PDF that has either the info page removed and/or watermarks removed.
Process a directory
remove_google_wm.py '/Downloaded PDFs'
- Remove prop pages on older Google PDFs
- Test more HathiTrust PDFs
- Test more archive.org/Microsoft PDFs