Hello,
We currently use Spire.PDF 5.8.16 in .NET Core. We use it to clean incoming PDFs for potentially malicious content by:
1) Load the input PdfDocument into memory and create a new output PdfDocument
2) Processing each page of the input file, saving it as a Drawing.Image in memory (200 dpi)
3) Loading a PdfBitMap from that image and calculating whether it needs to be resized
4) Using PdfPageBase.Canvas.DrawImage to draw the image loaded into memory
5) Save the output PdfDocument to a file
I have a scanned ~600KB PDF with about 50 pages that when cleaned this way, the output file size explodes to ~50MB. Way too much. When processing this file as described above, no resizing is done - the images are simply drawn in the original size of the document.
Now, I've looked into the ways you suggest compression, i.e. following your "C# How to compress PDF images" article on MSDN, but...:
* TryCompressImage doesn't seem available in either the newest version of Spire or in the .NET Core version.
* Reducing the quality of the images to 20 compromises the image quality quite a bit.
I've also looked at doc.CompressionLevel = PdfCompressionLevel.Best; - but that doesn't do much significant.
My question is, can this be done without compromising image quality too much? Or do you have any other suggestions for cleaning PDFs?
Thanks in advance.