We have a problem when trying to extract image from a PDF.
The images that we get from the page method are not complete : it's as if the image was translated outside of the canvas.
The extracted images have a transparent part, and some of the original image translated.
The page is rotated 270 degrees
Original PDF attached, as well as first and last image as extracted
we are using Spire Office for java, version 7.9.6
Code :
- Code: Select all
PdfPageCollection pages = document.getPages();
for (int i = 0; i < pages.getCount(); i++) {
PdfPageBase page = pages.get(i);
if (page != null) {
double pageWidth = page.getCanvas().getClientSize().getWidth();
double pageHeight = page.getCanvas().getClientSize().getHeight();
PdfPageRotateAngle rotation = page.getRotation();
// get infos on imges in page
PdfImageInfo[] infos = page.getImagesInfo();
if (infos != null) {
for (PdfImageInfo pdfImageInfo : infos) {
BufferedImage sourceImage = pdfImageInfo.getImage();
File output = new File("C:\\Users\\guipatry\\Documents\\Development\\ImageExtractor\\datas\\"
+ String.format("Image_%d_%d.png", i, pdfImageInfo.getIndex() ));
try {
ImageIO.write(sourceImage, "PNG", output);
} catch (IOException e) {
}
}
}
}
}
Regards
Guillaume PATRY