Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Wed Dec 27, 2023 10:20 am

Hello,
I'm using Spire.Office 8.12.0 version to extract data from a PDF file.
I just want to clarify 2 things about extracting data from a PDF file.

1. Table object
A file contains multiple tables of data. Using PdfDocument, I can extract only the data. But, not the table object.
Is it possible to get table-wise data?
For example in the image below, Key Credit Metrics, Capital Structure (1) and Yield Calibration Analysis are different tables. Can I have separate table data objects for these?
PDF-Tables.png


2. Scanned content
A file contains text data and scanned content. Is it possible to extract the scanned content using PdfDocument?

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Thu Dec 28, 2023 2:51 am

Hello,

Thank you for your inquiry.
Regarding your mentioned questions, here are the answers:
1. Extracting table content from PDF: We provide support for extracting the content of tables from PDF files. Please refer to the following sample code as an example:
Code: Select all
// Create a new instance of PdfDocument
PdfDocument pdfDocument = new PdfDocument();
// Load the PDF file
pdfDocument.loadFromFile("tableSample.pdf");
// Create a StringBuilder to store the extracted table data
StringBuilder builder = new StringBuilder();
// Create a PdfTableExtractor object with the PdfDocument
PdfTableExtractor extractor = new PdfTableExtractor(pdfDocument);
// Declare an array to store the extracted tables
PdfTable[] tableLists = null;
// Iterate through each page in the document
for (int pageIndex = 0; pageIndex < pdfDocument.getPages().getCount(); pageIndex++) {
    // Extract tables from the current page
    tableLists = extractor.extractTable(pageIndex);
    // Check if any tables were extracted
    if (tableLists != null && tableLists.length > 0) {
        // Iterate through each extracted table
        for (PdfTable table : tableLists) {
            // Get the number of rows and columns in the table
            int row = table.getRowCount();
            int column = table.getColumnCount();
            // Iterate through each cell in the table
            for (int i = 0; i < row; i++) {
                for (int j = 0; j < column; j++) {
                    // Get the text content of the current cell
                    String text = table.getText(i, j);
                    // Append the text to the StringBuilder
                    builder.append(text).append("  ");
                }
                builder.append("\r\n");
            }
        }
    }
}
// Create a FileWriter to write the extracted table data to a file
FileWriter fileWriter = new FileWriter("result.txt");
// Write the extracted table data from the StringBuilder to the file
fileWriter.write(builder.toString());
// Flush and close the FileWriter
fileWriter.flush();
fileWriter.close();

2. Extracting content from scanned PDF documents: Unfortunately, at this time, we do not directly support extracting content from scanned PDF documents. However, we offer a separate product called Spire.OCR for Java (Version: 1.9.2) that can help you extract text from images. You can download this product and refer to the tutorial "How to Scan and Recognize Text from Images in Java Projects" for testing and implementation.
If you have any further questions or need assistance, please feel free to ask.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Thu Dec 28, 2023 7:02 am

Hello,
Thank you for the information.

1. Table Extraction
I tried the sample code that you have shared to extract table content. It works for most of the cases. But in the image below, "Projected Cash Flows" are not identified as a table. Only the table with the borders is extracted as a table object. Could you please guide me on how can I get this table object?
Table.png


2. Scan OCR
I followed the steps mentioned in the tutorial and ended up in error on the code
Code: Select all
OcrScanner scanner = new OcrScanner();
. Will this product come along with Spire Office? Even I tried to request a free temporary license for OCR product for Java. But, this is not listed in the product name drop-down.

The actual error I'm getting is
Exception in thread "main" java.lang.NoSuchMethodError: 'com.spire.ocr.packages.sprrll com.spire.license.LicenseProvider.sprㆁ┩(java.lang.String)'
at com.spire.ocr.packages.sprytd.spr↮↮(Unknown Source)
at com.spire.ocr.OcrScanner.<init>(Unknown Source)
at com.cs.extract.ScanImageExtract.extractScannedContentFromPdf

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Thu Dec 28, 2023 10:04 am

Hello,

Thank you for your feedback.
Regarding your first question about table extraction from PDF files, we would like to clarify that PDF files themselves do not have a native concept of tables. In our product, we extract tables by first identifying the lines in the document, both horizontal and vertical. If multiple lines intersect, we consider it a "table," and the areas between adjacent vertical and horizontal lines are treated as "cells." We then compare the positions of the extracted text with the positions of these "cells" and place the text within the corresponding cell, thereby creating a table structure. However, in the case of the table shown in your provided screenshot, the absence of solid lines prevents proper recognition, and unfortunately, this issue cannot be currently resolved.
Regarding your second question about the error message you encountered, after investigating the matter, we found that Spire.OCR is currently incompatible with Spire.Office, which means you cannot use both components simultaneously in the same project. I've reported this issue to our development team, and it has been assigned the reference number SPIREOCR-47. Please note that Spire.OCR is an entirely separate component and is not attached to Spire.Office. However, we are working on implementing compatibility between Spire.OCR and Spire.Office, allowing you to use both components in the same project.
We apologize for any inconvenience caused by these limitations, and we appreciate your understanding.
Thank you for your patience and cooperation.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Thu Dec 28, 2023 1:39 pm

Hello,
Thank you for the detailed information.

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Fri Dec 29, 2023 1:35 am

Hello,

You're welcome!
Rest assured that once this issue is resolved, we will notify you immediately.
If you have any further questions or concerns, please don't hesitate to reach out to us.
Thank you for your understanding and support. Have a great day!

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Wed Jan 17, 2024 8:21 am

Hello,

Thank you for your patience.
We are pleased to inform you that we have just released Spire.Office for Java Version: 9.1.4, which addresses the issue SPIREOCR-47.
In the past, there was an incompatibility between Spire.OCR and Spire.Office. However, in this new version of Spire.Office, we have included Spire.OCR within it. As a result, the licensing for Spire.Office now covers Spire.OCR as well.
We invite you to download and test this latest version.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Wed Jan 17, 2024 10:21 am

Hello,
Thank you so much for the update.
I tested the Spire Office for Java's latest version for OCR scanning in a single project. It worked.

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Thu Jan 18, 2024 1:17 am

Hello,

You're welcome.
If you encounter other issues related to our products in the future, please feel free to contact us.
Have a nice day.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Thu Jan 18, 2024 11:07 am

Hello,
I give it a deep test with the latest version of Spire.Office 9.1.4 Many features are working as expected.

Previously the bug SPIREXLS-4996 has been fixed in version 8.12.0. But, this bug is reopened in the latest release.
Please look into this as soon as possible.
Thanks for your support.

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Fri Jan 19, 2024 1:19 am

Hello,

Thank you for your feedback.
I have tested the issue SPIREXLS-4996 using Spire.Office for Java Version: 9.1.4, but I was unable to reproduce the problem. In order to further investigate this matter, could you please provide the Excel file that you used for testing? You could attach it here or send it to us via email ([email protected]). Thanks in advance.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Fri Jan 19, 2024 6:09 am

Hello,

Thank you for providing the documents via email.
I have tested the Excel file you provided, and the program ran without any issues. I have sent you my testing results for your reference. Please check it.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Mon Jan 22, 2024 5:15 am

Hi,
Yes, it is working for the normal case

However, I have a requirement to copy the worksheet with another workbook. When I try to do that, I get the following error at the line
Code: Select all
ws.copyFrom(worksheet);


Exception in thread "main" java.lang.NullPointerException: Cannot invoke "com.spire.ms.System.Collections.Hashtable.get(Object)" because "copySetting.spr※" is null

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Mon Jan 22, 2024 6:32 am

Hello,

Thank you for your feedback.
Based on your description, I have retested the Excel file you provided using the following code and was able to reproduce the issue you mentioned. I have logged the issue into our bug tracking system with the ticket number SPIREXLS-5095. Our development team will investigate and fix it. Once it is resolved, I will inform you in time. Sorry for the inconvenience caused.
Code: Select all
Workbook workbook = new Workbook();
workbook.loadFromFile("Compass Health Financials March 2023 Distribution.xlsx");
Workbook workbook1 = new Workbook();
workbook1.getWorksheets().clear();
for (int i = 0; i < workbook.getWorksheets().getCount(); i++) {
    System.out.println(i);
    Worksheet worksheet = workbook.getWorksheets().get(i);
    workbook1.getWorksheets().add(worksheet.getName()).copyFrom(worksheet);
}
workbook1.saveToFile("result.xlsx");

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Mon Feb 26, 2024 12:12 pm

Hi,
Please update the status of this bug SPIREXLS-5095.
I tried with the latest version of Spire.Office (9.1.10). Still the same issue.

Usha.Thavasiappan
 
Posts: 40
Joined: Mon Nov 06, 2023 8:16 am

Return to Spire.PDF

cron