Thu Dec 28, 2023 10:04 am
Hello,
Thank you for your feedback.
Regarding your first question about table extraction from PDF files, we would like to clarify that PDF files themselves do not have a native concept of tables. In our product, we extract tables by first identifying the lines in the document, both horizontal and vertical. If multiple lines intersect, we consider it a "table," and the areas between adjacent vertical and horizontal lines are treated as "cells." We then compare the positions of the extracted text with the positions of these "cells" and place the text within the corresponding cell, thereby creating a table structure. However, in the case of the table shown in your provided screenshot, the absence of solid lines prevents proper recognition, and unfortunately, this issue cannot be currently resolved.
Regarding your second question about the error message you encountered, after investigating the matter, we found that Spire.OCR is currently incompatible with Spire.Office, which means you cannot use both components simultaneously in the same project. I've reported this issue to our development team, and it has been assigned the reference number SPIREOCR-47. Please note that Spire.OCR is an entirely separate component and is not attached to Spire.Office. However, we are working on implementing compatibility between Spire.OCR and Spire.Office, allowing you to use both components in the same project.
We apologize for any inconvenience caused by these limitations, and we appreciate your understanding.
Thank you for your patience and cooperation.
Sincerely,
Annika
E-iceblue support team