Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Sun Feb 23, 2020 6:26 am

Hello,

I am trying to get a unified inline text of each page of a pdf that has both text and images. Is there any way I can get the position of an image in the text so that I can insert the results of my ocr job inline for the text?

Thanks.

nlivni3846
 
Posts: 3
Joined: Thu Feb 20, 2020 10:54 am

Mon Feb 24, 2020 3:11 am

Hi,

Thanks for your inquiry.
Do you want to get the position of image that is relative to the text?
To help us investigate further, please provide your input PDF and the position information you want as well as the desired PDF result. You could upload it here or send to us([email protected]) via email.

Sincerely,
Betsy
E-iceblue support team
User avatar

Betsy.jiang
 
Posts: 3099
Joined: Tue Sep 06, 2016 8:30 am

Wed Feb 26, 2020 6:37 am

Attached is a good example, the first part of the page has text, then there is an image with Lorem ipsum, and then there is text after the image. I am doing ocr on the image itself but I want the extracted text to appear in between the top and bottom text. Is there a way to get a "cursor" of where a particular image breaks the text so that I can insert the extracted text into it?

nlivni3846
 
Posts: 3
Joined: Thu Feb 20, 2020 10:54 am

Wed Feb 26, 2020 9:38 am

Hi,

Thanks for your information.
Sorry that Spire.PDF doesn't support getting the position where a particular image breaks the text. We only support getting the position of the picture in the page.
Code: Select all
var location =  page.ImagesInfo[0].Bounds;


Sincerely,
Betsy
E-iceblue support team
User avatar

Betsy.jiang
 
Posts: 3099
Joined: Tue Sep 06, 2016 8:30 am

Wed Feb 26, 2020 10:41 am

Is there a way to get the text as paragraphs and their bounding boxes so that I can arrange everything in the order I want?

nlivni3846
 
Posts: 3
Joined: Thu Feb 20, 2020 10:54 am

Thu Feb 27, 2020 2:03 am

Hi,

Yes. Please refer to following code. Then do the adjustment according to your requirement by yourself.
Code: Select all
            PdfDocument pdf = new PdfDocument("Test.pdf");
            PdfTextFindCollection allTextFind = pdf.Pages[0].FindAllText();
            foreach (PdfTextFind find in allTextFind.Finds)
            {
                RectangleF buonds = find.Bounds;
            }


Sincerely,
Betsy
E-iceblue support team
User avatar

Betsy.jiang
 
Posts: 3099
Joined: Tue Sep 06, 2016 8:30 am

Return to Spire.PDF

cron