Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Nov 08, 2022 8:42 pm

We are extracting the text from this PDF using the PdfPageBase.ExtractText() command but the return is being characters in  format.

Could you check the attached PDF? Thanks.

RuiBarbosa
 
Posts: 19
Joined: Thu Apr 18, 2019 12:23 pm

Wed Nov 09, 2022 2:44 am

Hello,

Thanks for your inquiry.
After investigation, I reproduced your issue and logged it into our bug tracking system with the ticket number SPIREPDF-5607. Our development team will investigate and fix it. Once it is resolved, I will inform you in time. Sorry for the inconvenience caused.

Sincerely
Abel
E-iceblue support team
User avatar

Abel.He
 
Posts: 1010
Joined: Tue Mar 08, 2022 2:02 am

Tue Nov 15, 2022 9:27 am

Hello,

Hope you are doing well.
For the issue with the number SPIREPDF-5607, I have some updates to inform you:
This issue is not a bug in our product. When extracting text from Pdf document, what is actually extracted is the unicode encoding of text. In addition, the text display effect of Word document or txt document is according to the unicode encoding of text. However, the font of your pdf document (Boletos vencimento 11.2022.txt) don’t support unicode encoding, therefore, the extracted text cannot display its glyphs correctly. This
If you have any other issue, just feel free to contact us.

Sincerely
Abel
E-iceblue support team
User avatar

Abel.He
 
Posts: 1010
Joined: Tue Mar 08, 2022 2:02 am

Thu Nov 17, 2022 6:04 pm

Abel, thanks for the information.

Is there any way to identify this situation so that we can warn our user that the document cannot be processed?

RuiBarbosa
 
Posts: 19
Joined: Thu Apr 18, 2019 12:23 pm

Fri Nov 18, 2022 3:21 am

Hello,

Thanks for your feedback.
Sorry that our product doesn’t support determining whether the font of Pdf document support unicode encoding. However, for this pdf document (Boletos vencimento 11.2022.pdf), you can using the following code to achieve your requirement.

Code: Select all
    PdfDocument doc = new PdfDocument();
            // Read a pdf file
            string output = @"..\..\data\Boletos vencimento 11.2022.pdf";
            doc.LoadFromFile(output);

          //Get the fonts used in PDF
            PdfUsedFont[] fonts = doc.UsedFonts;

            //Travel the fonts used in PDF
            foreach (PdfUsedFont font in fonts)
            {
                if(font.Name == "F1"||font.Name=="F2"||font.Name=="F3")
                {
                    Console.WriteLine("Sorry,this pdf document does not support text extraction ");
                    break;
                }
            }


Sincerely
Abel
E-iceblue support team
User avatar

Abel.He
 
Posts: 1010
Joined: Tue Mar 08, 2022 2:02 am

Wed Nov 23, 2022 9:04 pm

Thanks Abel.

RuiBarbosa
 
Posts: 19
Joined: Thu Apr 18, 2019 12:23 pm

Thu Nov 24, 2022 1:19 am

Hello,

You're welcome! If you have any issue, just feel free to contact us. Have a nice day! :D

Sincerely
Abel
E-iceblue support team
User avatar

Abel.He
 
Posts: 1010
Joined: Tue Mar 08, 2022 2:02 am

Return to Spire.PDF

cron