Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Thu Mar 11, 2021 12:12 pm

Hello,

I am trying to change the used fonts in a PDF file, but I am getting an error message saying that: {Spire.Pdf.Exceptions.PdfException: 'The font being replaced is not a standard font of Type 1 font or a non-embeded TrueType font'}

I am trying to change the fonts because my problem is that the extracted text is not the same as the original pdf text. However, when I use Acrobat Reader DC to change the text font (from 'Times New Roman' to 'Traditional Arabic'), the extracted text matches the original PDF text with no problems. So, I am trying to change the fonts in the entire PDF programmatically since I have a 1000 page PDF document of Arabic text.

abd_rix_0
 
Posts: 5
Joined: Thu Mar 11, 2021 11:55 am

Fri Mar 12, 2021 9:57 am

Hello,

Thanks for your inquiry.
I simulated a PDF file which used "Times New Roman" and replaced the font with "Traditional Arabic". But I didn't reproduce your issue. If you were using an old Spire.PDF version, I suggest that you try again with our latest Spire.PDF Pack(Hot Fix) Version:7.3.1. If your issue still exists, to help us investigate further, please provide your input files and full testing code for reference. Thanks in advance.

Sincerely,
Elena
E-iceblue support team
User avatar

Elena.Zhang
 
Posts: 279
Joined: Thu Jul 23, 2020 1:18 am

Sun Mar 14, 2021 3:19 pm

Hello Elena,

Thank you for your reply. The PDF file is attached.
The letters are extracted just fine. The problems is with the digits.
As for the code, I am using the exact same code from the tutorial. Here it is:
Code: Select all
Private Sub btnTest_Click (sender As Object, e As EventArgs) Handles btnTest.Click
        Dim doc As New PdfDocument()
        doc.LoadFromFile("D:\Tester\967.pdf")

        Dim fonts As PdfUsedFont() = doc.UsedFonts
        Dim newfont As New PdfFont(PdfFontFamily.Courier, 11.0F, PdfFontStyle.Italic Or PdfFontStyle.Bold)
     
        For Each fnt As PdfUsedFont In fonts
               fnt.Replace(newfont)
        Next

        doc.SaveToFile("D:\Tester\967_output.pdf")
End Sub

I noticed that when I use Acrobat DC to change the font from Times New Roman to any other font, and then back to Times New Roman, and then save the file, the problem disappears. There is something going on with the attached fonts that I do not quite understand.

abd_rix_0
 
Posts: 5
Joined: Thu Mar 11, 2021 11:55 am

Mon Mar 15, 2021 2:49 am

Hello,

Thanks for your sharing and sorry for the late reply as the weekend.
Sorry to tell you that our Spire.Pdf currently only supports replacing standard font of type 1 and non-emmbeded true type font. I checked your document and noticed that the font you want to replace is embedded. Since this embedded font contains ToUnicode information, if we force to replace the embedded font with another font, it is most likely to cause mess of characters. Therefore, replacing this embedded font is unreachable in our Spire.PDF. Hope you can understand.

Sincerely,
Elena
E-iceblue support team
User avatar

Elena.Zhang
 
Posts: 279
Joined: Thu Jul 23, 2020 1:18 am

Mon Mar 15, 2021 6:09 am

Thank you Elena,

I totally understand the issue with replacing customized fonts that contain ToUnicode segment.
Is there a way to read the Unicode stream as Unicode values and not as characters? This will help solve the issue.

Currently, I am using the following code to extract the text from a PDF page:
Code: Select all
    Dim extractedText as string = SpirePDFdoc.Pages(pageNumber).FindAllText()

abd_rix_0
 
Posts: 5
Joined: Thu Mar 11, 2021 11:55 am

Mon Mar 15, 2021 10:04 am

Hello,

Thanks for your feedback.
You can refer to the following code to extract the text and then convert the string to Unicode values.
Code: Select all
    Sub Main()
        Dim doc As PdfDocument = New PdfDocument
        doc.LoadFromFile("967.pdf")
        Dim fonts() As PdfUsedFont = doc.UsedFonts
        Dim textFindCollection As String = doc.Pages(0).ExtractText
        Dim unicode As String = String2Unicode(textFindCollection)
    End Sub

    Function String2Unicode(ByVal source As String) As String
        Dim bytes = Encoding.Unicode.GetBytes(source)
        Dim stringBuilder = New StringBuilder
        Dim i = 0
        Do While (i < bytes.Length)
            stringBuilder.AppendFormat("\u{0:x2}{1:x2}", bytes((i + 1)), bytes(i))
            i = (i + 2)
        Loop
        Return stringBuilder.ToString
    End Function


Or if I misunderstood, please describe your requirements in detail. Thanks in advance.

Sincerely
Elena
E-iceblue support team
User avatar

Elena.Zhang
 
Posts: 279
Joined: Thu Jul 23, 2020 1:18 am

Mon Mar 15, 2021 11:25 am

Hello Elena,

Thank you for the prompt reply. Actually, I was asking if I could know the actual Unicode values stored in the PDf file before they being mapped to characters and stored in a string variable. This way I would be able to create my own characters mapping to recreate the correct text. The code you provided, however, finds the Unicode values of the extracted messed-up text.

For example, in the PDf file I attached earlier, there is a number '0381659' in Hindi digits, which is extracted as '9561637'. This is probably because the embedded font has a customized segment that is not recognized by the OS, and it is mapped to other glyphs following a rule that I am not aware of. So, when this number is extracted, it is stored as '9561637' and not as '0381659'. As a result, the Unicode of the extracted text will be different than the original Unicode stored inside the PDF, and I won't be able to recreate the original text.

Kind regards,

abd_rix_0
 
Posts: 5
Joined: Thu Mar 11, 2021 11:55 am

Tue Mar 16, 2021 10:03 am

Hello,

Thanks for your feedback.
Regarding the number '0381659' in Hindi digits you mentioned, are you referring to the text in the image below? I opened your file using multiple PDF readers and then tried to copy this text and paste it, but found that all the values displayed were " 9561637" instead of " 0381659". I also tried to extract the text using Adobe and the text was also extracted as " 9561637". You could verify this on your side.
Screenshot.png

I’m afraid that this issue may be related to your PDF file itself, rather than our Spire.PDF extracting the messed-up text. Hope you can understand.

Sincerely,
Elena
E-iceblue support team
User avatar

Elena.Zhang
 
Posts: 279
Joined: Thu Jul 23, 2020 1:18 am

Tue Mar 16, 2021 2:31 pm

Thank you Elena,

You are being very helpful with your prompt and thorough replies.

I know that the problem is not from Spire.PDF at all. Actually, I tried all the ways you mentioned and it gave me the same results. So, when you copy the text from the PDF file and paste in any word processor, the text will be messed-up. However, if you try to open the PDF in edit mode and try to copy that text, it will paste just fine in any word processor. Moreover, if you open the PDF in edit mode and change the font of the mentioned text to any other font and save the file; then when you copy it, it will paste just fine as well. This lead me to believe that the correct text could be extracted by working around it.

I am 99% sure that the problem lies in the embedded fonts inside the PDF. That is why I am looking for a way to read the Unicode values directly from the PDF, and I was wondering if Spire.PDF allows this level of access to the developer.

Kind regards,

abd_rix_0
 
Posts: 5
Joined: Thu Mar 11, 2021 11:55 am

Wed Mar 17, 2021 9:08 am

Hello,

Thanks for your response.
I did a test based on your description, and I did notice the behavior you said. I discussed this with our developers, and the response I got was that Adobe may have different mechanisms for edit mode and view mode, so the copied text content is different. Sorry that it is not possible to get the Unicode value you expect with our Spire.PDF.

Sincerely,
Elena
E-iceblue support team
User avatar

Elena.Zhang
 
Posts: 279
Joined: Thu Jul 23, 2020 1:18 am

Return to Spire.PDF