Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Jun 16, 2020 11:53 pm

The Arabic characters are not exported correctly for the PDFs , I think extraction values depend on the PDF font because when we are using the function PDFpage.ExtractText()

For PDF with font Calibri (Please see Calebri Arabic.pdf attached) the results as follows:
---------------------------------------------------------------------------------------------------------------------------------------
" Untitled Document Page 1 of 1 12/2/2020 :ﺍﻟﺘﺎﺭ\0ـﺦ ﺇ\0ﺼﺎﻝ ﻧﻘﺪ\0ﺔ ﺍﻟﻤ\0ﻠﻎ: ﻣﻠﺤﻮﻇﺔ \0ﻌﺘﺪ ﺑﻬﺬﺍ ﺇ\0ﺼﺎﻝ ﺣﺎﻝ ﻭﺟﻮﺩ ﺷﻄﺐ \0ﻌﺘﺪ ﺑﻬﺬﺍ ﺇ\0ﺼﺎﻝ ﻣﺎﻟﻢ \0ﻜﻦ ﻣﻤﻬﻮﺭﺍ \0 \0 ﺨﺘﻢ ﺩﻓﻊ ﻭﻣﻮﻗﻊ ﻣﻦ ﺍﻟﻤﻮﻇﻒ ﺍﻟﻤﺴﺌﻮﻝ ﺍﻟﻌﻤ\0ﻞ ﺍﻟﺴ\0ﺪ ﺷﺤﺎﺗﺔ ﺍﻷﻟ\0 ﺍﻻﻟ\0 ﻻ\0 ﻌﺘﺪ \0ﻪ\0 ﺇ\0ﺼﺎﻝ ﺳﺤﺐ ﻧﻘﺪﻯ ﺗﻮﻗﻴﻊ: ‐‐‐‐‐‐‐‐‐‐‐‐
---------------------------------------------------------------------------------------------------------------------------------------


For the same PDF with font Arial (Please see Arial Arabic.pdf attached) the results as follows:
-----------------------------------------------------------------------------------------------------------------------------------------
" Untitled Document Page 1 of 1 12/2/2020 :ﺍﻟﺗﺎﺭﻳﺦ ﺇﻳﺻﺎﻝ ﻧﻘﺩﻳﺔ ﺍﻟﻣﺑﻠﻎ: ﻣﻠﺣﻭﻅﺔ ﻳﻌﺗﺩ ﺑﻬﺫﺍ ﺇﻳﺻﺎﻝ ﺣﺎﻝ ﻭﺟﻭﺩ ﺷﻁﺏ ﻳﻌﺗﺩ ﺑﻬﺫﺍ ﺇﻳﺻﺎﻝ ﻣﺎﻟﻡ ﻳﻛﻥ ﻣﻣﻬﻭﺭﺍ ﺑﺧﺗﻡ ﺩﻓﻊ ﻭﻣﻭﻗﻊ ﻣﻥ ﺍﻟﻣﻭﻅﻑ ﺍﻟﻣﺳﺋﻭﻝ ﺍﻟﻌﻣﻳﻝ ﺍﻟﺳﻳﺩ ﺷﺣﺎﺗﺔ ﺍﻷﻟﻔﻰ ﺍﻻﻟﻔﻰ ﻻ ﻳﻌﺗﺩ ﺑﻪ ﺇﻳﺻﺎﻝ ﺳﺣﺏ ﻧﻘﺩﻯ ﺗﻭﻗﻳﻊ: ­­­­­­­­­­­­
------------------------------------------------------------------------------------------------------------------------------------------

As you see in the Calibri extraction a lot of Arabic characters are not extracted correctly unlike the Arial extraction so we need to resolve this issue or find a way to correct the extraction for any PDF with any Font

Appreciate your support

massoud.mohamed
 
Posts: 3
Joined: Tue Jun 16, 2020 11:24 pm

Wed Jun 17, 2020 2:32 am

Hello,

Thanks for your inquiry.
I tested your case with the latest Spire.PDF Pack(Hot Fix) Version:6.5.15, but didn't encounter the issue you mentioned, please see the attached "Calibri Arabic.txt". If you are using an older version, please try again with the latest version.
If the issue still occurs, to help us investigate further, please provide your OS information (E.g. Windows 7, 64bit) and region setting (E.g. China, Chinese). Thanks in advance.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Wed Jun 17, 2020 2:41 pm

Thank you for your response
Please note that the extracted Arabic text that you have sent for Calibri font is not extracted correctly as there a lot of Arabic characters are missing from the extraction

calibri font issue in arabic.png


For example the word “ايصال نقدية” is extracted as “ﺇ ﺼﺎﻝ ﻧﻘﺪ ﺔ” (character “يـ” is missing) and word “الألفي” is extracted as “ ﺍﻷﻟ” (characters “ف ي” are missing)

To help you to investigate, please be informed that this type of problem is not happen in Arial font so we need to find a way to standardize the extraction of Arabic characters across all fonts
Appreciate your support

massoud.mohamed
 
Posts: 3
Joined: Tue Jun 16, 2020 11:24 pm

Thu Jun 18, 2020 1:38 am

Hello,

Thanks for your more information.
I did notice that some characters were missing. This issue has been posted to our Dev team with the ticket SPIREPDF-3354 for further investigation. If there is any update, we will let you know.
Sorry for the inconvenience caused.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Thu Jun 18, 2020 5:31 pm

Thanks Rachel,

Please note that we are in the middle of a project that is totally based on extracting Arabic Documents. So even if this is an issue in the current release, we can accept any fast workaround solution that you recommend, and we have the technical capabilities to implement it.

I will be so grateful if you communicated with the development team to put this issue as a priority and kindly provide us with a time estimate of their response.

I urge your understanding and support for your fast response, as we are in a critical situation because of this issue and its implications.


Sincerely,
Mohamed Massoud

massoud.mohamed
 
Posts: 3
Joined: Tue Jun 16, 2020 11:24 pm

Fri Jun 19, 2020 7:55 am

Hello,

Thanks for your following up.
Regarding the issue of SPIREPDF-3354, after further analyzing, we found this is because that the Unicode value of some characters corresponds to a blank (invisible) symbol.
For example, for the character "يـ", when using the Calibri font, its Unicode value is 0000, which corresponds to a blank symbol, so it cannot be extracted as expected. While using the Arial font, the Unicode value of the character "يـ" is FEF3, which is not a blank symbol and can be extracted correctly. Also, you can directly use Adobe to extract the text, and you will get a similar result. Hope you can understand.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Fri Jun 19, 2020 10:44 pm

I am facing the same problem and it seems that Calibri font is not Arabic friendly so I have tried to use spire methods to replace it with different standard font however I got the following error when using the code block below
PdfUsedFont[] fonts = pdf.UsedFonts;
Font newfont = new Font("Arial", 11f);
Spire.Pdf.Graphics.PdfTrueTypeFont pf = new Spire.Pdf.Graphics.PdfTrueTypeFont(newfont, true);
foreach (PdfUsedFont font in fonts)
{
try
{
font.Replace(pf);

Error:
Spire.Pdf.Exceptions.PdfException: 'The font being replaced is not a standard font of Type 1 font or a non-embeded TrueType font

Please advise if I can resolve this or advise any other way to resolve this problem with Calibri font

Appreciate your support’

NasserTohamy
 
Posts: 19
Joined: Fri Jun 19, 2020 10:40 pm

Mon Jun 22, 2020 2:26 am

Hi Nasser,

Thanks for your post and sorry for the late reply as weekend.
To help us investigate your issue more accurately and quickly, please share your input file with us. You could send it to us([email protected]) via email.
Thanks in advance.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Mon Jun 22, 2020 3:09 pm

Hello Rachel
Thank you for your response
Please find attached sample PDF documents written in Arabic language in Calibri font and when I try to extract it , the characters are not appearing correctly ("جمهورية مصر العربية" extracted as "جمهور ة م الع ة") so I have tried to replace the font with spire font replace method however I got the following error

Error:
Spire.Pdf.Exceptions.PdfException: 'The font being replaced is not a standard font of Type 1 font or a non-embeded TrueType font


So I find a way to extract the text correctly for the Arabic Calibri font using Spire PDF or to find a way to replace this bad font with a more common font such as Times new roman for any Arabic document to be able to extract the text correctly or t

NasserTohamy
 
Posts: 19
Joined: Fri Jun 19, 2020 10:40 pm

Tue Jun 23, 2020 7:14 am

Hello,

Thanks for your sharing.
After analysis, we found that some fonts in your document are the embedded fonts that contain ToUnicode information. If we force to replace the embedded fonts with another font, it is most likely to cause mess of characters. Therefore, replacing such embedded fonts is unreachable in our Spire.PDF. Hope you can understand.
And regarding extracting text, as I mentioned above, this is because when using the Calibri font, the Unicode value of some characters corresponds to a blank (invisible) symbol. Sorry there is no good solution for your issue. If there is anything else we can do for you, please feel free to contact us.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Thu Jun 25, 2020 10:41 pm

I have reached a solution to replace the font with times new roman however I have faced another problem that the font extracted without spaces so when I extract the attached PDF the text extracted without spaces
For example the text
“جمهورية مصر العربية المتحدة”
Extracted as
“ﺟﻤﻬﻮﺭﻳﺔﻣﺼﺮﺍﻟﻌﺮﺑﻴﺔﺍﻟﻤﺘﺤﺪﺓ” [WITHOUT SPACES]
So we need to find a way to extract the text with the spaces as it appears in the PDF
Please extract the text using "PDFpage.ExtractText()" in the attached PDF and you will understand what I mean

NasserTohamy
 
Posts: 19
Joined: Fri Jun 19, 2020 10:40 pm

Fri Jun 26, 2020 3:11 am

Hi Nasser,

Thanks for your response.
I did an initial test and did notice the issue you mentioned. This issue has been posted to our Dev team with the ticket SPIREPDF-3373 for further investigation. If there is any update, we will inform you. Sorry for the inconvenience caused.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Tue Jun 30, 2020 11:09 am

Hi Nasser,

Hope you are doing well.
After analyzing the internal data of your PDF document, we found that the space in our visual effect is not really a space, it is just the spacing between two characters. And this spacing is too small that our PDF can't treat it as a space when extracting text. We are very sorry that there is no good way to resolve this issue at present.
If there is anything else we can do for you, please feel free to let us know.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Fri Sep 04, 2020 1:15 am

Hi Nasser,

Hope you are doing great.
I'm glad to tell you that our developers have improved the algorithm for extracting text, now the extracted text can retain the spaces as you expected. Please download the newly released Spire.PDF Pack(Hot Fix) Version:6.9.0 from the following links to test.
Website link: https://www.e-iceblue.com/Download/down ... t-now.html
Nuget link: https://www.nuget.org/packages/Spire.PDF/

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Thu Sep 10, 2020 8:38 am

Hi Nasser,

Greetings from E-iceblue!
Have you tested the hotfix? Thanks in advance for your feedback.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Return to Spire.PDF