The Arabic characters are not exported correctly for the PDFs , I think extraction values depend on the PDF font because when we are using the function PDFpage.ExtractText()
For PDF with font Calibri (Please see Calebri Arabic.pdf attached) the results as follows:
---------------------------------------------------------------------------------------------------------------------------------------
" Untitled Document Page 1 of 1 12/2/2020 :ﺍﻟﺘﺎﺭ\0ـﺦ ﺇ\0ﺼﺎﻝ ﻧﻘﺪ\0ﺔ ﺍﻟﻤ\0ﻠﻎ: ﻣﻠﺤﻮﻇﺔ \0ﻌﺘﺪ ﺑﻬﺬﺍ ﺇ\0ﺼﺎﻝ ﺣﺎﻝ ﻭﺟﻮﺩ ﺷﻄﺐ \0ﻌﺘﺪ ﺑﻬﺬﺍ ﺇ\0ﺼﺎﻝ ﻣﺎﻟﻢ \0ﻜﻦ ﻣﻤﻬﻮﺭﺍ \0 \0 ﺨﺘﻢ ﺩﻓﻊ ﻭﻣﻮﻗﻊ ﻣﻦ ﺍﻟﻤﻮﻇﻒ ﺍﻟﻤﺴﺌﻮﻝ ﺍﻟﻌﻤ\0ﻞ ﺍﻟﺴ\0ﺪ ﺷﺤﺎﺗﺔ ﺍﻷﻟ\0 ﺍﻻﻟ\0 ﻻ\0 ﻌﺘﺪ \0ﻪ\0 ﺇ\0ﺼﺎﻝ ﺳﺤﺐ ﻧﻘﺪﻯ ﺗﻮﻗﻴﻊ: ‐‐‐‐‐‐‐‐‐‐‐‐
---------------------------------------------------------------------------------------------------------------------------------------
For the same PDF with font Arial (Please see Arial Arabic.pdf attached) the results as follows:
-----------------------------------------------------------------------------------------------------------------------------------------
" Untitled Document Page 1 of 1 12/2/2020 :ﺍﻟﺗﺎﺭﻳﺦ ﺇﻳﺻﺎﻝ ﻧﻘﺩﻳﺔ ﺍﻟﻣﺑﻠﻎ: ﻣﻠﺣﻭﻅﺔ ﻳﻌﺗﺩ ﺑﻬﺫﺍ ﺇﻳﺻﺎﻝ ﺣﺎﻝ ﻭﺟﻭﺩ ﺷﻁﺏ ﻳﻌﺗﺩ ﺑﻬﺫﺍ ﺇﻳﺻﺎﻝ ﻣﺎﻟﻡ ﻳﻛﻥ ﻣﻣﻬﻭﺭﺍ ﺑﺧﺗﻡ ﺩﻓﻊ ﻭﻣﻭﻗﻊ ﻣﻥ ﺍﻟﻣﻭﻅﻑ ﺍﻟﻣﺳﺋﻭﻝ ﺍﻟﻌﻣﻳﻝ ﺍﻟﺳﻳﺩ ﺷﺣﺎﺗﺔ ﺍﻷﻟﻔﻰ ﺍﻻﻟﻔﻰ ﻻ ﻳﻌﺗﺩ ﺑﻪ ﺇﻳﺻﺎﻝ ﺳﺣﺏ ﻧﻘﺩﻯ ﺗﻭﻗﻳﻊ:
------------------------------------------------------------------------------------------------------------------------------------------
As you see in the Calibri extraction a lot of Arabic characters are not extracted correctly unlike the Arial extraction so we need to resolve this issue or find a way to correct the extraction for any PDF with any Font
Appreciate your support