Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Sat May 23, 2020 11:49 am

I'm getting inconsistent results when searching the PDF version of a document in comparison with it's Word counterpart. The issue seems to be that Pdf.FindText specifying the whole word parameter doesn't recognise the apostrophe character as a terminator, where DocFindAllString does.

I think that the Pdf processing is incorrect and should match that of the Doc libraries.

I have attached a couple of screenshots to illustrate the issue.

PDF Example
PDF Document Example.png


Word Example
Word Document Example.png


Many thanks in advance

wraydc
 
Posts: 130
Joined: Wed Apr 11, 2018 5:14 am

Mon May 25, 2020 6:01 am

Hello,

Thanks for your inquiry and sorry for the late reply as weekend.
I simulated a Word file and a PDF file and tested your scenario, and did notice the issue you mentioned. However, I found when using the regular expression "\b(attorney)\b" to match the text, the finding result is consistent with the Doc libraries, like the following code. I recommend you use the regular expression instead.
Code: Select all
    PdfDocument pdf = new PdfDocument();
    pdf.LoadFromFile("test.pdf");
    PdfTextFind[] result = null;
    foreach (PdfPageBase page in pdf.Pages)
    {
        result = page.FindText(@"\b(attorney)\b", TextFindParameter.Regex).Finds;
        foreach (PdfTextFind find in result)
        {
            find.ApplyHighLight();
        }
    }
    string output = "FindAndHighlightText_out.pdf";
    pdf.SaveToFile(output, FileFormat.PDF);

Anyway, I will also pass this issue to our Dev team for further investigation. If there is any update, we will let you know.

Sincerely,
Rachel
E-iceblue support team
User avatar

rachel.lei
 
Posts: 1571
Joined: Tue Jul 09, 2019 2:22 am

Return to Spire.PDF

cron