Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Mon Sep 09, 2019 3:07 pm

I have a list of strings that I need to identify and add hyperlinks for inside an existing PDF:

RX-1
RX-2
RX-3
RX-10
RX-11
RX-20
RX-22 (and the list goes on)

When using the code below, the links work, but I am getting double links for RX-11, RX-20, and RX-22, etc. (e.g. RX-11 has a link for RX-1 underneath a link for RX-11 - see image below)

Code: Select all
            PdfDocument sourcePDF = new PdfDocument();
            sourcePDF.LoadFromFile(documentPath);

            PdfTextFind[] result = null;
                       
            foreach (var exhibit in exhibitInfo)
            {
                var searchString = exhibit.ExhibitName;
                var linkPath = exhibit.ExhibitPath;

                //Check each page individually
                foreach (PdfPageBase page in sourcePDF.Pages)
                {
                    result = page.FindText(searchString).Finds;
                    foreach(PdfTextFind find in result)
                    {
                        PdfLaunchAction launchAction = new PdfLaunchAction(linkPath, PdfFilePathType.Relative);
                        RectangleF rect = new RectangleF(find.Position.X, find.Position.Y, find.Size.Width, find.Size.Height);
                        PdfActionAnnotation annotation = new PdfActionAnnotation(rect, launchAction);
                        annotation.Border = new PdfAnnotationBorder(2);
                        annotation.Color = new Spire.Pdf.Graphics.PdfRGBColor(System.Drawing.Color.Blue);
                        (page as PdfPageWidget).AnnotationsWidget.Add(annotation);
                    }
                    result = null;
                }
            }


For completeness, exhibit.ExhibitName and exhibit.ExhibitPath are from a List<ExhibitInfo> where ExhibitInfo is:

Code: Select all
namespace HyperLinker
{
    public class ExhibitInfo
    {
        public string ExhibitName;
        public string ExhibitPath;

        public ExhibitInfo()
        {

        }
    }
}


Now, in the first code block, I have tried replacing

- result = page.FindText(searchString).Finds;

with

- result = page.FindText(searchString, true).Finds;

I get the same result, double boxes on the matching entries (RX-10, oddly, does not) as you can see from the image below:
MatchResult.png


I have also tried "result = page.FindText(searchString, true, false).Finds;" which results in no links being placed.

"result = page.FindText(searchString, false, false).Finds;" is not much help either, as the result is the same as leaving out the bool values all together.

If it is any help:

Spire.Pdf 5.8.16.2040
runtime version v4.0.30319

MatthewPierce
 
Posts: 22
Joined: Thu Jul 16, 2015 4:45 pm

Tue Sep 10, 2019 2:24 am

Hi,

Thanks for your inquiry.
To help us investigate your issue accurately, please offer us your input Pdf file.
You could upload them here or send us([email protected]) via email.

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Tue Sep 10, 2019 8:31 am

Hi Amber,

Two documents attached.

Pre-process: Exhibit List-NoData.pdf
Post-process: Exhibit List-NoData_Linked.pdf

Thank you.

MatthewPierce
 
Posts: 22
Joined: Thu Jul 16, 2015 4:45 pm

Tue Sep 10, 2019 12:09 pm

Hi,

Thanks for your information.
After investigation, I found the double links in the result Pdf is because some texts you try to find are included in other texts. For example, the text “RX-1” is included in the text “RX-11”, “RX-12” and “RX-13”and so on, when you try to add links to the source file, the text “RX-1” will be found several times. Please find the whole word then add link on it.
And kindly note that the method(page.FindText(searchString).Finds) you were using is out of data, please use this new method instead.

Code: Select all
result = page.FindText(searchString,TextFindParameter. WholeWord).Finds;


But when I tested your file with the new method, I found it didn’t work. Anyway, this issue has been logged into our bug tracking system, once there is any progress, we will inform you. Sorry for the inconvenience caused.

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Tue Sep 10, 2019 12:50 pm

Yes, I had tried the WholeWord option as well, forgot to include that in my original post.

Since you are moving this over to bug tracking, it is also worth noting that if we remove the dashes and use TextFindParameter.WholeWord the result is the same. I am going to have instances where there will be a prefix with a space, no dash or dot, and then a number. Perhaps you could incorporate Regular Expressions into the search feature so that we could have a more robust match operation.

Thank you, again.

~Matthew

MatthewPierce
 
Posts: 22
Joined: Thu Jul 16, 2015 4:45 pm

Wed Sep 11, 2019 4:10 am

Hi,

Thanks for your reply.
Our Spire.Pdf supports to search text based on a regular expression. Below is the code for your reference.

Code: Select all
                //Change the regular expression according to your actual case
                result = page.FindText(@"\[\s*\d{1,}\s*x\s*\d{1,}\s*\]",TextFindParameter.WholeWord).Finds;


Any question, welcome to contact us.

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Wed Sep 11, 2019 1:19 pm

Hi Amber,

Awesome, RegEx to the rescue! I changed my search parameter to this:

Code: Select all
var searchString = exhibit.ExhibitName + "\\s";


Now, as it cycles through my list, \s gets added to the end of each string so we have a whitespace match requirement that prevents "RX-1" from matching on the first four characters of "RX-10" and creating the double-link. I did have to change the text find parameter to none though. If I used WholeWord I still get a match count of zero:

Code: Select all
result = page.FindText(searchString, TextFindParameter.None).Finds;


Still, the result is what I was going for:

MatchResult.png

MatthewPierce
 
Posts: 22
Joined: Thu Jul 16, 2015 4:45 pm

Thu Sep 12, 2019 9:09 am

Hi,

Thanks for your reply and glad to hear that you have found a method to solve this issue by yourself.
As for the new issue with “WholeWord” and regular expression, I will also post to our Dev team. Once the new issues are fixed, we will inform you. Sorry for the inconvenience caused.

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Wed Sep 18, 2019 11:11 am

Hi,

Thanks for your patient waiting.
Glad to tell you that the previous issue has been resolved in Spire.PDF Pack(Hot Fix) Version:5.9.6. Welcome to download and test it from the following links:
Website link: https://www.e-iceblue.com/Download/download-pdf-for-net-now.html
NuGet link: https://www.nuget.org/packages/Spire.PDF/5.9.6

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Wed Sep 18, 2019 1:31 pm

This is why you folks are the best!

Thank you!

~Matthew

MatthewPierce
 
Posts: 22
Joined: Thu Jul 16, 2015 4:45 pm

Thu Sep 19, 2019 1:08 am

Hi Matthew,

Thanks for your feedback.
Any question, welcome to contact us. Have a nice day.

Best wishes,
Amber
E-iceblue support team
User avatar

Amber.Gu
 
Posts: 525
Joined: Tue Jun 04, 2019 3:16 am

Return to Spire.PDF