Hello,
Thanks for your post.
I did notice that your PDF document is not searchable in Adobe, when copying and pasting your document content into the search box, the content is inconsistent. I am afraid that this is related to Adobe itself mechanism. However, using our the latest version(
Spire.PDF (hot fix) version 6.7.8) can search successfully. The below is my testing code and I uploaded my output file for your reference.
- Code: Select all
PdfDocument pdf = new PdfDocument("6343fos20200801_356.pdf");
PdfTextFind[] result = null;
foreach (PdfPageBase page in pdf.Pages)
{
result = page.FindText("Rathmann", TextFindParameter.None).Finds;
foreach (PdfTextFind find in result)
{
find.ApplyHighLight();
}
}
pdf.SaveToFile("result.pdf");
At the same time, our Spire.PDF also supports extracting the text content. You also can try the following code to extract them to a .txt file, and then search the string in it. If there is any other question, just feel free to write back.
- Code: Select all
PdfDocument doc = new PdfDocument();
doc.LoadFromFile("6343fos20200801_356.pdf");
StringBuilder content = new StringBuilder();
foreach (PdfPageBase page in doc.Pages)
{
content.Append(page.ExtractText());
}
String fileName = "outPut.txt";
File.WriteAllText(fileName, content.ToString());
Sincerely,
Sofia
E-iceblue support team
Login to view the files attached to this post.