Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Mon Feb 12, 2018 10:08 pm

I'm converting an HTML text file to a PDF using LoadFromHTML() and it works great. But the PDF is not searchable using PdfTextFind. Is it possible to convert an HTML file to a PDF using Spire.PDF so that the resulting PDF can be searched for text?

Here's the code for making the PDF:

string htmlString = File.ReadAllText(fileName);
PdfDocument doc = new PdfDocument();
PdfHtmlLayoutFormat htmlLayoutFormat = new PdfHtmlLayoutFormat();
PdfPageSettings setting = new PdfPageSettings();
setting.Size = PdfPageSize.Letter;

Thread thread = new Thread(() => { doc.LoadFromHTML(htmlString, false, setting, htmlLayoutFormat); });
thread.SetApartmentState(ApartmentState.STA);
thread.Start();
thread.Join();
doc.SaveToFile(newFileName, FileFormat.PDF);

And for searching (which returns nothing):

PdfDocument doc = new PdfDocument(fileName);
PdfTextFind[] results = null;
results = doc.Pages[0].FindText(searchText).Finds;
foreach (PdfTextFind find in results)
{
find.ApplyHighLight(Color.Green);
}
doc.SaveToFile(newFileName2, FileFormat.PDF);

Thanks for your help!

Nick Solomon

nicksolomon1
 
Posts: 5
Joined: Tue Feb 06, 2018 8:46 pm

Tue Feb 13, 2018 2:56 am

Dear Nick,

Thanks for your inquiry.
The LoadFromHTML method is based on the IE kernel. If the version of your IE is 9 or above version, the converted pdf will be rendered in image way. It is impossible to search text using Spire.PDF if the result pdf doesn't have any text.
Please try to use new plugin to convert html string to PDF, which will generate text properly.
https://www.e-iceblue.com/Tutorials/Spi ... -in-C.html
If you still cannot search the text after using the new plugin, please provide us with following information for further investigation.
1. Your HTML string and the result PDF you got.
2. The text you want to search.

Sincerely,
Betsy
E-iceblue support team
User avatar

Betsy.jiang
 
Posts: 3099
Joined: Tue Sep 06, 2016 8:30 am

Thu Feb 15, 2018 9:44 am

Dear Nick,

Did the solution with new plugin resolve your issue?
Thanks in advance for your valuable feedback and time.

Sincerely,
Amy
E-iceblue support team
User avatar

amy.zhao
 
Posts: 3044
Joined: Wed Jun 27, 2012 8:50 am

Thu Feb 15, 2018 5:01 pm

If I modify the sample program (namespace SALESSUPPORT_2429_HtmlToPdf_) I am able to convert the html file to pdf. It works great, and is readable! Here's the code I've changed:

static void Main(string[] args)
{
string sFile = @"file:///C:/Temp/PACER/PDF/TestPACER2.htm";
string sNewFile = @"C:\Temp\PACER\PDF\TestPACER2.PDF";
HtmlConverter.Convert(sFile, sNewFile,
//enable javascript
false,
//load timeout
1 * 1000,
//page size
new SizeF(612, 792),
//page margins
new PdfMargins(0, 0));
// System.Diagnostics.Process.Start("HTMLtoPDF.pdf");
}

I'm having problems getting any other project to work with the addin. I've followed your instructions and get the following error:

System.AccessViolationException was unhandled
Message: An unhandled exception of type 'System.AccessViolationException' occurred in Spire.Pdf.dll
Additional information: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

Here's my code:

try
{
HtmlConverter.PluginPath = @"C:\VS2015\Testing\Windows\SpireTest1\SpireTest1\bin\Debug\plugins";
HtmlConverter.Convert(
textEditURL.Text,
textEditPDFFileName.Text,
false,
2 * 1000,
new Size(612, 792),
new PdfMargins(0, 0));
}
catch (Exception ex)
{
MessageBox.Show(ex.Message);
}

Spire.Pdf: runtime version v4.0.30319. Version: 3.10.0.2040
Here's some output from debugging:

...
HTMLConverter: Check converter delegate.
HTMLConverter: Getting converter delegate...
HTMLConverter: Converter delegate found.
The program '[9720] SpireTest1.vshost.exe' has exited with code 0 (0x0).

I've tried replacing the Spire.Pdf dll several times. Other Spire.PDF functions seem to be working fine (locate text, drawimage, etc.).

I can't seem to understand why the test program works, but the other does not. Should I try using the Spire.Pdf included in the test program in my SpireTest1 app?

Any suggestions would be great. Thanks for your quick response too!

Nick Solomon

nicksolomon1
 
Posts: 5
Joined: Tue Feb 06, 2018 8:46 pm

Fri Feb 16, 2018 6:39 am

Hello,

Thanks for your response. The issue should be no related to the Spire.pdf dll, which is the architecture of your project that has the issue? X64 or X86? And which plugin package do you download? We suggest you should just used the X86 plugin packgin.

Sincerely,
Gary
E-iceblue support team
User avatar

Gary.zhang
 
Posts: 1380
Joined: Thu Apr 04, 2013 1:30 am

Mon Feb 19, 2018 10:26 pm

As you suggested, I switched to using the x86 plugin files, rebuilt my project, and things work GREAT! Thank you for your quick response and all the help.

Nick Solomon

nicksolomon1
 
Posts: 5
Joined: Tue Feb 06, 2018 8:46 pm

Tue Feb 20, 2018 1:19 am

Hello Nick,

Thanks for your feedback. If any other question, welcome to write to us.

Best regards,
Simon
E-iceblue support team
User avatar

Simon.yang
 
Posts: 620
Joined: Wed Jan 11, 2017 2:03 am

Return to Spire.PDF

cron