Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Fri May 15, 2020 11:53 am

Hi,

I am evaluating Spire.PDF to convert PDF to html. First of all Thank you for providing this utility. The output looks good.

I have below issues/questions:

1. The PDF pages seems to be embedded as SVG in a Html document. Is there any other way to convert the document as HTML instead of SVG? We convert the PDF to html and then do some manipulations like adding background color or border to certain text. It is very complex to apply the same in SVG.
2. As it is SVG, the text in converted HTML is not searchable, as even a single word is splitted into multiple nodes. Is there any solutions available for that?
3. All resources like image, css, font, etc. are embedded into html/svg. Is there options available to store these resources in separate folders instead of embedding it into html?

Thanks,
Sri
Sri

sri_1972
 
Posts: 1
Joined: Fri May 15, 2020 11:43 am

Mon May 18, 2020 8:58 am

Hello,

Thanks for your inquiry.
I simulated a PDF file, tested the following code with our latest Spire.PDF Pack(Hot Fix) Version: 6.5.9 and found the output are parsed as HTML tags and could be searched.
Code: Select all
PdfDocument pdf = new PdfDocument();
pdf.LoadFromFile("test.pdf");
pdf.ConvertOptions.SetPdfToHtmlOptions(useEmbeddedSvg: false, useEmbeddedImg: false);
pdf.SaveToFile("Result.html", FileFormat.HTML)

I suggest that you try again with the latest version. If there is still any issue on your side, please provide your test PDF document as well as your output Html to help us further look into it. You can upload here or send them to us via email ([email protected]).


Sincerely,
Sara
E-iceblue support team
User avatar

Sara.Yang
 
Posts: 33
Joined: Wed May 06, 2020 1:05 am

Return to Spire.PDF