We have an issue with extracting text from a pdf.
Our system is based on using Docker containers, we receive pdf as a stream format and we would like to extract data from it.
We are using the function LoadFromStream from SpirePDF. We get the page count number but the method ExtractText doesnt return the string.
Here is the code sample:
- Code: Select all
using (MemoryStream ms = new MemoryStream(Attachment.Body))
{
PdfDocument doc = new PdfDocument();
doc.LoadFromStream(ms);
Console.WriteLine(doc.Pages.Count);
PdfPageBase page = doc.Pages[0];
SimpleTextExtractionStrategy strategy = new SimpleTextExtractionStrategy();
string text = page.ExtractText(strategy);
Console.WriteLine(text);
return text;
}
Thank you very much for you answer,
Have a good day!