Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Wed Apr 07, 2021 1:32 pm

I tried your little changed demo code to test extraction of text and images in evaluation version from .pdf file. In many cases it works, but from some documents (like from attached one), it doesn't extract all of it's contents.

My code in C# is as follows:
Code: Select all
    static void Test(string fName)
    {
      //Create a pdf document.
      if (!File.Exists(fName))
      {
        Console.WriteLine($"File {fName} does not exist.");
        return;
      }

      PdfDocument doc = new PdfDocument();
      doc.LoadFromFile(fName);

      StringBuilder buffer = new StringBuilder();
      IList <Image> images = new List<Image>();

      foreach (PdfPageBase page in doc.Pages)
      {
        buffer.Append(page.ExtractText());
        foreach (Image image in page.ExtractImages())
        {
          images.Add(image);
        }
      }

      doc.Close();

      //save text
      var fName_Txt = fName + ".txt";
      File.WriteAllText(fName_Txt, buffer.ToString());

      //save image
      int index = 0;
      foreach (Image image in images)
      {
        var fNameImg = $"{fName}_img{index++}.png";
        image.Save(fNameImg, ImageFormat.Png);
      }

      //Launching the Text file.
      //System.Diagnostics.Process.Start(fName_Txt);
    }


Thank you for any suggestions.

bbrodnik
 
Posts: 20
Joined: Sat Jan 09, 2021 10:38 am

Thu Apr 08, 2021 5:05 am

Hello,

Thanks for your inquiry.
I noticed the behavior that the text cannot be extracted from your PDF document. After investigation, I found the contents in your file actully are paths which are made up of lines rather than real text. If you use Adobe to extract, it will also return empty text. Hence, this behavior is caused by your PDF document itself, sorry our Spire.PDF can't deal with it. Hope you can understand.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Mon Jul 12, 2021 5:13 pm

I understand that there is no text inside document, but contents in neither extracted as images. How can I extract contents which is shown when opened with Adobe, even as image, to be able to perform OCR on it?

bbrodnik
 
Posts: 20
Joined: Sat Jan 09, 2021 10:38 am

Tue Jul 13, 2021 7:24 am

Hello,

Thanks for your feedback.
You could use our Spire.PDF to convert Pdf pages to images, then use OCR to extract the text from the images. Please refer to the code below.
Code: Select all
            PdfDocument pdf = new PdfDocument();
            pdf.LoadFromFile("test.pdf");
            for (int i = 0; i < pdf.Pages.Count; i++)
            {
                String fileName = String.Format("img-{0}.png", i);
                using (Image image = pdf.SaveAsImage(i, 300, 300))
                {
                    image.Save(fileName, System.Drawing.Imaging.ImageFormat.Png);
                }
            }
            pdf.Close();

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Return to Spire.PDF