Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Sun Feb 04, 2024 7:08 am

Dear Annika
Thank you for your response

Indeed you have been so kind to help

for a free temporary license valid for one month

What is then the situation after one Month ?

Secondly, the reason for copying the 6 DLL files to the Debug folder of your testing project is because these DLL files are essential external dependencies. By placing them in the Debug folder, you ensure that these dependencies are fully functional and accessible during the testing process.

Ok Got it. Thanks
However, if you are looking for a tool specifically for viewing PDF files, you can consider using Spire.PDFViewer, which is designed for this purpose. Please note that Spire.PDFViewer does not support the direct viewing of image files


I incorporated PictureBox on the Form to view files of .jpg, .png etc

Now I am specifically asking you How Can I view the below
1. PDF Files with images,
2. PDF Files with Image as Text,
3. PDF Files with Images with Text incorporated

with PDFDocumentViewer1and get its text from the above 3 scenarios.

Which are the type of PDF Files where you cannot Extract Text using PDFDocumentViewer1

Thanks
SamD
9

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Sun Feb 04, 2024 9:18 am

Hello,

Thank you for your feedback.

Regarding the first issue, the temporary license will automatically expire after one month. Once it expires, you will see an evaluation warning watermark when using Spire.OCR to extract text. To permanently remove the evaluation warning watermark, please contact our sales team at [email protected] to purchase a formal license.

As for the third issue, our Spire.PDFViewer supports viewing any PDF file. Please refer to the sample demo I provided earlier on extracting text and displaying it in a text box.

If you have any further questions or concerns, please feel free to let me know.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Mon Feb 05, 2024 5:12 am

Anikka

Thanks for the reply

As for the third issue, our Spire.PDFViewer supports viewing any PDF file. Please refer to the sample demo I provided earlier on extracting text and displaying it in a text box


Have implemented the code and as mentioned earlier One type of PDF file not able to extract text also informed its bit description.
(As i don't know the term for this description)

As ocr.Scanner.Scan(path) only accepts .jpg, .jpeg, .bmp , .png and .Tiff files so why ocr.Scanner.Text not able to extract text from PDF

So wanted to know Which type of PDF file that Spire will not be able to extract text.

I created one PDF in the below following manner for the type of PDF which is not seen in PDFDocumentViewer1


Open MS-WORD
Typed some Text
Changed its Font Size and its Style
Select above all Text and Ctrl + C to Copy
And Select Paste > Paste Special > Picture (Enhanced MetaFile) > OK
Right Click on the above Text pasted as Image and
Save As changed the File Name & Folder Name and its Type to .JPEG
Opening the above file .JPG
Right Click > Open With > Photos
Click on Printer Icon
Select Microsoft Print to PDF and Saved the above File with .PDF




Thanks
SamD
10

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Mon Feb 05, 2024 9:38 am

Hello,

Thanks for your feedback.
For Spire.OCR, its underlying logic code determines that it is parsing the content of the image, so Spire.OCR can only extract the text in the image file. For scanned PDF files, currently we don't have any component support for extracting text from this type of file. But this feature is on our upgrade list and may be implemented in future development.
In addition, you mentioned that the PDF file you created can not be seen in our PDFDocumentViewer1 this problem, my side in accordance with the steps you provided to do a preliminary test, but did not reproduce the problem. Please provide your test PDF for our further investigation. You could attach it here or send it to us via email ([email protected]). Thanks in advance.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Mon Feb 05, 2024 11:54 am

Dear Annika

For scanned PDF files, currently we don't have any component support for extracting text from this type of file. But this feature is on our upgrade list and may be implemented in future development.

By What time the upgrade will be available and when you can let us know. My suggestion to you would to provide us tool with more flexibility
and accessibility. So that Normal PDF with Text can be viewed with respective Text Extracted and PDF with Image, Scanned Image with Text also could be extracted.

in addition, you mentioned that the PDF file you created can not be seen in our PDFDocumentViewer1 this problem, my side in accordance with the steps you provided to do a preliminary test, but did not reproduce the problem. Please provide your test PDF for our further investigation. You could attach it here or send it to us via email ([email protected]). Thanks in advance.


I just checked the file with your above message. While creating It got Corrupted and for this reason it did not display.
But Created another File With Text Saved as Image and then saved as PDF. I can view clearly the PDF File in PDFDocumentViewer1
Now awaiting the Extraction part of it.

Indeed You have been of great help
Thank you very much.

SamD
11

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Tue Feb 06, 2024 1:41 am

Hello,

Thanks for the feedback again.
Regarding the feature of extracting text from scanned PDF files, our developers are still investigating the implementation. It is uncertain when this feature will be released. But please rest assured that we will let you know as soon as this feature is realized. Thank you again for your understanding and support.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Sat Mar 30, 2024 6:00 am

Hello,

Thanks for the feedback again.
Regarding the feature of extracting text from scanned PDF files, our developers are still investigating the implementation. It is uncertain when this feature will be released. But please rest assured that we will let you know as soon as this feature is realized. Thank you again for your understanding and support.

Sincerely,
Annika
E-iceblue support team


Dear Annika
Hope you are doing good
Any progress on Extracting text from scanned PDF files

Thanks
SamD
12

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Mon Apr 01, 2024 8:01 am

Hello,

Thank you for following up with us.
Regarding the feature to extract text from scanned PDFs, we regret to inform you that it has not been implemented yet. Our development team is still working diligently to bring this feature to fruition. We kindly ask for your patience and understanding as we continue to work on it. Rest assured that once the feature is implemented, we will notify you promptly.
Thank you once again for your patience and cooperation.

Sincerely,
Annika
E-iceblue support team
User avatar

Annika.Zhou
 
Posts: 1657
Joined: Wed Apr 07, 2021 2:50 am

Thu Aug 01, 2024 8:46 am

Dear Annika
Hope you are doing good
Any progress on Extracting text from scanned PDF files ?

It has been very long time that i've asked you the same. Almost 4 months have been passed

Thanks
SamD
13

SamDsouza

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Fri Aug 02, 2024 9:52 am

Hello,

Thanks for your following up.
Sorry, there has been no breakthrough in this issue so far. I will urge our development team to speed up the processing speed. I will inform you promptly once there is any progress.

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Mon Sep 16, 2024 2:43 pm

Any progress on Extracting text from scanned PDF files ?


SamDsouza
17

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Tue Sep 17, 2024 4:00 am

Dear SamDsouza,

This is a new feature in our product, and it involves relatively complex modules, so unfortunately, it is not currently supported. However, our Spire.PDF supports converting scanned PDF into images, after which you can use Spire.OCR to recognize the text within those images. Have you considered this approach?

Sincerely,
Nina
E-iceblue support team
User avatar

Nina.Tang
 
Posts: 1385
Joined: Tue Sep 27, 2016 1:06 am

Thu Sep 26, 2024 5:29 am

Nina.Tang
However, our Spire.PDF supports converting scanned PDF into images after which you can use Spire.OCR to recognize the text within those images. Have you considered this approach?


Why did you not let me know the above approach ? Anyways....
May i know how can i achieve this


Thanks
SamDsouza
18

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Thu Sep 26, 2024 9:55 am

Hi SamDsouza,

I apologize for not informing you about this indirect method earlier. Please download both Spire.PDF and Spire.OCR from the following links:
Nuget download link:
https://www.nuget.org/packages/Spire.PDF/10.9.0
https://www.nuget.org/packages/Spire.OCR/1.9.8

This article provides detailed guidance on using Spire.OCR to extract text from images:
Extract Text from Images using the New Model of Spire.OCR for .NET

Additionally, I have attached the related code snippet below for your reference.
Code: Select all
// Create a new PdfDocument object
PdfDocument doc = new PdfDocument();

// Load the PDF file from the specified path
doc.LoadFromFile("test.pdf");

// Initialize a list to store image streams
List<Stream> lists = new List<Stream>();

// Iterate through each page in the PDF document
for (int i = 0; i < doc.Pages.Count; i++)
{
    // Save the current page as an image
    Image image = doc.SaveAsImage(i);

    // Create a new memory stream
    var stream = new MemoryStream();

    // Save the image to the memory stream in PNG format
    image.Save(stream, System.Drawing.Imaging.ImageFormat.Png);

    // Add the memory stream to the list
    lists.Add(stream);
}

// Iterate through the list of image streams
for (int i = 0; i < lists.Count; i++)
{
    // Get the image stream from the list
    var stream = lists[i];

    // Create a new OCR scanner object
    OcrScanner scanner = new OcrScanner();

    // Perform OCR scanning with the input stream in PNG format
    scanner.Scan(stream, OCRImageFormat.Png);

    // Get the OCR recognized text
    IOCRText ocrText = scanner.Text;

    // Convert the recognition results to a string
    string text = ocrText.ToString();

    // Write the recognized text to a file with the name "result_indexNumber.txt"
    File.WriteAllText($"result_{i}.txt", text);
}

If you have any further questions or need more assistance, please don't hesitate to ask.

Sincerely,
Nina
E-iceblue support team
User avatar

Nina.Tang
 
Posts: 1385
Joined: Tue Sep 27, 2016 1:06 am

Wed Oct 02, 2024 7:03 am

Hi Nina
Hi SamDsouza,

I apologize for not informing you about this indirect method earlier. Please download both Spire.PDF and Spire.OCR from the following links:
Nuget download link:
https://www.nuget.org/packages/Spire.PDF/10.9.0
https://www.nuget.org/packages/Spire.OCR/1.9.8

Vb.net
To Install the above two
Project>Manage Nu Get Packages...>Browsed as per aboveLinks>Installed

But not able to get PDFViewer Tool/Object in Toolbox

Pl guide me to get above things resolved
Thanks
SamDsouza
19

SamDsouza
 
Posts: 23
Joined: Wed Jan 31, 2024 11:05 am

Return to Spire.PDF