Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Apr 15, 2025 4:21 pm

I am working on a C# console application on .NET 8 using Spire.PDF version 11.3.0 to extract text from different types of PDFs.

I get inconsistent text extraction results across multiple runs for the SAME multi-page PDFs, and I am unable to determine why:

2025-04-15_11-07-33.png

2025-04-15_11-02-44.png


Please note that in BOTH screenshots shown above, the actual amount of text to be extracted from the PDF should be well over 100,000 characters, but I get different results with each run, so it is hard to determine an accurate number :(

Here is the code I'm using (sample project and large PDF attached):

Code: Select all
using Spire.Pdf;
using Spire.Pdf.Texts;
using System.Text;

const string PDF_FILENAME = @"<path to large PDF>";
const string LICENSE_FILENAME = @"<path to license.elic.xml>";

Spire.Pdf.License.LicenseProvider.SetLicense(LICENSE_FILENAME);

// Testing across multiple iterations
for (int i = 1; i <= 10; i++)
{
    using var document = new PdfDocument(PDF_FILENAME);

    var pdfText = new StringBuilder();

    Console.WriteLine($"Spire PDF Test {i}:  Page Count = {document.Pages.Count}");

    for (int j = 0; j < document.Pages.Count; j++)
    {
        var page = document.Pages[j];

        var extractor = new PdfTextExtractor(page);
        string pageText = extractor.ExtractText(new PdfTextExtractOptions() { IsExtractAllText = true });

        //Console.WriteLine($"    Page {j}, Character Count = {pageText.Length}");

        pdfText.Append(pageText);
    }

    Console.WriteLine($"    Total Character Count = {pdfText.Length}");

    document.Close();
}

drmcclelland
 
Posts: 3
Joined: Tue Apr 15, 2025 1:02 pm

Wed Apr 16, 2025 7:29 am

Hello,

Thanks for your inquiry.
The reason for your problem may be that the license was not properly applied. Without a valid license, Spire.PDF will limit extraction to the first 10 pages. When we execute your code, we consistently get the same output, as demonstrated in the attached screenshot. Please use the following method to apply the license and retest. Looking forward to your feedback.
Code: Select all
Spire.Pdf.License.LicenseProvider.SetLicenseKey("license key");

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Wed Apr 16, 2025 3:05 pm

Hi William,

We're setting the license key explicitly now and the sample yields the same result. (as opposed to setting the path of the license file as before).

The license key is copied from the license file and set as the string parameter for the SetLicenseKey method, as suggested.

You can see from the output the behavior is the same as before.

cullaa11
 
Posts: 1
Joined: Tue Apr 15, 2025 6:33 pm

Thu Apr 17, 2025 3:14 am

Hello,

Thanks for your feedback.
Sorry, I still can't reproduce the issue you mentioned. To help us investigate further, please provide us with the following information. Thanks in advance.
1.Your operating system (e.g. Windows 10);
2. Your computer's regional settings (e.g. English (United Kingdom));
3. A simplified and directly runnable project.(You can upload it to the attachment, for privacy reasons, I suggest you send it directly to this email: [email protected]).

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Thu Apr 17, 2025 4:18 pm

William - thank you for your offer to help, I have sent you an email with the information and sample that you requested.

drmcclelland
 
Posts: 3
Joined: Tue Apr 15, 2025 1:02 pm

Fri Apr 18, 2025 2:01 am

Hello,

Thansk for your feedback.
I have received your email and replied to it. Please check it. If you don't mind, we can communicate further via email later.

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Wed Apr 30, 2025 10:09 pm

Thank you for your help William - we appreciate you and your colleagues identifying the licensing issue that we encountered. Now that we have the right licenses, we are able to extract text from the entire PDF consistently!

drmcclelland
 
Posts: 3
Joined: Tue Apr 15, 2025 1:02 pm

Thu May 01, 2025 2:45 am

Hello,

Thanks for your feedback.
Glad to hear that. If you encounter any issues regarding our product in the future, please feel free to write to back.

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Return to Spire.PDF