I get inconsistent text extraction results across multiple runs for the SAME multi-page PDFs, and I am unable to determine why:
Please note that in BOTH screenshots shown above, the actual amount of text to be extracted from the PDF should be well over 100,000 characters, but I get different results with each run, so it is hard to determine an accurate number
Here is the code I'm using (sample project and large PDF attached):
- Code: Select all
using Spire.Pdf;
using Spire.Pdf.Texts;
using System.Text;
const string PDF_FILENAME = @"<path to large PDF>";
const string LICENSE_FILENAME = @"<path to license.elic.xml>";
Spire.Pdf.License.LicenseProvider.SetLicense(LICENSE_FILENAME);
// Testing across multiple iterations
for (int i = 1; i <= 10; i++)
{
using var document = new PdfDocument(PDF_FILENAME);
var pdfText = new StringBuilder();
Console.WriteLine($"Spire PDF Test {i}: Page Count = {document.Pages.Count}");
for (int j = 0; j < document.Pages.Count; j++)
{
var page = document.Pages[j];
var extractor = new PdfTextExtractor(page);
string pageText = extractor.ExtractText(new PdfTextExtractOptions() { IsExtractAllText = true });
//Console.WriteLine($" Page {j}, Character Count = {pageText.Length}");
pdfText.Append(pageText);
}
Console.WriteLine($" Total Character Count = {pdfText.Length}");
document.Close();
}