Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Feb 01, 2022 11:01 pm

Hi,

We are trying to produce a PDF for our clients, however Spire PDF appears to be throwing an error when trying to load a very large file into it's constructor
Code: Select all
new PdfDocument()
.

The exception is:

Code: Select all
Spire stitching failed with error: Spire.Pdf.Exceptions.PdfDocumentException: Invalid/Unknown/Unsupported format     at spr㓄.ᜂ(Int32 A_0)     at spr㓄.ᜂ(String A_0)     at spr㒺.ᜀ(Dictionary`2 A_0, spr㒽 A_1)     at spr㕡.ᜀ(Stream A_0, Boolean A_1)


The file in question is 29,000 pages and 9GB in size and was generated using Spire PDF prior to being saved to the file system within the same process that is trying to load it back into the constructor so I am unsure as to why it was able to write the file in its entirety but not be able to read it back into the object graph.

Unfortunately I cannot provide the file in question as it contains sensitive material related to our clients, but I would like to know your thoughts as to why or how this exception would be thrown when trying to load a very large file into the constructor of PdfDocument.

Thank you,
Zac

ZCliff92
 
Posts: 31
Joined: Thu Jan 20, 2022 1:58 am

Wed Feb 02, 2022 7:36 am

Hello Zac,

Thanks for your inquiry.
Because your pdf file is very large, we suspect that the file may not be recognized correctly when loading the file through the constructor. We recommend you try to specify the format of the file as followss.
Code: Select all
            Spire.Pdf.PdfDocument pdf = new Spire.Pdf.PdfDocument();
            pdf.LoadFromFile("test.pdf", Spire.Pdf.FileFormat.PDF);

If the issue persists, to help us further analyze, please provide an pdf file that could reproduce your issue. Don't worry, we promise to keep your document confidential and we will not use it for any other purpose. You could also remove the sensitive data of your pdf as long as the modified file could reproduce your issue.

Sincerely,
Brian
E-iceblue support team
User avatar

Brian.Li
 
Posts: 1271
Joined: Mon Oct 19, 2020 3:04 am

Thu Feb 03, 2022 9:49 pm

Hi Brian,

Thank you for your reply and suggestion. We tried this out and unfortunately it was not able to resolve the issue and eventually threw the same exception after approx. 4 hours of processing.

When the LoadFromFile method loaded the 9GB file memory never went above 40K and the page file was the same, however the "I/O Read Bytes" in Task Manager was reading Giga Bytes over the course of 15 minutes. The process appeared to be stuck on this call and crashed after 4 hours:

Code: Select all
Interop+Kernel32.ReadFile(System.Runtime.InteropServices.SafeHandle, Byte*, Int32, Int32 ByRef, IntPtr)


Here is a memory dump just prior to the process crashing:

Code: Select all
DBG   ID     OSID ThreadOBJ           State GC Mode     GC Alloc Context                  Domain           Count Apt Exception
   0    1     568c 0000029D3AE82270    2a020 Preemptive  0000029D3D377E70:0000029D3D3780E0 0000029d3aeaca90 0     MTA
   3    2     5018 0000029D3C67F3A0    2b220 Preemptive  0000000000000000:0000000000000000 0000029d3aeaca90 0     MTA (Finalizer)
   6    3     3428 0000029D3AEB8860  102a220 Preemptive  0000000000000000:0000000000000000 0000029d3aeaca90 0     MTA (Threadpool Worker)
0:000> !clrstack
OS Thread Id: 0x568c (0)
        Child SP               IP Call Site
000000E2B497DE30 00007ffb3a6ece34 [InlinedCallFrame: 000000e2b497de30]
000000E2B497DE30 00007ffa874afc1d [InlinedCallFrame: 000000e2b497de30]
000000E2B497DDF0 00007ffa874afc1d Interop+Kernel32.ReadFile(System.Runtime.InteropServices.SafeHandle, Byte*, Int32, Int32 ByRef, IntPtr)
000000E2B497DF00 00007ffa2ce3a1fe System.IO.FileStream.ReadFileNative(Microsoft.Win32.SafeHandles.SafeFileHandle, System.Span`1, System.Threading.NativeOverlapped*, Int32 ByRef) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 1195]
000000E2B497DF50 00007ffa2ce3a000 System.IO.FileStream.ReadNative(System.Span`1) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 495]
000000E2B497DFA0 00007ffa2ce38e6b System.IO.FileStream.ReadSpan(System.Span`1) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 417]
000000E2B497E000 00007ffa2ce386e8 System.IO.FileStream.Read(Byte[], Int32, Int32) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.cs @ 303]
000000E2B497E060 00007ffa2d1dadd0
000000E2B497E0B0 00007ffa2d1bc2db
000000E2B497E110 00007ffa2d1da0d3
000000E2B497E1E0 00007ffa2d1baa4e
000000E2B497E240 00007ffa2d1ba67a
000000E2B497E2B0 00007ffa2d1b9db8
000000E2B497E2F0 00007ffa2d1b9d0c Spire.Pdf.PdfDocument.LoadFromFile(System.String)
000000E2B497E340 00007ffa2d1b9c6b Spire.Pdf.PdfDocument.LoadFromFile(System.String, Spire.Pdf.FileFormat)
000000E2B497E370 00007ffa2ce150d8 ConsoleApp2.Program.MergeWithSpire(System.String, System.String, System.String)
000000E2B497E3E0 00007ffa2ce110c9 ConsoleApp2.Program.Main(System.String[])
000000E2B497E648 00007ffa8c916c93 [GCFrame: 000000e2b497e648]
000000E2B497EBE0 00007ffa8c916c93 [GCFrame: 000000e2b497ebe0]


Comparing this to when we tried to load the file into the constructor of PdfDocument, that led to memory climbing all the way to ~18GB over 4 hours of processing, even more for the page file, before the process crashed with the Invalid/Unknown format exception.

Unfortunately I am unable to provide the document in question that is causing this problem and am loathed to have to crawl through 29,000 pages in order to redact sensitive information so I hope the above information is helpful in any way to potentially identifying why this may be happening.

If I can come across any further information that may help I will make sure to update or reply to this thread with it.

Thank you,
Zac

ZCliff92
 
Posts: 31
Joined: Thu Jan 20, 2022 1:58 am

Fri Feb 04, 2022 2:59 am

Dear Zac,

Thanks for your further information.
According to your description, I guessed that your PDF document itself has problems with content. You could use PDF reader like Adobe to open your PDF to check if there is something wrong with content. Sorry for the fact that I can't help you out based on the current information, we need your real PDF for an accurate investigation. You could upload your PDF file via DropBox, then share us with the download link. Please send it to our email ([email protected]). Thanks for your assistance.

Sincerely,
Nina
E-iceblue support team
User avatar

Nina.Tang
 
Posts: 1385
Joined: Tue Sep 27, 2016 1:06 am

Sun Feb 06, 2022 11:04 pm

Hi Nina,

Thanks for your response. The document in question opens without error in Adobe Reader. As such we are currently looking at other solutions for our PDF generation process.

Thanks again,
Zac

ZCliff92
 
Posts: 31
Joined: Thu Jan 20, 2022 1:58 am

Mon Feb 07, 2022 6:56 am

Dear Zac,

Thanks for your feedback and I apologize for the late reply due to the weekend.
We did initial tests but didn't encounter your issue. The memory dump you provided was incomplete, our Dev team cannot provide any usable leads based on this information. We are sorry it's hard to provide solution for you at this moment. We need your further assistance. Maybe you could provide full stack information to help us investigate if there is any breakthrough. And please also share your testing environment (such as win7, 64bit, 8GB) and application type (such as console app, .net framework 4.7.2, x64 platform)
Nevertheless, I think the best way to drive progress is sharing the PDF document. Thanks for your understand.

Sincerely,
Nina
E-iceblue support team
User avatar

Nina.Tang
 
Posts: 1385
Joined: Tue Sep 27, 2016 1:06 am

Return to Spire.PDF