Hi Brian,
Thank you for your reply and suggestion. We tried this out and unfortunately it was not able to resolve the issue and eventually threw the same exception after approx. 4 hours of processing.
When the LoadFromFile method loaded the 9GB file memory never went above 40K and the page file was the same, however the "I/O Read Bytes" in Task Manager was reading Giga Bytes over the course of 15 minutes. The process appeared to be stuck on this call and crashed after 4 hours:
- Code: Select all
Interop+Kernel32.ReadFile(System.Runtime.InteropServices.SafeHandle, Byte*, Int32, Int32 ByRef, IntPtr)
Here is a memory dump just prior to the process crashing:
- Code: Select all
DBG ID OSID ThreadOBJ State GC Mode GC Alloc Context Domain Count Apt Exception
0 1 568c 0000029D3AE82270 2a020 Preemptive 0000029D3D377E70:0000029D3D3780E0 0000029d3aeaca90 0 MTA
3 2 5018 0000029D3C67F3A0 2b220 Preemptive 0000000000000000:0000000000000000 0000029d3aeaca90 0 MTA (Finalizer)
6 3 3428 0000029D3AEB8860 102a220 Preemptive 0000000000000000:0000000000000000 0000029d3aeaca90 0 MTA (Threadpool Worker)
0:000> !clrstack
OS Thread Id: 0x568c (0)
Child SP IP Call Site
000000E2B497DE30 00007ffb3a6ece34 [InlinedCallFrame: 000000e2b497de30]
000000E2B497DE30 00007ffa874afc1d [InlinedCallFrame: 000000e2b497de30]
000000E2B497DDF0 00007ffa874afc1d Interop+Kernel32.ReadFile(System.Runtime.InteropServices.SafeHandle, Byte*, Int32, Int32 ByRef, IntPtr)
000000E2B497DF00 00007ffa2ce3a1fe System.IO.FileStream.ReadFileNative(Microsoft.Win32.SafeHandles.SafeFileHandle, System.Span`1, System.Threading.NativeOverlapped*, Int32 ByRef) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 1195]
000000E2B497DF50 00007ffa2ce3a000 System.IO.FileStream.ReadNative(System.Span`1) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 495]
000000E2B497DFA0 00007ffa2ce38e6b System.IO.FileStream.ReadSpan(System.Span`1) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.Windows.cs @ 417]
000000E2B497E000 00007ffa2ce386e8 System.IO.FileStream.Read(Byte[], Int32, Int32) [/_/src/System.Private.CoreLib/shared/System/IO/FileStream.cs @ 303]
000000E2B497E060 00007ffa2d1dadd0
000000E2B497E0B0 00007ffa2d1bc2db
000000E2B497E110 00007ffa2d1da0d3
000000E2B497E1E0 00007ffa2d1baa4e
000000E2B497E240 00007ffa2d1ba67a
000000E2B497E2B0 00007ffa2d1b9db8
000000E2B497E2F0 00007ffa2d1b9d0c Spire.Pdf.PdfDocument.LoadFromFile(System.String)
000000E2B497E340 00007ffa2d1b9c6b Spire.Pdf.PdfDocument.LoadFromFile(System.String, Spire.Pdf.FileFormat)
000000E2B497E370 00007ffa2ce150d8 ConsoleApp2.Program.MergeWithSpire(System.String, System.String, System.String)
000000E2B497E3E0 00007ffa2ce110c9 ConsoleApp2.Program.Main(System.String[])
000000E2B497E648 00007ffa8c916c93 [GCFrame: 000000e2b497e648]
000000E2B497EBE0 00007ffa8c916c93 [GCFrame: 000000e2b497ebe0]
Comparing this to when we tried to load the file into the constructor of PdfDocument, that led to memory climbing all the way to ~18GB over 4 hours of processing, even more for the page file, before the process crashed with the Invalid/Unknown format exception.
Unfortunately I am unable to provide the document in question that is causing this problem and am loathed to have to crawl through 29,000 pages in order to redact sensitive information so I hope the above information is helpful in any way to potentially identifying why this may be happening.
If I can come across any further information that may help I will make sure to update or reply to this thread with it.
Thank you,
Zac