Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Apr 11, 2023 12:35 am

I used the following code to extract and output contents from table in pdf (not scanned). But how can I deal with the merged cells? Currently all merged cells contents are null. Is there any method to identify vertical merge or horizontal merge?

Code: Select all
                string txtPath = Path.Combine(@"C:\Users\xj2ssf\Desktop\EDM_test", string.Format("{0}_bySpire_{1}.txt",
                Path.GetFileNameWithoutExtension(pdf), DateTime.Now.ToString("yyyy_MM_dd_hh_mm_fff")));
                string xlsPath = Path.ChangeExtension(txtPath, ".xlsx");

                Spire.Pdf.PdfDocument doc = new Spire.Pdf.PdfDocument();
                doc.LoadFromFile(pdf);

                Spire.Pdf.Utilities.PdfTableExtractor extractor = new Spire.Pdf.Utilities.PdfTableExtractor(doc);

                using (ExcelPackage packageNew = new ExcelPackage(new FileInfo(@"D:\From_Git\NEWEDM\GeneralUI\Template\CheryICDRetrieveTemplate.xlsx")))
                using (StreamWriter sw = new StreamWriter(txtPath, true))
                {
                    // Loop through the pages                     
                    for (int pageIndex = 0; pageIndex < doc.Pages.Count; pageIndex++)
                    {
                        sw.WriteLine("Page: {0}", pageIndex);

                        Spire.Pdf.Utilities.PdfTable[] tableList = extractor.ExtractTable(pageIndex);

                        if (tableList != null && tableList.Length > 0)
                        {

                            int count = 0;
                            //Loop through the table in the list
                            foreach (Spire.Pdf.Utilities.PdfTable table in tableList)
                            {
                                count++;
                                ExcelWorksheet excelWorksheet = packageNew.Workbook.Worksheets.Add(String.Format(@"Page_{0}_Table_{1}", pageIndex.ToString(), count.ToString()));

                                //Get row number and column number of a certain table
                                int row = table.GetRowCount();
                                int column = table.GetColumnCount();

                                //Loop though the row and colunm
                                for (int i = 0; i < row; i++)
                                {
                                    for (int j = 0; j < column; j++)
                                    {
                                        var sb = new StringBuilder();
                                        //Get text from the specific cell
                                        string text = table.GetText(i, j);

                                        //Add text to the string builder
                                        sb.Append(text + ";");
                                        sw.Write(sb.ToString());

                                        excelWorksheet.Cells[i + 1, j + 1].Value = sb.ToString();
                                        excelWorksheet.Cells[i + 1, j + 1].Style.WrapText = true;

                                    }
                                    sw.WriteLine();
                                }
                            }
                        }
                        sw.WriteLine("============================================");
                    }
                    packageNew.SaveAs(new FileInfo(xlsPath));
                }


msdos241
 
Posts: 3
Joined: Tue Apr 11, 2023 12:22 am

Tue Apr 11, 2023 10:07 am

Hi,

Thanks for your message.
I simulated a PDF document and tested with our latest Spire.PDF Pack (Hot Fix) Version: 9.4.0, but I did not reproduce your issue. the merged cell content could be extracted. If you were using an old version? I suggest that you switch to the latest version to test again. If your issue persists after testing, please provide us with your input document for further investigation. You can upload it here or provide to us by email ([email protected]).Theoretically, PDF does not contain the table properties. It only contains lines and character glyph which we tend to interpret as tables. It cannot distinguish whether the merged cells are vertical or horizontal. Hope you can understand.

Best Regards,
Herman
E-iceblue support team
User avatar

Herman.Yan
 
Posts: 115
Joined: Wed Mar 08, 2023 2:00 am

Wed Apr 12, 2023 12:50 am

Hello,

Sorry for my incorrect description in the main post. What I wanted to say is that only the first cell in one merged group cells will be extracted the data and all other cells will not.

Please check out my upload attachment. In the zip file there are two files which are original pdf file and output excel file.

I did test it with latest dll.

Thank you for the support!

msdos241
 
Posts: 3
Joined: Tue Apr 11, 2023 12:22 am

Wed Apr 12, 2023 4:06 am

Hi,

Thanks for your message.
As for the merged cell, only the first cell contains the data, thus, another cell is null. Sorry I am a little confused about your desired output effect. You can share a screenshot to show your desired effect for further investigation.Additionally, our Spire.PDF also supports directly converting PDF to XLSX. The attachment is my testing result after testing with the below code.
Code: Select all
            String result = @"test.xlsx";
            //Create a pdf document
            using (PdfDocument doc = new PdfDocument())
            {
                doc.LoadFromFile(@"01. T1DPHEV CDU单元电路图上传KMS版_5.pdf");

                //Save to XLSX
                doc.SaveToFile(result, FileFormat.XLSX);

            }


Best Regards,
Herman
E-iceblue support team
User avatar

Herman.Yan
 
Posts: 115
Joined: Wed Mar 08, 2023 2:00 am

Wed Apr 12, 2023 6:19 am

Hello,

My expectation is to identify each cell whether it is part of a merged cell or not, and if it is, then the blank cell would be replaced by the data.

You could see the new attachment, and focus on the red cells. This is what I really want, because I need to extract data row to row completely and when encountering merged cell, the shared content must be identified other than null value.

msdos241
 
Posts: 3
Joined: Tue Apr 11, 2023 12:22 am

Thu Apr 13, 2023 7:58 am

Hi,

Thank you for your message.
As mentioned before, there are no table elements or cells in the PDF specification. Our product identifies tables by the combined lines and character glyph. We apologize that the screenshot you provided cannot be directly achieved by our Spire.PDF. However, you can use our Spire.XLS to copy and set the other data for your current excel file.If you need to use our Spire.XLS and Spire.PDF in the same program, please directly use our Spire.office to do the tests.

Best Regards,
Herman
E-iceblue support team
User avatar

Herman.Yan
 
Posts: 115
Joined: Wed Mar 08, 2023 2:00 am

Return to Spire.PDF

cron