Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Aug 15, 2023 6:11 pm

for a pdf page that has data that produces a text string of "hello world", that is "hello" followed by multiple whitespace and then the word world. How do I write the regular expression for your parser to match multiple whitespace between the words. I have tried many variations including those that validate correctly on reg exp checkers on the internet and nothing seems to match the multiple whitespace. If I print out the page of text and capture the exact number of whitespaces your routine finds between the two words and manually put that in the search string it still doesn't find the two strings with multiple whitespace between.

The actual string was "(word1) (word2)" with parentheses. Your text output had 11 spaces between the words. Inputting the string just given with 11 spaces and trying to search for that exactly didn't work either.

It finds just the first string (no problem), just the second string (no problem), but doesn't find the two strings with multiple whitespace.


New Information:

I have found the reason the regular expression pattern match is not working but it now raises a new issue. In order to try to debug the issue I did the following 2 things.

1)I used the Spire function page.extractText(true) returns the entire page of text with line breaks. Line 21 yields the following

(Counts) (psig) (°F) (Mcf) Mcf) (Mcf) (Btu/scf) (MMBtu)

I am trying to find the string: "(Mcf) (Btu/scf)"



2) I added the following routine ( copied from one of your examples) to print out the contents of the PdfTextFind getFinds()

private void printFindCollectionResults(PdfTextFindCollection pdftextfindcollection) throws Exception {
//
for (PdfTextFind find : pdftextfindcollection.getFinds()) {
//

System.out.println("=====================================================================");
System.out.println(" Match Text: " + find.getMatchText());
System.out.println(" Text: " + find.getSearchText());
System.out.println(" Size: " + find.getSize());
System.out.println(" Position: " + find.getPosition());
System.out.println(" The index of page which is including the searched text : " + find.getSearchPageIndex());
System.out.println(" The line that contains the searched text : " + find.getLineText());
System.out.println(" Match Text: " + find.getMatchText());
}
}


Now, if I try to find the string with the regular expression of: "\\(Mcf\\)( )+\\(Btu/scf\\)" it doesn't find it. But if I modify this to the following: "( )+\\(Btu/scf\\)" it finds it and prints out the following from your routine above:

==================================================================================
Match Text: (Btu/scf)
Text: ( )+\(Btu/scf\)
Size: com.spire.office.packages.sprthja[width=28.18225,height=6.95]
Position: Point2D.Float[401.3, 255.19]
The index of page which is including the searched text : 0
The line that contains the searched text : (Btu/scf)
Match Text: (Btu/scf)

Please note the second to the last line "The line that contains the searched text : (Btu/scf). Your find function is not finding the expression "\\(Mcf\\)( )+\\(Btu/scf\\)" because your pdftextfindcollection.getFinds() routine for some reason believes the entire line is " (Btu/scf)". It is chopping Line 21 into smaller parts. Can you clarify what is going. I have checked the full page using a hex editor and there are only spaces between the (Mcf) and the (Btu/scf) strings.

Any help you can give is appreciated.
Last edited by schultjd on Wed Aug 16, 2023 6:51 pm, edited 2 times in total.

schultjd
 
Posts: 7
Joined: Tue Jan 26, 2016 3:58 am

Wed Aug 16, 2023 6:49 am

Hi

Thank you for your inquiry.
I tested the code you provided and encountered the same problem as you.
This old method of finding text cannot recognize regular expressions and it will be deprecated. We recommend using our new method.
I put the complete code below for your referenece:
Code: Select all
  PdfDocument newP = new PdfDocument("data/1.pdf");
        PdfPageBase page;
        for (int i = 0; i < newP.getPages().getCount(); i++) {
            page = newP.getPages().get(i);
            PdfTextFinder pdfTextFinder = new PdfTextFinder(page);
            PdfTextFindOptions pdfTextFindOptions = new PdfTextFindOptions();
            pdfTextFindOptions.setTextFindParameter(EnumSet.of(TextFindParameter.Regex));
            List<PdfTextFragment> textFragments = new ArrayList<>();
            //Find the hello world string
            textFragments = pdfTextFinder.find("hello(\\s*)+world", pdfTextFindOptions);
            for (int j=0;j<textFragments.size();j++){
                System.out.println(textFragments.get(j).getLineText());
            }

        }

If my code doesn’t help you, please offer your pdf file to help us do further investigation, you can attach here or send it to us via email ([email protected]). Thank you in advance.

Sincerely,
Ula
E-iceblue support team
User avatar

Ula.wang
 
Posts: 282
Joined: Mon Aug 07, 2023 1:38 am

Wed Aug 16, 2023 7:47 pm

Not found in the latest release

TextFindParameter.Regex


Also, do you have a workaround for Spire Office (lic: April 19, 2022). can you point me to version that is within my year maintenance that will give me access to the pdf methods that you used in your reply. I do not have them in the licensed jar from April 2022.

These following methods do not exist in my offfice.jar

PdfTextFinder pdfTextFinder = new PdfTextFinder(page);
PdfTextFindOptions pdfTextFindOptions = new PdfTextFindOptions();
pdfTextFindOptions.setTextFindParameter(EnumSet.of(TextFindParameter.Regex));

schultjd
 
Posts: 7
Joined: Tue Jan 26, 2016 3:58 am

Thu Aug 17, 2023 8:01 am

Hi

Thank you for your inquiry.
I simulated a Pdf file and tested the code you provided, but did not encounter the same problem as you.
These are my test screenshots:
图片2.png

图片1.png

We have found that you purchased Spire.Office for Java on April 19, 2022. Therefore, for your license, the latest version of Spire.Office for Java is V8.3.6, which supports a new method of finding text.
You can download it through maven, I put the maven code below.
Code: Select all
<repositories>
<repository>
<id>com.e-iceblue</id>
<name>e-iceblue</name>
<url>https://repo.e-iceblue.com/nexus/content/groups/public/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>e-iceblue</groupId>
<artifactId>spire.office</artifactId>
<version>8.3.6</version>
</dependency>
</dependencies>

I suggest you can use the new method of finding text to test again your scenario, if the issue still exists, please offer your pdf file to help us do further investigation, you can attach here or send it to us via email ([email protected] ). Thank you in advance.

Sincerely,
Ula
E-iceblue support team
Last edited by Ula.wang on Wed Oct 11, 2023 8:49 am, edited 1 time in total.
User avatar

Ula.wang
 
Posts: 282
Joined: Mon Aug 07, 2023 1:38 am

Thu Aug 17, 2023 4:13 pm

Ula,

Thanks for the reply. I have stripped down the file that still recreates the issue. I have upgraded to office version 8.3.6 as you suggested and the problem is the same. Please see my writeup above. The issue is not necessarily with the regular expression. It is because the following method returns 2 separate lines for volume - when if you print out the page text - the two volumes occur on the same line. I need to be able to find the string "(Mcf) (Btu/scf). The problem is the code that you sent to me finds each of these texts as separate lines even though they print out as a single line. I'm searching for this regular expression because I need the coordinates of the bounding box and since (Mcf) occurs elsewhere on the line I need to find the above string with ONLY WHITESPACE between.

Please see the file that I emailed to be able to recreate the issue.

Thanks,

Jerry

schultjd
 
Posts: 7
Joined: Tue Jan 26, 2016 3:58 am

Thu Aug 17, 2023 4:48 pm

Also, If I use 8.3.6 I get the message that the document was prepared with Spire. Obviously, 8.3.6 doesn't work with my license file. Can you clarify for me please. Thanks

8.3.5 has the same problem. Also, 8.2.0

Evaluation Warning : The document was created with Spire.PDF for Java.

schultjd
 
Posts: 7
Joined: Tue Jan 26, 2016 3:58 am

Fri Aug 18, 2023 10:30 am

Hi

Thank you for your inquiry.
I tested the Pdf file you sent through email and encountered the same problem as you.
Firstly, because the document you provided is displayed in two text editing boxes (Mcf) and (Btu/scf) in software Adobe, it is automatically recognized as two lines, while my simulated Pdf document (Mcf) and (Btu/scf) are in the same editing box, please refer to the following screenshots. This is the reason for the same regular expression can be retrieved in my simulated Pdf document, but cannot be retrieved in your Pdf document.
3.png

4.png

Secondly, you need the coordinates of the bounding box. We can provide you with the following code for your reference:
Code: Select all
PdfDocument newP = new PdfDocument("data/test.pdf");
        PdfPageBase page;
        for (int i = 0; i < newP.getPages().getCount(); i++) {
            page = newP.getPages().get(i);
            PdfTextFinder pdfTextFinder = new PdfTextFinder(page);
            PdfTextFindOptions pdfTextFindOptions = new PdfTextFindOptions();
            pdfTextFindOptions.setTextFindParameter(EnumSet.of(TextFindParameter.Regex));
            List<PdfTextFragment> textFragments = new ArrayList<>();
            List<PdfTextFragment> textFragments1 = new ArrayList<>();
            textFragments = pdfTextFinder.find("(Btu/scf)", pdfTextFindOptions);
            textFragments1 = pdfTextFinder.find("(Mcf)", pdfTextFindOptions);
            ArrayList<Integer> list = new ArrayList<>();
       //Place the data in a list to find the maximum value
            for (int i1 = 0; i1 < textFragments1.size(); i1++) {
                Rectangle2D[] bounds = textFragments1.get(i1).getBounds();
                double X = bounds[0].getX();
                int max = (int) X;
                list.add(max);
            }
              int max = Collections.max(list);
                System.out.println(max);
   //the coordinates of the bounding box
            double x=0,y=0,w=0,h=0;
                for (int j = 0; j < textFragments.size(); j++) {
                    Rectangle2D[] bounds = textFragments.get(j).getBounds();
                    x = bounds[0].getX();
                    y = bounds[0].getY();
                    w = bounds[0].getWidth();
                   h = bounds[0].getHeight();
                }
                String s = page.extractText(new Rectangle2D.Double(max, y, x-max+w, h));
                System.out.println(s);
            }


In addition, I found that you only can use the version before Spire.Office 5.4.5 according to your license file, such as the following sceenshot. But in this vesrion, there are no new method of finding text. Therefore, I suggest you contact our sales([email protected]) to upgrade your license.
5.png

Sincerely,
Ula
E-iceblue support team
User avatar

Ula.wang
 
Posts: 282
Joined: Mon Aug 07, 2023 1:38 am

Mon Sep 04, 2023 3:30 am

Hi

I would like to know if the solution we have provided has helped you solved the problem you have encountered, and our team expected to have more communication with you.
If my solution doesn’t help you, please offer the following information to help us do further investigation. Thank you in advance.
1)Your input test file, you can attach here or send it to us via email ([email protected]).
2) Your full test code that can reproduce your issue.
3) Application type, such as Console App, .NET Framework 4.8.
4) Your test environment, such as OS info (E.g. Windows 7, 64-bit).

Sincerely,
Ula
E-iceblue support team
User avatar

Ula.wang
 
Posts: 282
Joined: Mon Aug 07, 2023 1:38 am

Return to Spire.PDF

cron