Spire.PDF is a professional PDF library applied to creating, writing, editing, handling and reading PDF files without any external dependencies. Get free and professional technical support for Spire.PDF for .NET, Java, Android, C++, Python.

Tue Jul 16, 2024 11:19 pm

I'm using a Regular Expression to find the pages that match a pattern. finder.find(pattern) returns a list page PdfPageBase object, but I can't find a way to get the page with a specific matching string so I get get related data from that page.

pmaneely
 
Posts: 4
Joined: Mon Jul 15, 2024 6:34 pm

Wed Jul 17, 2024 2:38 am

Hello,

Thanks for your inquiry.
Please note that you need to specify the page number before searching for text. Therefore, there is no need for an additional method to get the page number. I have attached the complete code for searching text by regular expression for your reference.
Code: Select all
// Specify the file path
String input = "data/findWithRegularExpression.pdf";
// Create a pdf file
PdfDocument doc = new PdfDocument();
// Load the PDF document
doc.loadFromFile(input);
// Get the first page of the PDF file
PdfPageBase page = doc.getPages().get(0);
// Match the regex
String regex = "(?<=\\{)[^}]*(?=\\})";
// Create text find options
PdfTextFindOptions findOptions = new PdfTextFindOptions();
// Set search parameter to use regular expression
findOptions.setTextFindParameter(EnumSet.of(TextFindParameter.Regex));
// Create a text finder object for the page
PdfTextFinder textFinder = new PdfTextFinder(page);
// Find text fragments that match the regex
List<PdfTextFragment> finds = textFinder.find(regex, findOptions);
// Define a color
PdfRGBColor color = new PdfRGBColor(Color.blue);
// Create a brush with the defined color
PdfBrush brush = new PdfSolidBrush(color);
// Define a font
PdfTrueTypeFont font = new PdfTrueTypeFont(new Font("Arial", Font.BOLD, 10));
// Set text alignment
PdfStringFormat centerAlign = new PdfStringFormat(PdfTextAlignment.Center, PdfVerticalAlignment.Middle);
// Define a rec
Rectangle2D rec;
// Iterate the find results
for (PdfTextFragment find : finds) {
    // Get the bounds of the found text
    rec = find.getBounds()[0];
    // Draw a rectangle around the found text
    page.getCanvas().drawRectangle(PdfBrushes.getWhite(), rec);
    // Set new text
    String newText = "New Text";
    //  Replace the found text with new text
    page.getCanvas().drawString(newText, font, brush, rec, centerAlign);
}
// Define output path
String result = "output/findWithRegularExpression_out.pdf";
// Save the modified document
doc.saveToFile(result);
// Close the PDF document
doc.close();
// Dispose of the PDF document (frees up system resources)
doc.dispose();

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Thu Jul 18, 2024 3:58 pm

I need to know the page number of a given text fragment so that I can determine if its on the same page as another text fragment. In other pdf extraction libraries that I'm evaluating this is done by referencing the associated page with the text fragment. The associated page has an page number method.

I don't see a way to do this is Spire.PDF.

The latest API isn't available online, perhaps is was added?

pmaneely
 
Posts: 4
Joined: Mon Jul 15, 2024 6:34 pm

Fri Jul 19, 2024 8:43 am

Hello,

Thanks for your reply.
Please refer to the following code to get the page number of the keyword. If you have any questions, please feel free to write to us.
Code: Select all
// Create a new PDF document
PdfDocument pdf = new PdfDocument();
// Load an existing PDF file
pdf.loadFromFile("test.pdf");
// Initialize a variable to store search results
List<PdfTextFragment> results = null;
// Create text find options for searching
PdfTextFindOptions findOptions = new PdfTextFindOptions();
// Set search parameters to find whole words only
findOptions.setTextFindParameter(EnumSet.of(TextFindParameter.WholeWord));
// Loop through the pages in the PDF file
for (Object pageObj : pdf.getPages()) {
    // Get each page in the PDF document
    PdfPageBase page = (PdfPageBase) pageObj;
    // Create a text finder object for the page
    PdfTextFinder textFinder = new PdfTextFinder(page);
    // Search for the text "Keyword" on the page
    results = textFinder.find("Keyword", findOptions);
    // Get page number
    for (PdfTextFragment fragment : results) {
        System.out.println("Keyword page number:"+(pdf.getPages().indexOf(fragment.getPage())+1));
    }
}
// Close the PDF document
pdf.close();
// Dispose of the PDF document (frees up system resources)
pdf.dispose();

Sincerely,
William
E-iceblue support team
User avatar

William.Zhang
 
Posts: 757
Joined: Mon Dec 27, 2021 2:23 am

Return to Spire.PDF

cron