Cute Web Email Extractor Advance v2.1.22
Ahmad Software Technologies

1. Introduction

Cute Web Email Extractor Advance is a professional tool to harvest email addresses from search engines, websites, local files, and via an embedded chromium web browser. It includes advanced filters, proxy support, CAPTCHA handling, multi‑language UI, and auto‑save recovery.

Core Features

  • Search engine extraction (Google, Bing, Yahoo, Baidu, Yandex)
  • Batch crawl from custom website list
  • Local file scanner (PDF, DOCX, XLSX, TXT, HTML, CSV, LOG)
  • Built‑in browser mode for JavaScript‑heavy pages
  • Real‑time email filters, domain validation, proxy rotation
  • Auto‑save and session recovery

search engine screenshot

Data Grid Screenshot

2. Getting Started

  1. Installation:

    Installation

    Install Cute Web Email Extractor Advance and all required components.

    1. 1

      Run the Installer

      Launch the installation package to begin the setup process.

    2. 2

      Follow the Setup Wizard

      Complete the on-screen installation wizard by following the provided instructions.

    3. 3

      Automatic Runtime Installation

      If the required Chromium browser components are not already installed, the setup program will automatically download and install them.

    4. 4

      Launch the Application

      After installation is complete, launch the application from the Start Menu or desktop shortcut.

    Important Note

    Cute Web Email Extractor Advance requires a Chromium-based browser runtime to properly render modern websites.

    Supported Runtimes

    • Microsoft Edge WebView2 Runtime Recommended
    • Google Chrome

    The installer automatically detects and installs any missing components. An active internet connection may be required during this process.

  2. First launch: Define an auto‑save directory (File → Set Auto‑save Directory) – essential for crash recovery.
    See Section 10 for more detail
  3. Activation: Click Activate Full version on the toolbar and enter your Purchase Order number.
    See Section 8 for more detail
  4. Language: Use the language dropdown on the toolbar (11 languages, RTL supported). description

3. Main Window Overview

The Main Window is the central workspace where you configure extraction methods, manage searches, view results, and export collected email addresses.

Software Toolbar

Main application window.

Main Components

The main window is divided into four primary areas:

1. Extraction Panel

Select the extraction method such as Search Engine, Websites List, Computer Files, or Browser Mode.

2. Toolbar

Provides quick access to search, validation, settings, filtering, and data management functions.

3. Results Grid

Displays extracted email addresses, page titles, and website URLs.

4. Statistics Bar

Shows crawl progress, extracted contacts, parsed pages, and queue information.

Extraction Methods

Extract from Search Engine

Searches Google, Bing, Yahoo, Baidu, and other supported search engines using keywords and automatically extracts email addresses from discovered websites.

Extract from Websites List

Processes a list of URLs and extracts email addresses directly from the specified websites.

Extract from Computer Files

Extracts email addresses from local files including DOC, Excel, TXT, HTML, PDF, CSV, and other supported formats.

Extract with Browser

Uses Chromium browser rendering to process JavaScript websites and dynamic web pages.

Toolbar Functions

Software Toolbar

Figure 4.2. Software Toolbar.

Search Now

Starts the extraction process.

Validate

Verifies extracted email addresses.

Save

Exports extracted data to supported formats.

Settings

Opens software configuration options.

Filter

Applies advanced email and domain filters.

Clear Grid

Removes all displayed results.

Clear History

Clears stored extraction history.

Registered

Displays license information and subscription status.

Results Grid

The Results Grid displays all extracted contacts and related information.

Column Description
Email Address Extracted email address.
Page Title Title of the page where the email was found.
Website Source URL containing the email address.

Statistics Panel

The statistics panel at the bottom of the window provides real-time information about extraction progress.

  • Domain Count: Number of unique domains discovered.
  • Fetched Contacts: Total email addresses extracted.
  • Pages Parsed: Number of processed web pages.
  • Fetched Pages in Queue: Pages waiting to be processed.
  • URL Queue: Remaining URLs scheduled for crawling.
  • Time Elapsed: Duration of the current extraction session.

📄 4. File Menu – Complete Reference

The File menu contains essential commands for saving your data, managing backups, restoring previous sessions, and exporting logs. Each option is described below.

File Menu expanded of Cute Web Email Extractorf

File Menu

💾 4.1 Save Email Addresses

Exports the currently displayed email grid to a file on your local disk.

Supported formats:
  • Microsoft Excel (.xlsx)
  • CSV (comma‑separated values)
  • Plain text (.txt)
Export options dialog:
  • Choose which columns to save: Serial number, Email address, Page Title, Website URL
  • Select text encoding: Unicode (UTF‑8) or ANSI
Save dialog (Excel, CSV, TXT options)

Save dialog (Excel, CSV, TXT options)

📂 4.2 Open Auto‑Save Directory Structure

Opens the auto‑save folder in Windows Explorer. This folder contains:

  • GridData.csv – backup of extracted emails
  • ErrorURLs.txt – list of failed pages
  • stats.dat – last session statistics
  • CuteWebExtractor.adb – SQLite database with crawl queue
  • History Folder – It contains previous searches extracted data. Before starting new search software copies GridData.csv in the directory and GridData is appended with current date and time.
  • HistoryDB Folder – It contains previous searches database containing previous search web URLs. Before starting new search software copies CuteWebExtractor.csv in the directory and CuteWebExtractor is appended with current date and time.
Auto Save directory structure

Auto Save directory structure

⚙️ 4.3 Set Auto‑save Directory

Opens the Auto Save Directories management dialog. Here you can:

  • See existing backup folders (e.g., DefaultAutoSave, DefaultAutoSave-1)
  • Create a new auto‑save folder by typing a name and clicking Add new directory
  • Delete a selected directory
  • Check “Never show again and use the default auto‑save directory” to skip future prompts
⚠️ Critical: When running multiple instances of the software, each instance must use a different auto‑save folder. Sharing the same folder will corrupt session data.
Auto Save directory settings

Auto‑save directory content

🔄 4.4 Load Last Search Results

Restores the previous extraction session from the currently selected auto‑save directory. The following data is reloaded:

  • All extracted emails (reconstructs the grid)
  • Processing queue (crawled pages, waiting URLs)
  • Error logs and parsed URL lists

Use this feature after a crash, accidental close, or when you want to continue where you left off.

ℹ️ Note: You must load the last search results (File → Load Last Search Results) before clicking the Search Now button, otherwise the extraction will start from scratch and may overwrite previous data.

📜 4.5 Show Parsed URL List in Last Search

Opens a text file containing all URLs that were successfully processed (status = 2) during the last search. Each URL is listed on a separate line. This is useful for auditing which pages were actually crawled.

📄 File format: plain text, one URL per line

💿 4.6 Save the last search URLs list in file – Export Options

When you click this menu item, a dialog titled “Select save options” appears. It allows you to choose which subset of URLs to export.

🔘 Save all the URLs during search

Exports all URLs that have been discovered or processed during the current search session, regardless of their status (pending, crawling, parsed, or failed). This gives you a complete history of every URL the crawler has encountered.

🔘 Save all the URLs to be crawled

Exports only the URLs that are still waiting in the crawl queue (status = pending). Useful if you want to resume the crawl later or inspect the remaining work.

🔘 Save all the URLs has been crawled

Exports only the URLs that have already been successfully processed (status = completed). This gives you a clean list of pages that have been fully parsed.

💾 File format & destination: After selecting one of the three options, click Save Now. A standard “Save As” dialog will appear, allowing you to choose between CSV (*.csv) or Text (*.txt) format and select the output folder.

Typical use cases

  • “Save all the URLs during search” – Audit all discovered links, debug crawler behaviour.
  • “Save all the URLs to be crawled” – Backup the remaining queue before stopping a long extraction.
  • “Save all the URLs has been crawled” – Keep a permanent record of successfully visited pages for later reference.
📸 Screenshot: “Select save options” dialog with the three radio choices
OptionExported URL StatusFile Content Example
Save all the URLs during searchAll discovered (pending + processing + completed + failed)https://example.com/page1
https://example.com/page2
https://example.com/error
Save all the URLs to be crawledOnly pending / queuedhttps://example.com/deep-link-1
https://example.com/deep-link-2
Save all the URLs has been crawledOnly successfully processed (status = 2)https://example.com/home
https://example.com/contact
⚙️ Note: The exported file is saved in plain text (UTF‑8) with one URL per line. The CSV option uses the same format but with a `.csv` extension.

⚠️ 4.7 Show last search Error pages URLs list

Opens a text file listing all URLs that failed to load during the last extraction. Error types include:

  • HTTP errors (404, 500, etc.)
  • Timeouts
  • Connection refused
  • CAPTCHA or blocking pages

Use this list to manually check problematic domains or adjust crawling settings (delay, proxy, browser rendering).

❌ 4.8 Quit

Exits the application. If an extraction is running, a confirmation dialog will appear. Partial results are saved only if auto‑save is enabled.

💡 Shortcut: You can also close the main window using the standard system close button.

Quick Reference – File Menu Actions

Menu ItemPrimary Use
Save Email AddressesExport grid to file
Open Auto‑Save DirectoryOpen backup folder
Set Auto‑save DirectoryManage backup folders
Load Last Search ResultsRestore previous session
Show Parsed URL ListView successful URLs
Save last search URLs listExport crawl list
Show Error pages URLs listView failed URLs
QuitExit software

5. Tools Menu

  • Settings: Opens the full configuration dialog (see section 12).
  • Allow Unsafe HTTP Header Parsing: Enables non‑standard headers (advanced).
  • Stop Search: Immediately halts extraction, saves partial results.
Tools Menu options of Cute Web Email Extractorf

Toots Menu of Cute Web Email Extractorf

6. Help Menu

  • About Cute Email Extractor: Version, contact aslogger@ahmadsoftware.com.
  • Help: Opens online documentation.
  • Check For Updates: Auto‑download newer version.
Help Menu expanded of Cute Web Email Extractorf

Helop Menu

🔍 7. Email Filters Dialog (Post‑Extraction)

This dialog allows you to filter already extracted email addresses without re‑running the extraction. Access it by clicking the Filter button on the main toolbar

Apply Email Filters dialog

Apply Email Filters dialog as shown in the software

⚙️ Filter Options

📝 Keywords (one per line)

Enter keywords or phrases, one per line. The filter is case‑insensitive. Any row that contains at least one of these keywords (in the selected columns) will be processed according to the chosen action below.

Example: @example.com, abuse, noreply

📌 Apply the filters on
  • Email Address column – check to filter based on the email itself.
  • Page Title column – check to filter based on the webpage title.
  • Website column – check to filter based on the source URL.

You can select one, two, or all three columns. If multiple columns are selected, a match in any of them will trigger the action.

⚡ Select an option to perform the delete action
🔹
Remove only the matched text from the email address

Deletes the matched keyword(s) from the email string, but keeps the row. For example, if the keyword is “noreply” and the email is noreply@example.com, it becomes @example.com. The row remains in the grid.

🔹
Delete the entire row (this action cannot be undone)

Completely removes any row that matches the keywords in the selected columns. This is irreversible – the data is lost unless you have saved the grid before filtering.

[Apply] [Cancel]

📖 Example: Removing all emails from a specific domain

  • Add the keyword: @example.com
  • Check “Apply the filters on Email Address column”
  • Select “Delete the entire row”
  • Click Apply

All rows where the email contains @example.com will be permanently deleted from the grid.

Field Reference

Field / ControlDescription
Remove email address if any of the below keywords matchMulti‑line text box. Each line is a keyword. The filter matches if any keyword is found (partial match, case‑insensitive).
Apply the filters on Email Address columnWhen checked, the email column is included in the search.
Apply the filters on Page Title columnWhen checked, the page title column is included.
Apply filters on Website columnWhen checked, the website (URL) column is included.
Remove only the matched text from the email addressAction that removes the matching substring(s) from the email address but keeps the row.
Delete the entire row (this action cannot be undone)Action that removes the whole row from the grid permanently.
ApplyExecutes the filter with the current settings.
CancelCloses the dialog without making any changes.
💡 Tip: To avoid accidental data loss, save your grid (Save Email Addresses) before applying irreversible filters.

🔑 8. Software Registration Dialog

The Registration dialog is accessed by clicking the Activate Full version button on the main toolbar. It displays your license status and allows you to activate or renew the software.

Software Registration dialog (Cute Web Email Extractor Advance)

Software Registration dialog

📜 License Information & Activation

✅ License Status:

Shows whether you are using a registered version. Example: “You are using the registered version of the software.” If unregistered, a trial message appears.

📧 Support Contact:

Email address for assistance: aslogger@ahmadsoftware.com – clickable link available in the actual dialog.

⏳ Expiration Info:

Displays the number of days remaining on your current license (if subscription‑based). Example: “Your current license is expiring in 790 days”.

🆔 Enter Purchase Order number:

Text field where you paste or type the order ID received after purchasing the software. This is required for online activation.

🔐 Software Serial Key:

A hardware‑based unique identifier (automatically generated). It is displayed here and can be used for offline activation if internet is unavailable.

🛒 Action Buttons:
  • Buy Subscription – Opens the purchase page in your default web browser to renew or upgrade your license.
  • Register Now – Validates the entered Purchase Order number online. If valid, the software unlocks all features and the dialog shows “Registered”.
  • Close – Closes the registration dialog without making any changes.
⚠️ Note: The serial key is read‑only. It is used by Ahmad Software Technologies for offline license generation if you cannot activate online.
📘 Typical activation workflow:
  1. Click Activate Full version on the main toolbar.
  2. In the Registration dialog, enter your Purchase Order number.
  3. Click Register Now (requires internet connection).
  4. After successful validation, the software becomes fully registered and the button changes to “Registered”.

📋 Field Reference – Registration Dialog

Field / ControlDescription
License Status textIndicates registration state (registered / trial / expired).
Support email linkaslogger@ahmadsoftware.com – opens default email client.
Expiring in X daysDays left on subscription (only shown for registered users).
Enter Purchase Order numberInput field for the order ID (alphanumeric).
Software Serial KeyRead‑only hardware ID, used for offline activation.
Buy SubscriptionOpens purchase website.
Register NowTriggers online license validation.
CloseCloses the dialog.
💡 Tip: If you lose your registration details, contact support at the email address above with your purchase order number.

9. Extraction Modes

🌐 9.1 Search Engine Extraction Mode

This mode extracts email addresses by submitting keywords to one or more search engines (Google, Bing, Yahoo, Baidu, Yandex, etc.) and then crawling the resulting pages and linked websites.

⚙️ Configuration Options

📋 Select Search Engine(s)

A list of available search engines with checkboxes. You can select one or multiple. The software will query each selected engine separately and combine results.

  • Bing – Microsoft Bing search engine
  • Google – Google search (requires browser rendering for best results)
  • Yahoo – Yahoo search
  • Additional engines (Baidu, Yandex, etc.) may appear based on your configuration.

⚡ Performance Settings

Crawl Depth Level

Specifies how many levels of links the software will follow after opening a search result. Higher values allow deeper website exploration and may discover additional email addresses, but will increase crawling time.

0 – Search results page only.
1 – Search results page + homepage.
2 – Homepage + links found on the homepage.
3 – Continue crawling up to three levels deep.
Recommended: 3
Linked Site Crawl Depth Level

Controls how deeply the software crawls pages after leaving the original search result website and following links to other domains.

Higher values increase coverage but may significantly increase crawl duration.

Recommended: 3
Crawler Threads

Specifies the number of simultaneous connections used to download web pages. Increasing this value can improve extraction speed but may cause websites to block or rate-limit requests.

Default: 30
Parser Threads

Determines how many threads are used to analyze downloaded pages and extract email addresses and contact information.

For best performance, this value should normally be lower than the Crawler Threads setting.

Default: 8

🎛️ Advanced Filter Options

Extract from Websites on Search Engine Results Only

Limits extraction to the websites that appear directly in the search engine results. The software will visit each result page but will not follow additional links within those websites.

Best for fast searches and targeted extraction.
Enable Email Search Filters ↗ see Section 14

Opens the Email Search Filters dialog where you can define keywords that must appear on a page before email addresses are extracted.

Useful for extracting emails only from pages related to specific topics.
Enable Page Link Filters ↗ see Section 14

Opens the Page Link Filters dialog where you can specify keywords that must be present on a page before the crawler follows links found on that page.

Helps focus crawling on relevant sections of a website.
Extract from Search Engine Pages Only

Restricts extraction to the search engine results pages themselves. The software will not visit any external websites.

May result in fewer email addresses being found.
Recommendation

For most email extraction projects, leave Extract from Search Engine Pages Only disabled and use a crawl depth of 3. This provides the best balance between search coverage and extraction accuracy.

🌐 Regional Settings

Opens the Regional Settings dialog, allowing you to target searches to specific geographic locations. Regional targeting helps improve search relevance and discover websites associated with a particular country, region, city, or local area.

Available Location Filters
Country

Restrict searches to a specific country.

State / Province

Focus searches within a selected state or province.

City Selection

Target up to three cities simultaneously.

Area / Town

Further narrow searches to a specific local area.

Additional Options
Force Regional Matching

Instructs the search engine to prioritize results matching the selected regional settings whenever possible.

Country-Level Domains Only

Extract emails only from websites using country-specific top-level domains such as .us, .uk, .ca, .au, or .de.

Example

To find business email addresses in Sydney, Australia, select: Country = Australia, State = New South Wales, and City = Sydney. This helps return more locally relevant websites and contacts.

📝 Keywords List

The Keywords List contains all search queries that will be submitted to the selected search engines. Each row represents a separate search task and can include optional location and domain restrictions to improve targeting accuracy.

Keywords List

Figure 1. Keywords List containing search phrases and optional targeting information.

Available Columns
Keywords

The primary search phrase used when querying search engines. This field is required.

Example: dentist in london
Top Domain

Restricts results to a specific country top level domain.

Example: linkedin.com
Region

Specifies a geographic region to improve search relevance.

City

Restricts searches to a specific city or metropolitan area.

Province / State

Narrows searches to a selected province or state.

Country

Restricts searches to websites associated with a specific country.

⚙️ Managing Keywords
Add

Create a new search query by entering keywords and optional targeting information.

Delete

Remove the selected keyword entry from the list.

Example
KeywordsCityCountry
dentistLondonUnited Kingdom

This query will prioritize websites related to dentists located in London, United Kingdom.

💡 Tip

Using location fields together with Regional Settings can significantly improve search accuracy and reduce unrelated results.

🔘 Action Buttons
Search Now Show Grid Selected Keyword Clear keywords list
  • Search Now – starts the extraction using the current configuration (changes to Stop Search while running).
  • Show Grid – switches the right panel to the email results grid.
  • Selected Keyword – applies the currently highlighted keyword row to the search engine query (useful for testing one keyword at a time).
  • Clear keywords list – removes all keyword rows.

📋 Complete Field Reference – Search Engine Mode

Field / ControlDescription
Search Engine checkboxesSelect which engines to query (Bing, Google, Yahoo, etc.)
Country dropdown (per engine)Localise search results to a specific country.
Crawl depth levelHow many link levels to follow inside each discovered website.
Linked Site Crawl Depth LevelDepth for external domains (different from the starting domain).
Crawler ThreadsConcurrent HTTP request threads (30 recommended).
Parser ThreadsConcurrent HTML parsing threads (8 recommended).
Extract from websites on Search Engine Results OnlyDo not follow internal links; only process the result URLs themselves.
Enable Email Search FiltersOpens a dialog to set page‑level email extraction filters – see Section 14.
Enable Page Link FiltersOpens a dialog to set page‑level link‑following filters – see Section 14.
Extract from Search Engine Pages OnlyNever leave the search engine result pages.
Click to do Regional SettingsAdvanced location targeting (country, state, city, area).
Keywords list (6 columns)Search phrases with optional domain, region, city, province, country.
Search Now / Stop SearchStart or stop the extraction process.
Show GridSwitch to the email results view.
Selected KeywordExecute only the selected keyword row.
Clear keywords listDelete all keyword rows.
💡 Tip: For search engines that rely heavily on JavaScript (e.g., Google), make sure “Use Browser Rendering for Search Engines” is enabled in Settings → Crawl Engine.

💾 9.2 Extract from Websites List

Learn how to extract email addresses directly from a targeted list of URLs or domains using custom crawling preferences as shown in screenshot.

Keywords List

Figure 1. website list based search screen.

1

Access the Feature

On the left-hand navigation panel, click on Extract from Websites List. This will open the primary configuration dashboard.

2

Configure Performance Settings

Optimize how the software handles resources and web crawling:

Crawl Depth Level

Defines how many levels deep the crawler navigates within the initial domain.

Linked Site Crawl Depth

Defines how deep the crawler navigates into external linked sites.

Crawler & Parser Threads

Adjusts parallel processing speeds. Higher numbers increase extraction speed but demand more system resources.

3

Set Advance Filter Options

Choose how strictly the crawler follows web paths via the configuration radio buttons:

  • Extract from Added Websites Only: Restricts searches strictly to your provided root domains.
  • Extract from Added and Linked Websites: Expands scanning to external links originating from target domains.
  • Extract from Added URLs Only: Targets the exact URL path without jumping deeper.
4

Input Your Target Websites

Populate your targeting grid using one of two methods:

Method A: Manual Entry

Type or paste line-separated URLs into the text box, then click the blue arrow button () to push them to the grid.

Method B: Bulk Import

Click the Import button above the right-hand grid to upload an external URL file (.txt/.csv).

5

Start the Extraction

Click the Search Now button on the top toolbar or below the grid to initialize extraction. Toggle Show Grid to view results in real-time.

Pro Tip: Grid Management

Use the Selected button to target and drop broken/unwanted entries mid-session, or use Clear keywords list to reset your space cleanly.

💾 9.3 Computer Files Extraction Mode

This mode extracts email addresses (or website URLs) directly from local files and folders without crawling the web. It supports a wide range of file types, including text files, PDFs, Office documents, and more.

⚙️ Configuration Options

📄 Files List

The Files List contains all files and folders that will be processed during extraction. Files are scanned in the order shown in the list.

  • Double‑click a row to edit the file path directly.
  • Ctrl+C to copy the selected row(s) to clipboard.
  • Files are processed in the order they appear.
  • Folders are scanned recursively (all supported files inside).
🔘 Extraction Mode
Extracted Email Addresses Only

Searches for email addresses within the file content. This is the default mode.

Extracted Website Links Only

Extracts all website URLs (http, https, ftp, www) found in the files, instead of email addresses.

🔘 Action Buttons
Browse Files Browse Directory Clear List Search Now Show Grid
  • Browse Files – Opens a file picker to select individual files (supports multi‑selection).
  • Browse Directory – Opens a folder picker; all supported files inside the folder (and subfolders) will be added recursively.
  • Clear List – Removes all entries from the Files List.
  • Search Now – Starts the extraction process on the selected files/folders (button changes to Stop while running).
  • Show Grid – Switches the right panel to the email results grid (or URL results, depending on mode).
📂 Supported File Types
TXT HTML XML CSV PDF DOCX XLSX PHP ASP JS CSS RSS

Office files (DOCX, XLSX, PDF) are parsed using internal libraries; large files are processed in chunks to avoid memory issues.

📋 Complete Field Reference – Computer Files Mode

Field / ControlDescription
Files ListDisplays all selected files/folders. Double‑click to edit path, Ctrl+C to copy.
Browse FilesOpen file picker – select one or more files.
Browse DirectoryOpen folder picker – recursively add all supported files.
Clear ListRemove all entries from the Files List.
Extracted Email Addresses OnlySearch for email addresses in the files.
Extracted Website Links OnlyExtract website URLs instead of emails.
Search Now / StopStart or stop the extraction process.
Show GridSwitch to the email/URL results grid.

📖 Example: Extract emails from a folder of PDFs and Excel files

  1. Click Browse Directory and select a folder containing PDF and XLSX files.
  2. Select “Extracted Email Addresses Only”.
  3. Click Search Now.
  4. The software will scan all supported files in the folder (recursively) and display any found email addresses in the grid.
Computer Files Extraction Mode configuration 📸 Screenshot: Computer Files configuration panel (as shown in the software)
💡 Tip: For very large files (e.g., multi‑gigabyte text files), the software reads them in chunks to avoid memory exhaustion. The “Pages Parsed” counter increments for each chunk, not necessarily each file.

🌍 9.4 Extract with Browser Mode

This mode uses an embedded Web browser (Microsoft Edge or Chromium) to render web pages exactly as a real user would see them. It is ideal for:

  • Websites that rely heavily on JavaScript (Single‑Page Applications, infinite scroll, dynamic content).
  • Pages that require login, form submission, or user interaction before content is revealed.
  • Scraping data behind authentication walls (e.g., internal company portals).
  • Bypassing simple anti‑bot measures that block standard HTTP crawlers.

⚙️ Browser Interface & Controls

🌐 Address Bar & Navigation

At the top of the browser panel, you will find:

  • URL text box – type or paste any web address (e.g., https://example.com).
  • Go button – click to navigate to the entered URL. While a page is loading, the button changes to Stop (cancels navigation).
  • Back / Forward buttons (if present in the toolbar) – navigate through browsing history.

After navigating, you can interact with the page normally: click links, fill forms, log in, scroll, and even solve CAPTCHAs manually.

🖼️ Web Page Rendering Area

The main area displays the fully rendered web page, including:

  • JavaScript execution (Angular, React, Vue, etc.)
  • CSS styling, images, videos, and interactive elements.
  • Infinite scroll – you can manually scroll to load more content, then let the software extract from the fully loaded page.
  • Popup dialogs (alerts, confirmations) are automatically suppressed to avoid interruption.
🔍 Extraction Workflow
  1. Navigate to the desired starting page (e.g., a search results page after logging in).
  2. Perform any necessary actions: click “Load more”, fill a form, scroll to reveal content, etc.
  3. Click the main toolbar’s Search Now button (not the “Go” button).
  4. The software will extract emails from the current page and then follow links according to the crawl depth settings that correspond to the type of page currently displayed in the browser:
    • If the current page is a search engine results page (e.g., Google, Bing, Yahoo), the software automatically applies the Search Engine mode crawl depth settings (from the “Crawl depth level” and “Linked Site Crawl Depth Level” under Performance Settings in the Search Engine tab).
    • If the current page is a normal (non‑search‑engine) website, the software automatically applies the Websites List mode crawl depth settings (from the “Crawl depth level” and the selected crawl type – “Scan Whole Website”, “Added and Linked Websites”, etc. – in the Websites List tab).

    No manual selection is required – the detection happens automatically based on the URL and domain of the page you are viewing.

  5. Results appear in the email grid. You can stop the extraction at any time by clicking Stop Search.
⚠️ Important: To extract only the current page (no link following), set the crawl depth to 0 in the Performance Settings of the corresponding extraction mode (Search Engine or Websites List) – the software will respect that setting once the page type is detected.
⚡ Performance & Resource Usage

Browser‑based extraction consumes more CPU and memory than standard HTTP crawling. Recommendations:

  • Reduce the number of concurrent browser pages (Settings → Crawl Engine) to 2‑4 on older machines.
  • Enable “Do not load images” and “Do not load videos” in the Crawl Engine tab to speed up page loading.
  • For large sites, combine browser mode with a shallow crawl depth (1 or 2) to avoid excessive resource consumption.

📋 Complete Field Reference – Browser Mode

Control / ElementDescription
URL address barEnter any web address (http:// or https://). Auto‑completes from history.
Go buttonLoad the entered URL. While loading, the button changes to “Stop” to cancel navigation.
Page rendering areaDisplays the fully interactive web page (JavaScript, CSS, forms, etc.).
Back / Forward (if present)Navigate through the browser history.
Search Now (main toolbar)Starts email extraction from the current page. The software automatically detects whether the page is a search engine results page or a normal website and applies the corresponding crawl depth settings.
Stop SearchHalts the extraction process (button toggles from “Search Now”).
Show GridSwitches the right panel to the email results grid.

📖 Example: Extracting emails from a LinkedIn search after login

  1. Click Extract with Browser to open the browser panel.
  2. Navigate to https://www.linkedin.com/login and log in manually.
  3. After login, perform a search for “marketing managers”.
  4. Scroll down the results page to load all visible profiles.
  5. Click the main toolbar’s Search Now button.
  6. The software automatically detects that the current page is a search engine (LinkedIn search) and applies the Search Engine mode crawl depth settings. It will extract all email addresses visible on the current page and follow profile links according to the configured depth.
  7. Emails appear in the grid for further processing or export.

This approach works for any website that requires user interaction before data can be scraped, and the software adapts its crawling behaviour automatically.

Extract with Browser mode interface 📸 Screenshot: Browser mode panel (as shown in the software)
💡 Tip: For websites that detect automation, use the proxy settings and alternative browser type (Chrome) in Settings → Crawl Engine to reduce detection risks.

10. Auto‑Save Directory Management

From File → Set Auto‑save Directory you can create, delete, or select backup folders. Each software instance must use a distinct folder.

Auto-Save Directories

What are Auto-Save Directories?

Auto-Save Directories are used to store the software's internal database while extraction sessions are running. The database contains discovered URLs, extracted email addresses, crawl progress, and other information required to restore searches later.

Benefits

  • ✓ Automatically saves extraction progress
  • ✓ Restores interrupted searches
  • ✓ Prevents data loss after system crashes
  • ✓ Stores discovered URLs and extracted emails
  • ✓ Supports multiple extraction projects

Information Stored

URLs

Discovered URLs and pending URLs waiting to be processed.

Email Addresses

Extracted email addresses collected during searches.

Search Progress

Current crawl status and extraction progress.

Recovery Data

Information required to resume interrupted sessions.

Managing Auto-Save Directories

  1. 1

    Create a Directory

    Enter a new directory name and click Add New Directory.

  2. 2

    Select a Directory

    Select the directory you want the software to use for storing extraction data.

  3. 3

    Delete Unused Directories

    Select an unused directory and click Delete Selected Directory.

  4. 4

    Confirm Settings

    Click OK to save your changes.

Important

The software requires at least one Auto-Save Directory to store its working database. Without an Auto-Save Directory, extraction sessions cannot be restored after the program is closed or interrupted.

If you run multiple instances of the software simultaneously, configure a different Auto-Save Directory for each instance to avoid database conflicts.

11. Statistics Panel (Bottom Bar)

  • Domain Count – unique domains with extracted emails.
  • Fetched Contacts – total emails in grid.
  • Pages Parsed – successfully processed pages.
  • Fetched pages in Q – pages waiting in parsing queue.
  • URL in Queue – URLs waiting to be crawled.
  • Time Elapse – HH:MM:SS running time.
Extract with Browser mode interface 📸 Screenshot: statistics panel (as shown in the software)

12. Complete Settings Reference (Tools → Settings)

⚙️ Settings Dialog – Common Buttons

The following buttons are present at the bottom of every tab in the Settings dialog. They control how changes are applied and saved.

ButtonDescription
Restore Defaults Resets all settings on the currently active tab to their original factory values.
Fonts Opens a font selection dialog. Allows you to change the font family, size, and style used throughout the Settings window. Does not affect crawling behaviour.
Save All Settings Permanently Writes the current configuration (across all tabs) to disk. Settings will persist after the software is restarted. Use this after making permanent changes.
Apply changes to session only Applies the current settings for the current extraction session only. When the software is closed, changes are discarded. Useful for temporary testing without altering saved defaults.
Close Closes the Settings dialog.
Settings Common Buttons (Cute Web Email Extractor Advance)

Settings Common Buttons

💡 Tip: Use Apply changes to session only when experimenting with new settings. Once you are satisfied, click Save All Settings Permanently to make them permanent.

🚫 12.1 Domain and URL Filters

The Domain and URL Filters tab allows you to control which URLs are crawled and how discovered links are prioritized. These filters are applied before pages are added to the crawl queue, helping reduce unnecessary requests, improve performance, and keep extraction focused on relevant content.

Domain and URL Filters tab (Cute Web Email Extractor Advance)

Domain and URL Filters tab.

🚫 URLs containing any of the following keywords will be excluded:

Enter one keyword, domain, or URL fragment per line. Any URL containing one of these values will be skipped completely and will not be crawled or parsed.

  • Useful for excluding advertising networks, analytics services, content delivery networks (CDNs), social media widgets, or other irrelevant URLs.
  • Example entries: gstatic.com, amazonaws.com, googleapis.com, google-analytics.com, schema.org.
📌 Tip: Adding google.com will exclude all Google URLs, including search result pages.
⭐ URLs containing any of the following keywords will be crawled first:

Enter one keyword per line. URLs containing one of these keywords are given higher priority and are processed before other URLs in the queue.

  • Useful for prioritizing important pages such as /contact, /about, /team, /support, or /services.
  • Example: entering contact prioritizes URLs such as https://example.com/contact.
✅ Enable this option to crawl only URLs containing the filters below

When enabled, only URLs matching at least one keyword from the priority list above will be crawled. URLs that do not match are ignored.

ℹ️ This option converts the priority list into a URL whitelist.
🔗 Follow page links only when the page content contains the following:

Controls whether links found on a page should be added to the crawl queue. If the page does not contain one of the specified keywords, its links will not be followed.

  • Enter one keyword per line.
  • Email extraction still occurs on the current page even when its links are not followed.
  • Useful for keeping the crawler focused on specific topics, industries, products, or services.
⚙️ Options (apply to the “Follow page links” filter):
  • Must match all the keywords on the page – Requires every keyword in the filter list to be present before links from the page are followed.
  • Match on page title only – Searches only the page title instead of the full page content. This improves performance but may reduce accuracy.
💡 Example Workflow: To focus crawling on product pages, add /product/ to the priority URL list and enable “Enable this option to crawl only URLs containing the filters below” . Then add keywords such as product, buy, or pricing to “Follow page links only when the page content contains the following” . This keeps the crawler focused on product-related content while avoiding unrelated sections of the website.

✅ 12.2 Content Validation

This tab allows you to filter which pages are processed for email extraction based on their textual content. You can also pre‑process the page HTML (replace or remove strings) before the extraction starts.

📌 Note: Enter custom filters to control which email addresses are extracted. Each line is treated as a separate rule.
Content Validation tab in Settings

Figure 12.2. Content Validation tab – whitelist/blacklist page content before email extraction.

🔄 Replace text on the page before extracting email addresses:

Use the format text to replace = replacement text. One rule per line. This is useful for decoding obfuscated emails (e.g., [at] = @, (dot) = .). The replacement happens before any keyword matching or email extraction.

Example: AT = @ will turn “user AT example.com” into “user@example.com”.
🗑️ Remove text before extracting emails:

One string per line. Any occurrence of these strings will be completely deleted from the page HTML before any extraction or filtering. Useful for stripping out noise like HTML comments, scripts, or irrelevant boilerplate.

Example: <!-- noemail --> will remove those comment blocks entirely.
✅ Extracted email addresses only if the page contains any of the following keywords:

This is a whitelist. Only pages that contain at least one of these keywords (or all, if “Match all keywords” is checked) will have their emails extracted. Leave empty to disable this filter.

  • Match all keywords – requires that the page contains every keyword in the list (AND logic).
  • Match page title only – instead of scanning the full page content, only the <title> tag is examined.
Example: If you add “contact” and “email”, with Match all keywords unchecked, any page containing either “contact” or “email” will be processed. With it checked, the page must contain both.
🚫 Do not extract email addresses if the page contains any of the following keywords:

This is a blacklist. Pages that contain any of these keywords (or all, if “Match all keywords” is checked) will be skipped entirely – no emails extracted.

  • Match all keywords – requires that the page contains every keyword in the list to be skipped (AND logic).
  • Match page title only – restricts the search to the page title.
Example: Add “login”, “captcha”, “terms of service”. Pages with those words will be ignored.

📋 Field Summary

Field / ControlEffect
Replace text before extractionPre‑process HTML: replace specified strings (e.g., obfuscation symbols).
Remove text before extractionDelete specified strings from the page HTML.
Extracted email addresses only if page contains keywordsWhitelist – page must contain at least one (or all) of the keywords.
Do not extract if page contains keywordsBlacklist – page will be skipped if it contains any (or all) of the keywords.
Match all keywordsRequires every keyword to be present (AND logic) instead of any (OR).
Match page title onlyRestricts keyword search to the HTML <title> element only.
💡 Typical use case: To extract emails only from “Contact” or “About” pages, add “contact” and “about” to the whitelist, leave “Match all keywords” unchecked, and leave the blacklist empty. To avoid “noreply” emails, add “noreply” to the blacklist.

✉️ 12.3 Email Filters (Detailed)

This tab provides fine‑grained control over which email addresses are accepted or rejected, how obfuscated symbols are decoded, and whether to validate domains or clean the page content.

Email Filters tab in Settings

Figure 12.3. Email Filters tab – control which emails are extracted, decode obfuscation, and enable validation.

🚫 Exclude emails containing (blacklist):

One keyword, domain, or pattern per line. Any email that contains any of these strings will be rejected. This is useful for filtering out file extensions, common spam domains, or role accounts.

Example entries: png, exe, google.com, @example.com, abuse, antispam.
🎯 Match Right Side of Email (domain whitelist):

Only emails whose domain part (after @) matches one of the given patterns will be kept. Leave empty to disable.

Example: @gmail.com will keep only Gmail addresses. @company.com will keep only internal addresses.
✅ Email must contain (positive filter):

Email must include at least one of these substrings. This is a whitelist for the email address itself (not the page content).

Example: sales will keep only emails containing “sales” (e.g., sales@domain.com).
🔄 Replace Email Symbols:

List of strings to be replaced with @ (for “at” symbols) or . (for dots) before extraction. Each line is a separate pattern. This decodes common obfuscation techniques.

Examples: [at], AT, snabel-a → become @; [dot], DOT, punkt → become ..

Note: The screenshot shows two separate lists – one for dot variants and one for at variants. The software combines them to normalise obfuscated emails like user[at]example[dot]com into user@example.com.

⚙️ Validation & Cleaning Options:
  • Extract only email addresses that belong to the website's domain – Only keep emails where the domain part matches the crawled website’s domain (e.g., john@example.com from https://example.com).
  • Ensure the domain part (after @) is valid during extraction – Performs a DNS lookup (MX record) to verify the domain can receive email. Invalid domains are rejected.
  • Remove Javascript before extracting emails – Strips all <script> tags and their contents from the page HTML before parsing.
  • Remove CSS before extracting emails – Strips all <style> tags and CSS content.

Removing JS and CSS reduces false positives (e.g., emails inside code snippets) and improves parsing speed.

📋 Field Summary

Field / ControlEffect
Exclude emails containing (blacklist)Rejects any email matching any of the listed substrings.
Match Right Side of Email (whitelist)Accepts only emails whose domain matches the given patterns.
Email must contain (positive filter)Email must include at least one listed substring.
Replace Email SymbolsDecodes obfuscated “at” and “dot” symbols before extraction.
Extract only emails belonging to the website's domainRestricts extraction to same‑domain emails.
Ensure domain part is valid (DNS)Performs real‑time MX record validation.
Remove Javascript / Remove CSSStrips script/style tags to clean page content.
💡 Typical use case: To extract only professional business emails from a specific domain, add the domain to “Match Right Side of Email”, enable “Extract only emails belonging to the website's domain”, and add common spam words (e.g., “abuse”, “noreply”) to the exclude list.

🏎️ 12.4 Crawl Engine

This tab configures how the software retrieves and processes web pages, including browser behaviour, delays, scrolling, and resource usage.

Crawl Engine tab in Settings

Figure 12.4. Crawl Engine tab – configure browser behaviour, delays, scrolling, and deep scan.

🆔 User Agent String:

The browser identity sent to web servers. Changing this can help avoid blocking. The screenshot shows a modern Edge/Chrome agent string.

Example: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 ...
🌐 Browser Type:

Note: Changing this option requires a software restart.

  • MS Edge – Recommended for normal scraping tasks without proxies. More stable, lighter, and faster.
  • Chrome – Recommended if you need proxies, proxy rotation, or advanced browser automation features.
⏱️ AJAX delay per page load in milliseconds:

Wait time after a page loads before extracting content. Increase for heavy JavaScript sites to allow dynamic content to render. Default: 300 ms.

🔍 Search Engine Pages Loading Options:
  • Use Browser Rendering for Search Engines – Enables full browser rendering for search engine result pages (needed for Google, Bing, etc.). Recommended for engines that require JavaScript.
  • Scroll Search Engine Pages to load all the pages – Simulates scrolling to load infinite results or “load more” buttons.
  • Auto adjust search engine delay timing to avoid blocking – Randomises delays between requests to reduce the risk of IP bans.
🧠 Full Deep Scan Mode (Browser-Based):

Load JavaScript and dynamic content. This improves compatibility with modern websites but reduces scraping speed.

  • Scrape all the websites using browser – Enables browser rendering for every page, not just search engines.
  • Do not load images in crawling the page – Speeds up rendering and reduces bandwidth.
  • Do not load videos in crawling the page – Similarly, blocks video content.
  • Scroll page to load all the contents / pages – Simulates scrolling to trigger lazy‑loaded content.
  • Number of concurrent browser pages – How many browser instances run in parallel (default 6). Increase with caution.
  • Maximum pages to scroll – How many “scroll steps” to perform on a single page (default 10).
  • Scroll step points – Pixels to scroll each time (default 200).
  • Scroll step delay in milliseconds – Pause between scroll steps (default 300 ms).

📋 Field Summary

Field / ControlEffect
User Agent StringIdentifies the browser to web servers; can bypass simple blocks.
Browser Type (Edge / Chrome)Edge is lighter/faster; Chrome supports proxy rotation and advanced automation.
AJAX delay (ms)Wait time for dynamic content to load after page ready.
Use Browser Rendering for Search EnginesEnables JS rendering for search result pages (required for Google, Bing).
Scroll Search Engine PagesSimulates scrolling to load all results.
Auto adjust delay to avoid blockingRandomises delays to avoid pattern detection.
Full Deep Scan (browser for all sites)Renders every page with a browser – slower but necessary for SPAs.
Do not load images / videosIncreases speed by skipping media.
Scroll page to load all contentsActivates infinite scroll simulation.
Concurrent browser pagesNumber of simultaneous browser instances.
Maximum pages to scroll, step points, step delayFine‑tunes the auto‑scrolling behaviour.
💡 Best practice: For normal websites, use MS Edge browser type, disable Full Deep Scan, and keep AJAX delay low (300 ms). For JavaScript‑heavy sites (React, Angular) or when using proxies, switch to Chrome and enable Full Deep Scan, but reduce concurrent browser pages to 2‑4 to avoid high memory usage.

📏 12.5 Crawling Rules

This tab controls the crawler’s behaviour, limits, and adherence to web standards. If you are not familiar with technical details, it is recommended to leave the default values unchanged.

Crawling Rules tab in Settings

Figure 12.5. Crawling Rules tab – limits, authorisation, and crawler behaviour flags.

📊 Crawl Limits
  • Maximum Pages to Crawl: Global limit across all domains. Default 10,000,000 (practically unlimited).
  • Max Search engine Pages per keyword: How many result pages to fetch per search keyword. Default 100.
  • Maximum pages search per domain: Stop crawling a single domain after N pages. Default 10,000,000.
  • Maximum Exhaustive Pages Without Emails per domain: Deprecated / similar to the next setting. (Kept for compatibility.)
  • Crawl Delay per Domain (milliseconds): Minimum time between requests to the same domain (1 ms default). Increase to avoid overloading servers or being blocked.
  • Maximum Emails to Extract per Domain: Stop extracting from a domain after N emails. Default 10,000.
  • Max Consecutive Pages Without Emails per domain: If a domain returns no emails for N consecutive pages, it is abandoned. Default 10,000.
⚙️ Advanced Limits (set to 0 for maximum value)
  • Maximum links to crawl per page: Limits the number of links extracted from a single page. 0 = unlimited.
  • Maximum amount of memory to use (MB): Soft memory limit. 0 = no limit.
  • Maximum page request time-out (Seconds): HTTP timeout. 0 = default (~30 seconds).
🔐 HTTP request authorization

Check “Each HTTP request should be authorized via login?” to enable basic authentication. Then provide a Username and Password. This is required for password‑protected websites.

⚙️ Standard Fast Crawler Engine Behaviour Settings
  • Enabled Cookies – Maintains cookies across requests (essential for sessions).
  • Don't allow crawling of already crawled pages? – Skips visited URLs to prevent loops.
  • Enable SSL certificate validation – Verifies HTTPS certificates (recommended ON).
  • Extract from Excel, word pages – Parses .docx, .xlsx files (slower).
  • Automatically Handle Gzip/Deflate Compression – Reduces bandwidth (ON).
  • Ignore Robots.txt file if Root URL Disallowed – Bypasses robots.txt restrictions (use with caution).
  • Ignore Links with rel='nofollow' – If ON, still crawl nofollow links. If OFF, respect them.
  • Delete previous logs on every new search – Starts a fresh log file each extraction.
  • Respect X-Robots-Tag 'nofollow' Links – Obeys the HTTP header `X-Robots-Tag: nofollow`.
  • Find links within text and Javascript also – Extracts URLs from inline JavaScript (may increase false positives).
  • Ignore Links on Pages with Meta Robots 'nofollow' – If ON, ignores meta robots directives.
  • Continuous Connectivity Monitoring – Pauses crawling when internet is lost, resumes automatically.
  • Respect robots.txt Rules – Standard compliance (recommended ON).

📋 Field Summary

Field / ControlEffect
Maximum Pages to CrawlGlobal crawl limit (0 = unlimited).
Max Search engine Pages per keywordNumber of search result pages per keyword.
Maximum pages per domainStop after N pages on a single domain.
Crawl Delay per Domain (ms)Minimum time between requests to the same domain.
Maximum Emails per DomainStop extracting from a domain after N emails.
Max Consecutive Pages Without EmailsAbandon domain after N empty pages.
HTTP request authorizationEnable basic authentication (username/password).
Behaviour checkboxesVarious options to control cookie handling, robots.txt respect, nofollow, link extraction, logging, etc.
💡 Best practice: To avoid being blocked, set a crawl delay of 1000–3000 ms, enable “Continuous Connectivity Monitoring”, and respect robots.txt. For deep crawls, increase “Maximum Pages to Crawl” but keep “Maximum pages per domain” reasonable (e.g., 10,000).

🌍 12.6 Country Filter

This tab allows you to restrict crawling to specific countries based on top‑level domains (TLDs). It is useful for targeting local businesses or avoiding irrelevant international results.

📌 Note: Choose one or more countries to restrict extraction to their domain extensions (e.g., .uk, .au). Leave empty to include all countries.
Country Filter tab in Settings

Figure 12.6. Country Filter tab – restrict crawling to specific TLDs.

📋 Country / TLD Table

A list of countries with their corresponding top‑level domains. Double‑click a row to select or deselect that country. Selected countries will be used to filter crawled domains.

Country (Double Click Row to Edit)Top Level Domain
1 International commercial.com
2 International organization.org
3 International network.net
4 International organizations.int
5 International Education.edu
6 U.S. national and state government agencies.gov
7 U.S. military.mil
Afghanistan.af
Åland Islands.ax
Albania.al
Algeria.dz
American Samoa.as
Andorra.ad
Angola.ao
Anguilla.ai
Antarctica.aq
Antigua and Barbuda.ag
Argentina.ar
Armenia.am
Aruba.aw
Ascension Island.ac
Australia.au
Austria.at
Azerbaijan.az
Bahamas.bs
Bahrain.bh
Bangladesh.bd
Barbados.bb

The list continues beyond what is shown; scroll down to see all countries.

🎛️ Selection Controls
  • Country Name: (label) – the currently highlighted country in the table.
  • Top Level Domain: (label) – the corresponding TLD of the highlighted country.
  • Save the TLD: Button – adds the currently highlighted TLD to the filter list. Effectively selects that country.
  • Selected Countries: (label) – shows the list of countries (or TLDs) that have been chosen.
  • Clear checked countries: Button – removes all selected countries, returning to “include all countries” mode.
📖 How to use: Double‑click a country row to select it (the row highlights). Click “Save the TLD” to add that TLD to the filter. To remove a selection, double‑click again or use “Clear checked countries” to remove all.

📋 Field Summary

Field / ControlEffect
Country / TLD tableDisplays all available countries and their TLDs. Double‑click to select/deselect.
Save the TLDAdds the TLD of the highlighted country to the filter list.
Selected Countries (label)Shows the currently selected TLDs (or countries).
Clear checked countriesRemoves all selections, disables country filtering.
💡 Example: To extract only Australian and New Zealand websites, double‑click “Australia” (.au) and “New Zealand” (.nz) in the table, then click “Save the TLD” for each. The crawler will then ignore domains that do not end with .au or .nz.

🎞️ 12.7 MIME Filters & Auto‑Save

This tab controls which file types the crawler processes and how automatic backups behave. Filtering out unnecessary MIME types improves extraction speed and reduces bandwidth.

MIME Filters and Auto-Save tab in Settings

Figure 12.7. MIME Filters & Auto‑Save tab – control which file types are processed and backup behaviour.

💾 Auto‑Save Last Search

When enabled, extracted data and links are automatically saved. If the software or computer shuts down unexpectedly, your results will be restored when you restart.

  • Enable Auto‑Save – Master switch for automatic backup.
  • Save fetched data after every [X] items – Number of emails after which the grid is saved (default 100).
  • Automatically load the last saved search data on startup – Restores previous session when the software launches.
  • Save processed page links – Also stores the list of crawled URLs (previous logs will be overwritten when starting a new search).
📄 MIME Types

Specify which file formats the crawler should process (e.g., text/html, text/plain). Filtering non‑essential types improves extraction speed and reduces data usage.

Mime TypeDescription
text/htmlHTML pages
text/plainPlain text files
text/cssCSS stylesheets
text/csvCSV files
text/xmlXML documents
application/xhtml+xmlXHTML documents
application/javascriptJavaScript files
application/jsonJSON data
application/rss+xmlRSS feeds
application/xmlGeneric XML
application/pdfPDF documents
application/vnd.mozilla.xul+xmlXUL interfaces
application/vnd.oasis.opendocument.textOpenDocument text
application/vnd.openxmlformats-officedocument.wordprocessingml.documentMicrosoft Word (DOCX)

Add new Mime Type: Type a custom MIME string (e.g., application/vnd.ms-excel) and click the add button (not shown in screenshot, but present in the actual interface).

📋 Field Summary

Field / ControlEffect
Enable Auto‑SaveTurns automatic backup on/off.
Save fetched data after every X itemsFrequency of auto‑save (number of new emails).
Automatically load last saved search on startupRestores previous session automatically.
Save processed page linksBacks up crawled URL list.
MIME Types tableLists allowed content types; crawler only processes these.
Add new Mime TypeAdds a custom MIME to the allowed list.
💡 Best practice: Remove MIME types you do not need (e.g., images, videos, CSS) to significantly speed up crawling. Keep text/html, text/plain, application/pdf if you need those formats. Enable Auto‑Save with a low threshold (e.g., 100) to protect against data loss.

🔒 12.8 Proxy Settings

This tab allows you to configure HTTP proxies for anonymous crawling, rotating IP addresses, or accessing geo‑restricted content. Only HTTP proxies are supported (not SOCKS).

Proxy Settings tab in Settings

Figure 12.8. Proxy Settings tab – configure HTTP proxies for anonymous crawling.

➕ Add / Edit HTTP Proxy Settings

Each field can contain multiple lines. The number of lines in each field must match (or be 1 if the same value applies to all proxies).

  • HTTP Proxy Address: One IP or domain per line. Example: 192.168.1.100 or proxy.example.com. The number of addresses should match the number of ports, usernames, and passwords.
  • Port Number: One port per line, or a single port if all proxies use the same port. Example: 8080.
  • Username (Proxy Authentication): One username per line, or a single username for all proxies. Leave blank if no authentication required.
  • Password: One password per line, or a single password for all proxies.
📌 Matching rule: If you have 5 proxy addresses, you must provide either 5 port numbers (one per proxy) or 1 port number (same for all). The same applies to usernames and passwords.
📋 Proxy List Table

Displays the currently configured proxies with their server address, port, and username. Double‑click a row to edit.

Proxy ServerPortUsername
(empty – add proxies using the fields above)

Delete Selected: Highlight a row in the table and click this button to remove it.

⚠️ Important: Only HTTP proxies are supported. For browser‑based crawling with proxies, you must also select the Chrome browser type in the Crawl Engine tab (see section 12.4).

📋 Field Summary

Field / ControlEffect
HTTP Proxy AddressOne IP or domain per line. Matched by line number to ports/credentials.
Port NumberOne port per line, or single port for all proxies.
Username / PasswordOptional authentication credentials. One per line or single for all.
Proxy List TableShows configured proxies; double‑click to edit; Delete Selected removes rows.
Restore DefaultsClears all proxy settings.
Save / Apply buttonsPersist or temporarily apply proxy configuration.
💡 Example configuration: To rotate through three proxies, enter three lines in HTTP Proxy Address, three lines in Port Number (or one port if same), and matching usernames/passwords if required. Click Save All Settings Permanently. Then enable “Use the selected proxies” (checkbox is located below the table in the actual software – not fully shown in screenshot, but present). The crawler will then randomly assign a proxy to each request.

🤖 12.9 Captcha Settings

This tab allows you to define keywords that indicate a CAPTCHA page and choose how the software reacts when one is detected. Proper configuration helps avoid getting stuck on blocked pages.

Captcha Settings tab in Settings

Figure 12.9. Captcha Settings tab – define CAPTCHA indicators and choose handling behaviour.

🔍 Captcha Page Keywords

Captcha keywords may include HTML class names, tag attributes, or any unique strings found in the page source that identify a captcha page. Multiple keywords can be added for each website.

📌 How it works: For each domain (e.g., google.com), you enter a list of keywords. If the page source contains any of those keywords, the software recognises it as a CAPTCHA page and triggers the selected action.
Domain NameAssociated Keywords (examples)
googlename="captcha", CheckboxCaptcha-Button, smartcaptcha, id="checkbox-captcha-form", data-testid="checkbox-captcha", class="captcha", id="turnstile-widget"
rambler.runova.rambler.ru/services/antifrod/captcha/, class="passMod_slide-control"
(other domains)Add your own keywords based on HTML inspection.

To add a new keyword: select a domain from the dropdown (or type a new domain), enter a keyword, and click the add button (interface not fully shown in screenshot but present in the software).

⚙️ Captcha Options
  • Pause extraction when a CAPTCHA is detected until it is resolved. – Completely stops all crawling activity and waits for the user to manually solve the CAPTCHA (e.g., in the built‑in browser). Extraction resumes only after the CAPTCHA is solved.
  • If a CAPTCHA appears, display it for manual resolution while continuing extraction. – Shows the CAPTCHA in a separate window, but other browser instances keep working. This is the recommended option for large‑scale scraping.
  • Continue extraction even if a CAPTCHA appears during search. – Ignores CAPTCHA pages. The crawler will skip the blocked page and move on, but no data will be extracted from that domain.

📋 Field Summary

Field / ControlEffect
Domain Name dropdown/listSelect or type the website domain for which you want to add CAPTCHA keywords.
Captcha keywords (per domain)List of strings that identify a CAPTCHA page (HTML IDs, class names, URLs, etc.).
Pause extraction until resolvedStops all activity, waits for manual CAPTCHA solving.
Display CAPTCHA while continuingShows CAPTCHA in a pop‑up but other browsers keep crawling.
Continue extraction (ignore CAPTCHA)Skips the blocked page, no extraction from that domain.
💡 Best practice: For search engines (Google, Bing, etc.), add keywords like class="g-recaptcha", id="captcha", name="captcha". Choose the second option (“display it for manual resolution while continuing extraction”) so that you can solve CAPTCHAs without stopping the entire crawl.

13. Regional Settings Dialog

Available in Search Engine mode via Click to do Regional Settings. Allows selecting country, state/province, up to three cities, and specific area. Options:

Opens the Regional Settings dialog, allowing you to target searches to specific geographic locations. Regional targeting helps improve search relevance and discover websites associated with a particular country, region, city, or local area.

Keywords List

Figure 1. Regional Settings window

Available Location Filters
Country

Restrict searches to a specific country.

State / Province

Focus searches within a selected state or province.

City Selection

Target up to three cities simultaneously.

Area / Town

Further narrow searches to a specific local area.

Additional Options
Force Regional Matching

Instructs the search engine to prioritize results matching the selected regional settings whenever possible.

Country-Level Domains Only

Extract emails only from websites using country-specific top-level domains such as .us, .uk, .ca, .au, or .de.

14. Additional Filter Dialogs

📧 Enable Email Search Filters Dialog

This dialog appears when you check “Enable Email Search Filters” in the Search Engine advanced options. It lets you define which pages are considered valid for email extraction based on their content and optional geographic fields.

Enable Email Search Filters dialog

Figure 14.1. Enable Email Search Filters dialog – define page‑level content and location filters.

📌 How it works: Add specific keywords or phrases that must be present on a searched webpage for the software to extract email addresses from it. If none of the defined keywords or phrases are found, the software will skip that page and move to the next one.
🔑 Keywords / Phrases (one per line)

Enter one keyword or phrase per line. Leave blank to skip filtering. The software checks the page’s full HTML (or title, depending on other settings) for any of these keywords.

Examples: marketing, software, contact us, sales@
⚠️ Tip: Using inaccurate or overly broad keywords may reduce extraction accuracy by including irrelevant pages, while overly restrictive keywords may cause the software to skip valid results.
📍 Select one or more location fields to match on the crawled page

These checkboxes tell the software to also check the page for specific location information (if present in the HTML, e.g., from the keyword row’s Region/City/Province/Country columns). When checked, the page must match the corresponding location value in addition to the keyword filter.

  • Region – match the keyword row’s region field.
  • City – match the keyword row’s city field.
  • State/Province – match the keyword row’s province/state field.
  • Country – match the keyword row’s country field.
Example: If you have a keyword row with Country = “Germany” and you check the “Country” box, the page will only be processed if it contains both the keyword and “Germany” (or a .de domain, depending on your settings).
Apply Settings Cancel

Apply Settings – saves the current keyword and location field selections for the current session.
Cancel – closes the dialog without saving any changes.

📋 Field Summary

Field / ControlEffect
Keywords / Phrases (multi‑line)Only pages containing at least one of these strings will yield email extraction. Leave empty to disable.
Region checkboxRequire the page to match the keyword row’s region value.
City checkboxRequire the page to match the keyword row’s city value.
State/Province checkboxRequire the page to match the keyword row’s province/state.
Country checkboxRequire the page to match the keyword row’s country.
Apply SettingsSave and use the current filter configuration.
CancelDiscard changes and close the dialog.
💡 Example: To extract emails only from German pages that mention “marketing”, add “marketing” as a keyword, check the Country location field, and ensure your keyword row has “Germany” in the Country column. The software will then process only pages that contain “marketing” and are associated with Germany.

💡 15. Best Practices & Troubleshooting

🚀 Performance Optimization

  • Adjust concurrency wisely: Start with 30 crawler threads and 8 parser threads (defaults). Increase crawler threads to 50–80 on high‑bandwidth connections, but keep parser threads lower (10–15) to avoid CPU overload.
  • Disable unnecessary assets: In Crawl Engine → Full Deep Scan Mode, check “Do not load images” and “Do not load videos”. This dramatically speeds up browser‑based crawling.
  • Choose the right browser type: Use MS Edge for normal HTTP crawling (lightweight, fast). Switch to Chrome only when you need proxy rotation or advanced automation.
  • Reduce AJAX delay for fast sites: Lower “AJAX delay per page load” to 100–200 ms for static pages. Increase to 1000–3000 ms for heavy JavaScript sites.
  • Limit concurrent browser pages: When using “Scrape all websites using browser”, set “Number of concurrent browser pages” to 2‑4 on older machines to prevent memory exhaustion.
  • Filter MIME types: In MIME Filters, remove types you don’t need (e.g., images, CSS, JSON) to reduce bandwidth and processing time.

🛡️ Avoiding IP Bans & CAPTCHAs

  • Use proxy rotation: Configure multiple HTTP proxies in Proxy Settings and enable “Use the selected proxies”. Also select Chrome browser type for full proxy support.
  • Enable auto‑adjust delay: In Crawl Engine, check “Auto adjust search engine delay timing” – this randomises request intervals to avoid pattern detection.
  • Increase crawl delay: In Crawling Rules, set “Crawl Delay per Domain” to 1000–3000 ms when targeting aggressive sites.
  • Respect robots.txt: Keep “Respect robots.txt Rules” enabled (default) to avoid being blocked by site policies.
  • Rotate user agents: Use the built‑in random user‑agent feature (under Crawl Engine) to avoid fingerprinting.
  • Handle CAPTCHAs gracefully: In Captcha Settings, choose “If a CAPTCHA appears, display it for manual resolution while continuing extraction”. This lets you solve CAPTCHAs without stopping the entire crawl.

💾 Data Safety & Recovery

  • Enable Auto‑Save: In MIME Filters tab, check “Enable Auto‑Save” and set a reasonable interval (e.g., after every 100 emails). This protects against crashes.
  • Use separate auto‑save folders for multiple instances: Each running copy of the software must have its own auto‑save directory (configured via File → Set Auto‑save Directory). Sharing a folder corrupts session data.
  • Load last search after crash: Use File → Load Last Search Results to restore your grid, queue, and logs.
  • Regularly export your grid: Click Save Email Addresses periodically to external Excel/CSV files as an additional backup.
  • Save URL lists: Use File → Save the last search URLs list in file to export discovered URLs for later resumption.

🐞 Common Errors and Solutions

Error / SymptomLikely Cause & Solution
Empty email gridTarget site uses heavy JavaScript. Enable Full Deep Scan Mode (Crawl Engine) and/or use Extract with Browser mode. Also check your Content Validation whitelist – maybe too restrictive.
“WebView2 runtime not found”Microsoft Edge WebView2 is not installed. Download from Microsoft and install the Evergreen Runtime.
CAPTCHA appears repeatedlyIncrease crawl delay, enable proxy rotation, and set CAPTCHA handling to “display and continue”. Solve manually when prompted.
License validation failsCheck your internet connection. If persistent, copy the Software Serial Key from the Registration dialog and your purchase order number and email aslogger@ahmadsoftware.com.
Slow crawling or high memory usageReduce “Number of concurrent browser pages” to 2‑4. Lower “Maximum pages to scroll” to 5. Disable image/video loading. Set a memory limit in Crawling Rules.
Extracted emails contain [at] or (dot)Add those strings to Replace Email Symbols in the Email Filters tab (e.g., “[at] = @”, “(dot) = .”).
Crawler stops after few pagesYou may be blocked. Enable proxy rotation, increase crawl delay, and enable auto‑adjust delay. Also check “Continuous Connectivity Monitoring” in Crawling Rules.
No results from search engines (Google, Bing)Enable Use Browser Rendering for Search Engines in Crawl Engine. Also ensure you have not excluded “google.com” in Domain/URL Filters.

⚙️ Recommended Settings for Different Scenarios

ScenarioRecommended Configuration
Fast scraping of static sitesBrowser type: MS Edge, Full Deep Scan: OFF, Crawler threads: 50, Parser threads: 10, AJAX delay: 0, Crawl delay: 100 ms, Disable images/videos.
JavaScript-heavy sites (SPAs)Browser type: Chrome, enable Full Deep Scan with “Scrape all websites using browser”, Concurrent browser pages: 4, Scroll steps: 300, AJAX delay: 1000 ms.
Anonymous scraping (avoid blocks)Use proxy rotation (HTTP proxies + Chrome browser type), enable auto‑adjust delay, crawl delay: 2000‑5000 ms, enable “Respect robots.txt”, CAPTCHA handling: “display and continue”.
Local file extraction (PDF, DOCX, XLSX)No special settings – just add files/folders in Computer Files mode. For large files, ensure memory limit is set to 1024 MB or higher.
Targeted extraction (specific countries/TLDs)Use Country Filter to select desired TLDs (e.g., .de, .fr). Also use Regional Settings in Search Engine mode with location fields.
High‑volume extraction (millions of pages)Increase “Maximum Pages to Crawl” to 0 (unlimited). Use fast crawling settings (MS Edge, no browser rendering). Enable auto‑save every 1000 items. Use a powerful machine with 16+ GB RAM.

✅ Quick Troubleshooting Checklist

  • Are you using the latest version? Check Help → Check For Updates.
  • Is WebView2 runtime installed? (Required for browser mode and search engine rendering).
  • Is your firewall blocking the software? Add exceptions for the executable.
  • Have you set a valid auto‑save directory? (Without it, session recovery is impossible).
  • Are your proxy settings correct? Test them with Test Proxies button.
  • Are your email filters too strict? Temporarily disable them to see if emails appear.
  • Is the target website still online? Try opening it in a normal browser.

For further assistance, refer to the online help (Help → Help) or contact support at aslogger@ahmadsoftware.com.

16. Support & Contact

Website: www.ahmadsoftware.com
Email: aslogger@ahmadsoftware.com
Video tutorials: Click Watch button in the software.