SkillAgentSearch skills...

Advanced Sitemap Parser

XML sitemap parser designed to extract and process millions of URLs while bypassing most modern anti-bot protections. Supports plain and compressed XML, unlimited nested sitemaps, multi-threading, multiple inputs, CloudScraper integration, fingerprint randomization, proxy/user agent rotation, auto stealth mode, and detailed monitoring.

Install / Use

npx skills add phase3dev/advanced-sitemap-parser

Installs into whichever agent you are using.

README

Advanced Sitemap Parser

An XML sitemap parser designed for large-scale URL extraction, capable of processing millions of URLs while bypassing most modern anti-bot protections. Supports both plain XML and compressed XML files (.xml.gz), as well as unlimited levels of nested child sitemaps. Sitemaps can be loaded directly from URLs, from a file containing multiple sitemap URLs, or from a local directory of .xml and .xml.gz files.

Also features a number of optional and advanced settings such as dynamic proxy and user agent rotation, CloudScraper integration, fingerprint randomization, auto stealth mode, and includes detailed logging and monitoring.

Key Features

Anti-Detection System

  • Human-like browsing patterns with randomized delays (configurable 3-30+ seconds)
  • Intelligent retry logic with exponential backoff for 403/429 errors
  • Real-time proxy rotation - switches IP for every single request
  • Dynamic user agent rotation - changes browser fingerprint per request
  • CloudScraper integration with browser fingerprint randomization
  • Enhanced HTTP headers mimicking real browser behavior (Sec-CH-UA, DNT, etc.)
  • Request timing variation with occasional longer "human-like" pauses
  • Stealth mode for maximum evasion with extended delays
  • Advanced fingerprint randomization without browser overhead

Multi-Threading with Intelligent Staggering

  • Configurable concurrency from 1 (stealth) to unlimited workers
  • Automatic request staggering (0.5-2 second delays between thread starts)
  • Thread-safe processing with proper interrupt handling
  • Sequential fallback for maximum stealth operations

Proxy System

  • Multiple proxy formats supported:
    • ip:port (basic HTTP proxy)
    • ip:port:username:password (authenticated proxy)
    • http://ip:port (full URL format)
    • http://username:password@ip:port (authenticated URL format)
  • Automatic proxy rotation - never reuses the same IP consecutively
  • Proxy health monitoring with real-time IP display
  • External proxy file support - load thousands of proxies from text file
  • Fallback to direct connection if no proxies provided

User Agent Management

  • Built-in user agent database with current browser versions
  • External user agent file support for custom UA lists
  • Per-request rotation ensuring maximum diversity
  • Real-time UA display showing current browser fingerprint
  • Browser-specific header matching (Chrome, Firefox, Safari, Edge)
  • CloudScraper override forcing custom user agents

Sitemap Processing Engine

  • Unlimited nested sitemap support - processes sitemap indexes recursively
  • Compressed file handling - automatic .gz decompression
  • XML namespace awareness - proper sitemap standard compliance
  • Duplicate URL prevention - intelligent deduplication system
  • Progress monitoring with real-time queue size and URL counts
  • Error recovery - continues processing despite individual failures
  • Failed URL tracking with detailed error reporting

Input/Output Management

  • Multiple input methods:
    • Single sitemap URL (--url)
    • Batch processing from file (--file)
    • Directory scanning for local .xml and .xml.gz files (--directory)
  • Configurable output directory (--save-dir)
  • Smart filename generation with readable source hints plus a short stability hash
  • Organized output files with metadata headers
  • Failed URL recovery files for easy reprocessing
  • Summary reporting with processing statistics
  • Comprehensive logging to sitemap_processing.log

Monitoring & Error Handling

  • Graceful interrupt handling - Ctrl+C stops processing cleanly
  • Session persistence through connection errors
  • Automatic retry mechanisms with configurable limits
  • Memory efficient processing of large sitemaps
  • Real-time progress reporting with timestamps
  • Detailed error tracking and retry statistics
  • Cross-platform compatibility (Linux, macOS, Windows)

Installation

Supported Python: 3.9+

# Clone the repository
git clone https://github.com/phase3dev/advanced-sitemap-parser.git
cd advanced-sitemap-parser

# Install dependencies
pip install -r requirements.txt

Quick Start

Basic Usage

# Simple sitemap extraction
python3 sitemap_extract.py --url https://example.com/sitemap.xml

# With output directory
python3 sitemap_extract.py --url https://example.com/sitemap.xml --save-dir ./results

Stealth Mode

# Maximum stealth with longer delays
python3 sitemap_extract.py --url https://example.com/sitemap.xml --stealth

# Stealth with proxy rotation
python3 sitemap_extract.py --url https://example.com/sitemap.xml --stealth --proxy-file proxies.txt

Proxy Rotation

# Basic proxy rotation
python3 sitemap_extract.py --url https://example.com/sitemap.xml --proxy-file proxies.txt

# Proxy + custom user agents
python3 sitemap_extract.py --url https://example.com/sitemap.xml --proxy-file proxies.txt --user-agent-file user_agents.txt

Multi-Threading

# Moderate concurrency (3 threads with staggering)
python3 sitemap_extract.py --url https://example.com/sitemap.xml --max-workers 3 --proxy-file proxies.txt

# High-speed processing (use with caution)
python3 sitemap_extract.py --url https://example.com/sitemap.xml --max-workers 10 --min-delay 1 --max-delay 3

Batch Processing

# Process multiple sitemaps from file
python3 sitemap_extract.py --file sitemap_urls.txt --proxy-file proxies.txt --stealth

# Reprocess failed URLs
python3 sitemap_extract.py --file failed_sitemap_urls.txt --proxy-file proxies.txt --max-retries 5

Timing Controls

# Custom delays for specific sites
python3 sitemap_extract.py --url https://example.com/sitemap.xml --min-delay 5 --max-delay 15 --max-retries 5

# Fast processing (less stealthy)
python3 sitemap_extract.py --url https://example.com/sitemap.xml --min-delay 0.5 --max-delay 2 --max-workers 5

Complete Command Reference

python3 sitemap_extract.py [OPTIONS]

# Input Options
--url URL                    Direct URL of sitemap file
--file FILE                  File containing list of sitemap URLs
--directory DIR              Directory containing .xml and .xml.gz files

# Output Options
--save-dir DIR               Directory to save all output files (default: current)

# Anti-Detection Options
--proxy-file FILE            File containing proxy list (see format below)
--user-agent-file FILE       File containing user agent list
--stealth                    Maximum evasion mode (5-12s delays, forces --max-workers=1)
--no-cloudscraper            Use standard requests instead of CloudScraper

# Performance Options
--max-workers N              Concurrent workers (default: 1 for max stealth)
--min-delay SECONDS          Minimum delay between requests (default: 3.0)
--max-delay SECONDS          Maximum delay between requests (default: 8.0)
--max-retries N              Maximum retry attempts per URL (default: 3)

File Formats

Proxy File Format (proxies.txt)

# Basic format
192.168.1.100:8080
proxy.example.com:3128

# With authentication
192.168.1.101:8080:username:password

# Full URL format
http://proxy.example.com:8080
http://user:pass@proxy.example.com:8080

User Agent File Format (user_agents.txt)

# One user agent per line
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36
Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/121.0

Sitemap URL File Format (sitemap_urls.txt)

# One URL per line
https://www.example.com/sitemap.xml
https://www.example.com/sitemap_index.xml
https://www.example.com/sitemaps/sitemap.xml.gz

Output Files

Individual Sitemap Files

  • Format: domain_com_hint_<short-hash>.txt
  • Contains: Deduplicated page URLs from that specific sitemap source only
  • Metadata: Source URL, generation timestamp, URL count
  • Readability: Remote files include the host plus path/query hints when available, and local files include the parent directory plus file stem
  • Uniqueness: The short hash is derived from the full original source string, so query-distinct child sitemap URLs do not overwrite each other
  • Example: www_example_com_sitemap_sitemap1_<short-hash>.txt

Merged URL File

  • all_extracted_urls.txt: Deduplicated union of all extracted page URLs from the run, with the same metadata header format as per-sitemap files

Summary Files

  • all_sitemaps_summary.log: Complete inventory of all sitemaps processed
  • failed_sitemap_urls.txt: Clean list of failed URLs for reprocessing
  • sitemap_processing.log: Detailed processing log with timestamps

Failed URL Recovery

When sitemaps fail to process, they're automatically saved to failed_sitemap_urls.txt for easy reprocessing:

# Reprocess failed URLs with different settings
python3 sitemap_extract.py --file failed_sitemap_urls.txt --proxy-file proxies.txt --max-retries 5

Real-Time Monitoring

The script provides comprehensive real-time feedback:

[2025-08-29 05:52:39] Loaded 1000 proxies from proxies.txt
[2025-08-29 05:52:39] Loaded 1000 user agents from user_agents.txt
[2025-08-29 05:52:39] Max concurrent workers: 3
[2025-08-29 05:52:39] Press Ctrl+C to stop gracefully...
[2025-08-29 05:52:39] Fetching (attempt 1): https://example.com/sitemap.xml
[2025-08-29 05:52:39] Using IP: 192.168.1.100
[2025-08-29 05:52:39] Using User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64)...
[2025-08-29 05:52:40] SUCCESS with 192.168.1.100
[2025-08-29 05:52:40] Processed: 15 nested sitemaps, 1,250 pages
[2025-08-29 05:52:40] Saved 1,250 URLs to exam

Related Skills

View on GitHub
GitHub Stars92
CategoryOperations
Updated9d ago
Forks7

Languages

Python

Security Score

100/100

Audited on Jul 30, 2026

No findings