Petco Knowledge Base Scraper
Advanced headless Chrome scraper that classifies, sections, and generates Q&A pairs for internal Petco knowledge base content, producing multiple training datasets with LLM-powered content analysis and progress tracking for AI model training.
Status: completed · 2024-03-01
Overview
The Petco Knowledge Base Scraper is a sophisticated toolset for automating the extraction and processing of internal knowledge base content. Using headless Chrome with Puppeteer, it connects to authenticated browser sessions, classifies content using LLM analysis, breaks documents into logical sections, and generates high-quality Q&A pairs for AI training data, with robust progress tracking and resumable operations.
Technologies
Puppeteer, Python, Node.js, JavaScript, Headless Chrome, LLM Integration, Data Processing, Content Classification, Q&A Generation, Progress Tracking, Concurrent Scraping, JSON Processing, Browser Automation, Error Handling, Resume Capability, Training Data Generation, Content Sectioning, Web Scraping, Automation, Data Pipeline, Chrome Debugging, LLM-Powered Analysis, Hierarchical Crawling, Data Cleaning, File Management
- Scraping Concurrency
- 5 Parallel Requests
- Data Types Generated
- 4 Formats
- Progress Tracking
- Resume-Capable
- Error Recovery
- Comprehensive
Petco Knowledge Base Scraper
Developed by the AI and Automation Team
This repository contains tools for scraping and processing Petco's internal knowledge base content to generate training data for AI models.
Prerequisites
- Chrome/Chromium browser
- Python 3.x or Node.js (depending on preferred scraping method)
Browser Setup
Before running the scraper, you need to launch Chrome/Chromium with remote debugging enabled. Use the appropriate command for your operating system:
Windows
```cmd "C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222 --no-first-run --no-default-browser-check --user-data-dir="%TEMP%\chrome_debug_profile" ```
macOS
```bash "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" --remote-debugging-port=9222 --no-first-run --no-default-browser-check --user-data-dir="/tmp/chrome_debug_profile" ```
Linux
```bash chromium --remote-debugging-port=9222 --no-first-run --no-default-browser-check --user-data-dir=/tmp/chrome_debug_profile ``` or ```bash google-chrome --remote-debugging-port=9222 --no-first-run --no-default-browser-check --user-data-dir=/tmp/chrome_debug_profile ```
Dependencies
Python Dependencies
``` pyppeteer pyzerox requests asyncio aiofiles uuid dataclasses ```
JavaScript Dependencies
``` puppeteer-core axios p-limit uuid ```
Scraping Process
- Launch Chrome with debugging enabled using the commands above
- Log in to the Petco Knowledge Base using your credentials
- Keep the browser window open
- Run one of the scraping scripts (Python or JavaScript)
Available Scripts
Primary Scraper (Recommended)
- `scraperV2.js`: Advanced JavaScript implementation that:
- Uses puppeteer to connect to Chrome
- Implements smart concurrent scraping (limited to 5 parallel requests)
- Analyzes document content using LLM to classify content types
- Breaks documents into logical sections for better processing
- Generates high-quality Q&A pairs for AI training data
- Includes robust progress tracking to resume interrupted scraping
- Saves multiple data types:
- Raw pre-training data
- Analyzed content with classifications
- Sectioned content for fine-tuning
- Q&A pairs for model training
- Features comprehensive error handling and recovery
Alternative Scrapers
Python Scraper
- `main.py`: Python implementation that mirrors scraperV2.js functionality:
- Connects to Chrome's debug port
- Implements hierarchical crawling of the knowledge base
- Uses LLM for content analysis and Q&A generation
- Supports concurrent processing with ThreadPoolExecutor
- Includes progress tracking for resumable operations
- Saves multiple data formats in `training_data_v2`
To test the Python scraper:
- Install additional dependencies: ```bash pip install aiofiles requests uuid dataclasses pyppeteer ```
- Start Chrome with debugging enabled
- Run the scraper: ```bash python main.py ```
- Monitor the `training_data_v2` directory for output files:
- `stock-pretraining-data-*.json`: Raw document content
- `analyzed-pretraining-data-*.json`: Documents with content analysis
- `pretraining-data-sections-*.json`: Sectioned content
- `fine-tuning-data-*.json`: Generated Q&A pairs
Legacy JavaScript Scraper
- `converter.js`: Original JavaScript implementation focused on document conversion:
- Connects to Chrome's debug port
- Extracts content from knowledge base articles
- Processes documents into sections using a local LLM
- Generates Q&A pairs with custom grammar rules
- Saves output to `training_data` directory:
- `documents_[date].json`: Raw document content
- `qa_[date]_petco_knowledgeowl_help.json`: Generated Q&A pairs
- Note: This script has been superseded by scraperV2.js
To run the legacy converter: ```bash node converter.js ```
- `scraper.js`: Original JavaScript implementation (deprecated)
Progress Tracking and Data Management
The scraper maintains progress using a `progress.json` file in the output directory (`training_data_v2` by default). This file:
- Tracks which categories and URLs have been processed
- Prevents duplicate processing of already scraped content
- Allows for resumable operations if the scraper is interrupted
To start a fresh scraping run:
- Create a new output directory (e.g., `training_data_v3`)
- Update the `OUTPUT_DIR` variable in `main.py` to point to your new directory
- Run the scraper - it will create a new `progress.json` in the new directory
To reprocess specific categories:
- Either delete the existing `progress.json` file to reprocess everything
- Or create a new output directory for a fresh run
Note: It's recommended to create a new output directory for each major scraping run to preserve previous data and avoid conflicts.
Data Processing Scripts
-
`removeblanks.py`: Cleans the scraped data by removing:
- Empty files
- Files containing only empty quotes
- Invalid JSON files
-
`GenerateDataFromTraining.py`: Processes raw scraped data by:
- Reading files from the `training_data_v2` directory
- Removing empty or invalid files
- Converting raw data into training format
-
`CombineData.py`: Merges multiple data files by:
- Processing files with 'stock-pretraining' prefix
- Combining title and text fields
- Outputting a consolidated `PreTuning.json` file
Usage Workflow
-
Start Chrome with debugging enabled (see Browser Setup section)
-
Log in to the Petco Knowledge Base
-
Run the recommended scraper: ```bash node scraperV2.js ```
The script will:
- Create a `training_data_v2` directory for output
- Track progress in `progress.json`
- Generate multiple types of training data
- Handle errors and resume capability automatically
-
Clean the data: ```bash python removeblanks.py ```
-
Generate training data: ```bash python GenerateDataFromTraining.py ```
-
Combine the data (if needed): ```bash python CombineData.py ```
Existing Training Data
The repository includes pre-existing training data in the training data directories. These can be used as reference or combined with newly scraped data using the processing scripts.
Note
Ensure you have the necessary permissions and are following company policies when accessing and scraping internal knowledge base content.