---
name: build-scrapers
description: "Set up competitor price scrapers on Scrapewise through the Scrapewise MCP server: pick the best source per shop, build the scrapers, test them cheaply, run them one at a time, turn raw prices into comparable prices (per unit, in EUR) with after-scrape rules, check the result and report the cost. Use when the user says 'build scrapers', 'competitor prices', 'track competitor prices', 'price monitoring setup', 'new customer setup', 'scrape these product links', or 'compare price per 100 g / per piece'."
---

# Build competitor price scrapers with Scrapewise

You drive Scrapewise through its MCP tools (server `https://mcp.scrapewise.ai/mcp`, API key with scope
`LLM_FULL`). All tool names below start with `scrapewise_`. Work in this order and do not skip steps.

Goal: **the smallest set of scrapers that gives correct, comparable prices every run.**
Pick every source by: **1. reliable over time, 2. scales to many products, 3. cost.**

## What the platform handles (plan for this size)

A typical setup has dozens of competitor shops (there is no limit on shops per group) and thousands to tens of
thousands of product links.
Build for that size from the start:

| Need | How | Size |
|---|---|---|
| Very big link list | Upload the links as a file (`scrapewise_create_file_scraper`), then feed a scraper's link list from it | Files up to 200 MB (CSV or JSON Lines). A fed link list of a product-page scraper holds up to about 700,000 links |
| Link list from another scraper | `scrapewise_create_scraper_site` with `linksSourceScraperId`, `urlFieldName`, `updateLinksBySourceScraperAutomatically: true` | Up to about 700,000 links for a product-page scraper, refreshed after each source run |
| Paste a list in one call | `scrapewise_run_scraper_url_list` | Up to 5,000 addresses per call |
| Bulk list by REST | JSON Lines upload (REST only) | Up to 100,000 addresses per upload |

Never create one scraper per link. One scraper per shop holds all links of that shop.

## Rules you always follow

- Ask the user before any step that costs money: full runs, group runs, schedules. Tests of one page are fine.
- Run **one scraper at a time** for first runs. Never start many big runs at once.
- Never delete scrapers, groups or data unless the user asks. Keep old scrapers as a fallback.
- Keep scrapers manual until the user asks for a schedule.
- The AI extractor copies values **as printed**. It never calculates. All maths is done by after-scrape rules.
- Use the same column names on every scraper (section 4).
- Report results in short, plain sentences: what works, what is missing and why, what it costs.

## 1. Intake (ask once)

1. Which shops, and which markets (countries)?
2. Does the user have **product links** for each competitor product, or only **shop names**?
3. What should be compared: price per item, per 50 g, per 100 g, per litre, per piece? In which currency?
4. Is there an own product list (file with SKU, EAN, name, price)? Upload it as the master (section 5).
5. How often: once, weekly, daily?

## 2. Choose the mode

| The user has | Mode | Layout |
|---|---|---|
| A list of product links (often thousands) | **Given links**: scrape only those links, no crawling | One group. One scraper per shop. Put market, own SKU and shop in each link title: `DE \| SKU-123 \| shop.com` |
| Only shop names, wants the whole catalogue | **Catalogue**: every product of each shop (often tens of thousands of products), then match to the master | One group per market |

Find or create the group: `scrapewise_get_scraper_group_list`, `scrapewise_create_scraper_group`.
In link titles, avoid numbers with 8 or more digits (they look like barcodes).

## 3. Pick the source for each shop (before you build)

Look at one product page and one category page. Use `scrapewise_get_scraper_load_site_content` or
`scrapewise_preview_scraper_from_url` (both fetch a page and are charged per page). Take the first source in
this list that gives price, a stable id, and the pack size or stock you need:

| Order | The shop has | Scraper | Why |
|---|---|---|---|
| 1 | A JSON API behind the listing or search page | API (from a cURL: `scrapewise_preview_scraper_from_curl`) | Many products per request. Use the largest page size that works |
| 2 | Shopify | API on `https://<shop>/variants/<variantId>.js?country=<XX>` | Price (in cents), before-price, weight in g, stock. `country=` sets the currency |
| 3 | WooCommerce | API on `/wp-json/wc/store/v1/products?slug=<slug>` | Price in minor units, regular price, stock |
| 4 | Magento or another shop REST API | API | Regular and final price |
| 5 | One JSON-LD `Product` block | JSON-LD scraper | Standard page data |
| 6 | `itemprop` microdata or clear page fields | HTML scraper with selectors | Stable page layer |
| 7 | Nothing structured | AI scraper with a custom schema (`scrapewise_create_customer_schema`) | Last choice. Ask for raw values only |

Checks while you look:
- **Past the end:** ask for the page after the last page. If it is not empty, use one link per page.
- **Plain or render:** if a plain fetch gives the same data as a rendered one, use plain (5 times cheaper).
- **Variants:** fetch 3 variant URLs. If they return the same page, use one URL per product.
- **Market:** set the market with the URL, a parameter or a header, not with the visitor location.
- **Price check:** compare the API price with the page price on a few products per market.
- Shopify weight is the shipping weight. Use it as the pack size only when it matches the real pack.

## 4. Columns (same names on every scraper)

Given-links price compare:
`name`, `price` (regular price), `sale_price` (today's price, only when lower), `on_sale`, `currency`,
`pack_grams` or `pack_pieces`, `price_per_50g` (or your unit), `sale_price_per_50g`, `price_eur`,
`sale_price_eur`, `price_per_50g_eur`, `availability`, `availability_norm`, `shop_sku`, `url`.

Catalogue and matching:
`name`, `brand`, `sku`, `productId`, `variantId`, `size`, `colour`, `price`, `url`, `image`, `ean`, `mpn`,
constant `competitor` / `market` / `currency`, `priceEur`, `eanClean`, `mpnClean`, `availabilityNorm`.
Always keep `image` and every EAN you can find. Read EANs as text, not numbers.

## 5. Build

- Create or update a scraper: `scrapewise_create_scraper` (read the allowed values first with
  `scrapewise_get_scraper_config_parameters`; read an existing one with `scrapewise_get_scraper`).
- Save the link list: `scrapewise_create_scraper_site` with `scraperId` + `links: [{url, title}]`.
  **This replaces the whole list.** To add links, send the old list plus the new links. Check the stored list
  with `scrapewise_get_scraper_site_links`.
- Thousands of links: do not paste them into one tool call. Upload them as a CSV (columns `url`, `title`,
  one row per link) with `scrapewise_create_file_scraper`, then save the scraper's link list with
  `linksSourceScraperId` = the file scraper, `urlFieldName: "url"`, `titleFieldName: "title"`. To change the
  list later, upload a new file with `scrapewise_update_file_scraper_file`.
- One scraper can hold a URL only once. If one URL is compared with two of your SKUs, add a harmless
  parameter to the second one (for example `&sku=SKU-456`).
- Upload the user's own product file as the master: `scrapewise_create_file_scraper` (CSV or JSON Lines up to
  200 MB, Excel up to 36 MB, first sheet only). Upload the raw file, do not clean it by hand.

## 6. Rules: make prices comparable

Rule kinds: `QUANTITY_NORMALIZE` (price per unit), `NUMBER_EXTRACT` (take a number from text),
`CURRENCY_CONVERT` (to EUR at the daily ECB rate), `ENUM_MAP` (map values, for example stock text),
`REGEX_CLEAN` (keep one kind of character, for example digits). Test **every** rule first with
`scrapewise_preview_scraper_preview_rule`.

Recipes:
- Price per 50 g: `QUANTITY_NORMALIZE` source `price`, `quantityField` `pack_grams`, `targetQuantity` 50
  → `price_per_50g`. Same for `sale_price`.
- Price in cents: `QUANTITY_NORMALIZE` source `price_cents`, `quantityValue` 100, `targetQuantity` 1 → `price`.
- Pack size only in text ("Ball 100 g", "9 pcs"): `NUMBER_EXTRACT` with `units` `WEIGHT_G` or `PIECES`
  → `pack_grams` / `pack_pieces`. Then the per-unit rule reads it.
- EUR: `CURRENCY_CONVERT` from `price_per_50g` → `price_per_50g_eur`. A convert rule may read the output of an
  earlier per-unit or number rule.
- Stock: `ENUM_MAP` from `availability` → `availability_norm` with `IN_STOCK`, `OUT_OF_STOCK`, `BACKORDER`,
  and default `UNKNOWN`.
- A rule that cannot compute (missing price or pack) writes **nothing**. Empty means unknown, never 0.

## 7. Test each scraper once (cheap)

`scrapewise_get_scraper_sample_data` returns a small sample and saves nothing. Check 5 rows against the live
page: price, before-price, pack size, currency, name, URL. If a rule column is missing in the sample, test the
rule with `scrapewise_preview_scraper_preview_rule`.

## 8. First runs, one at a time

Ask the user before the first full runs. Then:
1. `scrapewise_run_scraper` for one scraper. Wait until it ends: `scrapewise_get_scraper_load_history`.
2. Read errors: `scrapewise_get_scraper_job_errors`.
3. Read rows: `scrapewise_get_scraper_data_group_client` (page size 20 or less).
4. Check: rows vs expected, fill rate of every key column over all rows, duplicates, per-unit value =
   value ÷ pack × unit.
5. Fix and go to the next scraper. When all pass, `scrapewise_run_scraper_group` runs the whole group.

## 9. Audit against the user's list

Count links with a full row out of all links. Name each miss with its cause: product removed, shop closed,
no weight on the page, collection page instead of a product page. Ask the user for new links where needed.

## 10. Matching (catalogue mode)

Match by codes first, text last: EAN, then any exact code (cleaned MPN, the shop SKU without a brand prefix,
a code in the URL), then brand + model, then size and colour, then name. Set up matching with
`scrapewise_save_matcher_job` and read results with `scrapewise_get_matched_data`.

## 11. Cost

Each run in `scrapewise_get_scraper_load_history` has `costMicros` (1,000,000 = €1) and `totalRequests`.
Prices per 1,000 pages: plain €0.15, render €0.75, Super (residential proxy) €1.50, render + Super €3.75.
Imports of files are free.

## 12. Schedule (only when the user asks)

`scrapewise_update_scraper_group_schedule`, for example weekly on Monday. The uploaded master file is not
re-run by a schedule. A scheduled run starts only if the wallet has balance.

## 13. Report to the user

- Done: scrapers, links covered (for example "11,640 of 12,000 links complete"), columns delivered.
- Missing: each missing link and why.
- Cost: € per full run and per month at the chosen schedule.
- Next: what you need from the user (new links, a go for the schedule).

## Mistakes to avoid

1. Many first runs at the same time. Run one at a time.
2. Asking the AI to calculate per-unit prices. It copies numbers; rules calculate.
3. Taking the pack size from AI text reading. Prefer a structured field or `NUMBER_EXTRACT` on a spec line.
4. One URL per variant when all variants show the same page. Check 3 first.
5. Trusting a search page for a full catalogue. Deep pages often repeat. Compare ids.
6. Assuming the API price equals the page price. Check a few products per market.
7. Cleaning the customer file by hand ("100%" can become 1). Upload the raw file.
8. Different column names on different scrapers. Use section 4 from the first build.
9. Dividing by Shopify weight when it is shipping weight or 0.
10. A collection page as a product link. It can pick the wrong product.
11. Rendering when a plain fetch gives the same data.
12. Reading large data pages in one call. Use small pages.
13. One scraper per link, or thousands of links pasted into one call. Use one scraper per shop and a file or a feed.
