Skip to Content
REST APIUpload a file

Upload a file

An upload scraper is a scraper whose data comes from a file you send, not from a website. You give it one file. Scrapewise reads the file and writes every row into the group’s data, like the rows of a normal run.

After that the rows work like any other scraped data. You can see them under Data, export them, merge them, use Show latest run, feed another scraper from them, and use them in matching (as the master or as a scope).

An upload scraper never fetches a page. It is never scheduled, and Run group skips it. It only gets new data when you send a new file.

Uploading is free. An import costs nothing and works with a €0 wallet.

In the portal

  1. On the canvas, open the group’s ⋮ menu and pick Upload.
  2. Type a name for the scraper and pick your file.
  3. Press upload. A new scraper box appears under the group. It shows Running while the file is read, then Completed.

To send a new file later, press the play button on the box. The same dialog opens.

Data retention of an upload scraper is 1 by default: it keeps the latest file only. When a new file is fully read, the rows of the previous file are removed. You can change this in Data retention. When the file is used in matching, the rules are different — see When the file is used in matching.

Formats and limits

FormatFileSize limit (default)
CSV.csv — comma , or semicolon ; between values, UTF-8 text200 MB
JSON Lines.jsonl or .ndjson — one JSON object per line200 MB
Excel.xlsx — the first sheet only36 MB
  • For Excel files over 36 MB please save as CSV. Old .xls files, .json files and other formats are not read.
  • In CSV and Excel, the first row gives the column names. In JSON Lines, the keys of the objects do (in the order they first appear). A nested JSON value is stored as compact JSON text.
  • All values are stored as text. Nothing is turned into a number, so an EAN like 0012345678905 keeps its leading zeros. Long numbers in Excel keep all their digits.
  • Empty cells are left out. Rows that are completely empty are skipped.
  • In CSV, a line break ends a row, also a lone carriage return (CR) outside quotes. So a stray CR in the middle of a line splits it into two rows.
  • At most 500 columns. A column name is at most 256 characters, one cell at most 32 KB, and one row at most 1 MB.
  • 1 MB here means 1,000,000 bytes.
  • By default only 1 Excel import or add-only upload (see When the file is used in matching) runs at a time on the server. Others wait in line and start by themselves. The admin can raise this number (setting “Max Excel imports / file appends at once”, default 1).
  • Big files need a connection of at least about 0.6 Mbit/s. An upload that is slower than that for a full minute, or that takes longer than the time limit (60 minutes by default), is stopped with 408.

The size limits are set by the admin and can be lower on your account. Read the current values from GET /api/scraper/config/parameters before you send a big file. 200 MB is the highest possible limit.

Create an upload scraper — POST /api/scraper/group/{groupId}/file-scraper

Creates the upload scraper in the group and starts its first import. The file itself is the request body. Send it in one request — do not split it.

Mac or Linux (bash):

curl -X POST -T 'products.csv' \ -H 'Authorization: Bearer <YOUR_API_KEY>' \ -H 'Content-Type: text/csv' \ -H 'Expect:' \ 'https://portal.scrapewise.ai/api/scraper-api/api/scraper/group/65f1c2a9e4b0a1b2c3d4e5f6/file-scraper?name=Customer+feed&format=csv'

Windows (PowerShell or cmd):

curl.exe -X POST -T "products.csv" -H "Authorization: Bearer <YOUR_API_KEY>" -H "Content-Type: text/csv" -H "Expect:" "https://portal.scrapewise.ai/api/scraper-api/api/scraper/group/65f1c2a9e4b0a1b2c3d4e5f6/file-scraper?name=Customer+feed&format=csv"

Use -T. It streams the file from disk. Do not use --data-binary @file: that loads the whole file into memory first.

Keep the empty Expect: header (-H 'Expect:', on Windows -H "Expect:"). It stops curl from waiting for a “go ahead” before it sends the file. Without it, a file over the limit often ends in “connection reset” instead of a clear answer. With it, an upload over the limit usually answers 413 with a clear message.

Query parameterRequiredDescription
nameyesThe scraper’s name, 1-200 characters. URL-encode it (a space is + or %20). Spaces at the start and end are removed
formatyescsv, jsonl or xlsx. It must match the Content-Type
checknotrue = only check, send nothing — see Check first
formatContent-Type
csvtext/csv
jsonlapplication/x-ndjson (or application/jsonl)
xlsxapplication/vnd.openxmlformats-officedocument.spreadsheetml.sheet

The request must carry a Content-Length (curl -T sends it). A body sent in chunks, without a length, is refused with 413.

Response (201):

{ "scraper": { "id": "66a0c1d2e3f4a5b6c7d8e9f0", "name": "Customer feed", "type": "FILE", "...": "the rest of the scraper" }, "jobId": "66a0c1d2e3f4a5b6c7d8e9f1" }

Retries are safe. A name can be used only once in a group. If the group already has a scraper with that name, the call is refused with 409 NAME_TAKEN, and the answer carries that scraper’s scraperId and type. If it is an upload scraper, send the new file to it with the next route instead.

Send a new file — POST /api/scraper/{scraperId}/file

Sends a new file to an upload scraper you already have and starts a new import. The body and headers are the same as above; only format is needed in the query.

curl -X POST -T 'products.csv' \ -H 'Authorization: Bearer <YOUR_API_KEY>' \ -H 'Content-Type: text/csv' \ -H 'Expect:' \ 'https://portal.scrapewise.ai/api/scraper-api/api/scraper/66a0c1d2e3f4a5b6c7d8e9f0/file?format=csv'
Query parameterRequiredDescription
formatyescsv, jsonl or xlsx
keyColumnnoOnly when the file is used in matching and no key column is set yet — see When the file is used in matching
checknotrue = only check, send nothing

Response:

{ "jobId": "66a0c1d2e3f4a5b6c7d8e9f2", "keyColumnUsed": null }

keyColumnUsed is set only when the file is used in matching. It names the key column the new rows are matched by (for example "EAN"). Otherwise it is null.

The scraper must be an upload scraper (else 400), and it must not be importing already (else 409 ALREADY_RUNNING).

Small files as JSON

For small files (up to 2 MB of JSON) both routes also take a JSON body. This is the form the MCP tools use. At most 8 JSON uploads can run at once in total; a 9th gets 429 until one ends.

POST /api/scraper/group/65f1c2a9e4b0a1b2c3d4e5f6/file-scraper Authorization: Bearer <YOUR_API_KEY> Content-Type: application/json { "name": "Customer feed", "format": "csv", "content": "ean;name;price\n0012345678905;Blue mug;4.99\n" }
FieldRequiredDescription
namecreate onlyThe scraper’s name
formatyescsv, jsonl or xlsx
contentyesFor csv and jsonl: the file text as it is. For xlsx: the file as base64
keyColumnnoSend a new file only; see When the file is used in matching

Bad base64 is refused with 400 BAD_FORMAT, and nothing is created.

Check first

Before you send a big file, you can ask whether the upload would be accepted. Call the same route with the same query parameters plus check=true, the Content-Type of your format, and an empty body:

curl -X POST \ -H 'Authorization: Bearer <YOUR_API_KEY>' \ -H 'Content-Type: text/csv' \ -H 'Content-Length: 0' \ 'https://portal.scrapewise.ai/api/scraper-api/api/scraper/group/65f1c2a9e4b0a1b2c3d4e5f6/file-scraper?name=Customer+feed&format=csv&check=true'

The answer is 204 when all is fine. Otherwise it is the same error a real upload would get (for example 409 NAME_TAKEN). A check creates nothing. A check with a body is refused with 400.

Follow the import

The import runs in the background. Each upload is one run (a job), with the jobId from the answer. Follow it like any other run:

Then read the rows like any scraped data — see Scraped data.

A file import is all or nothing:

End stateWhat it means
COMPLETEDThe whole file was read. Its rows are the scraper’s latest data
FAILEDThe file could not be read (for example a broken file, a row that is too large, or too many columns). The job’s error message says why. A file with no data rows (empty, or only the header) fails with “The file has no data rows — nothing was imported.” Nothing from this file is kept. The previous file stays the latest data. The run stays in the run history with its reason, and you get the usual “run failed” email
STOPPEDYou stopped the import. Nothing from this file is kept, and the previous file is still used. A stopped import is removed from the run history. No email is sent

If Scrapewise is updated while your file is being imported, the import ends as FAILED with “The uploaded file is no longer available — please upload it again”. Send the same file again.

Errors

StatusCodeMeaning
400BAD_FORMATformat is missing or unknown, does not match the Content-Type, or the base64 is bad
400KEY_COLUMN_REQUIREDThe file is used in matching and no key column is known yet. Send keyColumn
400KEY_COLUMN_INVALIDkeyColumn is not a column of the stored data
400—The scraper is not an upload scraper, the name is empty or too long, or check=true came with a body
401—No API key, or the key is not valid
403—A read-only API key
404—The group or the scraper was not found in your account
408—The upload was too slow, or took longer than the time limit. Send it again from a faster connection
409NAME_TAKENThe group already has a scraper with this name. The answer carries its scraperId and type
409ALREADY_RUNNINGThis upload scraper is importing a file right now. Wait until it ends. When the file is used in matching, the message reads: “An upload is still being finished. This can take up to about N minutes after a restart — or press Stop on the box, wait about a minute, then upload again.” (N is set by the admin, 30 minutes by default)
413—Usually the answer when the file is over the size limit, the Excel file is over the Excel limit (“please save as CSV”), the JSON body is over 2 MB, or there is no Content-Length
415—The Content-Type is not one of the three above (or application/json)
429INGEST_CAPACITY, INGEST_CAPACITY_GLOBALToo many uploads at once (for you, or on the server in total). Wait a moment and send it again
500—Server error. Send it again later
503—The server is short of disk space right now. Try again later

A 400 answer carries the code in its body: { "message": "...", "errors": [ { "code": "KEY_COLUMN_REQUIRED", "message": "..." } ] }.

When the file is used in matching

Once any matching setup of the group uses the upload scraper’s data (as the master or as a scope), the scraper is protected. This stays on for good, even if the matching setup is deleted later.

A new file for a protected upload scraper never replaces or removes anything. Instead:

  • Rows are matched by a key column.
  • A row whose key is new is added.
  • A row whose key is already there gets its empty cells filled and its new columns added. Values that are already there are never changed.
  • Rows that are not in the new file stay as they are.
  • A row with an empty key is skipped. If the new file has the same key twice, the first row is used.
  • The rows stay in the same run, so the matching results keep pointing at them.
  • By default only 1 such upload or Excel import runs at a time (the admin can raise it); others wait in line.

The key column is set at the first protected upload and then stays fixed. It is found in this order:

  1. The column the matching setup uses to identify this data.
  2. Otherwise the first column named EAN, GTIN, EAN13, barcode, SKU, MPN, product_id, id or url (upper or lower case does not matter).
  3. Otherwise you pick it: send keyColumn (until then the upload is refused with 400 KEY_COLUMN_REQUIRED).

Once the key column is fixed, a keyColumn you send is ignored. The answer’s keyColumnUsed always names the key that was really used.

The key column must also be in the new file. If it is missing, the import ends as FAILED and nothing is changed; the job’s error message ends with (KEY_COLUMN_MISSING). If the file became used in matching while it waited, and no key column can be found, the import ends as FAILED with nothing changed; the error message ends with (KEY_COLUMN_REQUIRED). Send the file again with keyColumn. Key values are compared exactly, after spaces at the start and end are removed.

If such an import is stopped, fails, or is cut by a restart halfway, the rows added and cells filled so far stay. Upload the same file again to finish it.

After a restart, an unfinished upload can hold the scraper for up to about 30 minutes (the admin’s setting). During that time a new upload gets 409 ALREADY_RUNNING. To free it sooner, press Stop on the box, wait about a minute, then upload again.

Matching is not run again by itself after such an upload — also not when only empty cells were filled. Re-run matching to refresh the matched results.

What’s next