Create a Web Scraper task by selecting a scraping template and submitting parameters such as the target URL. You can either wait for the results to be returned or retrieve the task ID to query the results later.
POST /v1/web-scraper/task/create
Quick Start #
Before making the call, please prepare your API Key and obtain the template code and input parameter specifications from the product’s template page. Ensure that your current plan supports the template and that your account has sufficient balance.
The following example demonstrates scraping Amazon product data. Please replace <API_KEY> and the example product URL with your actual API endpoint, API Key, and target URL. The data used here is for demonstration purposes only.
curl --request POST 'https://data.clawoxy.com/v1/web-scraper/task/create' \
--header 'Authorization: Bearer <API_KEY>' \
--header 'Content-Type: application/json' \
--data '{
"template_code": "amazon_product_by_url",
"mode": "async",
"input": {
"url": "https://www.amazon.com/dp/B0EXAMPLE1"
}
}'
After the task is created, HTTP 202 is returned:
{
"code": 0,
"message": "success",
"data": {
"task_id": "01a0576790007a81b21db0d8b61eaabc",
"template_code": "amazon_product_by_url",
"mode": "async",
"status": "queued",
"result_status": "not_ready",
"error": null,
"result": null,
"result_url": null,
"result_url_expires_at": null,
"created_at": 1788172800,
"completed_at": null,
"result_expires_at": null
}
}
Save the task_id and query the task details to retrieve progress and results. queued indicates that the task is currently in the queue and has not yet been completed.
Request Parameters #
Request Headers #
| Header | Required | Description |
|---|---|---|
Authorization | Yes | Enter Bearer <API_KEY> |
Content-Type | Yes | Enter application/json |
Request Body #
Submit the following parameters using a JSON object:
| Parameter | Type | Required | Description |
|---|---|---|---|
template_code | string | Yes | Template code, e.g., amazon_product_by_url |
mode | string | Yes | async: Asynchronous scrapingsync: Synchronous scraping |
input | object | Yes | Scraping parameters; fill in according to the selected template’s instructions, e.g., target URL url |
The input field may vary depending on the template.
Choosing a Mode #
| Mode | Use Case | How to Get Results |
|---|---|---|
async | Submit a task now and retrieve the result later | Returns HTTP 202; save the task ID and query for details |
sync | Wait for the result within the same request | Returns HTTP 200 upon completion; returns HTTP 202 if still processing (continue querying for details) |
To use synchronous mode, simply change "mode": "async" in the request to "mode": "sync"; all other parameters remain the same.
Synchronous mode may also require follow-up queries. The termination of the request wait or a connection disconnection does not cancel the task that has already been created; please retain the task ID to avoid duplicate submissions.
Successful Response #
Please check data.status to determine if the task collection was successful. A code=0 indicates that the request itself was successful, but it does not mean the collection process is complete.
Upon completion of synchronous data collection, the result may be returned directly with an HTTP 200 status code:
{
"code": 0,
"message": "success",
"data": {
"task_id": "01a0576790007a81b21db0d8b61eaabc",
"template_code": "amazon_product_by_url",
"mode": "sync",
"status": "succeeded",
"result_status": "available",
"error": null,
"result": {
"url": "https://www.amazon.com/dp/B0EXAMPLE1",
"asin": "B0EXAMPLE1",
"title": "Sample Product"
},
"result_url": null,
"result_url_expires_at": null,
"created_at": 1788172800,
"completed_at": 1788172812,
"result_expires_at": 1788432012
}
}
The fields within result are determined by the template; the example above shows only a subset of product fields. If the result is large, result will be null; please use result_url to download the JSON result.
Response Fields #
| Field | Type | Description |
|---|---|---|
code | integer | 0 indicates a successful request; for other values, refer to the error code descriptions. |
message | string | Returns success on success; returns an English error description on failure. |
data | object or null | Task information; returns null on failure. |
The data object contains the following fields:
| Field | Type | Description |
|---|---|---|
task_id | string | Task ID, used to query progress and results. |
template_code | string | The template code used. |
mode | string | The mode selected upon creation: sync or async. |
status | string | Task status; see the table below. |
result_status | string | Indicates whether the result is available; see the table below. |
error | object or null | Reason for task failure, containing code and message; returns null if the task did not fail. |
result | object, array, or null | Directly returned scraping results; returns null if no result is returned directly. |
result_url | string or null | JSON Result download link; will not have a value at the same time as result |
result_url_expires_at | integer or null | Download link expiration time; null if no link exists |
created_at | integer | Task creation time |
completed_at | integer or null | Task completion time; null if not yet completed |
result_expires_at | integer or null | Result retention expiration time; null if not yet successful or if the task failed |
All time fields are Unix timestamps in seconds.
Proceeding based on status #
status | result_status | Meaning | Next Step |
|---|---|---|---|
queued | not_ready | Queued | Check details later |
running | not_ready | Collecting | Check details later |
succeeded | available | Collection successful | Read result or download result_url |
succeeded | expired | Result expired | Result no longer available |
failed | unavailable | Collection failed | Check the reason in error |
Retrieving Collection Results #
If result is returned, you can read the data directly. If result_url is returned, use the full URL to download the file:
curl '<RESULT_URL>' --output result.json
Replace <RESULT_URL> with the URL from the response. No API Key is required for the download; please keep the URL secure and do not share it publicly.
While the task is in progress, it is recommended to poll for details every 2–5 seconds and stop polling once the task completes or fails.
Please save the result before the result_expires_at timestamp.
Error Handling #
Unsuccessful Request #
When the code is not 0, please handle the error based on the error code. For example, insufficient balance returns an HTTP 402 status:
{
"code": 120301,
"message": "balance is insufficient",
"data": null
}
| Error Code | HTTP | Reason | Recommended Action |
|---|---|---|---|
100301 | 400 | Invalid request parameters | Check JSON format, required fields, and template parameter requirements |
120102 | 401 | API Key missing or invalid | Check the API Key and Authorization request header |
120103 | 403 | API Key disabled | Verify the API Key status |
120104 | 500 | API Key error |
120205500120301402120402429120405503120406400120501404120502403120503404120507429Task Collection Failed #
If the request is successful but data.status is failed, please check data.error. For example, this field might return:
{
"code": 120403,
"message": "The task failed during data collection."
}
data.error.code | Meaning | Recommended Action |
|---|---|---|
120403 | Collection failed or no valid results obtained | Check the target URL and template parameters before deciding whether to retry the collection |
120404 | Collection timed out | Retry the collection later as needed |
When a synchronous request returns a failed task, the HTTP status remains 200; please determine the collection result by checking the status and error fields.