Query the scraping progress, failure reasons, and result download link for a task. This endpoint supports queries for tasks created in either synchronous or asynchronous mode.
GET /v1/web-scraper/task/detail
Quick Start #
Initiate a query using the task_id returned when the task was created. Please replace the API URL, API Key, and task ID in the example.
curl --request GET 'https://data.clawoxy.com/v1/web-scraper/task/detail?task_id=<TASK_ID>' \
--header 'Authorization: Bearer <API_KEY>'
Upon successful scraping, an HTTP 200 response is returned, including the download link.
{
"code": 0,
"message": "success",
"data": {
"task_id": "01a0576790007a81b21db0d8b61eaabc",
"template_code": "amazon_product_by_url",
"mode": "async",
"status": "succeeded",
"result_status": "available",
"error": null,
"result": null,
"result_url": "https://example.com/results/example-task.json?signature=SIGNATURE_PLACEHOLDER",
"result_url_expires_at": 1788176412,
"created_at": 1788172800,
"completed_at": 1788172812,
"result_expires_at": 1788432012
}
}
Use the result_url to download the JSON result:
curl '<RESULT_URL>' --output result.json
Replace <RESULT_URL> with the full link from the response. You do not need to include your API Key when downloading; please do not share the link publicly.
Request Parameters #
Request Headers #
| Header | Required | Description |
|---|---|---|
Authorization | Yes | Enter Bearer <API_KEY> |
Query Parameters #
| Parameter | Type | Required | Description |
|---|---|---|---|
task_id | string | Yes | The task ID returned by the creation API |
Append the task_id after ?task_id= in the request URL; no request body is required.
Only tasks belonging to the current account can be queried. Other valid API keys under the same account may also be used; API keys from other accounts cannot query this task.
Response Fields #
An HTTP 200 status code is returned upon a successful query. This indicates only that the task information has been retrieved; please check data.status to see if the data collection is complete.
| Field | Type | Description |
|---|---|---|
code | integer | 0 indicates a successful query; for other values, refer to the error code descriptions. |
message | string | success upon a successful query; an English error description upon failure. |
data | object or null | Task information; null upon query failure. |
data contains the following fields:
| Field | Type | Description |
|---|---|---|
task_id | string | Task ID |
template_code | string | Template code used for the task |
mode | string | sync or async mode selected when creating the task |
status | string | Task status; see the table below |
result_status | string | Availability of the result; see the table below |
error | object or null | Reason for task failure, containing code and message; null if the task did not fail |
result | null | This API does not return the scraped data directly; please use the download link to retrieve it |
result_url | string or null | Download link for the JSON result; null if not yet successful or if the result has expired |
result_url_expires_at | integer or null | Expiration time of the current download link; null if no link exists |
created_at | integer | Task creation time |
completed_at | integer or null | Task completion time; null if not yet completed |
result_expires_at | integer or null | Expiration time for result retention; null if not yet successful or if the task failed |
All time fields are Unix timestamps in seconds.
Retrieving Results Based on Task Status #
status | result_status | Meaning | Next Step |
|---|---|---|---|
queued | not_ready | Queued | Check again later |
running | not_ready | Collecting | Check again later |
succeeded | available | Collection successful | Download results from result_url |
succeeded | expired | Result expired | Result no longer available |
failed | unavailable | Collection failed | Check the reason in error |
failed | unavailable | Collection failed | Check the reason in error |
It is recommended to poll every 2–5 seconds and stop polling once collection succeeds or fails. There is no need to create a new task just to refresh the progress.
The downloaded file contains the JSON data corresponding to the template; please refer to the template documentation for specific fields. The result field in the details API response is always null; even if the task was created in synchronous mode, you must use the result_url to download the data.
What if the download link has expired? #
Check the two expiration times:
result_url_expires_atis the validity period of the current link. If the link expires but the result is still within the retention period, you can query the details again to obtain a new link.result_expires_atis the deadline for result retention. Once this time passes, theresult_statusbecomesexpired, and the result cannot be recovered by refreshing the link.
Please rely on the expiration times provided in the response and download/save the data promptly. Repeated queries will not extend the result retention period. After the result expires, the task will still show as succeeded, indicating that the collection was previously successful.
Error Handling #
Unsuccessful Query #
When the code is not 0, please handle the error based on the error code. For example, if the task does not exist or does not belong to the current account, an HTTP 404 is returned:
{
"code": 120503,
"message": "Task not found.",
"data": null
}
| Error Code | HTTP | Reason | Recommended Action |
|---|---|---|---|
100301 | 400 | Task ID missing or incorrectly formatted | Use the full Task ID returned by the creation API |
120102 | 401 | API Key missing or invalid | Check the API Key and the Authorization request header |
120103 | 403 | API Key deactivated | Verify the API Key status |
120104 | 500 | API Key status abnormal | Contact customer support for assistance |
120405 | 503 | Service temporarily unavailable | Retry the query later using the same Task ID |
120503 | 404 | Task does not exist or access is denied | Check the Task ID and the account associated with the API Key |
A failed query does not mean the data collection failed, nor does it require recreating the task. If you need assistance, please provide customer support with the Task ID, request time, and the X-Request-Id from the response header; do not provide the full API Key or download links.
Query successful, but task collection failed #
In this case, the HTTP status remains 200 and the code remains 0, but data.status is failed. The reason for the failure is found in data.error, for example:
{
"code": 120404,
"message": "Task execution timed out."
}
data.error.code | Meaning | Recommended Action |
|---|---|---|
120403 | Collection failed or no valid results obtained | Check the target URL and template parameters before deciding whether to retry the collection |
120404 | Collection timed out | Retry the collection later as needed |
There is no need to continue checking the progress after a task fails. You can create a new task if you need to retry the collection.