Skip to content

Scraper API: send waitFor as an object, read the target status from http_code - #40

Merged
jehrr merged 1 commit into
mainfrom
fix/scraper-api-waitfor-object
Sep 23, 2026
Merged

jehrr merged 1 commit into
mainfrom
fix/scraper-api-waitfor-object

Conversation

@jehrr

@jehrr jehrr commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Two defects in scraper_api_client.py, both measured 2026-09-23 against https://scraper.2captcha.com/tasks/sync.

1. waitFor was sent as a JSON-encoded string. The live API now answers that with HTTP 422 params.waitFor must be an object — and still bills the task ($0.0005). So every run with --wait-text / --wait-element / --wait-state ended in exit 5. Sent as an object it answers 200. _build_wait_for now returns a dict (Optional[dict]); it is logged with json.dumps. The module and function docstrings that said it had to be a double-encoded string now say what was measured.

2. The target status was read from status. The response is {"status": "success", "http_code": <int>, "headers": …, "body": …}: status is the API's own verdict string, http_code is the target site's HTTP code. The client logged "Upstream page status success" and a target 403/503 was never visible. New _target_status() reads http_code (int), falling back to status only if that is an int.

Regression check — check_scraper_api_waitfor_object_and_http_code in smoke_test.py: drives the real fetch_html with requests.post monkeypatched (no network), asserts the payload's waitFor is a dict for --wait-text, and that the logged upstream status is the int 403 from a fake {"status":"success","http_code":403} response.

Control: with scraper_api_client.py restored to origin/main in a scratch copy (edit asserted to change the file), the suite went RED with both of this check's messages and nothing else; with the fix it is green.

Live: one call through the fixed client, --wait-text '$', https://www.farfetch.com/shopping/kids/girls-clothing-4/items.aspx: API HTTP 200, upstream status 200 (int), 1,016,665 bytes, cost $0.0005. The returned HTML carried no challenge marker and parsed to 96 products.

Not changed: the module docstring and --help still say this engine does not work against farfetch.com's Akamai-protected category pages. The single live call above returned the full catalogue, but one call is an anecdote, not a re-measurement, so that claim is left for a separate change with more runs behind it. No version bump; CHANGELOG entry under [Unreleased].

🤖 Generated with Claude Code

…ttp_code

Measured 2026-09-23 against scraper.2captcha.com/tasks/sync: waitFor sent
as a JSON-encoded string is refused with HTTP 422 ("params.waitFor must be
an object") and still billed; as an object it answers 200. The response's
status field is the API's verdict ("success"); the target's HTTP code is
http_code. Adds a no-network regression check, controlled red against the
unfixed client.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jehrr
jehrr merged commit e6f6213 into main Sep 23, 2026
8 checks passed
@jehrr
jehrr deleted the fix/scraper-api-waitfor-object branch September 23, 2026 15:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant