Skip to content

refactor(browser-pool): move remote connection logic into dedicated remote plugins#3868

Draft
l2ysho wants to merge 126 commits into
v4from
3849-browser-pool-move-remote-browser-connection-logic-out-of-the-abstract-browserplugin
Draft

refactor(browser-pool): move remote connection logic into dedicated remote plugins#3868
l2ysho wants to merge 126 commits into
v4from
3849-browser-pool-move-remote-browser-connection-logic-out-of-the-abstract-browserplugin

Conversation

@l2ysho

@l2ysho l2ysho commented Jul 16, 2026

Copy link
Copy Markdown
Contributor

BrowserPlugin no longer knows about remote connections. RemoteBrowserPool
swaps the supplied plugins for RemotePlaywrightPlugin/RemotePuppeteerPlugin,
which own the library connect() call via RemoteBrowserConnection.

Closes #3849

B4nan and others added 30 commits March 5, 2026 15:46
BREAKING CHANGE:

The project is now native ESM without a CJS alternative. This is fine since all supported node versions allow `require(esm)`.

Also all the dependencies are updated to the latest versions, including cheerio v1.
BREAKING CHANGE:

The crawler following options are removed:

- `handleRequestFunction` -> `requestHandler`
- `handlePageFunction` -> `requestHandler`
- `handleRequestTimeoutSecs` -> `requestHandlerTimeoutSecs`
- `handleFailedRequestFunction` -> `failedRequestHandler`
BREAKING CHANGE:

The crawling context no longer includes the `Error` object for failed requests. Use the second parameter of the `errorHandler` or `failedRequestHandler` callbacks to access the error.

Previously, the crawling context extended a `Record` type, allowing to access any property. This was changed to a strict type, which means that you can only access properties that are defined in the context.
….retireOnBlockedStatusCodes`

BREAKING CHANGE:

`additionalBlockedStatusCodes` parameter of `Session.retireOnBlockedStatusCodes` method is removed. Use the `blockedStatusCodes` crawler option instead.
….retireOnBlockedStatusCodes`

BREAKING CHANGE:

`additionalBlockedStatusCodes` parameter of `Session.retireOnBlockedStatusCodes` method is removed. Use the `blockedStatusCodes` crawler option instead.
also tries to bump better-sqlite3 to latest version to have prebuilds for node 22
- closes #2479
- closes #3106
- closes #3107
- closes #3078

In my opinion, it makes a lot of sense to do the remaining changes in a
separate PR.

- [x] Introduce a `ContextPipeline` abstraction
- [x] Update crawlers to use it
- [x] Make sure that existing tests pass
- [ ] Refine the `ContextPipeline.compose` signature and the semantics
of `BasicCrawlerOptions.contextPipelineEnhancer` to maximize DX
- [x] Write tests for the `contextPipelineEnhancer`
- [x] Resolve added TODO comments (fix immediately or make issues)
- [ ] Update documentation

The `context-pipeline` branch introduces a fundamental architectural
change to how Crawlee crawlers build and enhance the crawling context
passed to request handlers. The core motivation is to fix the
composition and extensibility nightmare in the current crawler
hierarchy.

1. **Rigid inheritance hierarchy**: Crawlers were stuck in a brittle
inheritance chain where each layer manipulated the context object while
assuming that it already satisfied its final type. Multiple overrides of
`BasicCrawler` lifecycle methods made the execution flow even harder to
follow.

2. **Context enhancement via monkey-patching**: Manual property
assignment (`crawlingContext.page = page`, `crawlingContext.$ = $`)
scattered everywhere. It was a mess to follow and impossible to reason
about.

3. **Cleanup coordination**: Resource cleanup was handled by separate
`_cleanupContext` methods that were not co-located with the
initialization.

4. **Extension mechanism was broken**: The `CrawlerExtension.use()` API
tried to let you extend crawlers (the ones based on `HttpCrawler`) by
overwriting properties - completely type-unsafe and fragile as hell.

Introduces `ContextPipeline` - a **middleware-based composition
pattern** where:

- Each crawler layer defines how it enhances the context through
explicit `action` functions
- Cleanup logic is co-located with initialization via optional `cleanup`
functions
- Type safety is maintained through TypeScript generics that track
context transformations
- The pipeline executes middleware sequentially with proper error
handling and guaranteed cleanup

Declarative middleware composition with co-located cleanup:

```typescript
contextPipeline.compose({
  action: async (context) => ({ page, $ }),
  cleanup: async (context) => { await page.close(); }
})
```

The `ContextPipeline<TBase, TFinal>` tracks type transformations through
the chain:

```typescript
ContextPipeline<CrawlingContext, CrawlingContext>
  .compose<{ page: Page }>(...) // ContextPipeline<CrawlingContext, CrawlingContext & { page: Page }>
  .compose<{ $: CheerioAPI }>(...) // ContextPipeline<CrawlingContext, CrawlingContext & { page: Page, $: CheerioAPI }>
```

The `CrawlerExtension.use()` is gone. New approach via
`contextPipelineEnhancer`:

```typescript
new BasicCrawler({
  contextPipelineEnhancer: (pipeline) =>
    pipeline.compose({
      action: async (context) => ({ myCustomProp: ... })
    })
})
```

The current way to express a context pipeline middleware has some
shortcomings (`ContextPipeline.compose`,
`BasicCrawlerOptions.contextPipelineEnhancer`). I suggest resolving this
in another PR.

For most legitimate use cases, this should be non-breaking. Those who
extend the Crawler classes in non-trivial ways may need to adjust their
code though - the non-public interface of `BasicCrawler` and
`HttpCrawler` changed quite a bit.

The pipeline uses `Object.defineProperties` for each middleware. Is this
a serious performance consideration?

---------

Co-authored-by: Martin Adámek <banan23@gmail.com>
Extracts `ProxyConfiguration` to `BasicCrawler` (related to discussion under #2917).

Pass the `ProxyConfiguration` instance to the `SessionPool` for new `Session` object creation.

Store and read the `ProxyInfo` from the `Session` instance instead of calling the `ProxyConfiguration` methods in the crawlers.

closes #3198
Phasing out `got-scraping`-specific interfaces in favour of native
`fetch` API.

Related to #3071
Fixes build toolchain errors caused by the recent rebase onto the
current `master` ([more details
here](https://apify.slack.com/archives/C02JQSN79V4/p1764373034961859)).

The largest thing is probably updating the dependency versions in
`package.json` - if `turborepo` doesn't find the matching version in the
local workspace, it will build against the package pulled from `npm`
(which doesn't match the v4 API at this point).
…client (#3286)

Removes incorrect implementation `KVS.getPublicUrl()` implementation from `@crawlee/core` and proxies the call to the storage client.

Closes #3272 
Closes #3076
…rfaces (#3295)

Works towards removing `got-scraping` as a direct Crawlee dependency.

Related to #3275
Related to #3071
Related to #3275

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
janbuchar and others added 29 commits June 1, 2026 13:36
Adds a `fingerprint` field to `Session` so repeated requests with the
same session can stay consistent on user-agent, headers, and TLS profile
across HTTP clients and browser crawlers.

Closes #3628.
)

`preNavigationHooks` and `postNavigationHooks` may now return a
`Partial<Context>`; returned own properties are merged into the crawling
context.

Closes #3660
Makes `@crawlee/impit-client` the default HTTP client for
`@crawlee/basic`'s `BasicCrawler`.
Adds recent changes to the `SessionPool` and related classes to the
existing guides.

Closes #796
Drops the `SDK_` prefix (replaces w/ `CRAWLEE_`) on multiple spots in
the codebase.

Closes #1729
Utilizes the byte-stream prescan
(https://html.spec.whatwg.org/multipage/parsing.html#prescan-a-byte-stream-to-determine-its-encoding)
to determine the HTML document encoding, if not specified in HTTP
headers (or explicitly forced by user).

Closes #2317
Due to the nature of the concurrency model in JS, we cannot reliably
abort the `requestHandler` execution on timeouts.

This can lead to "race conditions" with, e.g., the `requestHandler`
modifying the storages while the `errorHandler` / `failedRequestHandler`
is already running.

This PR adds a `tryCancel()` check throwing on elapsed timeouts before
each storage method call.

Closes #2889
Avoid double-counting requests in `RequestManagerTandem`. 

Closes #3787
…)` (#3831)

Serializes the passed data in `pushData` calls before storing it.
Closes #3396

---------

Co-authored-by: Jindřich Bär <jindrichbar@gmail.com>
…ite (#3844)

`scripts/copy.ts` rewrites the version of every `@crawlee/*` dependency
(and `crawlee`) to the monorepo version during canary publish and
`pin-versions`. But `@crawlee/fs-storage-native` is an externally
published package (pinned to `0.1.5-beta.18` in `@crawlee/fs-storage`)
that is **not** versioned in lockstep with the monorepo, so rewriting it
produced a non-existent version and broke CI publishing.

This excludes `@crawlee/fs-storage-native` from the version-rewrite loop
in both the `canary` and `pin-versions` branches.
…emote plugins

BrowserPlugin no longer knows about remote connections. RemoteBrowserPool
swaps the supplied plugins for RemotePlaywrightPlugin/RemotePuppeteerPlugin,
which own the library connect() call via RemoteBrowserConnection.

Closes #3849
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants