Skip to content

[Draft] Master catalog spi review 16#65809

Closed
morningman wants to merge 86 commits into
apache:masterfrom
morningman:master-catalog-spi-review-16
Closed

[Draft] Master catalog spi review 16#65809
morningman wants to merge 86 commits into
apache:masterfrom
morningman:master-catalog-spi-review-16

Conversation

@morningman

Copy link
Copy Markdown
Contributor

Only for testing #65785

@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@morningman

Copy link
Copy Markdown
Contributor Author

run buildall

morningman and others added 16 commits July 20, 2026 16:42
This multi-month refactor needs persistent state for progress, decisions,
risks, and cross-session agent handoff. Establishes a file-based tracking
system including dashboard, ADR decision log, deviation log, risk register,
per-stage task files, per-connector tracking, and an agent collaboration
playbook covering context budget / subagent usage / handoff norms. Closes
18 design decisions (D-001..D-018) and registers 14 risks (R-001..R-014).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…T27) (apache#63582)

## Summary

Lands the P0 SPI baseline for the catalog-SPI migration (master plan
§3.1 / RFC §2.1), with zero impact on the already-migrated JDBC + ES
connectors.

- **Batch 0** (commits 1-2): SPI types + fe-core bridges —
`ConnectorMetaInvalidator`, `ConnectorTransaction`,
`ConnectorMvccSnapshot`, `ExternalMetaCacheInvalidator`,
`ConnectorMvccSnapshotAdapter`, `PluginDrivenTransactionManager`
generalization.
- **Batch 1** (commit 3): DDL + Partition SPI —
`ConnectorCreateTableRequest` + 4 spec POJOs, 4 new defaults on
`ConnectorTableOps`, 3 new fields on `ConnectorPartitionInfo`, fe-core
converter, `PluginDrivenExternalCatalog.createTable` routing.
- **Batch 2** (commit 4): Import-gate + unit tests —
`tools/check-connector-imports.sh` wired through exec-maven-plugin;
`FakeConnectorPlugin` covering every default fall-through; routing tests
for the invalidator; converter tests for all 4 partition styles + 2
bucket flavors.

## Commits

- `[feat](connector) add P0 batch 0 SPI baseline: MetaInvalidator /
Transaction / MvccSnapshot` (T03-T08)
- `[feat](connector) wire P0 batch 0 SPI into fe-core` (T09-T12)
- `[feat](connector) add P0 batch 1 SPI: CreateTableRequest +
listPartitions` (T13-T20)
- `[feat](connector) add P0 batch 2 gate + unit tests` (T21-T23,
T26-T27)

## Test plan

- [x] `mvn -pl
fe-connector/fe-connector-api,fe-connector/fe-connector-spi -am compile`
— SPI modules compile
- [x] `mvn -pl fe-core -am compile -Dmaven.build.cache.enabled=false` —
fe-core compile
- [x] `mvn -pl fe-core checkstyle:check` — 0 violations
- [x] `mvn -pl fe-connector validate` — import gate runs and passes
(baseline clean)
- [x] `mvn -pl fe-core -am test
-Dtest='FakeConnectorPluginTest,ExternalMetaCacheInvalidatorTest,CreateTableInfoToConnectorRequestConverterTest,ConnectorPluginManagerTest,ConnectorSessionImplTest'`
— 39/39 green
- [x] `mvn -pl
fe-connector/fe-connector-jdbc,fe-connector/fe-connector-es -am compile`
— downstream connectors compile unchanged
- [ ] JDBC regression-test suite (T24) — to be exercised by this PR's CI
pipeline
- [ ] ES regression-test suite (T25) — to be exercised by this PR's CI
pipeline

## Tracking

Full plan, decisions, and risk log live under `plan-doc/` in the repo
(introduced by 6315983, already on the base branch). Per-task
status: `plan-doc/tasks/P0-spi-foundation.md`.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…pache#63641)

P1 batch A — close out scan-node SPI consolidation while keeping
migration-period fallbacks in place. Three surgical changes route
`PluginDrivenExternalTable` first in the nereids translator hot paths so
already-migrated SPI connectors (JDBC, ES) take the SPI route, while the
existing `instanceof XExternalTable` chains remain as fallbacks for
connectors still pending migration (P3–P7).

- **T3** — `PhysicalPlanTranslator.visitPhysicalFileScan`: move the
existing `PluginDrivenExternalTable` branch from position 8 to position
1; the 7 connector-specific branches (HMS / Iceberg / Paimon / Trino /
MaxCompute / LakeSoul / RemoteDoris) stay in place as migration-period
fallbacks
- **T4** — `PhysicalPlanTranslator.visitPhysicalHudiScan`: add a
`PluginDrivenExternalTable` branch routed to
`PluginDrivenScanNode.create(...)`, threading `tableSnapshot` +
`scanParams` through `FileQueryScanNode` setters; `incrementalRelation`
flagged as a P3 Hudi SPI extension TODO. The new branch is unreachable
today (`PhysicalHudiScan` is only built for `HMSExternalTable +
DLAType.HUDI`), so this is groundwork for P3 with zero current-day
runtime impact
- **T5** — `LogicalFileScan`: in `computeOutput()`, add a
`PluginDrivenExternalTable` branch calling new helper
`computePluginDrivenOutput()` — same shape as `computeIcebergOutput`,
using `getFullSchema()` + virtualColumns; in
`supportPruneNestedColumn()`, add an explicit `PluginDrivenExternalTable
→ false` branch. Both behaviorally equivalent for JDBC/ES today since
they have no hidden cols and no virtualColumns

P1 batch B (T1 — delete 13 legacy `Jdbc*Client` + `JdbcFieldSchema`) is
deferred to P8 because the 3 fe-core callers —
`PostgresResourceValidator`, `StreamingJobUtils`,
`CdcStreamTableValuedFunction` — are live CDC streaming code that
requires SPI extension for `getPrimaryKeys` / `getColumnsFromJdbc` /
`listTables`, which is out of P1 surgical scope.

Background and tracking docs live in `plan-doc/` (Master Plan §3.2 P1,
tasks/P1-scan-node-cleanup.md, decisions log).

- [x] `mvn -pl fe-core -am compile -Dmaven.build.cache.enabled=false` →
BUILD SUCCESS
- [x] `mvn -pl fe-core checkstyle:check` → 0 violations
- [x] JDBC + ES regression-test passing — baseline established in P0 /
PR apache#63582
- [ ] PR CI green on this PR
- [ ] Manual scan-node smoke for an SPI connector — JDBC `SELECT *`
should fall into the new `PluginDrivenExternalTable` branch first

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
…apache#64096)

### What problem does this PR solve?

Related PR: apache#63582 (P0 — SPI baseline), apache#63641 (P1 — nereids
plugin-driven routing)

Problem Summary:

This is **P2** of the catalog SPI migration and targets the
`branch-catalog-spi` feature branch (continuing P0 apache#63582 and P1
apache#63641). It fully migrates `trino-connector` off the legacy in-tree
`fe-core/datasource/trinoconnector/` implementation and onto the
connector SPI module `fe-connector-trino`, making `trino-connector` the
first connector to complete the SPI consumption playbook that later
connectors will reuse as a template.

All five batches land together so there is no intermediate state where a
newly-created trino catalog cannot be serialized.

**Batch A — complete the SPI surface (`fe-connector-trino` only, no
fe-core changes)**
- `TrinoConnectorProvider.validateProperties`: enforce the required
`trino.connector.name` property at `CREATE CATALOG` time (ported from
the legacy `checkProperties`).
- `TrinoDorisConnector.preCreateValidation`: call `ensureInitialized()`
so plugin loading + connector-factory resolution happen at catalog
creation instead of being deferred to the first `SELECT`.
- `TrinoConnectorDorisMetadata.applyFilter` / `applyProjection`: bridge
Trino native filter/projection pushdown, reusing
`TrinoPredicateConverter` to translate a Doris `ConnectorExpression`
into a Trino `TupleDomain`. `remainingFilter` is conservatively returned
as the original expression to match legacy behavior (conjuncts are not
stripped; BE re-evaluates them).

**Batch B — fe-core bridge for image compatibility**
- `GsonUtils`: atomically replace the three legacy `registerSubtype`
entries (`TrinoConnectorExternalCatalog` / `Database` / `Table`) with
`registerCompatibleSubtype` redirects onto the `PluginDrivenExternal*`
hierarchy. This must be atomic — `RuntimeTypeAdapterFactory` rejects
duplicate labels, so keeping both bindings would throw at static init.
Mirrors what ES/JDBC already did.
- `PluginDrivenExternalCatalog.gsonPostProcess`: extract a
`legacyLogTypeToCatalogType()` helper that maps `Type.TRINO_CONNECTOR` →
`"trino-connector"`; the generic `name().toLowerCase()` would otherwise
produce the wrong `"trino_connector"` (underscore) that `CatalogFactory`
does not recognize.
- `PluginDrivenExternalTable.getEngine()` / `getEngineTableTypeName()`:
add `trino-connector` branches that preserve the legacy engine-name /
table-type display across `SHOW TABLE STATUS` and `information_schema`.

**Batch C — flip the switch**
- Add `"trino-connector"` to `CatalogFactory.SPI_READY_TYPES` so catalog
creation routes through the SPI path.

**Batch D — remove legacy code**
- Drop the `instanceof TrinoConnectorExternalTable` scan branch in
`PhysicalPlanTranslator` (the `PluginDrivenExternalTable` SPI branch
already handles it).
- Drop `case "trino-connector"` in `CatalogFactory`.
- Delete `fe-core/datasource/trinoconnector/` (10 files) and the
now-dead legacy `TrinoConnectorPredicateTest`.
- Route the `TRINO_CONNECTOR` db-build case in `ExternalCatalog` to
`PluginDrivenExternalDatabase` (mirrors the migrated JDBC case).
- **Retained for image compatibility**: the
`InitCatalogLog.Type.TRINO_CONNECTOR` and
`TableType.TRINO_CONNECTOR_EXTERNAL_TABLE` enums, the GsonUtils
redirects, and the `MetastoreProperties` trino-connector entry.

**Batch E — tests + tracking docs**
- 29 JUnit 5 unit tests over the plugin-free converters:
- `TrinoPredicateConverterTest` — `ConnectorExpression` pushdown trees →
Trino `TupleDomain` (EQ / range / NE / IN / IS [NOT] NULL / AND / OR,
Slice encoding), plus graceful degradation to `TupleDomain.all()` on
null/unsupported input.
- `TrinoTypeMappingTest` — Trino SPI type → Doris `ConnectorType`
(scalars, decimal precision/scale, timestamp precision clamp,
array/map/struct, unsupported-type failure).
- `TrinoConnectorProviderTest` — `validateProperties` fast-fails when
`trino.connector.name` is missing/empty.
- No Trino plugin/cluster required; plugin-dependent paths remain
covered by the existing `external_table_p0/p2` `trino_connector`
regression suites.
- Sync the migration tracking docs under `plan-doc/` (already carried on
this feature branch since P0).

**Net effect**: 28 files, +1025 / −2681 (~1656 LOC net removed). Old FE
images holding legacy trino catalogs / databases / tables deserialize
onto the `PluginDrivenExternal*` hierarchy through the GsonUtils
string-name redirect, with engine-name display preserved.

**Deferred (follow-ups, not in this PR)**:
- `trino_connector_migration_compat` regression test (old-image
deserialization) — requires a running cluster + Trino plugin + docker,
unavailable in this dev environment; tracked as a CI/cluster follow-up.
- The plugin-install documentation update lives in the `doris-website`
repo and is handled separately.

### Release note

None

### Check List (For Author)

- Test
- [x] Unit Test — 29 new tests in `fe-connector-trino` (predicate
converter / type mapping / property validation).
- [ ] Regression test — existing `trino_connector` suites cover plugin
paths; the new old-image compat regression is deferred to a CI/cluster
follow-up.
    - [ ] Manual test (add detailed scripts or steps below)
    - [ ] No need to test or manual test. Explain why:
- [ ] This is a refactor/code format and no logic has been changed.
        - [ ] Previous test can cover this change.
        - [ ] No code files have been changed.
        - [ ] Other reason

- Behavior changed:
- [x] No. Internal routing moves from the legacy fe-core path to the SPI
path; image compatibility, engine-name display, and pushdown semantics
all mirror the legacy behavior. All batches land together, so there is
no serialization-gap window.

- Does this need documentation?
- [x] Yes. The trino-connector plugin-install doc update is a follow-up
in the `doris-website` repo.

### Check List (For Reviewer who merge this PR)

- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label
…tch design (hybrid, T02-T08) (apache#64143)

## Proposed changes

testing with apache#64146

P3 of the catalog-SPI migration (base: `branch-catalog-spi`). Migrates
the **hudi** connector following the **hybrid** strategy (D-019): harden
the dormant HMS-over-SPI hudi connector to correctness parity, build a
test baseline, and write the per-table dispatch design — **all behind
the closed gate** (`SPI_READY_TYPES` unchanged).

> ⚠️ **No user-visible behavior change.** The SPI hudi path stays
dormant (gate closed); hudi queries continue to use the legacy
`HMSExternalTable.dlaType=HUDI` path. This PR removes correctness
blockers ahead of the live cutover (deferred to P7 / batch E).

### What's included

**Correctness fixes (hardening dormant code, behind gate):**
- **T02** — fix hudi JNI `column_types` double bug: emit full Hive type
strings (was Doris bare type names, losing precision/scale/subtypes) and
send `column_names`/`column_types`/`delta_logs` as typed lists
end-to-end (was comma join/split, which shattered `decimal(10,2)` /
`struct<...>`). Matches the BE `hudi_jni_reader.cpp` contract (names `,`
/ types `#` / delta `,`).
- **T04** — fail loud on time-travel / incremental read in the SPI
`visitPhysicalHudiScan` branch (was silently returning the latest
snapshot / silently full-scanning).
- **T05** — real EQ/IN partition pruning in
`HudiConnectorMetadata.applyFilter` (was a placeholder that ignored
predicates and unconditionally switched the partition source from
Hudi-metadata to HMS); faithfully mirrors
`HiveConnectorMetadata.applyFilter`.
- **T07** — column-name casing fix in `avroSchemaToColumns` (top-level
lowercase, mirroring legacy `HMSExternalTable`).

**Test baseline (all three connector modules started P3 with 0 tests):**
- `fe-connector-hudi` (33): type-mapping / schema-parity (COW/MOR
golden) / table-type / partition-pruning / scan-range.
- `fe-connector-hms` (12): shared Hive-type-string parser tests.
- `fe-connector-hive` (14): file-format / partition-pruning (mirrors
T05).
- COW/MOR schema is **type-agnostic** (golden parity vs legacy
`initHudiSchema`); table type only affects scan planning.

**Decisions / design (code-grounded, design-only):**
- **T03** — defer `schema_id`/`history_schema_info` field-id evolution
to batch E (DV-006; not a model-agnostic SPI fix).
- **T06** — keep MVCC/snapshot SPI defaults (opt-out) + document
(DV-007).
- **T08** — `tableFormatType` dispatch design memo + **D-020**: single
`hms` catalog per-table routing via a new backward-compatible
`ConnectorMetadata.getScanPlanProvider(handle)` (per-table provider
seam); refines D-005. The keystone gap is split into M1 (identity
consumption, fe-core reads `tableFormatType` as an opaque string) and M2
(scan routing).

### Deferred to batch E / P7 (not in this PR)
Gate flip (`SPI_READY_TYPES += hms/hudi`), fe-core `tableFormatType`
consumption (M1+M2 implementation), live cutover, delete legacy
`datasource/hudi/`, full incremental/time-travel/MVCC, Iceberg-on-hms
via SPI (needs P6 `IcebergScanPlanProvider`), cluster/runtime
validation.

### Verification
Per task tracking, each code batch landed with: per-module compile +
checkstyle 0 (incl. test sources) + connector import-gate pass + new
unit tests green. The two most recent commits are docs-only
(`plan-doc/`); the code is unchanged since the last green batch. Gate
stays closed → the dormant SPI path is unreachable at runtime → zero
live-path risk. CI re-verifies.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
…core + make fe-core odps-free (T07-T09) (apache#64300)

Follow-up to apache#64253 (the MaxCompute catalog-SPI cutover). After the
cutover a `max_compute` catalog deserializes to
`PluginDrivenExternalCatalog` and no legacy `MaxComputeExternal*` object
is ever instantiated, so the legacy MaxCompute subsystem in fe-core is
dead code. This removes it and makes fe-core's dependency tree fully
odps-free.

**1. Remove legacy subsystem** (`7a4db351100`)
- Delete 20 fe-core files: `datasource/maxcompute/*` (incl.
`MCTransaction`, `MaxComputeScanNode`/`Split`), the MaxCompute
sink/insert/txn plumbing, and 2 legacy-only tests.
- Clean ~21 reverse-reference sites (imports + dead
`instanceof`/visitor/rule branches), keeping every
`PluginDriven`/connector sibling branch and the image/replay keep-set
(GsonUtils compat strings;
`TableType`/`TransactionType`/`TableFormatType`/`InitCatalogLog.Type`
`MAX_COMPUTE` enums; block-id thrift).
- Rewire 3 tests; e.g. `FrontendServiceImplTest`'s block-id RPC test now
mocks the generic `Transaction` SPI, since `getMaxComputeBlockIdRange`
reads the PluginDriven connector transaction.

**2. Make fe-core odps-free** (`409300a75b8`)
- Drop the two odps deps from `fe-core/pom.xml`.
- Move `MCUtils` from fe-common into
`be-java-extensions/max-compute-connector` (its only consumer after the
removal); keep `MCProperties` (odps-free constants) in fe-common.
- Drop `odps-sdk-core` from fe-common — it was also leaking
netty/protobuf transitively to fe-common's own
`DorisHttpException`/`GsonUtilsBase`, so declare `netty-all` +
`protobuf-java` directly (proper dependency hygiene).

**3. Doc-sync** (`f8c305765e8`) — plan-doc
PROGRESS/HANDOFF/deviations/design tracking notes.

- `mvn -pl :fe-core -am test-compile` (main+test) passes; checkstyle 0
violations; connector import-gate passes.
- `grep -rn com.aliyun.odps fe/fe-core/src` → empty.
- `mvn -pl :fe-core dependency:tree | grep odps` → empty (no odps,
direct or transitive).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…he#64446) (apache#64446)

Migrates the legacy in-tree Paimon catalog (fe-core `datasource/paimon`,
`metacache/paimon`, `property/metastore/*Paimon*` and their reverse
references) onto the catalog SPI as a self-contained
`fe-connector-paimon`
plugin, following the maxcompute full-adopter + cutover template.
`paimon`
is added to `SPI_READY_TYPES`, so a Paimon catalog now deserializes to
`PluginDrivenExternalCatalog` and all five functional areas —
normal-table
read, system-table read, DDL, MTMV/MVCC/time-travel (procedures are a
doc-only no-op: fe-core has none) — flow through the generic
`PluginDriven*`
bridge with no Paimon-specific branch in fe-core.

Paimon is the first lakehouse full-adopter to exercise the MVCC /
sys-table /
MTMV SPI surfaces, which become the reusable template for the upcoming
iceberg/hudi migrations, so the cutover is held to byte-parity with the
legacy path. The legacy `datasource/paimon/*` classes are left in place
(now
dead on the cutover path) and removed in a follow-up, mirroring the
maxcompute cutover -> P4 removal.

DefaultConnectorContext)

Source-agnostic seams so any connector can express these capabilities;
every
new method is a default no-op/identity, so hive/iceberg/jdbc/maxcompute/
trino/es are byte-for-byte unaffected.

- MVCC/time-travel: `resolveTimeTravel(session, handle,
ConnectorTimeTravelSpec)`
+ `applySnapshot(...)` (threads a pinned snapshot into the handle before
planScan); new immutable `ConnectorTimeTravelSpec` (SNAPSHOT_ID /
TIMESTAMP /
TAG / BRANCH / INCREMENTAL) that fe-core extracts from SQL;
`ConnectorMvccSnapshot`
  carries `schemaId` for schema-evolution-aware time travel.
- System tables: `listSupportedSysTables` / `getSysTableHandle`; fe-core
composes
  the `{base}${sys}` reference name.
- Scan: proportional split-weight
(`getSelfSplitWeight`/`getTargetSplitSize`),
COUNT(*) pushdown (`getPushDownRowCount`, 7-arg `planScan`), native-read
accounting (`isNativeReadRange`), `getDeleteFiles` (VERBOSE
merge-on-read),
`ignorePartitionPruneShortCircuit` (predicate-driven connectors keep
genuine-null
  partitions).
- Metadata: `ConnectorPartitionInfo.fileCount` +
`SUPPORTS_PARTITION_STATS` (rich
SHOW PARTITIONS); `ConnectorColumn.withTimeZone` (WITH_TIMEZONE DESC
marker,
  independent of the timestamp_tz mapping flag); connector cache control
  (`invalidateTable`/`invalidateAll`, `schemaCacheTtlSecondOverride`).
- `ConnectorContext`: six engine-side hooks the connector must not
implement itself
(`loadHiveConfResources`, `vendStorageCredentials`,
`getBackendStorageProperties`,
`normalizeStorageUri` x2, `getStorageProperties`), bound in
`DefaultConnectorContext`
over fe-core's single-source-of-truth credential/URI/HiveConf utilities
(no
  re-ported logic that could drift).

The metastore-side counterpart of the fe-filesystem StorageProperties
SPI: a
Paimon catalog resolves its backend (hive/dlf/rest/jdbc/filesystem) by
ServiceLoader dispatch on `paimon.catalog.type` instead of a fe-core
switch.

- `fe-connector-metastore-api`: hadoop/SDK-free `MetaStoreProperties`
contracts
  (HMS/DLF/REST/JDBC/FileSystem) carrying only Map/scalar facts.
- `fe-connector-metastore-spi`: `MetaStoreProvider<P>` + first-hit
`MetaStoreProviders`
dispatcher + five impls, registered via META-INF/services.
`HmsMetaStorePropertiesImpl`
reproduces the legacy `buildHmsHiveConf` key set and parity-critical
ordering
(storage overlay before the kerberos block; username/sasl last). Adding
a backend =
  one provider class + one services line (no central enum/switch).
- ServiceLoader is loaded with the SPI interface's own (plugin)
classloader, NOT the
thread-context CL: at CREATE CATALOG the static init first fires on an
FE worker
thread whose TCCL is the FE app loader, so a 1-arg load would cache an
empty
  provider list process-wide.
- fe-core bridges add
`initExecutionAuthenticator(List<StorageProperties>)` so
HDFS-Kerberos doAs is wired on the plugin path (legacy
`initializeCatalog` is dead).

Full read + DDL + partitions + statistics + system tables +
MVCC/time-travel across
all five metastore flavors, planning native (ORC/Parquet) vs JNI reads.

- `PaimonCatalogFactory` (pure flavor switch: Options + Hadoop
Configuration + HiveConf,
classloader-pinned for child-first loading), `PaimonCatalogOps`
(injection seam over
the remote Catalog for offline-testable metadata),
`PaimonConnectorMetadata` (~9 -> 28
methods: schema latest/at-snapshot, sys tables, MVCC/time-travel, DDL,
partitions,
  statistics, table descriptor).
- `PaimonScanPlanProvider` (~6 -> 30 methods): predicate/projection
pushdown,
native-vs-JNI routing, large-file sub-splitting with deletion vectors,
COUNT(*)
collapse, schema-evolution field-id dictionary, `$ro` unwrap to base
FileStoreTable.
- New: `PaimonSchemaBuilder` (CREATE TABLE),
`PaimonIncrementalScanParams` (@incr),
`PaimonTableResolver` (sys/branch-aware handle -> Table),
`PaimonLatestSnapshotCache`
+ `PaimonSchemaAtMemo` (per-catalog caches on the long-lived connector).
- `PaimonTypeMapping` gains the write direction (Doris -> Paimon) for
CREATE TABLE.

- NEW `PluginDrivenMvccExternalTable` (generic MvccTable + MTMV host:
the runtime class
for paimon tables), `PluginDrivenMvccSnapshot`,
`PluginDrivenSysExternalTable` +
`systable/PluginDrivenSysTable`. Behavior is selected by
`ConnectorCapability`
(SUPPORTS_MVCC_SNAPSHOT, SUPPORTS_PARTITION_STATS) — no paimon branch in
fe-core.
- `PluginDrivenScanNode`: MVCC pin threaded at every handle-consumption
site, sys-table
time-travel guard, COUNT pushdown, connector-agnostic EXPLAIN
re-emission.
- Cutover: `CatalogFactory` adds `paimon` to `SPI_READY_TYPES` and drops
the built-in
case; `GsonUtils` removes the 7 built-in `Paimon*` subtype registrations
and adds
replay compat so persisted `PaimonExternalCatalog` (all 5 flavors) /
`...Database`
-> `PluginDriven*` and `PaimonExternalTable` ->
`PluginDrivenMvccExternalTable` on
  deserialization (existing clusters upgrade without losing catalogs).
- Nereids/DDL: `PhysicalPlanTranslator` drops the `PaimonScanNode`
branch;
`UserAuthentication`, `CreateTableInfo`, `ShowPartitionsCommand`
generalized;
`Env.getDdlStmt` renders LOCATION/PROPERTIES for SHOW CREATE (paimon
engine, no
  credential leak).

- fe-filesystem: NEW `HdfsFileSystemProperties` / `HdfsConfigFileLoader`
(typed HDFS;
defaults-laden BE map vs defaults-free FE Hadoop-config map so it cannot
clobber a
co-bound object store's fs.s3a.* during a multi-backend merge); S3 MinIO
support +
  legacy tuning defaults; OSS/OBS/COS ANONYMOUS provider-type; new
`FileSystemFactory.bindAllStorageProperties` /
`FileSystemPluginManager.bindAll`.
- NEW `fe-kerberos` leaf module: dependency-free `AuthType` /
`KerberosAuthSpec` facts
  shared across modules without pulling in Hadoop.
- Packaging: NEW `fe-connector-paimon-hive-shade` relocates
`org.apache.thrift` for
HMS 2.3.7; the paimon plugin-zip excludes fe-thrift/libthrift (provided
parent-first)
to avoid a split `TBase` across loaders; self-contained bundle
(hadoop-aws, AWS-SDK,
hadoop-hdfs-client). `paimon.version` (1.3.1) pinned as the single
property across
FE connector / BE paimon-scanner / preload-extensions for FE<->BE
Table/Split serde.
- BE `PaimonJniScanner.getPredicates()` null-predicate backstop
(no-filter scan; pairs
  with the FE producer now always emitting an empty predicate list).

Parity fixes surfaced by the SPI cutover (each restores legacy
semantics):

- Storage/credentials: canonical `s3.*`/`oss.*` keys ->
`fs.s3a.*`/`fs.oss.*`
(STORAGE-CREDS); canonical `AWS_*` BE creds (STATIC-CREDS-BE); per-table
vended
DLF/STS token (REST-VENDED); `oss/cos/obs/s3a` -> `s3://` for data +
deletion-vector
paths (URI-NORMALIZE); external `hive.conf.resources` hive-site.xml
(HMS-CONFRES);
  real doAs on filesystem/jdbc + wrapped HMS read RPCs (KERBEROS-DOAS).
- Reads/types: paimon columns forced nullable (READ-NOTNULL);
`VARCHAR(65533)` not
widened to STRING (VARCHAR-BOUNDARY); dotted `enable.mapping.varbinary`
/
`.timestamp_tz` keys (MAPPING-FLAG-KEYS); per-type partition-value
rendering
(NATIVE-PARTVAL); genuine `\N` not coerced to NULL
(PARTITION-NULL-SENTINEL);
field-id history dict for rename-safe native reads (SCHEMA-EVOLUTION);
paimon-native
  split encode under `enable_paimon_cpp_reader` (CPP-READER).
- Planning/perf: restore the `force_jni_scanner` escape hatch
(FORCE-JNI-SCANNER);
intra-file native sub-splitting (NATIVE-SUBSPLIT); precomputed COUNT(*)
row count
(COUNT-PUSHDOWN); table row count for CBO / SHOW TABLE STATUS
(TABLE-STATS).
- Time travel / DDL: `FOR TIME AS OF` under CST/PST session zones
(TZ-ALIAS); resolved
+ validated `driver_url` for jdbc catalogs (JDBC-DRIVER-URL); reject a
local-only
  name conflict on CREATE TABLE (CREATE-TABLE-LOCAL-CONFLICT).

plan-doc P5 design / task-list / HANDOFF / deviation tracking notes.

- ~70 new FE unit suites: paimon connector (offline recording fakes:
`RecordingPaimonCatalogOps` / `FakePaimonTable`), fe-core PluginDriven*
MVCC /
sys-table / scan-node / split-weight, metastore-spi dispatch +
per-backend
  properties, fe-filesystem HDFS/MinIO/anonymous, fe-kerberos.
- checkstyle 0 violations; connector import-gate passes; FE main+test
compile clean.
- End-to-end paimon behavior is docker-gated (`enablePaimonTest=true`);
the five
functional areas are covered by the `external_table_p0/paimon`
regression suites,
  green on the latest CI run for this branch.
… make fe-core paimon-SDK-free (T29) (apache#64653)

Follow-up to apache#64446 (the Paimon catalog-SPI cutover). After the cutover
a `paimon`
catalog deserializes to `PluginDrivenExternalCatalog` and its tables to
`PluginDrivenMvccExternalTable`, so no legacy `Paimon*` catalog / table
/ scan-node
object is ever instantiated — the entire legacy Paimon subsystem in
fe-core is dead
code. This PR removes it and makes fe-core's dependency tree fully
paimon-SDK-free
(zero `org.apache.paimon` imports, the 5 paimon maven deps dropped).
Mirrors the P4
maxcompute cutover → legacy-removal (apache#64256).

**1. Remove legacy subsystem + sever reverse-refs** (`7632a074e4b`)
- Delete the dead fe-core paimon files: `datasource/paimon/*` (all 5
flavor catalogs,
`PaimonExternalTable`, `PaimonMetadataOps`, `PaimonUtil`,
`source/PaimonScanNode` /
`PaimonSplit` / `PaimonPredicateConverter` / `PaimonSource` /
`PaimonValueConverter`,
`profile/*`), `datasource/metacache/paimon/*`,
`systable/PaimonSysTable`, and the
  legacy-only tests.
- Sever the live reverse-references to the deleted classes, keeping
every `PluginDriven`
/ connector sibling branch and the image/replay keep-set: drop the
`PAIMON` db-creation
switch arm in `ExternalCatalog`; the paimon engine routing in
`ExternalMetaCacheMgr` /
`ExternalMetaCacheRouteResolver`; the dead `PAIMON_EXTERNAL_TABLE`
branch in
`Env.getDdlStmt` (keep the LIVE `PLUGIN_EXTERNAL_TABLE` SHOW CREATE
LOCATION/PROPERTIES
rendering); the `PaimonSysExternalTable` branch in `UserAuthentication`;
the legacy
clauses + `handleShowPaimonTablePartitions()` in `ShowPartitionsCommand`
(keep
`hasPartitionStatsCapability()` and the
`TableType.PAIMON_EXTERNAL_TABLE` enum).
Landed as one compiling unit (mirrors the P4 apache#64300 precedent:
`PaimonUtils` still
  called the removed `ExternalMetaCacheMgr.paimon()`).
- Inline the `getPaimonCatalogType()` string literals
(`hms`/`filesystem`/`dlf`/`rest`/
`jdbc`) to decouple the still-consumed thin metastore-property classes
from the deleted
  `PaimonExternalCatalog`.

**2. Make fe-core paimon-SDK-free** (`32de7d16702`)
- Strip the dead paimon-SDK catalog-building cluster from
`AbstractPaimonProperties` + the
5 flavors (`initializeCatalog` / `buildCatalogOptions` /
`appendCatalogOptions` /
`getMetastoreType` / `normalizeS3Config` / … — 0 live callers; the
plugin path builds
catalogs connector-side). Keep every LIVE SDK-free duty: `warehouse`
`@ConnectorProperty`,
the Kerberos `executionAuthenticator` doAs wiring (read by
`PluginDrivenExternalCatalog`),
property normalization/validation, `getPaimonCatalogType`,
`Type.PAIMON`.
- Vended credentials: delete `PaimonVendedCredentialsProvider` + its
test + the
`VendedCredentialsFactory` `case PAIMON`. Its SDK methods were dead —
reachable only via
the iceberg-only `getStoragePropertiesMapWithVendedCredentials`; the
real paimon vended
path is the connector's `PaimonScanPlanProvider.extractVendedToken`
(moved off fe-core at
cutover). Relocate its one LIVE duty (the REST "skip static storage map"
gate) to a new
SDK-free `MetastoreProperties.isVendedCredentialsEnabled()` (base=false,
`PaimonRestMetaStoreProperties`=true); `CatalogProperty` routes iceberg
through its
provider (byte-identical) and everything else through the
metastore-props method.
- pom: drop `paimon-core` / `-common` / `-format` / `-s3` / `-jindo`
from `fe-core/pom.xml`.
`fe/pom.xml` `paimon.version` is kept (fe-connector-paimon + BE
paimon-scanner still
  consume it).

**3. Doc-sync** (`58160cbcb56`, `895e214b2bd`) — plan-doc HANDOFF /
PROGRESS / connectors /
P5-T29 legacy-removal design tracking notes.

- `mvn -pl :fe-core -am test-compile` (main + test) passes; checkstyle 0
violations;
  connector import-gate (`tools/check-connector-imports.sh`) passes.
- `grep -rn org.apache.paimon fe/fe-core/src` → empty (main and test).
- `mvn -pl :fe-core dependency:tree -Dincludes=org.apache.paimon` →
empty (no paimon dep,
  direct or transitive).
- End-to-end paimon regression (`external_table_p0/paimon`, docker
`enablePaimonTest=true`)
  passed on this branch.
…kerberos (apache#64655)

**P3b: consolidate the drifted Kerberos/Hadoop authentication
implementations into the new top-level neutral leaf module
`fe-kerberos`** as the single source of truth.

Done as 3 commits:

1. **trino → JDK** (`4a740e1`) — replace the only external dependency in
the auth path, trino's `KerberosTicketUtils`, with a JDK-only
(`javax.security.auth.kerberos`) byte-for-byte equivalent, so the
kerberos path is trino-free.
2. **relocate** (`8898e15`) — move the 13 `fe-common`
`security.authentication.*` classes to `org.apache.doris.kerberos.*` in
`fe-kerberos`; retarget all consumer imports (fe-core + 3
be-java-extensions scanners); merge the duplicate `AuthType`.
3. **unify interface** (`5e3e896`) — merge the two competing
`HadoopAuthenticator` interfaces (fe-common's
`PrivilegedExceptionAction` variant vs fe-filesystem-spi's `IOCallable`
variant) into the single fe-kerberos one, and delete
fe-filesystem-hdfs's own
`KerberosHadoopAuthenticator`/`SimpleHadoopAuthenticator` copies (which
had drifted from the canonical impls). `DFSFileSystem` now routes
through the shared authenticators.

`fe-kerberos` remains a top-level neutral leaf (no dependency cycle).

HDFS filesystem access now uses the same authenticators as the HMS path
(restoring parity). Two intentional behavior changes in
fe-filesystem-hdfs: simple / no-`hadoop.username` now runs as remote
user `hadoop` (was: FE process user, direct); kerberos uses the shared
`LoginContext` + 80%-lifetime refresh.

fe-filesystem-hdfs 79/0/0 (+fe-kerberos/spi), checkstyle 0, connector
import-gate clean, whole-repo grep for the removed symbols = 0.

> ⚠️ docker kerberos e2e (HDFS kerberized + HMS) NOT yet run — the real
gate; UGI login can't be exercised in unit tests.
…move legacy fe-core subsystem (apache#64688)

Migrates the legacy in-tree Iceberg catalog (fe-core
`datasource/iceberg`, its
`property/metastore/*Iceberg*` clusters, and every reverse `instanceof
IcebergExternal*` / `case ICEBERG` coupling scattered across nereids /
planner /
alter) onto the catalog SPI as a self-contained `fe-connector-iceberg`
plugin,
following the paimon full-adopter + cutover template. `iceberg` is added
to
`SPI_READY_TYPES`, so an Iceberg catalog now deserializes to
`PluginDrivenExternalCatalog` and every functional area — read, write,
row-level
DML, `ALTER TABLE EXECUTE` procedures, system tables, DDL and
MVCC/time-travel —
flows through the generic `PluginDriven*` bridge with no
Iceberg-specific branch
in fe-core. (Iceberg tables surfaced *through* a `hive`/HMS catalog stay
on the
hive path until P7; this PR covers the dedicated `type=iceberg`
catalog.)

Iceberg is the widest lakehouse adopter migrated so far — 7 catalog
flavors,
merge-on-read row-level DML, 9 stored procedures, v3 row-lineage /
deletion
vectors — so the cutover is held to behavior-parity with the legacy
path. Unlike
the paimon migration (a cutover PR + a separate removal PR), this PR
does the
build-out, the live cutover **and** the full legacy-subsystem removal on
one
branch: fe-core `datasource/iceberg` and its property clusters are
deleted, not
just left dead.

Part of the catalog-SPI migration tracked in apache#65185 (phase **P6**).

Full read + write + row-level DML + procedures + system tables + DDL +
MVCC/time-travel, planning native (Parquet/ORC) vs JNI reads. Key seams:

- `IcebergCatalogFactory` (pure flavor switch, classloader-pinned for
child-first
loading), `IcebergCatalogOps` (metadata / DDL ops),
`IcebergConnectorMetadata`
(schema latest / at-snapshot, sys tables, MVCC/time-travel, DDL,
partitions,
statistics, table descriptor), `IcebergConnectorProvider` (ServiceLoader
entry).
- `IcebergScanPlanProvider` + `IcebergScanRange`: predicate/projection
pushdown
(`IcebergPredicateConverter` → Iceberg `Expression`), FileScanTask
byte-offset
sub-splitting, COUNT(\*) collapse, merge-on-read delete attachment
(position
delete / Puffin DV / equality delete), `IcebergColumnHandle` field-id
column
pruning, schema-evolution-safe reads (`IcebergSchemaUtils` field-id
dictionary),
  synthetic rowid + v3 row-lineage metadata columns.
- Write path: `IcebergWritePlanProvider` / `IcebergWriteContext` /
`IcebergWriterHelper` / `IcebergConnectorTransaction` — INSERT / INSERT
OVERWRITE
(dynamic + static partition), branch-targeted writes, WRITE ORDERED BY,
per-spec
distribution, optimistic-concurrency conflict detection, one SDK
transaction per
  statement, BE commit-fragment → DataFile/DeleteFile conversion.
- Row-level DML: DELETE / UPDATE / MERGE INTO over position-delete (v2)
and
  deletion-vector (v3) merge-on-read, baseSnapshot-anchored read.
- `action/`: the 9 `ALTER TABLE EXECUTE` procedures
(rollback_to_snapshot,
rollback_to_timestamp, set_current_snapshot, cherrypick_snapshot,
fast_forward,
expire_snapshots, publish_changes, rewrite_manifests,
rewrite_data_files)
dispatched by `IcebergExecuteActionFactory`; `rewrite/` hosts the
distributed
  bin-pack `rewrite_data_files` engine.
- `IcebergSchemaBuilder` (CREATE TABLE), `IcebergTypeMapping` (both
directions),
`IcebergColumnChange` / `IcebergComplexTypeDiff` (schema + partition
evolution),
`IcebergLatestSnapshotCache` / `IcebergManifestCache` (per-catalog
caches on the
  long-lived connector), `dlf/` (Aliyun DLF client).

(`fe-connector-metastore-iceberg`)

An Iceberg catalog resolves its backend by `MetaStoreProvider`
ServiceLoader
dispatch on `iceberg.catalog.type` (registered via `META-INF/services`),
instead
of a fe-core switch — **7 flavors**: REST, Hive Metastore, AWS Glue,
Hadoop /
filesystem, JDBC, Aliyun DLF, AWS S3Tables. Adding a backend = one
provider class
+ one services line. The connector self-builds the HMS-metastore
Kerberos
authenticator, so metastore auth is owned entirely by the plugin.

- `CatalogFactory` adds `iceberg` to `SPI_READY_TYPES` and drops the
built-in case.
- `GsonUtils` replay-compat: persisted `IcebergExternalCatalog` →
  `PluginDrivenExternalCatalog`, `IcebergExternalDatabase` →
  `PluginDrivenExternalDatabase`, `IcebergExternalTable` →
`PluginDrivenMvccExternalTable` on deserialization, so existing clusters
upgrade
  without losing catalogs (metadata backward-compat invariant).
- Nereids/DDL/alter generalized off the concrete Iceberg types; SHOW
CREATE
renders LOCATION / PARTITION BY / ORDER BY / PROPERTIES via the generic
path
  (iceberg engine name preserved, no credential leak).

Per the 2026-07 architecture decision (fe-core holds **no** property
parsing for
migrated connectors; storage parsing → `fe-filesystem`, metastore
parsing →
`fe-connector`, both plugin-side; the plugin assembles → BE thrift →
hands back to
fe-core), iceberg's property/credential/auth resolution moves entirely
into the
connector:

- Storage credentials bound to `fe-filesystem` on both scan and write
paths (one
source), retiring the second fe-core `StorageProperties.createAll`
parse.
- Vended temporary credentials, `warehouse` → `fs.defaultFS` derivation
(hadoop flavor only, via the neutral `Connector.deriveStorageProperties`
SPI),
  and Kerberos doAs are all connector-owned; the fe-core pre-execution
  authenticator is a no-op for plugin catalogs.
- fe-core's shared static property/metastore helpers are kept alive only
for the
connectors still on the legacy path (Hive / Hudi / LakeSoul /
HMS-iceberg).

Net **−21.7k** lines from fe-core. Removes the 4 dead entity classes
(`IcebergExternalTable` / `IcebergExternalDatabase` /
`IcebergSysExternalTable` /
`IcebergExternalCatalog`), the fe-core Iceberg `MetastoreProperties`
cluster + its
credential helpers, the 5 fe-core Iceberg connectivity probes and the
coordinator
branches, and un-registers the Iceberg property factories (fail-loud on
any stray
call). The now-dead paimon `MetastoreProperties` cluster left behind by
P5 is
swept in the same pass. `GsonUtils` retains only the three string compat
labels
for old-image upgrade.

Parity fixes surfaced by the SPI cutover (each restores legacy
semantics):

- Classloader: TCCL pinned to the plugin classloader at all three loci
where
fe-core crosses into bundled Iceberg (scan call thread, write/DDL/commit
engine
thread via `TcclPinningConnectorContext`, and Iceberg's own
manifest-writer
  worker pool), preventing a duplicate-copy `ClassCastException`.
- Nested schema dictionary field names lowercased so a mixed-case nested
struct
  (e.g. Iceberg `DROP_AND_ADD`) can't crash BE with `std::out_of_range`.
- Session time zone: lazy `ZoneId` resolution tolerant of Doris
CST/PST-style
  aliases in `FOR TIME AS OF` and datetime partition predicates.
- Storage / URI normalization (`oss/cos/obs/s3a` → `s3`), SHOW CREATE
secret
redaction, and `ALTER TABLE EXECUTE ... WHERE` rejection kept at legacy
wording.

- ~68 FE unit test classes: iceberg connector (offline recording fakes
for the
remote Catalog / Table), metastore-iceberg per-flavor dispatch +
properties, and
the fe-core `PluginDriven*` MVCC / sys-table / scan-node / GSON-replay
surfaces;
  many carry mutation checks.
- checkstyle 0 violations; connector import-gate passes (no
`fe-connector-iceberg → fe-core` import); FE main + test compile clean.
- End-to-end iceberg behavior is docker-gated
(`enableIcebergTest=true`); read,
write, row-level DML, procedures, system tables, DDL and time-travel are
covered
by the `external_table_p0/iceberg` (+ `external_table_p2/iceberg`)
regression
  suites.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…etaspace leak

JdbcConnectorClient cached JDBC-driver classloaders in a ref-counted map that
evicted the entry when refCount reached 0. A JDBC driver self-registers into
the static java.sql.DriverManager on class load, and close() never deregisters
it, so an evicted classloader stays pinned by DriverManager for the life of the
process. Serial CREATE -> DROP -> CREATE churn of the same driver URL therefore
rebuilt a fresh, separately-pinned classloader every cycle, leaking one driver's
worth of Metaspace per cycle.

Live FE evidence (external regression build 986696 OOM): only 4 jdbc catalogs
yet 38 FactoryURLClassLoader instances survived a full GC (jmap -histo:live),
tracking the DriverInfo count -> ~34 leaked, DriverManager-pinned driver
classloaders. FE Metaspace grew 165MB -> 1565MB over the single-concurrency 5.5h
run (multi-concurrency did not OOM: overlapping same-URL catalogs keep
refCount > 0 so the loader was reused instead of evicted-and-rebuilt).

Fix: revert to a keep-alive Map<URL, ClassLoader> (one loader per distinct
driver URL, shared across catalogs, never evicted), mirroring the pre-SPI
fe-core JdbcClient and the pattern IcebergConnector / PaimonConnector already
use. Bounded by the number of distinct driver jars. Adds a regression test
asserting repeated create/recreate cycles for one URL add exactly one cached
classloader.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKf5EzHW75qzgrXnkoBZcF
…TableTest after rebase

Rebase-integration fix, belongs with P4 (apache#64300 "remove legacy maxcompute
subsystem"); squash into it when finalizing that PR.

Upstream apache#64937 added MaxComputeExternalTableTest as a unit test for the
legacy MaxComputeExternalTable.parsePartitionValues. P4 deletes that class,
but the new test file rode into the rebase base with no modify/delete
conflict and then dangled: fe-core test-compile failed with
"cannot find symbol: MaxComputeExternalTable".

No coverage lost: the new connector's
MaxComputeConnectorMetadata.listPartitionValues already resolves partition
values by name via ODPS PartitionSpec.get(column) (structurally immune to
the out-of-order/invalid-spec bug apache#64937 fixed in the legacy path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es into SPI connectors

After rebasing branch-catalog-spi onto upstream master (4dfcb4f), an audit
of the 3 upstream commits touching fe-core datasource/ found behaviour that was
dropped when P5/P6 removed the legacy fe-core paimon/iceberg subsystems. Port the
bounded correctness gaps into the new fe-connector SPI code (main + tests):

- apache#65332 (paimon JNI IOManager): PaimonScanPlanProvider.getBackendPaimonOptions
  returned emptyMap() for non-jdbc catalogs, so the three JNI IOManager options
  (paimon.doris.enable_jni_io_manager / .tmp_dir / .impl_class) never reached BE
  and Paimon primary-key merge reads could not spill (OOM risk). Collect the three
  options (strip the "paimon." prefix) before the metastore-type check and forward
  them for all catalog flavours; jdbc-only driver logic stays in the jdbc branch.

- apache#65094 tier1 (build-time correctness): IcebergSchemaBuilder buildPartitionSpec /
  buildSortOrder and PaimonSchemaBuilder primary/partition keys used bare column
  names, so a partition/sort/key referenced with a different case than the (now
  case-preserving) schema column failed the engine's case-sensitive lookup. Add
  case-insensitive resolution back to the canonical column name (iceberg via
  Schema.caseInsensitiveFindField; paimon via a lower->canonical map), mirroring
  upstream getIcebergColumnName / getPaimonColumnNames.

- apache#65094 tier2 (consistency, worse than upstream): PaimonConnectorMetadata
  emitted "partition_columns" with case-preserving keys while column names are
  lowercased, so PluginDrivenExternalTable's case-sensitive byName lookup silently
  dropped mixed-case paimon partition columns (table treated as non-partitioned).
  Lower-case partition_columns with the same bare toLowerCase() the columns use.

Verified: mvn package -pl :fe-connector-paimon -am (95 paimon tests) + iceberg
23 tests, 0 failures. Read-path column-name case preservation (tier3) and the
apache#63068 OIDC session-credential port are tracked separately.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013KKXx3btbQc6w3MbZbEoRf
…column-name case on the read path

Align the iceberg + paimon catalog-SPI connectors with upstream apache#65094: external
table top-level column names now keep their remote case on the read path instead
of being lowercased. FE-only (BE untouched): the byte-match invariant "Doris scan
slot name == TSchema top-level TField name" is preserved by flipping BOTH the scan
slot names AND the field-id-dict / current-TSchema top-level names to case-preserving
together, so BE (which matches top-level by field-id) needs no change. Nested struct
children stay as-is (iceberg lowercase for BE StructNode.children.at; paimon
paimon-cased). Reverts tier2's lowercasing of paimon partition_columns (bce28b2)
now that the paimon columns are themselves case-preserving.

iceberg (IcebergConnectorMetadata): partition_columns / getColumnHandles / parseSchema.
iceberg (IcebergPartitionUtils): getIdentityPartitionColumns (path_partition_keys),
getIdentityPartitionInfoMap (scan partition-value map) and generateRawPartition -- the
listPartitions value-map key that PluginDrivenExternalTable.getNameToPartitionItems
looks up by the case-preserved partition_columns remote name; omitting it would
silently drop a mixed-case partition column's value in MTMV / partition pruning.
paimon (PaimonConnectorMetadata): partition_columns / getColumnHandles / mapFields.
paimon (PaimonScanPlanProvider): path_partition_keys / getPartitionInfoMap and the
current (-1) TSchema entry (buildSchemaInfo lowercaseTopLevelNames -> false at the
sole current-schema call site; the historical entries and nested names are unchanged).

Tests: flip the connector tests that pinned mixed-case -> lowercase output to assert
case preservation; add an IcebergSchemaUtils mixed-case byte-match case and an
IcebergPartitionUtils.listPartitions regression pinning the case-preserved
partition-value-map key.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013KKXx3btbQc6w3MbZbEoRf
…l DTO + user-session capability

Re-migrate the Iceberg REST OIDC per-user session-credential feature (upstream
apache#63068 / e545f1a, dropped by the 2026-07-09 P6 rebase) onto the catalog SPI.
This is the SPI layer (fe-connector-api):

- ConnectorDelegatedCredential: neutral, immutable DTO carrying a user's
  per-connection OIDC/JWT/SAML token across the fe-core -> connector boundary so
  the connector never imports a fe-core type. The token is connection-scoped and
  in-memory only; toString() redacts it (never edit-logged / persisted / SHOWn).
- ConnectorSession.getDelegatedCredential() (default empty) + getSessionId()
  (default = getQueryId): a connector reads the credential and the stable
  AuthSession key (preserved across FE observer->master forwarding) off the session.
- ConnectorCapability.SUPPORTS_USER_SESSION: gates FE credential injection
  (least-privilege -- a connector that never consumes the token never receives it).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013KKXx3btbQc6w3MbZbEoRf
morningman and others added 22 commits July 20, 2026 16:57
…P2-remainder recon; tracker/handoff updates

- Design + summary for the resolveScanProvider() per-handle memo (fix committed in 87ff73b).
- Records the session-4 recon of the three remaining P2 candidates (all evidence re-verified live,
  not stale): the streaming-pump micro-batch (fe-core framework, needs sign-off because micro-batching
  changes the backend split distribution - not byte-identical), the C15b/C13 wire-byte dedup (protocol
  evolution: thrift delete-file dictionary + BE reader + mixed-version compat, needs sign-off), and the
  C14 generic-node slices #1/#2/#4 (deferred: cross-connector Hadoop-Path byte risk / new SPI capability
  surface). User chose only C14 slice #3 (provider memo).
- tasklist/HANDOFF: mark C14 provider-memo done; HANDOFF next-step = the two sign-off-gated blocks only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…ing spaces

- plan-doc/master-todo/: a cross-task register of deferred, decision-gated big items
  (needs user sign-off or protocol evolution), so they are not forgotten. Seeds two entries
  with full background/current-call-stack/solution-call-stack/example writeups: the streaming
  split-dispatch micro-batch (fe-core; not byte-identical -> changes backend split distribution)
  and the wire-side delete-file dictionary (thrift + BE reader + mixed-version compat).

- plan-doc/per-statement-table-owner-port/: tracking space for porting iceberg's PERF-07
  per-statement table-load owner pattern to other connectors. Key finding recorded: the neutral
  fe-core/fe-connector-api scaffolding (ConnectorStatementScope + ConnectorSession.getStatementScope
  + StatementContext plumbing) is already in place, so porting is connector-side-only work.
  Candidate map (to be scoped next session): paimon (high) / hive-hms (mid-high) / hudi (mid);
  maxcompute/es/jdbc/trino likely excluded (no per-statement metastore loadTable fan-out).

No product code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…conclusions + Trino-refactor groundwork

Recon (recon + adversarial cross-check, double-signed) of every catalog-backed
connector for porting iceberg's per-statement table-load-owner pattern. Verdict:
no connector is worth porting now. iceberg is unique in being the only connector
with a migrated row-level write path (DELETE/MERGE) — the only capability that
triggers the multi-arm reload storm the pattern targets. Every other connector is
read-only (paimon/hudi/es/trino) or INSERT/append-only (hive/hms, jdbc, maxcompute),
and existing cross-query caches already collapse loads to ~1.

Also recorded the architecture-unification analysis: the uniform SPI standard
already exists (neutral getStatementScope() + ConnectorStatementScope); Trino's
per-transaction-metadata model is not directly portable (Doris connectors are
shared singletons; the only per-statement span is StatementContext). Altitude tiers
L0-L3 laid out; paimon's future row-level-write migration is the correct trigger to
extract a shared helper (L1) into fe-connector-api (no iron-rule-A breach).

Per user decision: this round only records conclusions to a standalone doc; the
Trino-architecture refactor (L2/L3) is deferred to a dedicated next-session
discussion (groundwork staged in the doc's section 7). No product code changed.

- designs/recon-findings-and-trino-refactor-groundwork.md (new, full findings)
- tasklist.md: PORT-01..04 all marked recon-judged-unnecessary (paimon = future candidate)
- progress.md / HANDOFF.md: session 1 recon + next-session Trino-refactor discussion

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…oundation)

Foundation for the per-statement ConnectorMetadata redesign: an engine-owned
funnel that memoizes one ConnectorMetadata per (statement, catalog) and can
close it deterministically at statement end. Read-side rerouting, close
wiring, HMS sibling, write sharing, and cache isolation land in later commits.

- ConnectorStatementScope (SPI): add getOrCreateMetadata(key, factory) typed
  memo over computeIfAbsent + closeAll() default no-op. NONE runs the factory
  every call (byte-identical to today).
- ConnectorStatementScopeImpl: idempotent best-effort closeAll() closing each
  AutoCloseable value once (close-once guard, log-and-continue).
- PluginDrivenMetadata: the single funnel - get(session, connector) memoizes
  connector.getMetadata(session) on the session's per-statement scope, keyed
  by catalog id.
- Tests: memo-once-per-statement, NONE-each-call, cross-catalog isolation,
  close-once idempotency.

No production seam routes through the funnel yet; this commit is byte-neutral
and independently compilable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…adata scope (close wiring)

Wire the per-statement ConnectorStatementScope to be closed deterministically at
statement end, so a later commit that gives ConnectorMetadata.close() real work
(releasing FileIO / tables) releases it exactly once, not on GC. P1 close is still
a no-op, so this commit is behavior-neutral; it establishes the lifecycle.

Two-tier close (a single register-at-scope-creation point was rejected: it would
leak the callback for DDL / SHOW / DESCRIBE / EXPLAIN / foreground ANALYZE, which
create a scope via Command.run but never reach unregisterQuery, and the callback
registry has no TTL):
- PRIMARY (coordinated scans/writes/cursor-fetch/internal): getSplits registers
  scope::closeAll on the existing query-finish callback, same query-id key as the
  read-transaction release beside it. Object-capture; skip NONE. Fires after
  off-thread pump quiescence; no registry leak.
- FALLBACK (non-coordinated Command statements): StatementContext.close() closes
  the scope, guarded by isReturnResultFromLocal so arrow-flight defers to its own
  query-finish close. Runs in ConnectProcessor's per-statement finally (direct)
  and a new finally in proxyExecute (master-side forwarded DDL/SHOW).
- resetConnectorStatementScope() closes-before-null, and each retry attempt
  resets, so a reused StatementContext (prepared EXECUTE / retry) never memoizes
  into an already-closed scope whose values would never close.

closeAll is idempotent, so the dual primary+fallback triggers on a coordinated
local query close the scope exactly once.

Tests: reset-closes-before-drop, close()-closes-locally, close()-defers-for-async.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…TEP 1 (C1 foundation + C2 close wiring)

- expanded-scope-phasing-and-security-decisions.md: user decisions (phased, not
  big-bang; the "propagate the instance to background threads" ask was refuted
  by recon and corrected to read-through; cache isolation confirmed as a REAL
  security fix since list != load) + verified 3-area findings + STEP 1-4
  sequencing + C1/C2 implementation progress + close-wiring residual risks.
- trino-parity-metadata-redesign-design.md, P1-implementation-design.md: target
  architecture + seam-by-seam blueprint (from the prior session). P1-design
  section 4 corrected: the "register at scope creation" close plan was found to
  leak (DDL/SHOW/EXPLAIN/ANALYZE via Command.run create a scope but never reach
  unregisterQuery, and the callback registry has no TTL); replaced with the
  two-tier close landed in the code commits.
- HANDOFF / progress / tasklist: C1 (5b7312f) + C2 (12f3e95) marked done;
  next session = STEP 1 C3 (read-side rerouting + background read-through + fix
  the ANALYZE row-count scope-binding hazard).

Docs only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…tement metadata funnel + force read-through for cross-statement loaders

Completes the read-side of the engine-held per-statement ConnectorMetadata
keystone (C1 funnel + C2 close wiring already landed). Four mechanical parts,
all in one commit so every intermediate state stays compilable and green
(the sub-parts interleave within two shared files):

1. buildCrossStatementSession() helper on PluginDrivenExternalCatalog: mirrors
   buildConnectorSession() (same delegated-credential handling) but forces the
   per-statement scope to ConnectorStatementScope.NONE. Under NONE the funnel
   memoizes nothing, so a metadata resolved with this session is built fresh and
   never bound into -- nor closed with -- a live statement's scope.

2. Force read-through for the 9 cross-statement background loaders (database /
   table name caches, schema cache, column-statistic cache, chunk-size and
   row-count loaders, name-mapping resolvers, and the BE-driven metadata TVF):
   each now builds its session via buildCrossStatementSession() and resolves
   through PluginDrivenMetadata.get(...). This also fixes a latent hazard where
   fetchRowCount, reached synchronously from AnalysisManager.buildAnalysisJobInfo
   on the ANALYZE execution thread, would otherwise capture that statement's live
   scope; forcing NONE makes read-through a contract rather than an accident.

3. Reroute the 49 on-thread read / DDL / command-TVF / MVCC seams from
   connector.getMetadata(session) to PluginDrivenMetadata.get(session, connector),
   so a statement resolves one memoized ConnectorMetadata per (catalog) instead
   of rebuilding it at each resolver.

4. PluginDrivenScanNode: add a volatile cachedMetadata field + metadata()
   accessor (mirroring the existing resolvedScanProvider holder); the 8 per-method
   resolvers share it, and the static create() factory routes through the funnel
   directly.

The write-path getMetadata seams (translator x2, BindSink x2, IcebergMergeSink,
IcebergRowLevelDmlTransform, InsertExecutor, and the write-only
resolveWriteTargetHandle) are deliberately left unchanged for the write-sharing
step.

Tests: adapt the plugin/scan/mvcc/command/TVF unit tests -- their mocked
ConnectorSession now flows through the funnel, so getStatementScope() must return
NONE (a Mockito mock returns null for the default method), and test doubles that
override buildConnectorSession() also override buildCrossStatementSession(). No
assertions or verify() call-counts were weakened. Adds
ConnectorSessionImplTest.explicitNoneStatementScopeWinsOverLiveContext pinning
that an explicit NONE scope wins over a live ConnectContext capture.

Verification: fe-core main compile BUILD SUCCESS; 265 targeted tests green;
checkstyle 0 violations. Byte-neutral under NONE (offline/no ConnectContext),
matching pre-funnel behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…st + progress/handoff

Adds the verified seam-by-seam checklist for the read-side rerouting step
(66 connector-metadata factory seams partitioned into 51 reroute / 7-9
force-NONE loaders / 8 write-deferred, cross-checked against current line
numbers by an adversarial recon), records the user's two decisions
(name-mapping seams -> force-NONE; loaders route through the funnel with a
NONE session for a zero-exception anti-drift gate), and updates progress /
tasklist / handoff to point the next session at the anti-drift gate that
closes the read keystone and then the HMS heterogeneous-gateway sibling
fan-out.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…e per-statement funnel

Close out the read keystone of the per-statement metadata redesign with a
build-time gate so the read/scan/DDL/MVCC seams cannot drift back to bare
connector.getMetadata(session) calls.

tools/check-fecore-metadata-funnel.sh (+ self-test) greps fe-core main sources
and fails the build on any Connector#getMetadata(session) call outside the
funnel PluginDrivenMetadata.get(...). Kept silent: the funnel file itself; the
no-arg getMetadata() (a different method); comment mentions; and write-path
seams carrying a `getMetadata-funnel-exempt` marker (on the call line or the
line above it). Wired as a validate-phase exec in fe-core/pom.xml, mirroring
tools/check-connector-imports.sh.

The 8 write-path seams (INSERT/DELETE/MERGE sink, bind, row-level DML, write
target handle) still call getMetadata directly and are marked exempt; the
marker is removed and the gate auto-tightens onto them when the write-sharing
step reroutes writes through the funnel.

Verified: self-test PASS; gate green on the current tree and red on the 8
unmarked write sites; fe-core checkstyle 0 violations; `mvn -pl fe-core validate`
fires the gate with BUILD SUCCESS.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…rift gate) + STEP 2 handoff

RD-1 (read keystone) marked complete: C1 foundation + C2 close wiring + C3
reroute + scan-node field + background read-through + the anti-drift gate.
Progress log records the gate landing (grep universe, rule set, verification);
handoff re-points the next session at the HMS heterogeneous-gateway sibling
fan-out (STEP 2 / RD-2), including the note that the sibling funnel lives in
fe-connector-hive and must go through session.getStatementScope().getOrCreateMetadata
(not fe-core's PluginDrivenMetadata) to stay off the connector-imports gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…ta per statement

The Hive heterogeneous gateway forwards iceberg/hudi-on-HMS operations to embedded
sibling connectors. It obtained a fresh sibling ConnectorMetadata on every forward,
so one statement rebuilt a table's sibling metadata a dozen-plus times. Route the
four sibling-metadata acquisition sites through the existing per-statement scope so
a statement reuses ONE sibling metadata per owner (mirroring fe-core's own funnel).

- Add SiblingOwner{connector,label}. HiveConnector.resolveSiblingOwnerLabeled yields
  the owner label from its matched ownsHandle arm (no force-build, no identity
  comparison). resolveSiblingOwner now delegates to it, so the three connector-level
  provider seams are unchanged.
- HiveConnectorMetadata.memoizedSiblingMetadata keys the scope as
  "metadata:"+catalogId+":"+label. The three connectors share one catalogId, so the
  label keeps their metadata entries distinct (a plain catalogId key would collapse
  them and misroute). The three helpers (icebergSiblingMetadata / hudiSiblingMetadata
  by TYPE, siblingMetadata by HANDLE) and the getTableSchema capability-reflection
  stray route through it; beginTransaction(session,handle) is covered for free because
  it forwards via siblingMetadata. By-TYPE and by-HANDLE mint the same label for one
  owner, so the getTableHandle divert and later per-handle forwards reuse one instance.
- Under NONE scope the factory runs on every call: byte-identical to before. Only
  fe-connector-api types are used (no fe-core import); the foreign sibling is never cast.

Tests: five funnel assertions (one build across many forwards incl. the write txn;
NONE rebuilds each call; by-TYPE == by-HANDLE key; cross-catalog isolation; iceberg /
hudi isolated by label). Existing sibling suites now pass a NONE-scope session (the
forward path dereferences the scope); no forwarding / routing / fail-loud assertion is
weakened. Copies TestStatementScope + a ScopeSession test double into the hive test
tree (connector tests cannot import fe-core / iceberg).

The heterogeneous-gateway e2e (INSERT/DELETE/MERGE/ALTER/EXECUTE on an iceberg-on-HMS
table vs a standalone iceberg catalog) is deferred to the later unified delegation-e2e
pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…oseout + STEP 3 handoff

Record the HMS heterogeneous-gateway sibling-metadata funnel landing (commit
5fd55d0): grounding correction (4 acquisition sites, not ~43), the design
confirmed with the user, e2e deferred to the unified delegation pass, 348 tests
green + gates green + adversarial review clean. Flip RD-2 to done in tasklist and
rewrite HANDOFF to point at STEP 3 (write-side sharing: reroute the 8 write seams
through the funnel, uphold the hive start-of-write refresh + read/write identity
gates, lift ConnectorTransaction ownership).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
… the per-statement funnel

The 8 write-arm getMetadata call sites -- the INSERT/DELETE/MERGE sink
translators, the insert executor's connector setup, the write-target handle
resolver, the two BindSink partition validators, the row-level DML mode check,
and the merge-sink partition-field derivation -- were the last direct
Connector#getMetadata calls in fe-core (each carried a funnel-exempt marker).
Reroute them through PluginDrivenMetadata.get so read and write share the ONE
memoized ConnectorMetadata instance per statement, aligning with Trino's
CatalogTransaction (which memoizes exactly one ConnectorMetadata per
transaction+catalog and shares it across read planning and writes), and delete
the 8 markers so the anti-drift gate now covers all of fe-core with no
exemptions.

Guard the shared reuse with a fail-loud identity check in the funnel: a
session=user connector bakes the querying user's delegated credential into the
metadata instance at build time, so reusing it under a different identity would
execute one user's write with another's credentials. A statement resolves a
catalog under exactly one identity (one statement = one user = one credential),
so the check never fires in practice; it turns any future violation of that
invariant into a hard error instead of a silent cross-user leak. It uses the
Doris principal (getUser), never a credential token, so fe-core still parses no
credentials; under NONE it stores nothing and the check is vacuous.

The hive begin-write refresh (its own getTable under the write auth context for
the ACID reject) is downstream of beginTransaction and untouched. Tests: two new
funnel identity assertions (same user reuses, different user fails loud); the
three write-path suites that now flow through the funnel get the NONE
statement-scope stub (the read-side convention).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…to a per-statement co-holder

Give the write transaction a first-class statement-level owner,
CatalogStatementTransaction, co-held on the per-statement scope next to the one
memoized ConnectorMetadata the read and write arms now share -- mirroring
Trino's CatalogTransaction (one metadata + one transaction per statement +
catalog). The insert executor opens the transaction through the co-holder's
begin(), which mints it from that shared write-ops facet (so the write inherits
the read arm's client/ops) and registers it with the manager exactly as before.
The tx<->session binding, the global registration for the BE block-allocation /
commit-data RPCs, and the executor's own commit/rollback at onComplete/onFail
are all unchanged.

The co-holder's own job is deterministic statement-end teardown. The scope's
closeAll now runs in two ordered passes: pass 1 finalizes the transaction
(finalizeAtStatementEnd rolls back a transaction the executor never committed --
only a mid-flight abort leaves one active), pass 2 closes the remaining
AutoCloseable values (the memoized metadata). A transaction is therefore always
finished before the instance it was minted from is closed. The backstop is
idempotent and can never undo a committed write: PluginDrivenTransactionManager
removes a transaction from its map on both commit and rollback, so the new
isActive(id) is the single source of truth -- on every normal path the executor
already finished the transaction, so the backstop finds nothing active and does
nothing.

Under NONE (offline / no live statement) the scope stores nothing, so the
co-holder is transient and the executor's own commit/rollback is the only
lifecycle -- byte-identical to the pre-co-holder path.

Tests: CatalogStatementTransaction (begin registers active; finalize rolls back
an orphan, never undoes a committed txn, is idempotent after rollback, no-ops
with no txn opened); the scope's two-pass ordering (txn finalized before the
metadata is closed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…+ STEP 4 handoff

Record the read/write-shares-one-metadata-instance step as done: the P3
implementation design + as-built (3a reroute + identity gate, 3b co-holder +
two-pass closeAll, side-car orthogonality), the RD-3 tasklist row (two commits),
the session-7 progress entry, and a rewritten HANDOFF pointing at the cache
isolation step (STEP 4, the independent security track: getIdentityShardKey()
SPI, iceberg projection-cache key sharding, fe-core schema-cache bypass, the
heterogeneous-gateway hole, anti-drift gate, and the required threat-model
sign-off).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…che leak fix)

Records the signed-off design for the cache-isolation security track, with the
three grounding corrections the adversarial verification surfaced:

- The only ACTIVE cross-user leaks under iceberg.rest.session=user are the
  latestSnapshotCache (beginQuerySnapshot reads it without a preceding per-user
  loadTable) and the fe-core name-keyed schema cache (no shouldBypassSchemaCache).
  partitionCache/formatCache are safe-but-fragile (guarded only by an upstream
  per-user loadTable that rests on tableCache being null).
- The heterogeneous HMS gateway is a LATENT, not active, hole: the embedded
  iceberg sibling is forced iceberg.catalog.type=hms and can never be
  session=user; the "propagate owner label to fe-core" fix-vehicle was refuted.

Signed-off route: DISABLE authorization-sensitive projection caches under
session=user (no new SPI, no stale-authz window, mirrors tableCache/commentCache),
all three projection caches unified; fe-core schema-cache bypass mirroring the
db/table-name-cache bypass; a fail-loud gateway guard; an anti-drift gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…ion caches under session=user

The latestSnapshotCache / partitionCache / formatCache were built unconditionally,
so under iceberg.rest.session=user (per-user delegated authorization) they were
shared across users with no user dimension in the key. beginQuerySnapshot reads
latestSnapshotCache WITHOUT a preceding per-user loadTable, so a "can-list-cannot-
load" principal hitting a cached entry received another user's snapshotId/schemaId
with the per-user authorization (which lives inside loadTable) bypassed. partition/
format were only safe because their readers happen to run a per-user loadTable
first (tableCache being null under session=user) -- a fragile dependency.

Gate all three null under isUserSessionEnabled() (mirroring the existing tableCache/
commentCache discipline), so session=user keeps NO live cross-query metadata cache
and every projection is re-loaded live per-user -- re-running authorization every
call with no stale-authz window. beginQuerySnapshot gains a null-cache fallback
(loadLatestSnapshotPin, mirroring resolveTableForRead's tableCache ternary); the
partition/format readers already tolerate a null cache. The invalidate* hooks gain
null-guards on the three (they were unconditional and would NPE on a nulled cache).
Non-session catalogs are byte-unchanged.

Tests: latest-snapshot/partition/format each null for session=user (non-null for
plain + vended); invalidate hooks no-throw with all three nulled. 22 cache-test +
29 mvcc + 51 metadata green; checkstyle 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…ser catalogs

The generic table-schema cache is keyed by table NAME only (SchemaCacheKey wraps a
NameMapping, no user/identity dimension), so under a SUPPORTS_USER_SESSION connector
whose remote loadTable enforces per-user authorization (Iceberg REST session=user) a
shared hit served one user's schema to another who could LIST the table but was NOT
authorized to LOAD it — the same "list != load" metadata disclosure the db/table-name
caches already bypass, but the schema cache was missed.

Add shouldBypassSchemaCache(SessionContext) (default false) as the schema-level analog
of shouldBypassTableNameCache/DbNameCache, overridden in PluginDrivenExternalCatalog
with the identical predicate (supportsUserSession() && ctx.hasDelegatedCredential()).
ExternalTable.getSchemaCacheValue() consults it and, on bypass, reads schema live via
initSchema (exactly what the cache miss-loader would do) instead of the shared cache —
so schema is re-read per user with no stale-authz window. A credential-less session
keeps the shared cache; the fail-closed rejection stays connector-side. This covers the
MVCC "latest" path too (it funnels through super.getSchemaCacheValue()).

Tests: predicate bypasses only for capable+credentialed sessions (empty/null keep the
cache; non-user-session never bypasses); read-site reads live via initSchema and never
touches the shared cache under bypass. 6 tests green; checkstyle 0; funnel gate exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…ling is session=user

fe-core keys its per-user schema/name cache bypass off the FRONT-DOOR connector's
capabilities. A hive gateway front door never declares SUPPORTS_USER_SESSION, so for a
delegated iceberg-on-HMS sibling the bypass would NOT trigger — a session=user sibling
would silently leak cross-user metadata through the shared caches. Today this cannot
happen: the sibling is built with iceberg.catalog.type forced to hms
(IcebergSiblingProperties.synthesize), which can never satisfy isUserSessionEnabled()
(REST-flavor only). Guard that invariant explicitly: getOrCreateIcebergSibling now
throws if the freshly built sibling declares SUPPORTS_USER_SESSION, converting a future
fail-open (someone letting the sibling be REST session=user) into a loud failure at
sibling-build time instead of a silent leak.

Tests: a plain (no-capability) sibling is accepted; a sibling declaring
SUPPORTS_USER_SESSION fails loud with a message naming the invariant. FakeSibling gains
an optional capability set (no-arg ctor unchanged). 16 tests green; checkstyle 0;
connector-import gate exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…on on IcebergConnector

A new cross-query cache field added to IcebergConnector without gating it for
iceberg.rest.session=user would silently reintroduce the "list != load" metadata
disclosure this track just fixed (a shared, un-partitioned cache bypasses the per-user
loadTable authorization). tools/check-authz-cache-sharding.sh turns that from a silent
leak into a build failure: every final cache-typed holder field on IcebergConnector must
carry, on its declaration line or the line directly above it, one of two markers —
authz-cache-session-user-disabled (null under isUserSessionEnabled()) or
authz-cache-exempt (justified as read-only-after-a-per-user-load, e.g. the default-off
manifest cache).

The field-declaration regex matches any "final <Type>Cache " / "final Cache<" holder
regardless of visibility or static-ness, so a future authz cache added as a raw
Caffeine/Guava Cache<...>/LoadingCache<...> is still caught (not only the shipped
"private final Iceberg*Cache" convention). Marker-based like check-fecore-metadata-funnel
.sh; a RED/GREEN mktemp self-test locks the behavior across all these forms and proves the
markers are load-bearing and independent.

Wired as a second exec-maven-plugin validate execution alongside check-connector-imports
in fe/fe-connector/pom.xml (inherited=false). The five projection/table/comment caches
carry the disabled marker; the manifest cache the exempt marker. The fe-core generic
schema cache is a different name-keyed cache protected by shouldBypassSchemaCache, not
by this gate.

Verified: self-test PASS (incl. raw Cache/LoadingCache/static/visibility forms); real gate
exit 0; both validate gates run (BUILD SUCCESS); checkstyle 0; iceberg cache test 22 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…main line done

RD-4 (cache isolation, the "list != load" authorization-leak fix) landed across four
commits: iceberg disables its three authz-sensitive projection caches under
session=user, fe-core bypasses the shared schema cache, the hive gateway fails loud on
a session=user sibling, and an anti-drift gate enforces the invariant. Grounding +
adversarial verification corrected the plan (only latestSnapshotCache + fe-core schema
cache are active leaks; the heterogeneous gateway is latent, not active), and the user
signed off the disable route (no getIdentityShardKey SPI, no stale-authz window).
Clean-room review returned 0 blocker/major.

Updates tasklist (RD-4 -> done), appends progress session 8, and rewrites HANDOFF: the
read/write refactor main line (RD-1..RD-4 / STEP 1-4) is complete; the only remaining
debt is the consolidated e2e (heterogeneous gateway DML/DDL + can-list-cannot-load
authorization) under external_table_p2/refactor_catalog_param.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
…rmalizer signature after rebase

The rebase onto upstream branch-catalog-spi replayed our 46 review commits
onto a base that already carries upstream's four "port fe-core external-table
fixes into the connector" commits. All four survived the 3-way auto-merge and
are functionally intact; the only breakage was a test API drift, not a code
conflict.

Upstream apache#65676's new test convertDeleteRejectsMalformedDeletionVector calls
convertDelete(delete, Collections.emptyMap()) -- the pre-refactor signature.
Our PERF-06 (635aee0) changed convertDelete's second parameter from a raw
Map to a UnaryOperator<String> uriNormalizer (derive the vended storage config
once per scan). The upstream test predates that refactor, so it no longer
compiles.

Align the test to the refactored signature with UnaryOperator.identity(),
matching the sibling DV tests (1669/1685/1698/1773). identity() is correct
here: the case asserts a malformed DV is rejected, where URI normalization is
immaterial. This is the ONLY test migration the rebase required; no connector
code changed. fe-connector-iceberg 488 tests green (IcebergScanPlanProviderTest
113/113, incl. upstream's projection-order + 6 DV-validation tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMtYwYyyubZZiC1odLZTG3
@morningman
morningman force-pushed the master-catalog-spi-review-16 branch from 4563f2d to 960bb33 Compare July 20, 2026 09:14
@morningman

Copy link
Copy Markdown
Contributor Author

run buildall

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 29849 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit 960bb33b7b5bbb908deb47b3a776a6bdadce511d, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17726	4067	4056	4056
q2	2040	341	206	206
q3	10391	1380	829	829
q4	4683	478	338	338
q5	7571	875	583	583
q6	186	176	137	137
q7	769	839	617	617
q8	9327	1609	1653	1609
q9	5626	4347	4344	4344
q10	6740	1759	1487	1487
q11	508	381	321	321
q12	720	583	456	456
q13	18104	3443	2752	2752
q14	263	260	243	243
q15	q16	794	786	718	718
q17	959	974	978	974
q18	7359	5738	5632	5632
q19	1157	1375	1128	1128
q20	793	648	575	575
q21	5903	2610	2551	2551
q22	442	359	293	293
Total cold run time: 102061 ms
Total hot run time: 29849 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	4425	4294	4347	4294
q2	286	321	209	209
q3	4627	4963	4380	4380
q4	2069	2155	1405	1405
q5	4380	4263	4291	4263
q6	257	177	127	127
q7	1734	2185	1716	1716
q8	2575	2199	2189	2189
q9	8070	7964	7828	7828
q10	4693	4619	4197	4197
q11	558	423	413	413
q12	775	783	548	548
q13	3324	3610	2988	2988
q14	311	311	272	272
q15	q16	717	746	661	661
q17	1376	1320	1452	1320
q18	7863	7433	7379	7379
q19	1195	1117	1143	1117
q20	2213	2209	1939	1939
q21	5221	4608	4411	4411
q22	507	441	390	390
Total cold run time: 57176 ms
Total hot run time: 52046 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 177845 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit 960bb33b7b5bbb908deb47b3a776a6bdadce511d, data reload: false

query5	4310	616	484	484
query6	465	218	209	209
query7	4919	612	313	313
query8	346	186	175	175
query9	8828	4119	4104	4104
query10	459	359	325	325
query11	5911	2336	2096	2096
query12	160	101	102	101
query13	1245	573	390	390
query14	6204	5278	4870	4870
query14_1	4226	4235	4230	4230
query15	205	201	183	183
query16	1054	488	466	466
query17	1127	716	583	583
query18	2658	483	361	361
query19	214	194	155	155
query20	115	108	106	106
query21	242	161	144	144
query22	13549	13518	13269	13269
query23	17394	16382	16124	16124
query23_1	16287	16290	16217	16217
query24	7671	1745	1267	1267
query24_1	1314	1305	1282	1282
query25	619	422	360	360
query26	1346	340	206	206
query27	2614	616	370	370
query28	4453	2011	1988	1988
query29	1084	603	462	462
query30	338	262	228	228
query31	1110	1099	973	973
query32	107	59	59	59
query33	527	310	244	244
query34	1182	1169	660	660
query35	767	783	672	672
query36	1196	1205	1062	1062
query37	148	106	93	93
query38	1874	1707	1644	1644
query39	880	882	882	882
query39_1	823	830	845	830
query40	244	155	137	137
query41	65	67	63	63
query42	95	95	90	90
query43	319	325	273	273
query44	1402	781	769	769
query45	189	202	178	178
query46	1069	1193	734	734
query47	2199	2105	1996	1996
query48	389	416	296	296
query49	587	422	308	308
query50	1086	457	335	335
query51	10742	11009	10584	10584
query52	91	90	74	74
query53	262	283	192	192
query54	278	231	208	208
query55	74	72	67	67
query56	293	293	293	293
query57	1329	1297	1196	1196
query58	292	258	263	258
query59	1617	1670	1464	1464
query60	310	284	265	265
query61	154	147	154	147
query62	559	492	434	434
query63	236	205	213	205
query64	2838	1022	845	845
query65	4775	4625	4577	4577
query66	1810	512	375	375
query67	29273	29204	29106	29106
query68	3233	1666	995	995
query69	409	306	265	265
query70	1076	969	990	969
query71	378	341	292	292
query72	3061	2713	2635	2635
query73	850	787	434	434
query74	5043	4896	4718	4718
query75	2574	2511	2140	2140
query76	2316	1145	796	796
query77	348	388	297	297
query78	11842	11817	11386	11386
query79	1389	1115	751	751
query80	1032	584	560	560
query81	508	338	292	292
query82	555	151	119	119
query83	400	318	301	301
query84	276	154	128	128
query85	1004	608	522	522
query86	409	286	286	286
query87	1856	1823	1751	1751
query88	3709	2790	2770	2770
query89	446	363	339	339
query90	1775	197	194	194
query91	203	188	163	163
query92	64	57	57	57
query93	1630	1553	1051	1051
query94	648	343	315	315
query95	784	489	554	489
query96	1134	798	334	334
query97	2620	2641	2491	2491
query98	218	207	198	198
query99	1108	1116	971	971
Total cold run time: 263907 ms
Total hot run time: 177845 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 25.69 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit 960bb33b7b5bbb908deb47b3a776a6bdadce511d, data reload: false

query1	0.01	0.01	0.00
query2	0.15	0.08	0.09
query3	0.39	0.24	0.25
query4	1.60	0.24	0.24
query5	0.34	0.32	0.32
query6	1.17	0.68	0.67
query7	0.04	0.01	0.01
query8	0.09	0.07	0.07
query9	0.51	0.39	0.40
query10	0.58	0.58	0.58
query11	0.31	0.18	0.18
query12	0.31	0.18	0.18
query13	0.52	0.52	0.53
query14	0.93	0.93	0.93
query15	0.68	0.59	0.59
query16	0.39	0.40	0.40
query17	1.01	1.04	1.01
query18	0.31	0.28	0.29
query19	1.95	1.87	1.79
query20	0.02	0.02	0.02
query21	15.44	0.36	0.31
query22	4.96	0.14	0.13
query23	15.85	0.49	0.30
query24	2.43	0.59	0.44
query25	0.15	0.10	0.10
query26	0.74	0.26	0.22
query27	0.11	0.10	0.10
query28	3.43	0.91	0.53
query29	12.48	4.29	3.38
query30	0.38	0.27	0.26
query31	2.78	0.64	0.32
query32	3.23	0.59	0.47
query33	3.02	2.92	2.96
query34	15.72	4.08	3.38
query35	3.24	3.28	3.26
query36	0.66	0.54	0.51
query37	0.12	0.10	0.09
query38	0.08	0.07	0.06
query39	0.07	0.06	0.05
query40	0.20	0.18	0.17
query41	0.13	0.08	0.08
query42	0.08	0.06	0.05
query43	0.07	0.07	0.06
Total cold run time: 96.68 s
Total hot run time: 25.69 s

@hello-stephen

Copy link
Copy Markdown
Contributor

FE UT Coverage Report

Increment line coverage 48.74% (3411/6999) 🎉
Increment coverage report
Complete coverage report

@hello-stephen

Copy link
Copy Markdown
Contributor

BE UT Coverage Report

Increment line coverage 🎉

Increment coverage report
Complete coverage report

Category Coverage
Function Coverage 57.51% (23994/41723)
Line Coverage 41.23% (236070/572539)
Region Coverage 37.06% (186605/503493)
Branch Coverage 38.27% (83844/219105)

@morningman morningman closed this Jul 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants