Skip to content

[Feature] Make DROP TABLE asynchronous and expose deletion progress - #18806

Open
Caideyipi wants to merge 3 commits into
apache:masterfrom
Caideyipi:async-drop-table-procedure
Open

Caideyipi wants to merge 3 commits into
apache:masterfrom
Caideyipi:async-drop-table-procedure

Conversation

@Caideyipi

Copy link
Copy Markdown
Collaborator

Description

DROP TABLE now acknowledges the request after the ConfigNode validates the table and persists its procedure. The deletion continues in the background, so the client does not wait for data and device removal.

Procedure execution

  • Route table deletion to a dedicated scheduler and worker, including procedure recovery and timeout rescheduling. Other procedures continue using the existing worker pool.
  • Preserve synchronous DROP VIEW behavior.
  • Return an error if the initial procedure persistence fails.

Visibility and progress

  • Keep the table visible in SHOW TABLES while its metadata is in PRE_DELETE; SHOW TABLES DETAILS shows that status.
  • Add information_schema.drop_table_procedures with database, table_name, procedure_id, state, progress, and error_message. progress reports the current deletion stage. Recently completed procedures remain visible until the existing procedure retention period expires.

Verification

  • ConfigNode unit tests: DropTableProcedureTest and ClusterSchemaInfoTest (9 tests passed).
  • Table model integration tests: IoTDBTableIT#testAsyncDropTableProgress, IoTDBTableIT#testManageTable, and IoTDBDatabaseIT#testInformationSchema passed in a test cluster.
  • Compiled the ConfigNode and DataNode reactor with both English and Chinese locales; the final branch also compiled after rebasing onto origin/master.

This PR has:

  • been self-reviewed.
  • added unit tests or modified existing tests to cover new code paths.
  • added integration tests.
  • been tested in a test IoTDB cluster.

Key changed/added classes (or packages if there are too many classes) in this PR
  • ProcedureManager, ProcedureExecutor, and DropTableProcedure: asynchronous submission, worker isolation, and stage reporting.
  • InformationSchema and InformationSchemaContentSupplierFactory: progress table and query results.
  • ConfigNode Thrift RPC and clients: transport procedure progress to DataNodes.

Comment on lines +1239 to +1254
statement.execute("create database async_drop_db");
statement.execute("use async_drop_db");
statement.execute("create table drop_target (device string tag, reading int32)");
statement.execute("drop table drop_target");

try (final ResultSet resultSet =
statement.executeQuery(
"select procedure_id, state, progress "
+ "from information_schema.drop_table_procedures "
+ "where database = 'async_drop_db' and table_name = 'drop_target'")) {
assertTrue(resultSet.next());
assertTrue(resultSet.getLong(1) >= 0);
assertFalse(resultSet.getString(2).isEmpty());
assertFalse(resultSet.getString(3).isEmpty());
assertFalse(resultSet.next());
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The procedure may finish before the query starts.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated the IT to wait for completion and verify the retained SUCCESS result, removing the redundant query immediately after DROP TABLE. This covers both a drop that finishes before the first query and one that is still running. The unit test checks the active deletion stage while execution is held by a CountDownLatch, so that assertion does not depend on deletion timing.

Comment on lines +1261 to +1272
try (final ResultSet resultSet =
statement.executeQuery(
"select procedure_id, state, progress, error_message "
+ "from information_schema.drop_table_procedures "
+ "where database = 'async_drop_db' and table_name = 'drop_target'")) {
assertTrue(resultSet.next());
assertTrue(resultSet.getLong(1) >= 0);
assertEquals("SUCCESS", resultSet.getString(2));
assertEquals("SUCCESS", resultSet.getString(3));
Assert.assertNull(resultSet.getString(4));
assertFalse(resultSet.next());
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are the results stored even after the procedure finishes? How long and how many?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. The view includes both active drops and completed results retained by the executor's existing completed-procedure cleaner. Retention uses procedure_completed_evict_ttl (60 seconds by default, measured from the procedure's last update); the cleaner runs every procedure_completed_clean_interval (30 seconds by default), so an expired record disappears on a subsequent cleanup. There is no separate count limit: all drop results within that retention window are included. This PR reuses the existing procedure store and cleanup policy rather than adding a separate history store. I also documented the shared TTL retention on getDropTableProcedures().

Comment on lines +145 to +147
dropTableWorkerThreads = new CopyOnWriteArrayList<>();
dropTableWorkerThreads.add(
new WorkerThread(threadGroup, dropTableScheduler, "DropTableProcedureWorker-"));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Better to delay the init until the first DropTable comes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changed this to lazy initialization: init()/startWorkers() no longer create a DROP TABLE worker when there is no drop to run. The first scheduled drop creates one worker; recovered runnable/failed drops create it during recovery and start it in startWorkers(). Initialization, startup, and shutdown share the lifecycle lock so concurrent first submissions cannot create or start it twice. Added regression coverage for concurrent first drops, recovery/restart, completed drops not creating a worker, DROP VIEW routing, and regular procedures progressing while a drop is blocked.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants