Company or project name
No response
Describe what's wrong
With spec.settings.enableDatabaseSync: true (default), cluster reconciliation gets stuck in a loop on CleanupDatabaseReplicas:
type: SchemaInSync
status: "False"
reason: ReplicasNotCleanedUp
message: "Some stale replicas are not cleaned up"
Because the reconcile loop keeps failing and requeuing on this step, new pods never start.
Setting enableDatabaseSync: false bypasses the issue and lets the pods start immediately.
Does it reproduce on the most recent release?
Yes
How to reproduce
- Deploy a standard KeeperCluster and ClickHouseCluster with default settings (enableDatabaseSync left enabled).
- Scale up replicas (or wait for reconciliation against an existing pod).
- The operator runs CleanupDatabaseReplicas and ClickHouse throws Code 32.
Expected behavior
Reconciliation succeeds and allows new pods to start
Error message and/or stacktrace
<Error> executeQuery: Code: 32. DB::Exception: Attempt to read after eof: Cannot parse Int32 from String, because value is too short: while executing 'FUNCTION toInt32(__table1.
database_shard_name :: 0) -> toInt32(__table1.database_shard_name) Int32 : 16'. (ATTEMPT_TO_READ_AFTER_EOF) (in query: SELECT `__table1`.`cluster` AS `database`, toInt32(`__table1`.
`database_shard_name`) AS `shard_id`, toInt32(`__table1`.`database_replica_name`) AS `replica_id`, `__table1`.`is_active` AS `is_active` FROM `system`.`clusters` AS `__table1` WHERE
(`__table1`.`database_replica_name` != '') AND ((toInt32(`__table1`.`database_shard_name`) >= 1) OR (toInt32(`__table1`.`database_replica_name`) >= 1)))
Additional context
In internal/controller/clickhouse/commands.go, listStaleDatabaseReplicasQuery runs:
toInt32(database_shard_name) AS shard_id,
toInt32(database_replica_name) AS replica_id,
system.clusters contains static clusters from remote_servers (such as default), where database_shard_name is "". ClickHouse's filter transform evaluates toInt32("") on those rows before filtering, which throws Code: 32.
Company or project name
No response
Describe what's wrong
With spec.settings.enableDatabaseSync: true (default), cluster reconciliation gets stuck in a loop on CleanupDatabaseReplicas:
Because the reconcile loop keeps failing and requeuing on this step, new pods never start.
Setting enableDatabaseSync: false bypasses the issue and lets the pods start immediately.
Does it reproduce on the most recent release?
Yes
How to reproduce
Expected behavior
Reconciliation succeeds and allows new pods to start
Error message and/or stacktrace
Additional context
In internal/controller/clickhouse/commands.go, listStaleDatabaseReplicasQuery runs:
system.clusters contains static clusters from remote_servers (such as default), where database_shard_name is
"". ClickHouse's filter transform evaluates toInt32("") on those rows before filtering, which throws Code: 32.