From 022403d94610652fb8597b6dbcca25f2309d690c Mon Sep 17 00:00:00 2001 From: Ricardo Zanini Date: Thu, 1 Oct 2026 15:20:23 -0400 Subject: [PATCH 1/2] Remove kubebuilder default from ApplicationSpec.Replicas to allow HPA override When users omit `spec.replicas`, the field stays nil and the controller doesn't enforce a value, allowing HorizontalPodAutoscaler to manage scaling freely without interference from reconciliation loops. Updates ApplicationSpec documentation to clarify this behavior and regenerates CRDs. Closes #13 --- api/v1/deployment_types.go | 3 +- ...ogic.kubesmarts.org_logicflowruntimes.yaml | 3 +- .../logic.kubesmarts.org_logicplatforms.yaml | 9 ++---- .../ROOT/pages/deployment/production.adoc | 31 +++++++++++++++++-- 4 files changed, 34 insertions(+), 12 deletions(-) diff --git a/api/v1/deployment_types.go b/api/v1/deployment_types.go index b1ddc3a..b3a48b5 100644 --- a/api/v1/deployment_types.go +++ b/api/v1/deployment_types.go @@ -17,10 +17,9 @@ type ApplicationSpec struct { ImagePullPolicy corev1.PullPolicy `json:"imagePullPolicy,omitempty"` // Replicas is the desired number of pod replicas. - // Ignored when HorizontalPodAutoscaler is configured. + // When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. // +optional // +kubebuilder:validation:Minimum=0 - // +kubebuilder:default=1 Replicas *int32 `json:"replicas,omitempty"` // Resources specifies compute resource requirements. diff --git a/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml b/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml index 2745edc..7d983e9 100644 --- a/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml +++ b/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml @@ -9474,10 +9474,9 @@ spec: type: array type: object replicas: - default: 1 description: |- Replicas is the desired number of pod replicas. - Ignored when HorizontalPodAutoscaler is configured. + When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. format: int32 minimum: 0 type: integer diff --git a/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml b/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml index 7697b98..352f8b5 100644 --- a/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml +++ b/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml @@ -9500,10 +9500,9 @@ spec: type: array type: object replicas: - default: 1 description: |- Replicas is the desired number of pod replicas. - Ignored when HorizontalPodAutoscaler is configured. + When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. format: int32 minimum: 0 type: integer @@ -19276,10 +19275,9 @@ spec: type: array type: object replicas: - default: 1 description: |- Replicas is the desired number of pod replicas. - Ignored when HorizontalPodAutoscaler is configured. + When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. format: int32 minimum: 0 type: integer @@ -28834,10 +28832,9 @@ spec: type: array type: object replicas: - default: 1 description: |- Replicas is the desired number of pod replicas. - Ignored when HorizontalPodAutoscaler is configured. + When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. format: int32 minimum: 0 type: integer diff --git a/docs/antora/modules/ROOT/pages/deployment/production.adoc b/docs/antora/modules/ROOT/pages/deployment/production.adoc index 5d981d6..85b0ba7 100644 --- a/docs/antora/modules/ROOT/pages/deployment/production.adoc +++ b/docs/antora/modules/ROOT/pages/deployment/production.adoc @@ -90,13 +90,40 @@ spec: == Horizontal Pod Autoscaling -The operator sets `spec.replicas` on the managed `Deployment`. -Create an `HorizontalPodAutoscaler` targeting that `Deployment` to scale based on CPU or custom metrics. +The operator creates a `Deployment` for the runtime and sets `spec.replicas` based on your LogicFlowRuntime spec. +To enable autoscaling, create a `HorizontalPodAutoscaler` targeting that `Deployment`. + +Example HPA for a LogicFlowRuntime named `hello-runtime`: + +[source,yaml] +---- +apiVersion: autoscaling/v2 +kind: HorizontalPodAutoscaler +metadata: + name: hello-runtime-autoscaler + namespace: default +spec: + scaleTargetRef: + apiVersion: apps/v1 + kind: Deployment + name: hello-runtime # Must match LogicFlowRuntime name + minReplicas: 1 + maxReplicas: 10 + metrics: + - type: Resource + resource: + name: cpu + target: + type: Utilization + averageUtilization: 70 +---- [NOTE] ==== When PostgreSQL persistence is enabled, scale-up is safe — leases ensure each workflow instance is owned by exactly one pod. Without persistence, scale-up creates duplicate in-memory state for in-flight workflows. + +The operator manages lease assignment/cleanup automatically when replica count changes. ==== == cert-manager From 9a1919ccf1a24d32cd90e491343ead6f887608d3 Mon Sep 17 00:00:00 2001 From: Ricardo Zanini Date: Thu, 1 Oct 2026 15:44:48 -0400 Subject: [PATCH 2/2] Use webhook to default LogicFlowRuntime.Replicas and clarify HPA targeting - Add mutating webhook to default LogicFlowRuntime.spec.replicas=1 - Remove kubebuilder default from ApplicationSpec to allow LogicPlatform.DataIndex HPA management - Register defaulter in webhook manager and test setup - Regenerate CRDs with updated ApplicationSpec (no default) Update HPA documentation: - Clarify that LogicFlowRuntime HPA must target the CR (not Deployment) for lease reconciliation - Add separate Data Index HPA section targeting the Deployment - Explain why: leases derive desired count from spec.replicas, so HPA must update the CR - Data Index can omit replicas to let HPA manage freely This enables HPA on both runtimes (via scale subresource) and Data Index (via nil replicas). --- api/v1/deployment_types.go | 2 +- api/v1/logicflowruntime_webhook.go | 19 +++++++ api/v1/webhook_integration_test.go | 1 + cmd/main.go | 1 + ...ogic.kubesmarts.org_logicflowruntimes.yaml | 2 +- .../logic.kubesmarts.org_logicplatforms.yaml | 6 +-- config/webhook/manifests.yaml | 20 ++++++++ .../ROOT/pages/deployment/production.adoc | 51 +++++++++++++++++-- 8 files changed, 92 insertions(+), 10 deletions(-) diff --git a/api/v1/deployment_types.go b/api/v1/deployment_types.go index b3a48b5..c660e56 100644 --- a/api/v1/deployment_types.go +++ b/api/v1/deployment_types.go @@ -17,7 +17,7 @@ type ApplicationSpec struct { ImagePullPolicy corev1.PullPolicy `json:"imagePullPolicy,omitempty"` // Replicas is the desired number of pod replicas. - // When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. + // Ignored when HorizontalPodAutoscaler is configured. // +optional // +kubebuilder:validation:Minimum=0 Replicas *int32 `json:"replicas,omitempty"` diff --git a/api/v1/logicflowruntime_webhook.go b/api/v1/logicflowruntime_webhook.go index 5e952e7..84f4e24 100644 --- a/api/v1/logicflowruntime_webhook.go +++ b/api/v1/logicflowruntime_webhook.go @@ -6,6 +6,25 @@ import ( "sigs.k8s.io/controller-runtime/pkg/webhook/admission" ) +// +kubebuilder:webhook:path=/mutate-logic-kubesmarts-org-v1-logicflowruntime,mutating=true,failurePolicy=fail,sideEffects=None,groups=logic.kubesmarts.org,resources=logicflowruntimes,verbs=create;update,versions=v1,name=mlogicflowruntime-v1.kb.io,admissionReviewVersions=v1 + +// +kubebuilder:object:generate=false +type LogicFlowRuntimeDefaulter struct{} + +var _ admission.Defaulter[*LogicFlowRuntime] = &LogicFlowRuntimeDefaulter{} + +// Default sets Replicas to 1 if not specified. +// We use a webhook instead of +kubebuilder:default=1 because ApplicationSpec is also used by LogicPlatform.DataIndex, +// which needs to support HPA management without operator enforcement. The webhook allows us to apply the default +// conditionally only for LogicFlowRuntime, leaving DataIndex.Application.Replicas nil for HPA to manage freely. +func (d *LogicFlowRuntimeDefaulter) Default(_ context.Context, obj *LogicFlowRuntime) error { + if obj.Spec.Replicas == nil { + one := int32(1) + obj.Spec.Replicas = &one + } + return nil +} + // +kubebuilder:webhook:path=/validate-logic-kubesmarts-org-v1-logicflowruntime,mutating=false,failurePolicy=fail,sideEffects=None,groups=logic.kubesmarts.org,resources=logicflowruntimes,verbs=create;update,versions=v1,name=vlogicflowruntime-v1.kb.io,admissionReviewVersions=v1 type LogicFlowRuntimeValidator struct{} diff --git a/api/v1/webhook_integration_test.go b/api/v1/webhook_integration_test.go index 7eba241..295f895 100644 --- a/api/v1/webhook_integration_test.go +++ b/api/v1/webhook_integration_test.go @@ -81,6 +81,7 @@ func TestWebhookIntegration(t *testing.T) { t.Fatalf("register LFD webhook: %v", err) } if err := builder.WebhookManagedBy(mgr, &LogicFlowRuntime{}). + WithDefaulter(&LogicFlowRuntimeDefaulter{}). WithValidator(&LogicFlowRuntimeValidator{}). Complete(); err != nil { t.Fatalf("register LFR webhook: %v", err) diff --git a/cmd/main.go b/cmd/main.go index 84bcc8c..709dbd1 100644 --- a/cmd/main.go +++ b/cmd/main.go @@ -248,6 +248,7 @@ func main() { os.Exit(1) } if err := builder.WebhookManagedBy(mgr, &logicv1.LogicFlowRuntime{}). + WithDefaulter(&logicv1.LogicFlowRuntimeDefaulter{}). WithValidator(&logicv1.LogicFlowRuntimeValidator{}). Complete(); err != nil { setupLog.Error(err, "unable to create webhook", "webhook", "LogicFlowRuntime") diff --git a/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml b/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml index 7d983e9..3a11da6 100644 --- a/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml +++ b/config/crd/bases/logic.kubesmarts.org_logicflowruntimes.yaml @@ -9476,7 +9476,7 @@ spec: replicas: description: |- Replicas is the desired number of pod replicas. - When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. + Ignored when HorizontalPodAutoscaler is configured. format: int32 minimum: 0 type: integer diff --git a/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml b/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml index 352f8b5..06e66d5 100644 --- a/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml +++ b/config/crd/bases/logic.kubesmarts.org_logicplatforms.yaml @@ -9502,7 +9502,7 @@ spec: replicas: description: |- Replicas is the desired number of pod replicas. - When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. + Ignored when HorizontalPodAutoscaler is configured. format: int32 minimum: 0 type: integer @@ -19277,7 +19277,7 @@ spec: replicas: description: |- Replicas is the desired number of pod replicas. - When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. + Ignored when HorizontalPodAutoscaler is configured. format: int32 minimum: 0 type: integer @@ -28834,7 +28834,7 @@ spec: replicas: description: |- Replicas is the desired number of pod replicas. - When omitted, defaults to 1. Leave unset to allow HorizontalPodAutoscaler to manage scaling. + Ignored when HorizontalPodAutoscaler is configured. format: int32 minimum: 0 type: integer diff --git a/config/webhook/manifests.yaml b/config/webhook/manifests.yaml index 49087ab..6a39946 100644 --- a/config/webhook/manifests.yaml +++ b/config/webhook/manifests.yaml @@ -4,6 +4,26 @@ kind: MutatingWebhookConfiguration metadata: name: mutating-webhook-configuration webhooks: +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /mutate-logic-kubesmarts-org-v1-logicflowruntime + failurePolicy: Fail + name: mlogicflowruntime-v1.kb.io + rules: + - apiGroups: + - logic.kubesmarts.org + apiVersions: + - v1 + operations: + - CREATE + - UPDATE + resources: + - logicflowruntimes + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/docs/antora/modules/ROOT/pages/deployment/production.adoc b/docs/antora/modules/ROOT/pages/deployment/production.adoc index 85b0ba7..de4be79 100644 --- a/docs/antora/modules/ROOT/pages/deployment/production.adoc +++ b/docs/antora/modules/ROOT/pages/deployment/production.adoc @@ -90,8 +90,13 @@ spec: == Horizontal Pod Autoscaling -The operator creates a `Deployment` for the runtime and sets `spec.replicas` based on your LogicFlowRuntime spec. -To enable autoscaling, create a `HorizontalPodAutoscaler` targeting that `Deployment`. +The operator exposes the Kubernetes scale subresource on LogicFlowRuntime and does not enforce default replicas on LogicPlatform.DataIndex, enabling HPA management for both. + +=== LogicFlowRuntime Autoscaling + +Target LogicFlowRuntime directly (not the Deployment). +HPA updates `spec.replicas` on the CR, which the operator uses to reconcile both the Deployment AND lease pool. +If you target the Deployment directly, leases won't track replica changes. Example HPA for a LogicFlowRuntime named `hello-runtime`: @@ -104,9 +109,9 @@ metadata: namespace: default spec: scaleTargetRef: - apiVersion: apps/v1 - kind: Deployment - name: hello-runtime # Must match LogicFlowRuntime name + apiVersion: logic.kubesmarts.org/v1 + kind: LogicFlowRuntime + name: hello-runtime minReplicas: 1 maxReplicas: 10 metrics: @@ -126,6 +131,42 @@ Without persistence, scale-up creates duplicate in-memory state for in-flight wo The operator manages lease assignment/cleanup automatically when replica count changes. ==== +=== Data Index Autoscaling + +Target the generated Data Index Deployment directly (`flow-platform-data-index` or similar). +Leave `spec.dataIndex.application.replicas` unset in LogicPlatform so the operator doesn't enforce a value. + +Example HPA for Data Index: + +[source,yaml] +---- +apiVersion: autoscaling/v2 +kind: HorizontalPodAutoscaler +metadata: + name: data-index-autoscaler + namespace: data-index-system +spec: + scaleTargetRef: + apiVersion: apps/v1 + kind: Deployment + name: flow-platform-data-index # Deployment created by operator + minReplicas: 1 + maxReplicas: 5 + metrics: + - type: Resource + resource: + name: cpu + target: + type: Utilization + averageUtilization: 75 +---- + +[NOTE] +==== +Do not set `spec.dataIndex.application.replicas` when using HPA. +The operator respects the value if set, but will override HPA scaling. +==== + == cert-manager The admission webhooks require a valid TLS certificate.