We should have full traceability and observability of our app, so that we are aware of every thing that correlates to problems. It would greatly help our agents to solve the problems and it's better to do it sooner. It's based on my discussion with Honza Skalla (infra / devops expert) and analysis from AI. Please take a look @redeyecz, later on we'll convert this into epic and split it into smaller separate tasks.
AI analysis
https://chatgpt.com/share/695d8213-a8b0-800e-94f3-d6ae14ac1e0d
Validation / corrections (your pasted notes)
eBPF tooling list
- bpftrace: ✅ correct; best for fast ad‑hoc tracing, not “always‑on” observability.
- libbpf + CO‑RE: ✅ correct; best for “productized” agents (portable via BTF/CO‑RE).
- BCC: ✅ exists / works, but ⚠️ increasingly “legacy” vs libbpf/CO‑RE (runtime LLVM, heavier).
- Cilium/Hubble: ✅ but K8s/CNI‑centric; not a fit unless you run K8s + Cilium.
- Pixie: ✅ but Kubernetes‑only by design (it’s “for Kubernetes applications”). ([Pixie Documentation]1)
bpftrace examples
tracepoint:syscalls:sys_enter_* { @[probe] = count(); }: ✅ valid, but ⚠️ very high‑volume; you’ll drown in events if left broad.
- “userspace function tracing (requires debug symbols)” via
uprobe:/path:func: ⚠️ not strictly true. You need a resolvable symbol (or an address). Debug symbols help for arg decoding + nicer symbolization, but not always required.
USDT / Node.js claims (most important correction)
USDT syntax in bpftrace (fix)
- The snippet omitted the namespace/provider. Correct forms include:
usdt:binary_path:[probe_namespace]:probe_name (namespace optional only if unique). ([bpftrace.org]4)
Pixie deployment note (fix)
- “Quick Pixie install …
px deploy … for Docker compose” → ⚠️ misleading. Pixie is built around cluster components (Vizier/PEMs) and is documented as for Kubernetes. ([Pixie Documentation]1)
“Docker: CAP_BPF + CAP_PERFMON”
“Beyla will connect everything with trace IDs”
Odigos claims (nuance)
- Odigos: ✅ real “zero‑code” approach, but primarily Kubernetes workflow (operator mounts instrumentation + env vars into pods). ([Odigos - Enterprise Grade OpenTelemetry]9)
- Also: it solves “app tracing without code changes”, not “kernel ↔ trace_id perfect fusion” by itself.
What “full traceability per request” actually takes (Node/Medusa on VMs + Docker)
1) Make OpenTelemetry the source of truth for “one request”
Because eBPF can’t magically know your business request context (trace_id) unless something propagates it. OTel exists exactly for this. ([OpenTelemetry]10)
Minimum bar in MedusaJS:
- OTel SDK for Node: auto‑instrument HTTP server/client, fetch/undici/axios, etc.
- Add manual spans around Medusa “business steps” (cart calc, inventory check, payment authorize, order create).
- Ensure W3C propagation end‑to‑end (
traceparent headers).
Logs:
- Structured logs must include
trace_id + span_id (log ↔ trace pivot).
DB:
2) Use eBPF for “outside-in” truth + gap filling (not primary correlation)
Best fit on VMs/Docker (no K8s):
-
Beyla / OpenTelemetry eBPF Instrumentation (OBI) as a host/sidecar agent:
- emits OTel spans (HTTP/S, gRPC) ([GitHub]7)
- can propagate trace context via
traceparent ([Grafana Labs]6)
- can run in Docker, but needs privileges + PID sharing. ([Grafana Labs]11)
-
Keep bpftrace/BCC for on-demand deep dives (syscalls, TCP retrans, file I/O) — not always on.
3) Reality check on “connecting kernel syscalls to a single request”
You can get usable correlation, but “perfect” is hard:
-
Strong correlation (easy): traces ↔ logs ↔ app metrics via trace_id/span_id (all inside OTel world).
-
Medium correlation: eBPF network telemetry ↔ traces by:
- traceparent injection (Beyla) ([Grafana Labs]6)
- plus resource attrs (host, container, pid) and timestamps.
-
Weak correlation: raw syscall events ↔ traces (usually by PID+time windows; good for investigations, not dashboards).
Concrete “do this” stack for your setup (VMs + Docker, self-hosted)
Core (request traceability):
- OTel SDK in MedusaJS (canonical trace_id)
- OTel Collector (receives OTLP, enriches, exports)
- Backend UI: Tempo/Jaeger for traces + Loki for logs + metrics backend (Prometheus/Mimir)
eBPF add-ons:
- Beyla/OBI on each VM to auto-capture HTTP/gRPC spans and context propagation when useful. ([GitHub]7)
- bpftrace only for targeted scripts (avoid wildcard syscall tracing in prod except short windows).
Docker permissions (don’t hand-wave this):
- Expect needing privileged or
SYS_ADMIN and shared PID namespace for “agent container instruments app container” pattern. ([Grafana Labs]11)
If you want, I’ll sanity-check next
Send (paste, no secrets):
- kernel version(s) on the VMs (
uname -r)
- Node version + how you run Medusa (single container? multiple? sidecars?)
- DB: Postgres? MySQL? hosted where? (same VM vs separate)
- which external calls matter most (payments/shipping/search) so we define “golden spans” first
We should have full traceability and observability of our app, so that we are aware of every thing that correlates to problems. It would greatly help our agents to solve the problems and it's better to do it sooner. It's based on my discussion with Honza Skalla (infra / devops expert) and analysis from AI. Please take a look @redeyecz, later on we'll convert this into epic and split it into smaller separate tasks.
AI analysis
https://chatgpt.com/share/695d8213-a8b0-800e-94f3-d6ae14ac1e0d
Validation / corrections (your pasted notes)
eBPF tooling list
bpftrace examples
tracepoint:syscalls:sys_enter_* { @[probe] = count(); }: ✅ valid, butuprobe:/path:func:USDT / Node.js claims (most important correction)
“Node.js has built‑in USDT probes (since v12)… enable with
--enable-dtrace-probes/ default” → ❌ wrong/outdated.usdt:/usr/bin/node:http__server__requestis very likely not usable on modern Node builds. ([Opeyemi Onikute]3)USDT syntax in bpftrace (fix)
usdt:binary_path:[probe_namespace]:probe_name(namespace optional only if unique). ([bpftrace.org]4)Pixie deployment note (fix)
px deploy… for Docker compose” →“Docker: CAP_BPF + CAP_PERFMON”
CAP_SYS_ADMINin container environments; Docker “CAP_BPF/CAP_PERFMON” support is not a safe assumption. ([eunomia.dev]5)“Beyla will connect everything with trace IDs”
Mostly ✅, with important boundaries:
Beyla does distributed tracing via W3C
traceparentpropagation; it can read incoming context and inject/propagate outgoing. ([Grafana Labs]6)Beyla is eBPF auto‑instrumentation that produces OpenTelemetry spans for HTTP/S and gRPC (that’s what it claims as core capture). ([GitHub]7)
The pasted claim “HTTP → MikroORM query → Postgres syscalls” is ❌ overstated:
Odigos claims (nuance)
What “full traceability per request” actually takes (Node/Medusa on VMs + Docker)
1) Make OpenTelemetry the source of truth for “one request”
Because eBPF can’t magically know your business request context (trace_id) unless something propagates it. OTel exists exactly for this. ([OpenTelemetry]10)
Minimum bar in MedusaJS:
traceparentheaders).Logs:
trace_id+span_id(log ↔ trace pivot).DB:
Since MikroORM lacks first‑class OTel instrumentation, do one of:
em.flush(), transaction boundaries, and “unit of work” operations with manual spans; add useful attrs (entity counts, flush reason). ([GitHub]8)2) Use eBPF for “outside-in” truth + gap filling (not primary correlation)
Best fit on VMs/Docker (no K8s):
Beyla / OpenTelemetry eBPF Instrumentation (OBI) as a host/sidecar agent:
traceparent([Grafana Labs]6)Keep bpftrace/BCC for on-demand deep dives (syscalls, TCP retrans, file I/O) — not always on.
3) Reality check on “connecting kernel syscalls to a single request”
You can get usable correlation, but “perfect” is hard:
Strong correlation (easy): traces ↔ logs ↔ app metrics via trace_id/span_id (all inside OTel world).
Medium correlation: eBPF network telemetry ↔ traces by:
Weak correlation: raw syscall events ↔ traces (usually by PID+time windows; good for investigations, not dashboards).
Concrete “do this” stack for your setup (VMs + Docker, self-hosted)
Core (request traceability):
eBPF add-ons:
Docker permissions (don’t hand-wave this):
SYS_ADMINand shared PID namespace for “agent container instruments app container” pattern. ([Grafana Labs]11)If you want, I’ll sanity-check next
Send (paste, no secrets):
uname -r)