Skip to content

feat(observability): deployment (OTel collector + Parca/eBPF) - #5377

Draft
Ma77Ball wants to merge 86 commits into
apache:mainfrom
Ma77Ball:obs/pr3/deployment
Draft

Ma77Ball wants to merge 86 commits into
apache:mainfrom
Ma77Ball:obs/pr3/deployment

Conversation

@Ma77Ball

@Ma77Ball Ma77Ball commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this PR?

Provides the local deployment wiring that receives and stores telemetry. Default-off and isolated from the application services.

  • Adds an OpenTelemetry Collector configuration and wires it into the single-node docker-compose stack.
  • Adds the Parca server configuration and the Parca eBPF agent for continuous profiling.
  • Updates the single-node up.sh and .env to start the observability backends.
  • Infrastructure only; the application runs unchanged whether or not these services are started.

Any related issues, documentation, or discussions?

Closes: #5369
Part of #4070. Stacked on #5376.

How was this PR tested?

  • Configuration-validation specs for the collector and Parca config.
  • sbt scalafmtCheckAll passes; the build runs in this PR's CI.

Was this PR authored or co-authored using generative AI tooling?

Co-authored with Claude Opus 4.8 in compliance with ASF

Ma77Ball and others added 3 commits June 5, 2026 04:49
…, SDK bootstrap (default-off)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ca/eBPF profiling

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tracing primitives

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added dependencies Pull requests that update a dependency file docs Changes related to documentations infra common labels Jun 5, 2026
@codecov-commenter

codecov-commenter commented Jun 5, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.66%. Comparing base (4d7fd49) to head (6d738cd).
⚠️ Report is 1 commits behind head on main.
✅ All tests successful. No failed tests found.

Additional details and impacted files
@@             Coverage Diff              @@
##               main    #5377      +/-   ##
============================================
+ Coverage     93.62%   93.66%   +0.03%     
  Complexity     4842     4842              
============================================
  Files          1211     1211              
  Lines         49991    49951      -40     
  Branches       6125     6124       -1     
============================================
- Hits          46805    46786      -19     
+ Misses         1675     1653      -22     
- Partials       1511     1512       +1     
Flag Coverage Δ *Carryforward flag
access-control-service 80.27% <100.00%> (+0.09%) ⬆️
agent-service 99.32% <ø> (ø)
amber 89.94% <ø> (+0.09%) ⬆️ Carriedforward from fd09f20
computing-unit-managing-service 77.20% <100.00%> (+0.05%) ⬆️
config-service 87.25% <100.00%> (+0.12%) ⬆️
file-service 83.65% <ø> (ø) Carriedforward from fd09f20
frontend 96.16% <ø> (+<0.01%) ⬆️ Carriedforward from fd09f20
notebook-migration-service 83.73% <ø> (ø) Carriedforward from fd09f20
pyamber 98.47% <ø> (+0.06%) ⬆️
workflow-compiling-service 74.25% <ø> (-0.75%) ⬇️

*This pull request uses carry forward flags. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions github-actions Bot added the platform Non-amber Scala service paths label Jun 5, 2026
Ma77Ball added 21 commits June 26, 2026 15:53
reduced the comments to include less design details and information not
needed in the codebase.
…WorkflowMetrics rename

Follow the OTel Java demo patterns instead of custom wrappers (zuozhiw
review):

- WorkflowService: start the run-level span with the standard OTel API,
name it WorkflowService.initExecutionService, and drop the
initExecutionServiceSpanned split so no span is passed as an argument.
The real execution failure is now recorded onto the span from
errorHandler, where it is actually caught.

- TexeraTracer: drop the withSpan wrapper, keeping only the tracer
accessor (single instrumentation scope) and currentContext.

- SpanAttrs: drop the awkward with*/set* setter helpers, keep the
standard label keys and sanitizeFreeText; callers set attributes via the
OTel API.

- Rename TexeraMetrics to WorkflowMetrics, document that this facade is
only for the workflow-execution cluster, and add an example of calling
the OTel meter API directly at a call site.

Update SpanAttrsSpec and WorkflowMetricsSpec to match.

@zuozhiw zuozhiw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

left some comments, please fix them, I clicked "approve" so that you can feel free to merge this PR after fixing these commments, because this is a env setup PR, please make sure to carefully make AI actually test these configurations and make sure they work

apart from that I have some reservations about ebpf, it's more sensitive than application level log/traces/metrics. Furthermore, ebpf is normally used for very advanced diagnostics and profiling, and seems a bit heavy for our use case. But again feel free to try it and see if we see it being useful.

Also one more missing piece of the telemetry data is host level metrics, e.g. cpu, memory being the most important ones, please check how we can collect them, e.g. I think otel has some host metrics receiver, please check that, also check if docker / k8s expose metrics and how we can get them.

Another missing piece is database related metrics. E.g. database load, it would also be nice to consider them, but maybe in future PRs.

sh -c 'apk add --no-cache curl jq bash > /dev/null 2>&1 && bash /examples/load-examples.sh'

# ========================================================================
# Part 5: Observability stack (PR 6).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove these descriptions about "PR 6", the code comments should be more concise and factual and it's pointless to carry such information, claude nowadays is not good at writing good and concise comments, ask claude to pay more attention to it

security_opt:
- no-new-privileges:true
volumes:
- ../observability/otel-collector/config.yaml:/etc/otelcol-contrib/config.yaml:ro,z

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path might not exist in Texera’s published Docker Compose bundle. The release workflow currently archives only bin/single-node/, sql/, and NOTICE, while this mount depends on bin/observability/otel-collector/config.yaml.

can you make sure to let AI actually run and test both single node and local dev release bundle workflows with these new files?

Comment thread bin/single-node/.env
# TEXERA_OBSERVABILITY_* env vars below into the right COMPOSE_PROFILES.
# * Or edit COMPOSE_PROFILES directly here.
#
# Disable env-var conventions (consumed by up.sh):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we should by default turn on the entire observability stack, it's more for deployments, and when we deploy, each deployment should override these configurations.

These profiles add the collector, three signal backends, Parca, and a privileged Parca agent to every single-node installation. Their configured memory limits alone total roughly 5.3 GB, while Texera documents 4 GB as the minimum for the complete single-node deployment. I think these services should be opt-in, or enabled through an explicit installation option. Remember we are open source and we might serve various users who might not need observability, unless they are hosting a service.

Comment thread bin/single-node/.env
# TEXERA_OBSERVABILITY_METRICS=disabled drops victoriametrics
# TEXERA_OBSERVABILITY_TRACES=disabled drops jaeger
# TEXERA_OBSERVABILITY_PROFILES=disabled drops parca + parca-agent
# TEXERA_OBSERVABILITY_COLLECTOR=disabled drops the otel-collector (rare)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These variables might not be consumed by Texera’s official bin/single-node.sh path. That entry point delegates to bin/single-node/main.sh, which does not inspect any TEXERA_OBSERVABILITY_* variables.

The referenced up.sh seems to no longer be the canonical launcher, so commands such as TEXERA_OBSERVABILITY_TRACES=disabled bin/single-node.sh up do not disable the profile as documented. Please integrate this behavior into the canonical launcher and test it there. Please double check, I'm not very familiar with the current launching process

# IntelliJ) can emit telemetry. In compose-only deploys, services
# talk to otel-collector:4317/:4318 via the bridge network.
ports:
- "127.0.0.1:4317:4317"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't local dev override config be in the local dev docker override file?

# posture. The privileged + bind-mount block here is the ONLY
# observability service that requires elevated permissions; the
# surface is documented and reviewed.
parca-agent:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ebpf is much more sensitive than logs/metrics/traces, because those are application level telemetry, ebpf is host level telemetry and involves much high privileges. What do you mean by "the surface is documented and reviewed"? documented where and how is it reviewed?

also we are turning it on by default, and the comment below explicitly say that it requires sys admin class privilege and macos/windows cannot run, so I really don't think we need to turn it on by default, this should be an opt-in feature and document the security implications before the user enables it.

also make sure we have proper default overrides in the local dev overrides

privileged: true
pid: "host"
# Bind-mount kernel state read-only — the agent reads /proc and
# /sys for stack-trace symbolization but cannot write to either.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you ask more AI, maybe different models to double check and fact check and think harder on this statement?

`parca-agent.env`:

- `deployment=texera`
- `cluster=local` (override per env)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make sure that parca really only collects texera process and not other processes in the host.

- `deployment=texera`
- `cluster=local` (override per env)

When the PR 7 Texera query gateway runs Parca queries, it filters on

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

again the readme file should not contain any thing about "pr5, pr6, pr7", make absolutely sure to press claude to carefully inspect the writing style of these readmes and code comments! this is an important global comment for all the PRs!

We do **not** label profiles with `workflow.id` / `execution.id`. As
with metrics, those are unbounded identifiers and would blow up
Parca's storage cardinality. Per-execution profile views are reached
by joining on `trace_id` at query time (the Parca query API supports

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please fact check this statement that the ebpf collections can join with trace_id, how does it know our application level trace id? make sure test it

@github-actions github-actions Bot added the ci changes related to CI label Sep 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci changes related to CI common dependencies Pull requests that update a dependency file docs Changes related to documentations engine infra platform Non-amber Scala service paths

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Observability] Deploy OpenTelemetry Collector and Parca/eBPF profiling backends

3 participants