otel-collector
Catalog date: 2026-06-24.
The single OTLP ingress for the cluster. Every workload pod sends traces, metrics, and logs to this gateway; the gateway fans out by signal type to the appropriate backend.
Identity
| Field | Value |
|---|---|
| Service id | otel-collector |
| Role | observability |
| Image | otel/opentelemetry-collector-contrib:0.120.0 — pinned in both the collector-gateway.yaml (gateway, Deployment mode) and collector-agent.yaml (agent, mode: daemonset) OpenTelemetryCollector CRs; correcting "gateway flavour" of the base (non-contrib) image, which does not match either manifest |
| Port | 4317 (OTLP gRPC) + 4318 (OTLP HTTP) |
| Health | no health_check extension or :13133 port was found in either checked-in manifest; the OTel Operator may inject a default one at admission — flagging as unverified rather than restating a specific port |
Deployment surfaces
| Surface | Where |
|---|---|
| Compose | service otel-collector in alphaswarm_platform/compose/docker-compose.yml |
| Kustomize | observability/opentelemetry-collector-gateway/ — gateway Deployment + DaemonSet agent (canonical) |
| Operator | observability/opentelemetry-operator/ — auto-instrumentation CRDs |
| Legacy | observability/otel-collector/ — rollback only; NOT wired to overlays |
Routing
| Signal | Destination |
|---|---|
traces.infrastructure | Jaeger — per a comment in grafana-datasources.yaml, "Tempo was referenced across the observability stack but never deployed," so the earlier "Jaeger (in-cell) / Tempo (cloud cells)" split is corrected to Jaeger-only |
traces.ai (OpenInference spans) | phoenix |
metrics | the traces/metrics/logs pipeline in collector-gateway.yaml exports metrics via prometheusremotewrite to the in-cluster Prometheus (prometheus-operated:9090/api/v1/write) only; VictoriaMetrics receives metrics separately, via vmagent's own Prometheus-style scrape configs (including scraping the otel-gateway's own :8888 self-metrics) rather than as an otel-collector export target — correcting the earlier "VictoriaMetrics + Prometheus" routing claim |
logs | routed to otlp/data-prepper (OpenSearch, primary) and otlphttp/loki (secondary/compat path) per collector-gateway.yaml's own pipeline comment |
Correction: no connectors: block or routing-type connector exists
in collector-gateway.yaml.
The single traces pipeline runs a tail_sampling processor (with
ai_traces / llm_traces / agent_traces string-attribute policies
plus an errors policy and a 1% probabilistic_baseline) and then
sends every span retained by that processor to both the
otlp/jaeger and otlp/phoenix exporters — there is no
attribute-based split routing infra spans to Jaeger-only and AI spans
to Phoenix-only.
Dependencies
Upstream: every alphaswarm workload pod (auto-instrumentation
through the OTel operator + manual SDK init in alphaswarm/observability/).
Downstream: Jaeger, Phoenix, Prometheus, VictoriaMetrics, Loki.
Operations
- Sampling: tail-based for traces — keep 100% of error spans, 5% of healthy traffic. Tuned per cell.
- Resource tagging: every span carries
tenant_id,cell_id,service.id(matching topology), andexperiment_id/test_idwhen set. - Auto-instrumentation: Python via
opentelemetry-distro; Node via the OTel operator's auto-injected sidecar; Go services use manual SDK.
See also
observability.md— observability concept doc.observability-stack.md— stack composition + dashboards.phoenix— AI / LLM observability upstream.