已合并
feat(observability): 在 Grafana 可视化展示 metrics 指标以及 ms_service_metric 指标 #202
feat(observability): 在 Grafana 可视化展示 metrics 指标以及 ms_service_metric 指标 #202
已合并
LinWei100创建于 6月1日
共 34 个文件变更+5865-1
@@ -109,4 +109,9 @@ venv.bak/
109**/*_pb2_grpc.py109**/*_pb2_grpc.py
110 110 
111# deployer111# deployer
112-examples/deployer/output_yamls/*112+deployer/output/*
113+ 
114+# observability stack runtime artifacts
115+examples/features/observability/stack/generated/
116+examples/features/observability/stack/.native-runtime/
117+examples/features/observability/stack/.env
@@ -0,0 +1,64 @@
1+# pymotor observability stack — environment overrides
2+#
3+# Copy to `.env` and edit to taste. All values shown are defaults.
4+ 
5+# -------- Image registry --------
6+# Prefix prepended to every image name. Leave empty to pull from Docker Hub.
7+# Example for internal Harbor: REGISTRY_PREFIX=harbor.example.com/library/
8+REGISTRY_PREFIX=
9+ 
10+# -------- Image versions --------
11+GRAFANA_VERSION=11.3.0
12+PROMETHEUS_VERSION=v2.55.1
13+TEMPO_VERSION=2.6.1
14+LOKI_VERSION=3.3.0
15+OTEL_COLLECTOR_VERSION=0.115.1
16+NODE_EXPORTER_VERSION=v1.8.2
17+CADVISOR_VERSION=v0.49.1
18+NPU_EXPORTER_IMAGE=swr.cn-south-1.myhuaweicloud.com/ascendhub/npu-exporter:v6.0.0
19+ 
20+# -------- Grafana admin --------
21+GF_SECURITY_ADMIN_USER=motor
22+GF_SECURITY_ADMIN_PASSWORD=motor
23+ 
24+# -------- Discovery / launch (used by launch.sh) --------
25+# Kubernetes namespace / job_id
26+MOTOR_NAMESPACE=
27+# Optional NodePort access IP override
28+MOTOR_NODE_IP=
29+# Optional pyMotor user_config.json path
30+MOTOR_USER_CONFIG=
31+# Engine /metrics management port
32+MOTOR_ENGINE_MGMT_PORT=10001
33+# pyMotor 上报 tracing 使用的观测主机
34+OBS_HOST=
35+# Docker stack mode: minimal or full
36+OBS_STACK_MODE=full
37+# First host port used when Docker needs PodIP bridge forwards
38+MOTOR_PORT_FORWARD_BASE=19000
39+# -------- Proxy(可选)--------
40+# PROXY_SH:仅 native runtime 从 GitHub/Grafana CDN 下载 Prometheus/Grafana/Tempo 等二进制时使用。
41+# - 留空(默认):不加载任何代理文件。
42+# - 填写路径:指向本机 dotenv 文件(每行 KEY=VALUE,例如 http_proxy=、HTTPS_PROXY=),
43+# 不是 Docker 拉镜像用的配置;Docker 请在 launch 前对当前 shell source 代理或 export HTTP_PROXY。
44+# - 示例文件内容见 SERVICE_GUIDE.md §2.4.2
45+# PROXY_SH=/home/you/pymotor-proxy.env
46+PROXY_SH=
47+ 
48+# -------- Service ports (host-side) --------
49+GRAFANA_PORT=3000
50+PROMETHEUS_PORT=9090
51+TEMPO_QUERY_PORT=3200
52+LOKI_PORT=3100
53+OTEL_GRPC_PORT=4317
54+OTEL_HTTP_PORT=4318
55+NODE_EXPORTER_PORT=9100
56+CADVISOR_PORT=8088
57+NPU_EXPORTER_PORT=8082
58+ 
59+# -------- Prometheus config --------
60+# launch.sh 会自动生成 ./generated/prometheus.yml 并在启动时注入。
61+# start.sh 直接启动时默认使用模板文件。
62+PROMETHEUS_CONFIG_FILE=./prometheus/prometheus.yml
63+OTEL_CONFIG_FILE=./otel-collector/otel-collector.yaml
64+GRAFANA_PROVISIONING_DIR=./grafana/provisioning
@@ -0,0 +1,229 @@
1+# pyMotor 可观测性栈 · Grafana 使用指导
2+ 
3+本指导说明栈内 Grafana 的页面设计、内置看板与数据源,并提供「如何在看板中新增其他 metrics 指标」的可复现步骤。
4+ 
5+> 服务拉起 / 停止操作见 [SERVICE_GUIDE.md](SERVICE_GUIDE.md)。
6+ 
7+前提:已按 [SERVICE_GUIDE.md](SERVICE_GUIDE.md) 拉起栈,且 `http://localhost:3000` 可访问。
8+ 
9+---
10+ 
11+## 1. 登录与访问
12+ 
13+| 项 | 默认值 | 说明 |
14+|----|--------|------|
15+| 地址 | `http://localhost:3000` | 端口由 `.env` 的 `GRAFANA_PORT` 控制 |
16+| 账号 | `motor` | `.env` 的 `GF_SECURITY_ADMIN_USER` |
17+| 密码 | `motor` | `.env` 的 `GF_SECURITY_ADMIN_PASSWORD` |
18+ 
19+登录后进入 **Dashboards** 即可看到下文的三个内置看板。
20+ 
21+---
22+ 
23+## 2. 页面设计总览
24+ 
25+### 2.1 数据源(Datasources)
26+ 
27+数据源由 `grafana/provisioning/datasources/datasources.yml` 预置,无需手动添加:
28+ 
29+| 数据源 | UID | 地址 | 用途 |
30+|--------|-----|------|------|
31+| **Prometheus** | `prometheus` | `http://prometheus:9090` | 指标查询(默认数据源) |
32+| **Tempo** | `tempo` | `http://tempo:3200` | 分布式追踪(Trace) |
33+| **Loki** | `loki` | `http://loki:3100` | 日志(仅 full 模式拉起) |
34+ 
35+并已配置三者间的联动跳转:
36+ 
37+- **Trace → Log**(Tempo `tracesToLogsV2`):按 `service.name` / `x_request_id` 关联,时间窗口前后 5 分钟。
38+- **Trace → Metrics**(Tempo `tracesToMetrics`):按 `service.name` 关联 Prometheus。
39+- **Log → Trace**(Loki `derivedFields`):从日志中正则提取 `trace_id` / `x_request_id`,一键跳转 Tempo。
40+ 
41+### 2.2 看板(Dashboards)
42+ 
43+看板 JSON 位于 `grafana/dashboards/`,由 `grafana/provisioning/dashboards/dashboard-providers.yml` 自动加载(平铺、无文件夹层级,`foldersFromFilesStructure: false`,`updateIntervalSeconds: 30`)。当前仅保留三个:
44+ 
45+| 看板 | UID | 文件 | 内容 |
46+|------|-----|------|------|
47+| **pyMotor Metrics · 指标总览** | `motor-all-metrics` | `motor-all-metrics.json` | 集群概览、PD Role / Instance 分组、吞吐与延迟 |
48+| **KV 缓存** | `motor-kv-cache` | `motor-kv-cache.json` | vLLM KV cache 使用率、prefix cache 命中率 |
49+| **引擎性能剖析** | `motor-vllm-profiling` | `motor-vllm-profiling.json` | `vllm_profiling_*` 性能剖析(显存、forward/execute/scheduler 时延等) |
50+ 
51+> Grafana 容器以 **只读** 方式挂载 `grafana/dashboards`,因此在 UI 上的临时修改不会落盘;要长期保留改动需写回对应 JSON 文件(见第 4 节)。
52+ 
53+### 2.3 看板变量(Template Variables)
54+ 
55+以「指标总览」为例,顶部变量用于跨集群 / 角色 / 实例过滤,均为 `query` 类型并基于 Prometheus 标签动态生成:
56+ 
57+| 变量 | Label |
58+|------|-------|
59+| $cluster | Cluster |
60+| $motor_metric_scope | Metric Scope |
61+| $role | Role |
62+| $pd_role | PD Role |
63+| $dp_rank | DP Rank |
64+| $pod_ip | Pod IP |
65+| $instance_id | Instance ID |
66+| $model_name | Model |
67+ 
68+各变量在 Grafana 中的取值查询(`label_values`):
69+ 
70+- **$cluster**:`label_values({cluster!=""}, cluster)`
71+- **$motor_metric_scope**:`label_values({cluster=~"$cluster", motor_metric_scope!=""}, motor_metric_scope)`
72+- **$role**:`label_values({cluster=~"$cluster", role!=""}, role)`
73+- **$pd_role**:`label_values({cluster=~"$cluster", pd_role!=""}, pd_role)`
74+- **$dp_rank**:`label_values({cluster=~"$cluster", dp_rank!=""}, dp_rank)`
75+- **$pod_ip**:`label_values({cluster=~"$cluster", pod_ip!=""}, pod_ip)`
76+- **$instance_id**:`label_values({cluster=~"$cluster", instance_id!=""}, instance_id)`
77+- **$model_name**:`label_values(vllm:num_requests_running{cluster=~"$cluster"}, model_name)`
78+ 
79+新增面板时**应复用这些变量**做标签过滤,并将变量的 `allValue` 设为 `.*`,避免标签缺失导致 No Data。「引擎性能剖析」看板另有 `$source`、`$job`、`$phase`、`$dp` 等变量,含义类似。
80+ 
81+---
82+ 
83+## 3. 验证指标是否已被采集
84+ 
85+新增看板指标前,先确认 Prometheus 已抓到目标指标,避免在 Grafana 侧反复调试。
86+ 
87+```bash
88+# 列出所有指标名(确认指标存在)
89+curl -s http://localhost:9090/api/v1/label/__name__/values | tr ',' '\n' | grep -i <keyword>
90+ 
91+# 直接查询某指标当前值
92+curl -sG http://localhost:9090/api/v1/query --data-urlencode 'query=<metric_name>'
93+ 
94+# 查看抓取目标是否 UP
95+curl -s http://localhost:9090/api/v1/targets
96+```
97+ 
98+也可在 Grafana 左侧 **Explore** 选择 Prometheus 数据源,直接输入 PromQL 验证表达式。
99+ 
100+---
101+ 
102+## 4. 在看板中新增其他 metrics 指标
103+ 
104+有两种方式:**UI 编辑后写回 JSON**(推荐,可纳入版本库)或 **直接编辑 JSON 文件**。
105+ 
106+### 4.1 方式一:UI 编辑面板并写回 JSON
107+ 
108+1. 打开目标看板(如「指标总览」),点击右上角 **Edit**。
109+2. 点击 **Add → Visualization** 新增面板,选择数据源 **Prometheus**。
110+3. 在 **Query** 中输入 PromQL,复用看板变量做过滤,例如新增「每实例的等待请求数」:
111+ 
112+ ```promql
113+ sum by (instance_id) (
114+ vllm:num_requests_waiting{cluster=~"$cluster", instance_id=~"$instance_id", pd_role=~"$pd_role"}
115+ )
116+ ```
117+ 
118+4. 选择可视化类型(Time series / Stat / Bar gauge / Pie chart 等),设置标题、单位、阈值。
119+5. **Apply** 返回看板,调整面板位置与大小。
120+6. 写回源文件以长期保留:点击看板设置(齿轮)→ **JSON Model**,复制完整 JSON,覆盖写入对应文件,例如:
121+ - 指标总览 → `grafana/dashboards/motor-all-metrics.json`
122+ - KV 缓存 → `grafana/dashboards/motor-kv-cache.json`
123+ - 引擎性能剖析 → `grafana/dashboards/motor-vllm-profiling.json`
124+ 
125+ provisioner 每 30s 重新加载挂载目录,刷新页面即可看到生效(容器以只读挂载,必须写回文件才会持久化)。
126+ 
127+### 4.2 方式二:直接编辑看板 JSON 文件
128+ 
129+在 `grafana/dashboards/<dashboard>.json` 的 `panels` 数组中追加一个面板对象。可参考「指标总览」中现有 `stat` 面板的最小结构:
130+ 
131+```json
132+{
133+ "id": 100,
134+ "type": "timeseries",
135+ "title": "Waiting Requests by instance_id",
136+ "gridPos": { "h": 8, "w": 12, "x": 0, "y": 56 },
137+ "datasource": { "type": "prometheus", "uid": "prometheus" },
138+ "targets": [
139+ {
140+ "expr": "sum by (instance_id) (vllm:num_requests_waiting{cluster=~\"$cluster\", instance_id=~\"$instance_id\"})",
141+ "refId": "A",
142+ "legendFormat": "{{instance_id}}"
143+ }
144+ ],
145+ "fieldConfig": { "defaults": { "unit": "short" } }
146+}
147+```
148+ 
149+要点:
150+ 
151+- `id` 在同一看板内唯一;`gridPos` 的 `x/y/w/h` 决定布局(看板宽度 24 格,`x` 取 0–23)。
152+- `datasource.uid` 固定为 `prometheus`(或 `tempo` / `loki`)。
153+- `targets[].expr` 为 PromQL,**务必带上看板变量过滤**(`{cluster=~"$cluster", ...}`),否则切换变量时该面板不会随动。
154+- `legendFormat` 用 `{{label}}` 渲染图例。
155+ 
156+保存文件后,provisioner 自动重载(约 30s),刷新浏览器即可。若 JSON 改动较大,可执行 `docker compose restart grafana` 强制重载。
157+ 
158+### 4.3 新增需要先接入新数据源的指标
159+ 
160+如果指标来自尚未被抓取的新组件:
161+ 
162+1. 在生成 / 模板 `prometheus.yml` 中新增对应 scrape job(自动发现产物为 `generated/prometheus.yml`)。
163+2. 重新执行 `./launch.sh`(或 `curl -X POST http://localhost:9090/-/reload` 触发 Prometheus 热加载,需启用 lifecycle,本栈已开启 `--web.enable-lifecycle`)。
164+3. 确认 `targets` 为 UP、指标可查询后,再按 4.1 / 4.2 添加面板。
165+ 
166+### 4.4 性能剖析看板的自动生成(可选)
167+ 
168+`grafana/scripts/build-profiling-dashboard.py` 可从 Prometheus 拉取 `vllm_profiling_*` 指标族并自动生成 `motor-vllm-profiling.json`(核心面板常开、明细面板默认折叠):
169+ 
170+```bash
171+cd examples/features/observability/stack
172+python3 grafana/scripts/build-profiling-dashboard.py \
173+ --prometheus-url http://localhost:9090 \
174+ --output grafana/dashboards/motor-vllm-profiling.json
175+```
176+ 
177+适用于 Engine 暴露了新的 `vllm_profiling_*` 指标后,批量刷新剖析看板。
178+ 
179+---
180+ 
181+## 5. Trace 与 Profiling 数据接入(pyMotor 侧)
182+ 
183+要让 Tempo / profiling 面板有数据,需在 pyMotor 侧开启上报;**需要在拉起栈之前完成的 pyMotor 配置清单见 [SERVICE_GUIDE.md §1.4](SERVICE_GUIDE.md)**。本节为操作要点速查。
184+ 
185+> 基础指标(指标总览 / KV 缓存)无需改 pyMotor 配置即可生效;当前方案不使用 Controller metrics 接口,相关配置可忽略。
186+ 
187+### 5.1 Tracing
188+ 
189+pyMotor 建议配置(`<obs-host>` 为观测主机,参考 `config/tracing.example.json`):
190+ 
191+- Coordinator:`tracer_config.endpoint = http://<obs-host>:4318/v1/traces`
192+- Engine:`engine_config.otlp-traces-endpoint = http://<obs-host>:4318/v1/traces`
193+- `OTEL_SERVICE_NAME` 建议:`mindie-motor-coordinator`、`vllm-server-p`、`vllm-server-d`
194+ 
195+在 Grafana **Explore** 选 Tempo 时,若默认 Query type 为 TraceQL,可切到 **Search**,或在 TraceQL 输入 `{}` 后执行搜索。
196+ 
197+### 5.2 Profiling
198+ 
199+需在 Engine 侧安装 [`ms_service_metric`](https://gitcode.com/Ascend/msserviceprofiler/tree/master/ms_service_metric)(`pip install ms_service_metric`;依赖 Python >= 3.10、pyyaml、prometheus-client、posix_ipc)。详细步骤见 [SERVICE_GUIDE.md §1.4.2](SERVICE_GUIDE.md)。
200+ 
201+Engine 启动前:
202+ 
203+```bash
204+export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p "$PROMETHEUS_MULTIPROC_DIR"
205+# 可选:rm -rf $PROMETHEUS_MULTIPROC_DIR/*
206+```
207+ 
208+Engine ready 后开启指标采集:
209+ 
210+```bash
211+ms-service-metric on # 开启
212+ms-service-metric off # 关闭
213+ms-service-metric restart # 重启(重新加载配置)
214+ms-service-metric status # 查看状态
215+```
216+ 
217+随后 `vllm_profiling_*` 指标会被 Prometheus 抓取,「引擎性能剖析」看板即可显示数据。
218+ 
219+---
220+ 
221+## 6. 常见问题
222+ 
223+| 现象 | 处理建议 |
224+|------|----------|
225+| 看板变量下拉为空 | 对应标签未被任何指标暴露;先确认 Prometheus 已抓到带该标签的指标(第 3 节)。 |
226+| 新增面板 No Data | 检查 PromQL 是否带了看板变量过滤;变量 `allValue` 是否为 `.*`;指标名是否含冒号(如 `vllm:*`,本栈已开启 `--enable-feature=utf8-names`)。 |
227+| UI 改动刷新后丢失 | 容器只读挂载 `grafana/dashboards`,需将 JSON Model 写回源文件(第 4 节)。 |
228+| 看板报 500 / 504 | Grafana 容器经外网代理访问 `prometheus` / `tempo` 超时;确认容器内 `HTTP_PROXY` 为空(详见 [SERVICE_GUIDE.md](SERVICE_GUIDE.md) 第 6 节)。 |
229+| Loki 数据源不可用 | Loki 仅在 full 模式拉起,minimal 模式无 Loki。 |
@@ -0,0 +1,23 @@
1+# pyMotor 可观测性一键栈
2+ 
3+本目录提供 Prometheus、Grafana、Tempo、OTel Collector 等组件的一键发现与拉起能力。
4+ 
5+## 文档导航
6+ 
7+| 文档 | 说明 |
8+|------|------|
9+| [SERVICE_GUIDE.md](SERVICE_GUIDE.md) | 前提条件、镜像准备、`launch.sh` 拉起与停止、常见问题 |
10+| [GRAFANA_GUIDE.md](GRAFANA_GUIDE.md) | Grafana 页面、数据源、看板设计与新增指标步骤 |
11+ 
12+## 快速开始
13+ 
14+```bash
15+cd examples/features/observability/stack
16+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal
17+```
18+ 
19+Grafana 默认:<http://localhost:3000>(`motor` / `motor`)。
20+ 
21+## 代理 / 内网拉镜像
22+ 
23+内网需代理访问外网时,请区分 **Docker 拉镜像**(当前 shell 的 `HTTP_PROXY`)与 **Native 下载二进制**(`.env` 中的 `PROXY_SH`);发现阶段建议关闭代理。完整说明见 [SERVICE_GUIDE.md §2.4](SERVICE_GUIDE.md#24-代理配置)。
@@ -0,0 +1,459 @@
1+# pyMotor 可观测性栈 · 服务拉起与停止指导
2+ 
3+本指导面向需要在已部署 pyMotor 的节点上拉起 / 停止可观测性栈(Prometheus + Grafana + Tempo + OTel Collector + Loki)的使用者,提供可逐步复现的完整操作步骤。
4+ 
5+> 配套文档:Grafana 页面设计与看板指标扩展见 [GRAFANA_GUIDE.md](GRAFANA_GUIDE.md)。
6+ 
7+整体流程:
8+ 
9+```text
10+前提条件检查 → 准备镜像(联网拉取) → 拉起服务(launch.sh) → 验收 → 停止服务(stop.sh)
11+```
12+ 
13+---
14+ 
15+## 1. 前提条件
16+ 
17+### 1.1 运行环境
18+ 
19+| 项 | 要求 |
20+|----|------|
21+| 工作目录 | 进入仓库内 `examples/features/observability/stack` |
22+| Python | 已安装 `python3`(用于运行 `scripts/discover-targets.py`) |
23+| Kubernetes | 能访问目标集群 API,`kubectl get pods -n <namespace>` 可正常返回 |
24+| Docker(推荐) | 安装 Docker,且支持 Docker Compose **v2**(`docker compose version` 可用);无 Docker 时可用 `--native` 走原生二进制 |
25+| 网络 | 观测机到 pyMotor Coordinator / Engine 的 **NodePort**,或经主机端口转发的 **PodIP** 可达 |
26+ 
27+### 1.2 业务侧就绪
28+ 
29+- 目标 **namespace** 内 Coordinator、Engine(含 `vllm-p0` / `vllm-d0` 等命名)Pod 已处于 **Running**。
30+- 切换 `mindie-*` 等不同环境时,先执行 `./stop.sh` 再重新 `./launch.sh`,避免复用其他 namespace 的旧 `generated/discovered.env`。
31+- 可选:准备 pyMotor 的 `user_config.json` 路径,用于从 `motor_deploy_config.job_id` 推断 namespace。
32+ 
33+### 1.3 配置文件
34+ 
35+```bash
36+cd examples/features/observability/stack
37+cp -n .env.example .env # launch.sh 在无 .env 时也会自动从 .env.example 复制
38+```
39+ 
40+按需编辑 `.env`(所有值均有默认,见 `.env.example`):
41+ 
42+| 变量 | 说明 |
43+|------|------|
44+| `REGISTRY_PREFIX` | 镜像前缀,与内网 Harbor 一致;留空则从 Docker Hub 拉取 |
45+| `GRAFANA_VERSION` / `PROMETHEUS_VERSION` / `TEMPO_VERSION` / `OTEL_COLLECTOR_VERSION` / `LOKI_VERSION` | 各组件镜像版本 |
46+| `OBS_STACK_MODE` | 栈模式,默认 `full`;也可在命令行用 `--minimal` / `--full` 覆盖 |
47+| `GF_SECURITY_ADMIN_USER` / `GF_SECURITY_ADMIN_PASSWORD` | Grafana 管理员账号 / 密码(默认 `motor` / `motor`) |
48+| `GRAFANA_PORT` / `PROMETHEUS_PORT` / `TEMPO_QUERY_PORT` / `OTEL_GRPC_PORT` / `OTEL_HTTP_PORT` / `LOKI_PORT` | 主机侧服务端口 |
49+| `MOTOR_PORT_FORWARD_BASE` | Docker 需要 PodIP 桥接转发时使用的起始主机端口(默认 `19000`) |
50+| `PROXY_SH` | **可选**。Native runtime 从 GitHub / Grafana CDN 下载二进制时使用的代理配置文件路径;留空则不加载(见 [§2.4 代理配置](#24-代理配置)) |
51+ 
52+### 1.4 需要调整 pyMotor 配置才能生效的能力(重要,请提前配置)
53+ 
54+部分观测能力需要在 **pyMotor 侧**(`env.json` / `user_config.json`,或引擎运行环境)提前配置,否则观测栈拉起后对应看板会无数据。请在拉起栈**之前**对照下表完成配置:
55+ 
56+| 观测能力 | 是否需改 pyMotor 配置 | 需要的配置 |
57+|----------|----------------------|-----------|
58+| Coordinator 基础指标(指标总览 / KV 缓存的请求数、KV、吞吐、延迟等) | **否** | Coordinator 默认在管理端口暴露 `/metrics`、`/instance/metrics`,无需额外配置;只需保证该端口可被观测机或主机端口转发访问 |
59+| Engine / vLLM 指标 | **否**(默认开启) | Engine 在管理端口(默认 `10001`)暴露 `/metrics`;保证端口可达即可 |
60+| Tracing(Tempo 链路) | **是** | 见下方「1.4.1 Tracing 接入」 |
61+| 引擎性能剖析(`vllm_profiling_*`) | **是** | 需安装并开启 `ms_service_metric`,见下方「1.4.2 Profiling 接入」 |
62+ 
63+> 说明:当前方案**不使用 Controller 的 metrics 接口**,因此 Controller observability 相关配置(如 `observability_enable`、`1027` 端口)无需调整,可忽略。
64+ 
65+#### 1.4.1 Tracing 接入(让 Tempo 链路看板有数据)
66+ 
67+需修改 deploy 使用的 `env.json` 与 `user_config.json`(`<obs-host>` 为观测栈所在主机 IP),随后用 `deploy.py` 重新部署生效:
68+ 
69+- `env.json`:在 `motor_coordinator_env` / `motor_engine_prefill_env` / `motor_engine_decode_env` 下新增:
70+ - `OTEL_SERVICE_NAME`(建议 `mindie-motor-coordinator` / `vllm-server-p` / `vllm-server-d`)
71+ - `OTEL_EXPORTER_OTLP_TRACES_PROTOCOL=http/protobuf`
72+ - `OTEL_EXPORTER_OTLP_TRACES_INSECURE=true`
73+- `user_config.json`:
74+ - `motor_coordinator_config.tracer_config.endpoint = http://<obs-host>:4318/v1/traces`
75+ - `motor_engine_prefill_config.engine_config.otlp-traces-endpoint = http://<obs-host>:4318/v1/traces`
76+ - `motor_engine_decode_config.engine_config.otlp-traces-endpoint = http://<obs-host>:4318/v1/traces`
77+ 
78+> `4318` 为栈内 OTel Collector 的 OTLP HTTP 端口。完整片段见 `config/tracing.example.json`,详细部署见 `docs/zh/user_guide/tracing_deployment.md`。
79+ 
80+#### 1.4.2 Profiling 接入(让「引擎性能剖析」看板有数据)
81+ 
82+「引擎性能剖析」看板依赖 `ms_service_metric` 暴露的 `vllm_profiling_*` 指标,需在 **Engine 侧** 安装并开启采集。完整说明见上游文档:[ms_service_metric · Ascend/msserviceprofiler](https://gitcode.com/Ascend/msserviceprofiler/tree/master/ms_service_metric)。
83+ 
84+**安装**
85+ 
86+```bash
87+pip install ms_service_metric
88+```
89+ 
90+**依赖**
91+ 
92+- Python >= 3.10
93+- pyyaml
94+- prometheus-client
95+- posix_ipc(Linux 平台)
96+ 
97+**快速开始**
98+ 
99+1. **vLLM 集成**
100+ 
101+ vLLM 通过 `entry_points` 机制自动适配,无需额外代码:
102+ 
103+ - 安装 `ms_service_metric`
104+ - Engine 启动**前**设置多进程 metric 采集环境变量:
105+ 
106+ ```bash
107+ # 开启 vLLM 多进程 metric 采集环境变量
108+ export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p "$PROMETHEUS_MULTIPROC_DIR"
109+ 
110+ # 可选,清理上次的指标文件
111+ # rm -rf $PROMETHEUS_MULTIPROC_DIR/*
112+ 
113+ # 启动 vLLM / Engine
114+ # vllm serve --model your_model
115+ ```
116+ 
117+2. **控制指标采集**(Engine ready **后**执行)
118+ 
119+ ```bash
120+ # 开启指标采集
121+ ms-service-metric on
122+ 
123+ # 关闭指标采集
124+ ms-service-metric off
125+ 
126+ # 重启(重新加载配置)
127+ ms-service-metric restart
128+ 
129+ # 查看状态
130+ ms-service-metric status
131+ ```
132+ 
133+未执行上述步骤时,`vllm_profiling_*` 指标不会产生,「引擎性能剖析」看板将无数据。
134+ 
135+---
136+ 
137+## 2. 准备镜像(联网拉取)
138+ 
139+拉起过程会启动多个容器镜像。**首次本地无镜像时**,`start.sh` 默认执行:
140+ 
141+```text
142+docker compose up -d --pull missing --no-build
143+```
144+ 
145+含义:**本地已有镜像则不拉取,缺失时才 `docker pull`**。可通过环境变量覆盖:
146+ 
147+| 变量 | 取值 | 含义 |
148+|------|------|------|
149+| `OBS_COMPOSE_PULL` | `missing`(默认) | 缺镜像才拉取 |
150+| | `never` | 禁止拉取(离线 / 镜像已齐全) |
151+| | `always` | 每次启动都尝试拉取 |
152+| `OBS_COMPOSE_BUILD` | `0`(默认) | 不本地 build Grafana |
153+| | `1` | 允许 `compose up --build` |
154+ 
155+### 2.1 需要拉取的核心镜像(minimal 模式)
156+ 
157+| 镜像(默认 tag,见 `.env.example`) | 用途 |
158+|-----------------------------------|------|
159+| `grafana/grafana:11.3.0` | Grafana 看板 |
160+| `prom/prometheus:v2.55.1` | 指标存储与查询 |
161+| `grafana/tempo:2.6.1` | Trace 存储 |
162+| `otel/opentelemetry-collector-contrib:0.115.1` | OTLP 接入 |
163+ 
164+**full** 模式额外拉起(默认 `full`,对应 Compose `--profile full`):`grafana/loki`、`prom/node-exporter`、`gcr.io/cadvisor/cadvisor`;可选 `--profile npu` 使用 Ascend `npu-exporter` 镜像。
165+ 
166+### 2.2 代理分工(需区分阶段)
167+ 
168+内网需 HTTP 代理才能访问外网 registry 时,**不要**让 kubectl 发现、镜像拉取、容器内访问栈内服务共用同一套代理策略:
169+ 
170+| 阶段 | 主机 `HTTP_PROXY` | 说明 |
171+|------|-------------------|------|
172+| `kubectl` / `discover-targets.py` | **建议关闭** | 发现脚本内 `_kubectl_env()` 会剔除代理,避免 API Server 经代理超时;也可先 `unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY` |
173+| `docker pull` / `compose pull` / `up --pull missing` | **需要时开启** | 拉镜像时 Docker 客户端继承**当前 shell** 代理 |
174+| Grafana / Prometheus 等容器内 | **已禁用** | Compose 已为 Grafana 清空 `HTTP_PROXY`,`NO_PROXY` 含 `prometheus,tempo`,访问栈内数据源不走外网代理 |
175+ 
176+### 2.3 推荐拉取流程(代理环境 · 首次拉起)
177+ 
178+```bash
179+cd examples/features/observability/stack
180+ 
181+# ① 需要拉镜像时:在**当前 shell** 开启代理(见 §2.4;与 PROXY_SH 无关)
182+source /path/to/your-proxy.sh # 或手动 export HTTP_PROXY / HTTPS_PROXY
183+docker compose pull # 仅需一次;本地已有镜像可跳过
184+ 
185+# ② 发现与启动:建议关闭代理,避免 kubectl 异常
186+unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY all_proxy ALL_PROXY
187+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal
188+```
189+ 
190+**仅本地已有镜像、禁止任何拉取**(离线场景):
191+ 
192+```bash
193+export OBS_COMPOSE_PULL=never
194+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal
195+```
196+ 
197+### 2.4 代理配置
198+ 
199+内网环境访问外网 registry 或 GitHub 时常需 HTTP/HTTPS 代理。观测栈在**不同阶段**对代理的要求不同,请按场景配置,避免混用导致 kubectl 超时或 Grafana 看板 504。
200+ 
201+#### 2.4.1 三阶段分工(速查)
202+ 
203+| 阶段 | 配置方式 | 是否建议开代理 |
204+|------|----------|----------------|
205+| **目标发现**(`discover-targets.py` / `kubectl`) | 关闭 shell 代理;脚本内已对 kubectl 剔除代理变量 | **否** |
206+| **Docker 拉镜像**(`docker compose pull` / `launch.sh` → `start.sh`) | 在**启动前**对当前 shell `source` 代理脚本或 `export HTTP_PROXY=...` | **需要外网 registry 时是** |
207+| **Native 下载二进制**(`start-native.sh` / `launch.sh --native`) | `.env` 中设置 `PROXY_SH`,或启动前 export 同名环境变量 | **需要访问 GitHub / dl.grafana.com 时是** |
208+| **容器内访问 Prometheus / Tempo** | 无需配置;Compose 已清空 Grafana 的 `HTTP_PROXY` 并设置 `NO_PROXY` | **否**(已内置) |
209+ 
210+#### 2.4.2 `PROXY_SH`(Native runtime 专用)
211+ 
212+`PROXY_SH` 仅用于 **native runtime** 首次下载 Prometheus、Grafana、Tempo、OTel Collector 等二进制(`curl`/`wget` 访问 GitHub、Grafana CDN)。**不会**影响 Docker 镜像拉取,也不会自动作用于 `kubectl`。
213+ 
214+**配置步骤:**
215+ 
216+1. 复制并编辑环境文件:
217+ 
218+ ```bash
219+ cd examples/features/observability/stack
220+ cp -n .env.example .env
221+ ```
222+ 
223+2. 准备代理配置文件(**dotenv 格式**,每行 `KEY=VALUE`,与 `.env` 相同;不要用个人机器上的绝对路径提交到仓库):
224+ 
225+ ```bash
226+ # 示例:~/pymotor-proxy.env(路径自定)
227+ cat > ~/pymotor-proxy.env <<'EOF'
228+ http_proxy=http://proxy.example.com:8080
229+ https_proxy=http://proxy.example.com:8080
230+ HTTP_PROXY=http://proxy.example.com:8080
231+ HTTPS_PROXY=http://proxy.example.com:8080
232+ no_proxy=localhost,127.0.0.1,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16
233+ NO_PROXY=localhost,127.0.0.1,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16
234+ EOF
235+ ```
236+ 
237+3. 在 `.env` 中指向该文件(**留空表示不加载**,为默认值):
238+ 
239+ ```bash
240+ PROXY_SH=/home/you/pymotor-proxy.env
241+ ```
242+ 
243+4. 启动 native 栈:
244+ 
245+ ```bash
246+ unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY # 发现阶段仍建议关代理
247+ MOTOR_NAMESPACE=<namespace> ./launch.sh --native
248+ ```
249+ 
250+启动日志出现 `[native] loaded proxy config: ...` 表示已加载;若路径不存在或 `PROXY_SH` 为空,则跳过(下载失败时需检查网络或补全代理文件)。
251+ 
252+> **注意:** 请勿将他人开发机路径(如 `/mnt/<工号>/proxy.sh`)写入 `.env` 并提交;每台机器应使用本机可访问的代理文件路径,或保持 `PROXY_SH=` 为空。
253+ 
254+#### 2.4.3 Docker 模式拉镜像(shell 代理,非 `PROXY_SH`)
255+ 
256+Docker 客户端继承**当前 shell** 的 `HTTP_PROXY` / `HTTPS_PROXY`,不读取 `PROXY_SH`。推荐在**同一终端**按顺序执行:
257+ 
258+```bash
259+cd examples/features/observability/stack
260+source /path/to/your-proxy.sh # 或 export HTTP_PROXY=... HTTPS_PROXY=...
261+ 
262+# 可选:预拉镜像
263+docker compose --profile full pull
264+ 
265+# 发现前关闭代理,避免 kubectl 走代理
266+unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY all_proxy ALL_PROXY
267+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal
268+```
269+ 
270+也可在已 `source` 代理的 shell 中直接 `./launch.sh`(`start.sh` 在 `--pull missing` 时会用当前 shell 代理拉缺失镜像);若发现阶段报错,请按上表在 `launch` 前 `unset` 代理变量。
271+ 
272+#### 2.4.4 使用公司统一 shell 代理脚本
273+ 
274+若团队提供 `source proxy.sh`(bash `export` 形式),可用于 **Docker 拉镜像**;Native 下载请任选其一:
275+ 
276+- 启动前在同一 shell 执行 `source proxy.sh`,并**不要**设置 `PROXY_SH`(依赖当前环境变量);或
277+- 将相同变量写入 dotenv 文件,仅在 `.env` 中配置 `PROXY_SH=/path/to/pymotor-proxy.env`(推荐,与 `launch.sh` 解耦)。
278+ 
279+#### 2.4.5 常见问题
280+ 
281+| 现象 | 处理 |
282+|------|------|
283+| Native 下载 Prometheus/Grafana 超时 | 检查 `PROXY_SH` 路径、代理是否可达;`cat "$PROXY_SH"` 确认含 `https_proxy` |
284+| `kubectl` / 发现超时 | `unset` 全部代理后再 `./launch.sh`;勿对 API Server 走 HTTP 代理 |
285+| Grafana 看板 500 / 504 | 容器内代理问题,见本文 [§6](#6-常见问题总结) 与 `docker-compose.yml` 中 Grafana 的 `NO_PROXY` |
286+| `.env` 里 `PROXY_SH` 指向不存在文件 | 保持为空即可;错误路径不会加载,但 native 下载可能失败 |
287+ 
288+---
289+ 
290+## 3. 拉起服务(`launch.sh` 统一入口)
291+ 
292+**所有联调与验收请使用 `./launch.sh`**,它会先做目标发现,再启动栈;不要直接跳过发现步骤调用 `start.sh`(除非仅调试 Compose)。
293+ 
294+### 3.1 基本用法
295+ 
296+```bash
297+cd examples/features/observability/stack
298+ 
299+# 最常用:指定 namespace,minimal 栈(联调推荐)
300+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal
301+ 
302+# 完整栈(含 Loki、node-exporter、cAdvisor 等)
303+MOTOR_NAMESPACE=<namespace> ./launch.sh --full
304+ 
305+# 指定 NodePort 访问 IP(可选)
306+MOTOR_NAMESPACE=<namespace> ./launch.sh --minimal --node-ip <node-ip>
307+ 
308+# Docker 不可用 / 镜像拉取失败时,显式走 native runtime
309+MOTOR_NAMESPACE=<namespace> ./launch.sh --native
310+```
311+ 
312+### 3.2 命令行参数说明
313+ 
314+| 参数 | 说明 |
315+|------|------|
316+| `--namespace <ns>` | Kubernetes namespace / job_id,等同环境变量 `MOTOR_NAMESPACE` |
317+| `--node-ip <ip>` | NodePort 访问使用的节点 IP,等同 `MOTOR_NODE_IP` |
318+| `--user-config <path>` | pyMotor `user_config.json` 路径,等同 `MOTOR_USER_CONFIG` |
319+| `--minimal` | 启动 minimal Docker 栈(Prometheus / Grafana / Tempo / OTel) |
320+| `--full` | 启动 full Docker 栈(额外含 Loki / node-exporter / cAdvisor) |
321+| `--discover-only` | 只运行目标发现,写出 `generated/*`,不启动栈 |
322+| `--dry-run` | 发现并打印生成的 `generated/prometheus.yml`(前 240 行),不启动栈 |
323+| `--native` | 跳过 Docker Compose,直接运行原生二进制 runtime |
324+| `-h`, `--help` | 显示帮助 |
325+ 
326+### 3.3 环境变量(可与参数混用,参数优先)
327+ 
328+```bash
329+export MOTOR_NAMESPACE=<namespace> # K8s namespace / job_id
330+export MOTOR_NODE_IP=<node-ip> # NodePort 访问 IP
331+export MOTOR_USER_CONFIG=/path/user_config.json
332+export MOTOR_ENGINE_MGMT_PORT=10001 # Engine /metrics 管理端口,默认 10001
333+export OBS_HOST=<obs-host> # pyMotor 上报 tracing / OTLP 的观测主机
334+export OBS_STACK_MODE=minimal|full # 未传 --minimal/--full 时生效
335+export PROXY_SH=/path/to/pymotor-proxy.env # native runtime 下载二进制(dotenv 格式,可选;见 §2.4)
336+```
337+ 
338+### 3.4 `launch.sh` 模式一览
339+ 
340+| 模式 | 命令 / 参数 | 行为 |
341+|------|-------------|------|
342+| **默认 Docker · full** | `./launch.sh` 或 `./launch.sh --full` | 发现 → `start.sh --full` 启动完整 Compose profile |
343+| **Docker · minimal** | `./launch.sh --minimal` | 发现 → `start.sh --minimal`;生成 minimal provisioning / Prometheus / OTel;启动主机 port-forward helper |
344+| **仅发现** | `./launch.sh --discover-only` | 只运行 `discover-targets.py`,写出 `generated/*`,不启动栈 |
345+| **发现 + 预览配置** | `./launch.sh --dry-run` | 发现并打印 `generated/prometheus.yml`,不启动栈 |
346+| **强制 native** | `./launch.sh --native` | 跳过 Docker,直接 `scripts/start-native.sh`(本地下载二进制运行 Prometheus / Grafana / Tempo 等) |
347+| **Docker 失败回退** | `./launch.sh`(未加 `--native`) | Docker 启动非 0 退出时,**自动** fallback 到 native(日志会提示) |
348+ 
349+### 3.5 内部调用链(便于排障)
350+ 
351+```text
352+launch.sh
353+ ├─ discover-targets.py → generated/prometheus.yml, generated/discovered.env
354+ ├─ [--discover-only / --dry-run] → 结束
355+ ├─ [--native] → scripts/start-native.sh
356+ └─ [默认] start.sh --minimal|--full
357+ ├─ scripts/run-k8s-port-forwards-host.sh(存在 discovered.env 时)
358+ ├─ ensure_compose_images + docker compose up --pull missing
359+ └─ [失败] launch.sh 自动回退 scripts/start-native.sh
360+```
361+ 
362+兼容入口:`./start-real.sh [options]` 等价于 `./launch.sh [options]`,仅做转发。
363+ 
364+### 3.6 自动发现规则(参考)
365+ 
366+- **Namespace**:`--namespace` / `MOTOR_NAMESPACE` → `user_config.json` 的 `motor_deploy_config.job_id` → 扫描含 Coordinator observability NodePort 的 namespace。
367+- **Node IP**:`--node-ip` / `MOTOR_NODE_IP` → Coordinator Pod `hostIP` → Kubernetes Node `InternalIP`。
368+- **Coordinator**:自动发现 observability NodePort(默认服务端口 `1027`),生成 `/metrics` 及 `type=instance|role|dp|node` 等指标端点。
369+- **Engine**:优先 Engine metrics NodePort;无 NodePort 时回退 PodIP + `MOTOR_ENGINE_MGMT_PORT`(默认 `10001`),识别 `vllm-p0` / `vllm-d0` 等命名并推断 `pd_role` 与 `instance_id`。
370+- **Tracing**:写入 `OBS_HOST`、`OTLP_HTTP_ENDPOINT=http://<obs-host>:4318/v1/traces`、`OTLP_GRPC_ENDPOINT=http://<obs-host>:4317`。
371+ 
372+### 3.7 默认端口
373+ 
374+| 组件 | 默认端口 | 用途 |
375+|------|----------|------|
376+| Grafana | `3000` | 浏览器访问看板 |
377+| Prometheus | `9090` | 指标查询与 targets |
378+| Tempo | `3200` | Trace 查询 API |
379+| OTel Collector gRPC | `4317` | OTLP gRPC 上报 |
380+| OTel Collector HTTP | `4318` | OTLP HTTP `/v1/traces` |
381+| Loki(仅 full) | `3100` | 日志数据源 |
382+| Coordinator observability | `1027` | Coordinator typed metrics |
383+| Engine management metrics | `10001` | Engine `/metrics` |
384+ 
385+---
386+ 
387+## 4. 拉起后验收
388+ 
389+### 4.1 健康检查
390+ 
391+```bash
392+curl -s http://localhost:9090/-/healthy # Prometheus
393+curl -s -u motor:motor http://localhost:3000/api/health # Grafana
394+curl -s http://localhost:3200/ready # Tempo
395+curl -s http://localhost:9090/api/v1/targets # 抓取目标
396+curl -sG http://localhost:9090/api/v1/query \
397+ --data-urlencode 'query=count(up{motor_component=~"coordinator|engine"})'
398+```
399+ 
400+### 4.2 Grafana
401+ 
402+- 访问地址:`http://localhost:3000`(默认账号 `motor` / `motor`)。
403+- 看板变量:`source=real`,`cluster=<当前 namespace>`。
404+- Prometheus Targets 中 `motor-coordinator`、`motor-engine` 应为 **UP**。
405+- 页面布局与看板指标扩展详见 [GRAFANA_GUIDE.md](GRAFANA_GUIDE.md)。
406+ 
407+### 4.3 发现产物(排障用)
408+ 
409+| 文件 | 内容 |
410+|------|------|
411+| `generated/discovery-summary.txt` | 发现摘要 |
412+| `generated/discovered.env` | `OBS_HOST`、`PORT_FORWARD_*` 等 |
413+| `generated/prometheus.yml` | 自动生成的 scrape 配置 |
414+ 
415+---
416+ 
417+## 5. 停止服务
418+ 
419+```bash
420+cd examples/features/observability/stack
421+ 
422+# 停止 Docker Compose 与 native runtime,并清理相关主机端口转发
423+./stop.sh
424+ 
425+# 同时清空数据卷 / native 数据(彻底重置)
426+./stop.sh --purge
427+```
428+ 
429+`stop.sh` 行为:
430+ 
431+1. 先执行 `scripts/stop-k8s-port-forwards-host.sh` 清理主机侧 port-forward。
432+2. 若 `docker compose` 可用,执行 `docker compose --profile npu down`(带 `--purge` 时追加 `-v` 删除数据卷)。
433+3. 停止 native runtime 进程(Grafana / Prometheus / OTel / Tempo 的 pid),`--purge` 时删除 `.native-runtime/{data,logs,run}`。
434+ 
435+---
436+ 
437+## 6. 常见问题总结
438+ 
439+| 现象 | 根因 / 处理建议 |
440+|------|----------------|
441+| `kubectl is unavailable` / `kubectl` 超时 | `source proxy` 后 kubectl 经 HTTP 代理访问 API Server 超时。发现阶段 `unset` 代理(脚本内已对 kubectl 清代理),并确认 `MOTOR_NAMESPACE` 正确。 |
442+| `docker pull` / `compose pull` 超时 | 外网 registry 需代理。拉镜像前 `source` 代理脚本;或内网预拉镜像后设 `OBS_COMPOSE_PULL=never`。 |
443+| `docker compose up --build` 失败 | 默认不应 build Grafana。确认未设置 `OBS_COMPOSE_BUILD=1`,使用上游 `grafana/grafana` 镜像。 |
444+| Grafana 看板 500 / 504 | Grafana 容器继承了 `HTTP_PROXY`,访问 `prometheus:9090` / `tempo:3200` 走外网代理超时。确认镜像与 Compose 为最新,容器内 `HTTP_PROXY` 为空(`docker exec pymotor-grafana printenv HTTP_PROXY` 应为空)。 |
445+| Dashboard 无曲线 / No Data | Prometheus target 指向的节点 IP 不可达,或未识别 `vllm-p0`/`vllm-d0` Pod,或 namespace / Pod IP 变化后未重建转发。重新 `./launch.sh` 发现;切换环境先 `./stop.sh`。 |
446+| `motor-coordinator` / `motor-engine` 为 DOWN | 观测机到 Pod 网段不通,或 NodePort 不可达。检查 `generated/discovered.env` 的 `PORT_FORWARD_*` 与主机 `tcp-forward.py` 是否生效。 |
447+| 无 Docker 环境 | 使用 `./launch.sh --native` 走原生二进制 runtime。 |
448+| 误用其他 namespace 的旧发现结果 | 切换环境时先 `./stop.sh` 再 `./launch.sh`,勿复用旧 `generated/discovered.env`。 |
449+ 
450+---
451+ 
452+## 7. 运行产物与提交边界
453+ 
454+以下运行时文件不应进入版本库(已由 `.gitignore` 忽略):
455+ 
456+- `.env`
457+- `.native-runtime/`
458+- `generated/prometheus.yml`、`generated/discovered.env`、`generated/discovery-summary.txt`
459+- 本地下载的二进制、日志、pid、Tempo WAL、Prometheus TSDB、Grafana data
@@ -0,0 +1,47 @@
1+{
2+ "_comment": "Tracing 接入示例片段。复制到 deploy 使用的 env.json / user_config.json,将 <obs-host> 替换为 observability 栈所在主机 IP 或 Service 地址。",
3+ "env_json_snippet": {
4+ "motor_coordinator_env": {
5+ "OTEL_SERVICE_NAME": "mindie-motor-coordinator",
6+ "OTEL_EXPORTER_OTLP_TRACES_PROTOCOL": "http/protobuf",
7+ "OTEL_EXPORTER_OTLP_TRACES_INSECURE": "true"
8+ },
9+ "motor_engine_prefill_env": {
10+ "OTEL_SERVICE_NAME": "vllm-server-p",
11+ "OTEL_EXPORTER_OTLP_TRACES_PROTOCOL": "http/protobuf",
12+ "OTEL_EXPORTER_OTLP_TRACES_INSECURE": "true"
13+ },
14+ "motor_engine_decode_env": {
15+ "OTEL_SERVICE_NAME": "vllm-server-d",
16+ "OTEL_EXPORTER_OTLP_TRACES_PROTOCOL": "http/protobuf",
17+ "OTEL_EXPORTER_OTLP_TRACES_INSECURE": "true"
18+ }
19+ },
20+ "user_config_snippet": {
21+ "motor_coordinator_config": {
22+ "tracer_config": {
23+ "endpoint": "http://<obs-host>:4318/v1/traces",
24+ "root_sampling_rate": 1.0,
25+ "remote_parent_sampled": 1.0,
26+ "remote_parent_not_sampled": 1.0,
27+ "local_parent_sampled": 1.0,
28+ "local_parent_not_sampled": 1.0
29+ }
30+ },
31+ "motor_engine_prefill_config": {
32+ "engine_config": {
33+ "otlp-traces-endpoint": "http://<obs-host>:4318/v1/traces"
34+ }
35+ },
36+ "motor_engine_decode_config": {
37+ "engine_config": {
38+ "otlp-traces-endpoint": "http://<obs-host>:4318/v1/traces"
39+ }
40+ }
41+ },
42+ "recommended_protocol": {
43+ "coordinator_tracer_config_endpoint": "http://<obs-host>:4318/v1/traces",
44+ "engine_otlp_traces_endpoint": "http://<obs-host>:4318/v1/traces",
45+ "env_OTEL_EXPORTER_OTLP_TRACES_PROTOCOL": "http/protobuf"
46+ }
47+}
@@ -0,0 +1,190 @@
1+# pyMotor Observability Stack
2+# ----------------------------
3+# A self-contained Docker Compose deployment that mirrors NVIDIA Dynamo's
4+# local observability stack (Prometheus + Grafana + Tempo + Loki + OTel
5+# Collector + Exporters) for pyMotor.
6+#
7+# Profiles:
8+# - default(no profile flag) : core stack
9+# - npu : enable Ascend npu-exporter (requires Ascend drivers on host)
10+#
11+# Usage:
12+# docker compose up -d
13+# docker compose --profile npu up -d
14+ 
15+name: pymotor-observability
16+ 
17+x-image-prefix: &registry "${REGISTRY_PREFIX:-}"
18+ 
19+networks:
20+ obs:
21+ driver: bridge
22+ 
23+volumes:
24+ prometheus-data:
25+ grafana-data:
26+ tempo-data:
27+ loki-data:
28+ 
29+services:
30+ # ---------------------------------------------------------------
31+ # Core: Prometheus
32+ # ---------------------------------------------------------------
33+ prometheus:
34+ image: ${REGISTRY_PREFIX:-}prom/prometheus:${PROMETHEUS_VERSION:-v2.55.1}
35+ pull_policy: if_not_present
36+ container_name: pymotor-prometheus
37+ restart: unless-stopped
38+ networks: [obs]
39+ ports:
40+ - "${PROMETHEUS_PORT:-9090}:9090"
41+ command:
42+ - --config.file=/etc/prometheus/prometheus.yml
43+ - --storage.tsdb.path=/prometheus
44+ - --storage.tsdb.retention.time=72h
45+ - --web.enable-lifecycle
46+ - --web.enable-remote-write-receiver
47+ # 保留 vllm:* 等带冒号的指标名(与 prometheus.yml 中
48+ # metric_name_validation_scheme: utf8 配套)
49+ - --enable-feature=utf8-names
50+ volumes:
51+ - ${PROMETHEUS_CONFIG_FILE:-./prometheus/prometheus.yml}:/etc/prometheus/prometheus.yml:ro
52+ - prometheus-data:/prometheus
53+ extra_hosts:
54+ - "host.docker.internal:host-gateway"
55+ 
56+ # ---------------------------------------------------------------
57+ # Core: Tempo (traces)
58+ # ---------------------------------------------------------------
59+ tempo:
60+ image: ${REGISTRY_PREFIX:-}grafana/tempo:${TEMPO_VERSION:-2.6.1}
61+ pull_policy: if_not_present
62+ container_name: pymotor-tempo
63+ restart: unless-stopped
64+ networks: [obs]
65+ ports:
66+ - "${TEMPO_QUERY_PORT:-3200}:3200"
67+ command: ["-config.file=/etc/tempo/tempo.yaml"]
68+ volumes:
69+ - ./tempo/tempo.yaml:/etc/tempo/tempo.yaml:ro
70+ - tempo-data:/var/tempo
71+ 
72+ # ---------------------------------------------------------------
73+ # Core: Loki (logs)
74+ # ---------------------------------------------------------------
75+ loki:
76+ image: ${REGISTRY_PREFIX:-}grafana/loki:${LOKI_VERSION:-3.3.0}
77+ container_name: pymotor-loki
78+ restart: unless-stopped
79+ profiles: [full]
80+ networks: [obs]
81+ ports:
82+ - "${LOKI_PORT:-3100}:3100"
83+ command: ["-config.file=/etc/loki/loki.yaml"]
84+ volumes:
85+ - ./loki/loki.yaml:/etc/loki/loki.yaml:ro
86+ - loki-data:/var/loki
87+ 
88+ # ---------------------------------------------------------------
89+ # Core: OpenTelemetry Collector
90+ # ---------------------------------------------------------------
91+ otel-collector:
92+ image: ${REGISTRY_PREFIX:-}otel/opentelemetry-collector-contrib:${OTEL_COLLECTOR_VERSION:-0.115.1}
93+ pull_policy: if_not_present
94+ container_name: pymotor-otel-collector
95+ restart: unless-stopped
96+ networks: [obs]
97+ depends_on: [tempo]
98+ command: ["--config=/etc/otelcol/otel-collector.yaml"]
99+ volumes:
100+ - ${OTEL_CONFIG_FILE:-./otel-collector/otel-collector.yaml}:/etc/otelcol/otel-collector.yaml:ro
101+ ports:
102+ - "${OTEL_GRPC_PORT:-4317}:4317"
103+ - "${OTEL_HTTP_PORT:-4318}:4318"
104+ 
105+ # ---------------------------------------------------------------
106+ # Core: Grafana (visualization)
107+ # ---------------------------------------------------------------
108+ grafana:
109+ # Use upstream image + volume-mounted provisioning (no local build / buildx).
110+ image: ${REGISTRY_PREFIX:-}grafana/grafana:${GRAFANA_VERSION:-11.3.0}
111+ pull_policy: if_not_present
112+ container_name: pymotor-grafana
113+ restart: unless-stopped
114+ networks: [obs]
115+ depends_on: [prometheus, tempo]
116+ environment:
117+ GF_SECURITY_ADMIN_USER: ${GF_SECURITY_ADMIN_USER:-motor}
118+ GF_SECURITY_ADMIN_PASSWORD: ${GF_SECURITY_ADMIN_PASSWORD:-motor}
119+ GF_USERS_ALLOW_SIGN_UP: "false"
120+ GF_LOG_LEVEL: warn
121+ # Host proxy must not apply to in-compose datasources (prometheus/tempo).
122+ HTTP_PROXY: ""
123+ HTTPS_PROXY: ""
124+ http_proxy: ""
125+ https_proxy: ""
126+ NO_PROXY: prometheus,tempo,otel-collector,localhost,127.0.0.1,host.docker.internal,.local,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16
127+ no_proxy: prometheus,tempo,otel-collector,localhost,127.0.0.1,host.docker.internal,.local,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16
128+ ports:
129+ - "${GRAFANA_PORT:-3000}:3000"
130+ volumes:
131+ - grafana-data:/var/lib/grafana
132+ # Dashboards are baked into the image, but mount the source dir read-only
133+ # so the Grafana provisioner picks up live edits during development.
134+ - ./grafana/dashboards:/var/lib/grafana/dashboards:ro
135+ - ${GRAFANA_PROVISIONING_DIR:-./grafana/provisioning}:/etc/grafana/provisioning:ro
136+ 
137+ # ---------------------------------------------------------------
138+ # Infra exporters
139+ # ---------------------------------------------------------------
140+ node-exporter:
141+ image: ${REGISTRY_PREFIX:-}prom/node-exporter:${NODE_EXPORTER_VERSION:-v1.8.2}
142+ container_name: pymotor-node-exporter
143+ restart: unless-stopped
144+ profiles: [full]
145+ networks: [obs]
146+ pid: host
147+ command:
148+ - --path.rootfs=/host
149+ - --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc|var/lib/docker)($$|/)
150+ volumes:
151+ - /:/host:ro,rslave
152+ ports:
153+ - "${NODE_EXPORTER_PORT:-9100}:9100"
154+ 
155+ cadvisor:
156+ image: ${REGISTRY_PREFIX:-}gcr.io/cadvisor/cadvisor:${CADVISOR_VERSION:-v0.49.1}
157+ container_name: pymotor-cadvisor
158+ restart: unless-stopped
159+ profiles: [full]
160+ networks: [obs]
161+ privileged: true
162+ devices:
163+ - /dev/kmsg
164+ volumes:
165+ - /:/rootfs:ro
166+ - /var/run:/var/run:ro
167+ - /sys:/sys:ro
168+ - /var/lib/docker/:/var/lib/docker:ro
169+ ports:
170+ - "${CADVISOR_PORT:-8088}:8080"
171+ # ---------------------------------------------------------------
172+ # Ascend NPU exporter (profile: npu) — requires host drivers
173+ # ---------------------------------------------------------------
174+ ascend-npu-exporter:
175+ image: ${NPU_EXPORTER_IMAGE:-swr.cn-south-1.myhuaweicloud.com/ascendhub/npu-exporter:v6.0.0}
176+ container_name: pymotor-npu-exporter
177+ restart: unless-stopped
178+ profiles: [npu]
179+ # NPU exporter needs host network + privileged + DCMI / driver mounts.
180+ network_mode: host
181+ privileged: true
182+ command:
183+ - --listen=0.0.0.0:${NPU_EXPORTER_PORT:-8082}
184+ - --updateTime=5
185+ volumes:
186+ - /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro
187+ - /usr/local/dcmi:/usr/local/dcmi:ro
188+ - /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro
189+ - /var/log/Ascend:/var/log/Ascend
190+ - /etc/localtime:/etc/localtime:ro
@@ -0,0 +1,16 @@
1+ARG GRAFANA_VERSION=11.3.0
2+FROM grafana/grafana:${GRAFANA_VERSION}
3+ 
4+LABEL org.opencontainers.image.title="pymotor-grafana" \
5+ org.opencontainers.image.description="Grafana with pre-provisioned datasources & dashboards for pymotor observability stack."
6+ 
7+ENV GF_SECURITY_ADMIN_USER=motor \
8+ GF_SECURITY_ADMIN_PASSWORD=motor \
9+ GF_USERS_ALLOW_SIGN_UP=false \
10+ GF_INSTALL_PLUGINS="" \
11+ GF_LOG_LEVEL=warn
12+ 
13+COPY provisioning/ /etc/grafana/provisioning/
14+COPY dashboards/ /var/lib/grafana/dashboards/
15+ 
16+EXPOSE 3000
@@ -0,0 +1,145 @@
1+{
2+ "uid": "motor-kv-cache",
3+ "title": "KV 缓存",
4+ "tags": [
5+ "pymotor",
6+ "kv-cache"
7+ ],
8+ "timezone": "browser",
9+ "schemaVersion": 39,
10+ "version": 1,
11+ "refresh": "10s",
12+ "time": {
13+ "from": "now-30m",
14+ "to": "now"
15+ },
16+ "editable": true,
17+ "graphTooltip": 1,
18+ "templating": {
19+ "list": [
20+ {
21+ "name": "source",
22+ "label": "Source",
23+ "type": "query",
24+ "datasource": {
25+ "type": "prometheus",
26+ "uid": "prometheus"
27+ },
28+ "query": "label_values({source!=\"\"}, source)",
29+ "refresh": 2,
30+ "includeAll": true,
31+ "multi": true,
32+ "allValue": ".*",
33+ "current": {
34+ "text": "All",
35+ "value": "$__all"
36+ }
37+ },
38+ {
39+ "name": "instance_id",
40+ "label": "Instance",
41+ "type": "query",
42+ "datasource": {
43+ "type": "prometheus",
44+ "uid": "prometheus"
45+ },
46+ "query": "label_values(vllm:kv_cache_usage_perc{source=~\"$source\"}, instance_id)",
47+ "refresh": 2,
48+ "includeAll": true,
49+ "multi": true,
50+ "current": {
51+ "text": "All",
52+ "value": "$__all"
53+ }
54+ }
55+ ]
56+ },
57+ "panels": [
58+ {
59+ "id": 1,
60+ "type": "timeseries",
61+ "title": "vLLM KV cache usage %",
62+ "gridPos": {
63+ "h": 8,
64+ "w": 12,
65+ "x": 0,
66+ "y": 0
67+ },
68+ "datasource": {
69+ "type": "prometheus",
70+ "uid": "prometheus"
71+ },
72+ "targets": [
73+ {
74+ "expr": "avg by (model_name) (vllm:kv_cache_usage_perc{source=~\"$source\"}) * 100",
75+ "legendFormat": "{{model_name}}",
76+ "refId": "A"
77+ }
78+ ],
79+ "fieldConfig": {
80+ "defaults": {
81+ "unit": "percent",
82+ "custom": {
83+ "drawStyle": "line",
84+ "lineWidth": 1,
85+ "fillOpacity": 10
86+ },
87+ "min": 0,
88+ "max": 100
89+ }
90+ },
91+ "options": {
92+ "tooltip": {
93+ "mode": "multi"
94+ },
95+ "legend": {
96+ "displayMode": "table",
97+ "placement": "bottom"
98+ }
99+ }
100+ },
101+ {
102+ "id": 2,
103+ "type": "timeseries",
104+ "title": "vLLM prefix cache hit rate",
105+ "gridPos": {
106+ "h": 8,
107+ "w": 12,
108+ "x": 12,
109+ "y": 0
110+ },
111+ "datasource": {
112+ "type": "prometheus",
113+ "uid": "prometheus"
114+ },
115+ "targets": [
116+ {
117+ "expr": "(sum(rate(vllm:prefix_cache_hits_total{source=~\"$source\"}[5m])) / clamp_min(sum(rate(vllm:prefix_cache_queries_total{source=~\"$source\"}[5m])), 1))",
118+ "legendFormat": "hit_rate",
119+ "refId": "A"
120+ }
121+ ],
122+ "fieldConfig": {
123+ "defaults": {
124+ "unit": "percentunit",
125+ "custom": {
126+ "drawStyle": "line",
127+ "lineWidth": 1,
128+ "fillOpacity": 10
129+ },
130+ "min": 0,
131+ "max": 1
132+ }
133+ },
134+ "options": {
135+ "tooltip": {
136+ "mode": "multi"
137+ },
138+ "legend": {
139+ "displayMode": "table",
140+ "placement": "bottom"
141+ }
142+ }
143+ }
144+ ]
145+}
@@ -0,0 +1,14 @@
1+apiVersion: 1
2+ 
3+providers:
4+ - name: motor
5+ orgId: 1
6+ folder: ""
7+ type: file
8+ disableDeletion: false
9+ editable: true
10+ updateIntervalSeconds: 30
11+ allowUiUpdates: true
12+ options:
13+ path: /var/lib/grafana/dashboards
14+ foldersFromFilesStructure: false
@@ -0,0 +1,27 @@
1+# Grafana datasource provisioning for minimal mode.
2+apiVersion: 1
3+ 
4+datasources:
5+ - name: Prometheus
6+ type: prometheus
7+ uid: prometheus
8+ access: proxy
9+ url: http://prometheus:9090
10+ isDefault: true
11+ jsonData:
12+ timeInterval: 5s
13+ httpMethod: POST
14+ 
15+ - name: Tempo
16+ type: tempo
17+ uid: tempo
18+ access: proxy
19+ url: http://tempo:3200
20+ jsonData:
21+ httpMethod: GET
22+ nodeGraph:
23+ enabled: true
24+ search:
25+ hide: false
26+ serviceMap:
27+ datasourceUid: prometheus
@@ -0,0 +1,60 @@
1+# Grafana datasource provisioning.
2+# 预置 Prometheus / Tempo / Loki,并配置 Trace<->Log 双向跳转。
3+apiVersion: 1
4+ 
5+datasources:
6+ - name: Prometheus
7+ type: prometheus
8+ uid: prometheus
9+ access: proxy
10+ url: http://prometheus:9090
11+ isDefault: true
12+ jsonData:
13+ timeInterval: 5s
14+ httpMethod: POST
15+ 
16+ - name: Tempo
17+ type: tempo
18+ uid: tempo
19+ access: proxy
20+ url: http://tempo:3200
21+ jsonData:
22+ httpMethod: GET
23+ tracesToLogsV2:
24+ datasourceUid: loki
25+ tags:
26+ - { key: service.name, value: service_name }
27+ - { key: x_request_id, value: x_request_id }
28+ spanStartTimeShift: -5m
29+ spanEndTimeShift: 5m
30+ filterByTraceID: true
31+ filterBySpanID: false
32+ tracesToMetrics:
33+ datasourceUid: prometheus
34+ tags:
35+ - { key: service.name, value: service }
36+ nodeGraph:
37+ enabled: true
38+ search:
39+ hide: false
40+ serviceMap:
41+ datasourceUid: prometheus
42+ lokiSearch:
43+ datasourceUid: loki
44+ 
45+ - name: Loki
46+ type: loki
47+ uid: loki
48+ access: proxy
49+ url: http://loki:3100
50+ jsonData:
51+ derivedFields:
52+ - name: TraceID
53+ matcherRegex: '(?:trace_id|traceID|traceId)\s*[:=]\s*"?([A-Fa-f0-9]+)"?'
54+ url: '$${__value.raw}'
55+ datasourceUid: tempo
56+ - name: x_request_id
57+ matcherRegex: '(?:x_request_id|x-request-id)\s*[:=]\s*"?([A-Za-z0-9_-]+)"?'
58+ url: '$${__value.raw}'
59+ datasourceUid: tempo
60+ urlDisplayLabel: 'Find by request_id'
@@ -0,0 +1,239 @@
1+#!/usr/bin/env python3
2+# -*- coding: utf-8 -*-
3+ 
4+"""
5+Build `motor-vllm-profiling.json` from Prometheus `vllm_profiling_*` metric families.
6+"""
7+ 
8+from __future__ import annotations
9+ 
10+import argparse
11+import json
12+import logging
13+import sys
14+import urllib.error
15+import urllib.parse
16+import urllib.request
17+from pathlib import Path
18+from typing import Dict, List
19+ 
20+logger = logging.getLogger(__name__)
21+ 
22+CORE_METRICS = [
23+ "vllm_profiling_forward_duration_seconds_bucket",
24+ "vllm_profiling_execute_model_duration_seconds_bucket",
25+ "vllm_profiling_scheduler_duration_seconds_bucket",
26+ "vllm_profiling_batch_size",
27+ "vllm_profiling_running_queue_size",
28+]
29+ 
30+ 
31+def fetch_metric_names(prometheus_url: str) -> List[str]:
32+ endpoint = urllib.parse.urljoin(prometheus_url.rstrip("/") + "/", "api/v1/label/__name__/values")
33+ req = urllib.request.Request(endpoint, headers={"Accept": "application/json"})
34+ with urllib.request.urlopen(req, timeout=10) as resp:
35+ payload = json.loads(resp.read().decode("utf-8"))
36+ if payload.get("status") != "success":
37+ raise RuntimeError(f"prometheus response is not success: {payload}")
38+ values = payload.get("data", [])
39+ return sorted(name for name in values if isinstance(name, str) and name.startswith("vllm_profiling_"))
40+ 
41+ 
42+def infer_query(metric_name: str) -> str:
43+ if metric_name.endswith("_bucket"):
44+ metric = metric_name
45+ return (
46+ f"histogram_quantile(0.95, sum(rate({metric}[5m])) by (le, pd_role, role, instance_id))"
47+ )
48+ if metric_name.endswith("_total") or metric_name.endswith("_count"):
49+ return f"sum(rate({metric_name}[5m])) by (pd_role, role, instance_id)"
50+ if metric_name.endswith("_sum"):
51+ return f"sum(rate({metric_name}[5m])) by (pd_role, role, instance_id)"
52+ return f"avg({metric_name}) by (pd_role, role, instance_id)"
53+ 
54+ 
55+def make_timeseries_panel(panel_id: int, title: str, expr: str, y_unit: str = "short") -> Dict:
56+ return {
57+ "id": panel_id,
58+ "type": "timeseries",
59+ "title": title,
60+ "datasource": {"type": "prometheus", "uid": "prometheus"},
61+ "targets": [
62+ {
63+ "refId": "A",
64+ "expr": expr,
65+ "legendFormat": "{{pd_role}}/{{instance_id}}",
66+ }
67+ ],
68+ "fieldConfig": {
69+ "defaults": {
70+ "unit": y_unit,
71+ },
72+ "overrides": [],
73+ },
74+ "gridPos": {"h": 8, "w": 12, "x": 0, "y": 0},
75+ "options": {
76+ "legend": {
77+ "displayMode": "list",
78+ "placement": "bottom",
79+ }
80+ },
81+ }
82+ 
83+ 
84+def place_panels(panels: List[Dict], start_y: int) -> int:
85+ x = 0
86+ y = start_y
87+ for idx, panel in enumerate(panels):
88+ panel["gridPos"]["x"] = x
89+ panel["gridPos"]["y"] = y
90+ x += 12
91+ if idx % 2 == 1:
92+ x = 0
93+ y += 8
94+ if len(panels) % 2 == 1:
95+ y += 8
96+ return y
97+ 
98+ 
99+def build_dashboard(metric_names: List[str], title: str) -> Dict:
100+ core_names = [name for name in CORE_METRICS if name in metric_names]
101+ detail_names = [name for name in metric_names if name not in core_names]
102+ 
103+ panel_id = 1
104+ top_panels: List[Dict] = []
105+ detail_panels: List[Dict] = []
106+ 
107+ for metric in core_names:
108+ panel = make_timeseries_panel(
109+ panel_id=panel_id,
110+ title=f"{metric} (Core)",
111+ expr=infer_query(metric),
112+ )
113+ top_panels.append(panel)
114+ panel_id += 1
115+ 
116+ for metric in detail_names:
117+ panel = make_timeseries_panel(
118+ panel_id=panel_id,
119+ title=metric,
120+ expr=infer_query(metric),
121+ )
122+ detail_panels.append(panel)
123+ panel_id += 1
124+ 
125+ current_y = 0
126+ current_y = place_panels(top_panels, current_y)
127+ 
128+ detail_row = {
129+ "id": panel_id,
130+ "type": "row",
131+ "title": "指标明细 Metric Details",
132+ "collapsed": True,
133+ "gridPos": {"h": 1, "w": 24, "x": 0, "y": current_y},
134+ "panels": [],
135+ }
136+ panel_id += 1
137+ detail_y = current_y + 1
138+ place_panels(detail_panels, detail_y)
139+ detail_row["panels"] = detail_panels
140+ 
141+ templating_list = [
142+ {
143+ "name": "cluster",
144+ "type": "query",
145+ "datasource": {"type": "prometheus", "uid": "prometheus"},
146+ "query": "label_values(up, cluster)",
147+ "refresh": 1,
148+ "includeAll": True,
149+ "multi": True,
150+ "current": {"text": "All", "value": "$__all"},
151+ "label": "集群 Cluster",
152+ },
153+ {
154+ "name": "pd_role",
155+ "type": "query",
156+ "datasource": {"type": "prometheus", "uid": "prometheus"},
157+ "query": "label_values(up, pd_role)",
158+ "refresh": 1,
159+ "includeAll": True,
160+ "multi": True,
161+ "current": {"text": "All", "value": "$__all"},
162+ "label": "角色 Role",
163+ },
164+ {
165+ "name": "instance_id",
166+ "type": "query",
167+ "datasource": {"type": "prometheus", "uid": "prometheus"},
168+ "query": "label_values(up, instance_id)",
169+ "refresh": 1,
170+ "includeAll": True,
171+ "multi": True,
172+ "current": {"text": "All", "value": "$__all"},
173+ "label": "实例 Instance",
174+ },
175+ ]
176+ 
177+ return {
178+ "uid": "motor-vllm-profiling",
179+ "title": title,
180+ "tags": ["pymotor", "profiling", "ms_service_metric"],
181+ "timezone": "browser",
182+ "schemaVersion": 39,
183+ "version": 1,
184+ "refresh": "10s",
185+ "editable": True,
186+ "graphTooltip": 1,
187+ "time": {"from": "now-30m", "to": "now"},
188+ "templating": {"list": templating_list},
189+ "panels": [*top_panels, detail_row],
190+ }
191+ 
192+ 
193+def parse_args() -> argparse.Namespace:
194+ parser = argparse.ArgumentParser(description="Generate vLLM profiling dashboard JSON.")
195+ parser.add_argument(
196+ "--prometheus-url",
197+ default="http://localhost:9090",
198+ help="Prometheus base URL.",
199+ )
200+ parser.add_argument(
201+ "--output",
202+ default=str(Path(__file__).resolve().parents[1] / "dashboards" / "motor-vllm-profiling.json"),
203+ help="Output dashboard path.",
204+ )
205+ parser.add_argument(
206+ "--title",
207+ default="引擎性能剖析",
208+ help="Dashboard title.",
209+ )
210+ return parser.parse_args()
211+ 
212+ 
213+def main() -> int:
214+ logging.basicConfig(
215+ level=logging.INFO,
216+ format="[%(name)s] %(message)s",
217+ )
218+ args = parse_args()
219+ try:
220+ metric_names = fetch_metric_names(args.prometheus_url)
221+ except (urllib.error.URLError, RuntimeError, json.JSONDecodeError) as exc:
222+ logger.error("failed to query Prometheus: %s", exc)
223+ return 1
224+ 
225+ if not metric_names:
226+ logger.error("no vllm_profiling_* metrics found.")
227+ return 1
228+ 
229+ dashboard = build_dashboard(metric_names=metric_names, title=args.title)
230+ output = Path(args.output)
231+ output.parent.mkdir(parents=True, exist_ok=True)
232+ output.write_text(json.dumps(dashboard, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
233+ logger.info("generated: %s", output)
234+ logger.info("metrics: %d", len(metric_names))
235+ return 0
236+ 
237+ 
238+if __name__ == "__main__":
239+ sys.exit(main())
@@ -0,0 +1,164 @@
1+#!/usr/bin/env bash
2+ 
3+set -euo pipefail
4+ 
5+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
6+cd "${SCRIPT_DIR}"
7+. "${SCRIPT_DIR}/scripts/load-dotenv.sh"
8+ 
9+NAMESPACE="${MOTOR_NAMESPACE:-}"
10+NODE_IP="${MOTOR_NODE_IP:-}"
11+USER_CONFIG="${MOTOR_USER_CONFIG:-}"
12+ENGINE_MGMT_PORT="${MOTOR_ENGINE_MGMT_PORT:-10001}"
13+OBS_HOST_INPUT="${OBS_HOST:-}"
14+ 
15+FORCE_NATIVE=0
16+DISCOVER_ONLY=0
17+DRY_RUN=0
18+STACK_MODE="${OBS_STACK_MODE:-full}"
19+ 
20+usage() {
21+ cat <<'EOF'
22+Usage: ./launch.sh [options]
23+ 
24+Options:
25+ --namespace <namespace> Kubernetes namespace / job_id
26+ --node-ip <node-ip> Node IP used for NodePort access
27+ --user-config <path> pyMotor user_config.json path
28+ --minimal Start minimal Docker stack (Prometheus/Grafana/Tempo/OTel)
29+ --full Start full Docker stack (adds Loki/node-exporter/cAdvisor)
30+ --discover-only Only run discovery, do not start stack
31+ --dry-run Run discovery and print generated Prometheus config
32+ --native Skip Docker Compose and run native runtime
33+ -h, --help Show this help
34+ 
35+Environment:
36+ MOTOR_NAMESPACE
37+ MOTOR_NODE_IP
38+ MOTOR_USER_CONFIG
39+ MOTOR_ENGINE_MGMT_PORT
40+ OBS_HOST
41+ PROXY_SH dotenv file for native binary downloads only (see SERVICE_GUIDE.md §2.4)
42+ 
43+Proxy (see SERVICE_GUIDE.md §2.4):
44+ - Discovery/kubectl: unset shell proxy before launch (script also strips proxy for kubectl).
45+ - Docker image pull: export HTTP_PROXY in current shell or source your proxy.sh before launch.
46+ - Native runtime: set PROXY_SH=/path/to/dotenv in .env (optional; default empty).
47+ - Grafana container: HTTP_PROXY cleared for in-stack prometheus/tempo.
48+ - OBS_COMPOSE_PULL=never|missing|always OBS_COMPOSE_BUILD=0|1
49+EOF
50+}
51+ 
52+while [[ $# -gt 0 ]]; do
53+ case "$1" in
54+ --namespace)
55+ [[ $# -lt 2 ]] && { echo "[launch] missing value for --namespace" >&2; exit 1; }
56+ NAMESPACE="$2"
57+ shift 2
58+ ;;
59+ --node-ip)
60+ [[ $# -lt 2 ]] && { echo "[launch] missing value for --node-ip" >&2; exit 1; }
61+ NODE_IP="$2"
62+ shift 2
63+ ;;
64+ --user-config)
65+ [[ $# -lt 2 ]] && { echo "[launch] missing value for --user-config" >&2; exit 1; }
66+ USER_CONFIG="$2"
67+ shift 2
68+ ;;
69+ --minimal)
70+ STACK_MODE="minimal"
71+ shift
72+ ;;
73+ --full)
74+ STACK_MODE="full"
75+ shift
76+ ;;
77+ --discover-only)
78+ DISCOVER_ONLY=1
79+ shift
80+ ;;
81+ --dry-run)
82+ DRY_RUN=1
83+ shift
84+ ;;
85+ --native)
86+ FORCE_NATIVE=1
87+ shift
88+ ;;
89+ -h|--help)
90+ usage
91+ exit 0
92+ ;;
93+ *)
94+ echo "[launch] unknown option: $1" >&2
95+ usage
96+ exit 1
97+ ;;
98+ esac
99+done
100+ 
101+if [[ ! -f .env && -f .env.example ]]; then
102+ cp .env.example .env
103+ echo "[launch] created .env from .env.example"
104+fi
105+ 
106+DISCOVERY_RUNTIME="docker"
107+if [[ "${FORCE_NATIVE}" -eq 1 ]]; then
108+ DISCOVERY_RUNTIME="native"
109+fi
110+DISCOVERY_CMD=(python3 "./scripts/discover-targets.py" "--output-dir" "./generated" "--engine-mgmt-port" "${ENGINE_MGMT_PORT}" "--runtime" "${DISCOVERY_RUNTIME}")
111+[[ -n "${NAMESPACE}" ]] && DISCOVERY_CMD+=("--namespace" "${NAMESPACE}")
112+[[ -n "${NODE_IP}" ]] && DISCOVERY_CMD+=("--node-ip" "${NODE_IP}")
113+[[ -n "${USER_CONFIG}" ]] && DISCOVERY_CMD+=("--user-config" "${USER_CONFIG}")
114+[[ -n "${OBS_HOST_INPUT}" ]] && DISCOVERY_CMD+=("--obs-host" "${OBS_HOST_INPUT}")
115+ 
116+echo "[launch] discovering targets..."
117+"${DISCOVERY_CMD[@]}"
118+ 
119+load_dotenv "./generated/discovered.env"
120+ 
121+if [[ "${DRY_RUN}" -eq 1 ]]; then
122+ echo
123+ echo "========== generated/prometheus.yml =========="
124+ sed -n '1,240p' "./generated/prometheus.yml"
125+ echo "============================================="
126+fi
127+ 
128+if [[ "${DISCOVER_ONLY}" -eq 1 || "${DRY_RUN}" -eq 1 ]]; then
129+ echo "[launch] discovery completed."
130+ exit 0
131+fi
132+ 
133+run_native() {
134+ echo "[launch] starting native runtime..."
135+ echo "[launch] refreshing discovery for native runtime..."
136+ NATIVE_DISCOVERY_CMD=(python3 "./scripts/discover-targets.py" "--output-dir" "./generated" "--engine-mgmt-port" "${ENGINE_MGMT_PORT}" "--runtime" "native")
137+ [[ -n "${NAMESPACE}" ]] && NATIVE_DISCOVERY_CMD+=("--namespace" "${NAMESPACE}")
138+ [[ -n "${NODE_IP}" ]] && NATIVE_DISCOVERY_CMD+=("--node-ip" "${NODE_IP}")
139+ [[ -n "${USER_CONFIG}" ]] && NATIVE_DISCOVERY_CMD+=("--user-config" "${USER_CONFIG}")
140+ [[ -n "${OBS_HOST_INPUT}" ]] && NATIVE_DISCOVERY_CMD+=("--obs-host" "${OBS_HOST_INPUT}")
141+ "${NATIVE_DISCOVERY_CMD[@]}"
142+ ./scripts/start-native.sh \
143+ --env-file "./generated/discovered.env" \
144+ --prometheus-file "./generated/prometheus.yml"
145+}
146+ 
147+if [[ "${FORCE_NATIVE}" -eq 1 ]]; then
148+ run_native
149+ exit 0
150+fi
151+ 
152+echo "[launch] starting Docker Compose stack..."
153+set +e
154+PROMETHEUS_CONFIG_FILE="./generated/prometheus.yml" \
155+OBS_HOST="${OBS_HOST:-}" \
156+./start.sh "--${STACK_MODE}"
157+DOCKER_RC=$?
158+set -e
159+ 
160+if [[ "${DOCKER_RC}" -ne 0 ]]; then
161+ echo "[launch] Docker startup failed (exit=${DOCKER_RC}), cleaning partial Docker stack before native fallback."
162+ ./stop.sh || true
163+ run_native
164+fi
@@ -0,0 +1,55 @@
1+# Loki single-binary configuration for local development.
2+# All-in-one mode with filesystem storage; suitable for demo / PoC.
3+ 
4+auth_enabled: false
5+ 
6+server:
7+ http_listen_port: 3100
8+ grpc_listen_port: 9096
9+ log_level: warn
10+ 
11+common:
12+ instance_addr: 127.0.0.1
13+ path_prefix: /var/loki
14+ storage:
15+ filesystem:
16+ chunks_directory: /var/loki/chunks
17+ rules_directory: /var/loki/rules
18+ replication_factor: 1
19+ ring:
20+ kvstore:
21+ store: inmemory
22+ 
23+ingester:
24+ chunk_idle_period: 1m
25+ chunk_retain_period: 30s
26+ 
27+schema_config:
28+ configs:
29+ - from: 2024-01-01
30+ store: tsdb
31+ object_store: filesystem
32+ schema: v13
33+ index:
34+ prefix: index_
35+ period: 24h
36+ 
37+storage_config:
38+ tsdb_shipper:
39+ active_index_directory: /var/loki/index
40+ cache_location: /var/loki/index_cache
41+ filesystem:
42+ directory: /var/loki/chunks
43+ 
44+limits_config:
45+ reject_old_samples: false
46+ retention_period: 72h
47+ allow_structured_metadata: true
48+ 
49+compactor:
50+ working_directory: /var/loki/compactor
51+ retention_enabled: true
52+ delete_request_store: filesystem
53+ 
54+analytics:
55+ reporting_enabled: false
@@ -0,0 +1,37 @@
1+# OpenTelemetry Collector configuration for minimal mode.
2+ 
3+receivers:
4+ otlp:
5+ protocols:
6+ grpc:
7+ endpoint: 0.0.0.0:4317
8+ http:
9+ endpoint: 0.0.0.0:4318
10+ 
11+processors:
12+ batch:
13+ send_batch_size: 1024
14+ timeout: 5s
15+ resource:
16+ attributes:
17+ - key: deployment.environment
18+ value: pymotor-observability-minimal
19+ action: upsert
20+ 
21+exporters:
22+ otlp/tempo:
23+ endpoint: tempo:4317
24+ tls:
25+ insecure: true
26+ debug:
27+ verbosity: basic
28+ 
29+service:
30+ pipelines:
31+ traces:
32+ receivers: [otlp]
33+ processors: [batch, resource]
34+ exporters: [otlp/tempo, debug]
35+ telemetry:
36+ logs:
37+ level: info
@@ -0,0 +1,56 @@
1+# OpenTelemetry Collector configuration.
2+# Receives OTLP traces/logs from pyMotor and routes:
3+# traces → Tempo (otlp/tempo → tempo:4317)
4+# logs → Loki
5+#
6+# pyMotor Coordinator 需在 user_config.json 设置 tracer_config.endpoint,
7+# 并在 env.json 设置 OTEL_EXPORTER_OTLP_TRACES_PROTOCOL(见 config/tracing.example.json)。
8+ 
9+receivers:
10+ otlp:
11+ protocols:
12+ grpc:
13+ endpoint: 0.0.0.0:4317
14+ http:
15+ endpoint: 0.0.0.0:4318
16+ 
17+processors:
18+ batch:
19+ send_batch_size: 1024
20+ timeout: 5s
21+ resource:
22+ attributes:
23+ - key: deployment.environment
24+ value: pymotor-observability-local
25+ action: upsert
26+ # 保留上游 resource 中的 service.name(Coordinator / Engine 的 OTEL_SERVICE_NAME)
27+ # 把 trace_id/span_id 提升为 Loki label 以便 derived fields 跳转
28+ attributes/loki:
29+ actions:
30+ - key: loki.format
31+ value: json
32+ action: insert
33+ 
34+exporters:
35+ otlp/tempo:
36+ endpoint: tempo:4317
37+ tls:
38+ insecure: true
39+ loki:
40+ endpoint: http://loki:3100/loki/api/v1/push
41+ debug:
42+ verbosity: basic
43+ 
44+service:
45+ pipelines:
46+ traces:
47+ receivers: [otlp]
48+ processors: [batch, resource]
49+ exporters: [otlp/tempo, debug]
50+ logs:
51+ receivers: [otlp]
52+ processors: [batch, resource, attributes/loki]
53+ exporters: [loki, debug]
54+ telemetry:
55+ logs:
56+ level: info
@@ -0,0 +1,50 @@
1+# Minimal Prometheus config for pyMotor real-data observability.
2+ 
3+global:
4+ scrape_interval: 5s
5+ evaluation_interval: 15s
6+ metric_name_validation_scheme: utf8
7+ external_labels:
8+ monitor: motor-observability
9+ 
10+scrape_configs:
11+ - job_name: prometheus
12+ static_configs:
13+ - targets:
14+ - "localhost:9090"
15+ 
16+ - job_name: motor-coordinator
17+ metrics_path: /metrics
18+ static_configs:
19+ - targets:
20+ - "host.docker.internal:1027"
21+ labels:
22+ motor_component: coordinator
23+ motor_metric_scope: full
24+ cluster: example-namespace
25+ source: real
26+ 
27+ - job_name: motor-engine
28+ metrics_path: /metrics
29+ honor_labels: true
30+ static_configs:
31+ - targets:
32+ - "host.docker.internal:10001"
33+ labels:
34+ motor_component: engine
35+ pd_role: prefill
36+ role: prefill
37+ instance_id: p0
38+ cluster: example-namespace
39+ source: real
40+ 
41+ - job_name: vllm-profiling
42+ metrics_path: /metrics
43+ honor_labels: true
44+ static_configs:
45+ - targets:
46+ - "host.docker.internal:10001"
47+ labels:
48+ motor_component: vllm
49+ cluster: example-namespace
50+ source: real
@@ -0,0 +1,120 @@
1+# Prometheus 通用模板(建议通过 scripts/discover-targets.py 生成运行时配置)
2+#
3+# 正式一键启动入口:
4+# ./launch.sh
5+# 运行时会覆盖为:
6+# generated/prometheus.yml
7+#
8+# 模板用途:
9+# - 作为本地调试与配置说明参考
10+# - 说明 Coordinator / Engine 的标准 scrape 约定
11+ 
12+global:
13+ scrape_interval: 5s
14+ evaluation_interval: 15s
15+ metric_name_validation_scheme: utf8
16+ external_labels:
17+ monitor: motor-observability
18+ 
19+scrape_configs:
20+ - job_name: prometheus
21+ static_configs:
22+ - targets:
23+ - "localhost:9090"
24+ 
25+ # Coordinator observability / typed metrics
26+ - job_name: motor-coordinator
27+ metrics_path: /metrics
28+ static_configs:
29+ - targets:
30+ - "host.docker.internal:1027"
31+ labels:
32+ motor_component: coordinator
33+ motor_metric_scope: full
34+ cluster: example-namespace
35+ 
36+ - job_name: motor-coordinator-instance
37+ metrics_path: /metrics?type=instance
38+ static_configs:
39+ - targets:
40+ - "host.docker.internal:1027"
41+ labels:
42+ motor_component: coordinator
43+ motor_metric_scope: instance
44+ cluster: example-namespace
45+ 
46+ - job_name: motor-coordinator-role-prefill
47+ metrics_path: /metrics?type=role&role=prefill
48+ static_configs:
49+ - targets:
50+ - "host.docker.internal:1027"
51+ labels:
52+ motor_component: coordinator
53+ motor_metric_scope: role
54+ role: prefill
55+ pd_role: prefill
56+ cluster: example-namespace
57+ 
58+ - job_name: motor-coordinator-role-decode
59+ metrics_path: /metrics?type=role&role=decode
60+ static_configs:
61+ - targets:
62+ - "host.docker.internal:1027"
63+ labels:
64+ motor_component: coordinator
65+ motor_metric_scope: role
66+ role: decode
67+ pd_role: decode
68+ cluster: example-namespace
69+ 
70+ # Engine management metrics(honor_labels 打开,保留应用侧标签)
71+ - job_name: motor-engine
72+ metrics_path: /metrics
73+ honor_labels: true
74+ static_configs:
75+ - targets:
76+ - "host.docker.internal:10001"
77+ labels:
78+ motor_component: engine
79+ pd_role: prefill
80+ role: prefill
81+ instance_id: p0
82+ cluster: example-namespace
83+ - targets:
84+ - "host.docker.internal:10001"
85+ labels:
86+ motor_component: engine
87+ pd_role: prefill
88+ role: prefill
89+ instance_id: p1
90+ cluster: example-namespace
91+ - targets:
92+ - "host.docker.internal:10001"
93+ labels:
94+ motor_component: engine
95+ pd_role: decode
96+ role: decode
97+ instance_id: d0
98+ cluster: example-namespace
99+ 
100+ # vLLM Profiling(来自 ms_service_metric 的 vllm_profiling_*)
101+ - job_name: vllm-profiling
102+ metrics_path: /metrics
103+ honor_labels: true
104+ static_configs:
105+ - targets:
106+ - "host.docker.internal:10001"
107+ labels:
108+ motor_component: vllm
109+ cluster: example-namespace
110+ 
111+ # Docker Compose 基础 exporter
112+ - job_name: node-exporter
113+ static_configs:
114+ - targets:
115+ - "node-exporter:9100"
116+ 
117+ - job_name: cadvisor
118+ static_configs:
119+ - targets:
120+ - "cadvisor:8080"
@@ -0,0 +1,88 @@
1+# 示例:真实集群 Prometheus 配置模板
2+# 推荐方式:使用 ./launch.sh 自动发现并生成 generated/prometheus.yml
3+ 
4+global:
5+ scrape_interval: 5s
6+ evaluation_interval: 15s
7+ metric_name_validation_scheme: utf8
8+ external_labels:
9+ monitor: motor-observability
10+ 
11+scrape_configs:
12+ - job_name: prometheus
13+ static_configs:
14+ - targets:
15+ - "localhost:9090"
16+ 
17+ - job_name: motor-coordinator
18+ metrics_path: /metrics
19+ static_configs:
20+ - targets:
21+ - "host.docker.internal:1027"
22+ labels:
23+ motor_component: coordinator
24+ motor_metric_scope: full
25+ cluster: example-namespace
26+ 
27+ - job_name: motor-coordinator-instance
28+ metrics_path: /metrics?type=instance
29+ static_configs:
30+ - targets:
31+ - "host.docker.internal:1027"
32+ labels:
33+ motor_component: coordinator
34+ motor_metric_scope: instance
35+ cluster: example-namespace
36+ 
37+ - job_name: motor-coordinator-role-prefill
38+ metrics_path: /metrics?type=role&role=prefill
39+ static_configs:
40+ - targets:
41+ - "host.docker.internal:1027"
42+ labels:
43+ motor_component: coordinator
44+ motor_metric_scope: role
45+ role: prefill
46+ pd_role: prefill
47+ cluster: example-namespace
48+ 
49+ - job_name: motor-coordinator-role-decode
50+ metrics_path: /metrics?type=role&role=decode
51+ static_configs:
52+ - targets:
53+ - "host.docker.internal:1027"
54+ labels:
55+ motor_component: coordinator
56+ motor_metric_scope: role
57+ role: decode
58+ pd_role: decode
59+ cluster: example-namespace
60+ 
61+ - job_name: motor-engine
62+ metrics_path: /metrics
63+ honor_labels: true
64+ static_configs:
65+ - targets:
66+ - "host.docker.internal:10001"
67+ labels:
68+ motor_component: engine
69+ pd_role: prefill
70+ role: prefill
71+ instance_id: p0
72+ cluster: example-namespace
73+ - targets:
74+ - "host.docker.internal:10001"
75+ labels:
76+ motor_component: engine
77+ pd_role: prefill
78+ role: prefill
79+ instance_id: p1
80+ cluster: example-namespace
81+ - targets:
82+ - "host.docker.internal:10001"
83+ labels:
84+ motor_component: engine
85+ pd_role: decode
86+ role: decode
87+ instance_id: d0
88+ cluster: example-namespace
@@ -0,0 +1,120 @@
1+# Prometheus 通用模板(建议通过 scripts/discover-targets.py 生成运行时配置)
2+#
3+# 正式一键启动入口:
4+# ./launch.sh
5+# 运行时会覆盖为:
6+# generated/prometheus.yml
7+#
8+# 模板用途:
9+# - 作为本地调试与配置说明参考
10+# - 说明 Coordinator / Engine 的标准 scrape 约定
11+ 
12+global:
13+ scrape_interval: 5s
14+ evaluation_interval: 15s
15+ metric_name_validation_scheme: utf8
16+ external_labels:
17+ monitor: motor-observability
18+ 
19+scrape_configs:
20+ - job_name: prometheus
21+ static_configs:
22+ - targets:
23+ - "localhost:9090"
24+ 
25+ # Coordinator observability / typed metrics
26+ - job_name: motor-coordinator
27+ metrics_path: /metrics
28+ static_configs:
29+ - targets:
30+ - "host.docker.internal:1027"
31+ labels:
32+ motor_component: coordinator
33+ motor_metric_scope: full
34+ cluster: example-namespace
35+ 
36+ - job_name: motor-coordinator-instance
37+ metrics_path: /metrics?type=instance
38+ static_configs:
39+ - targets:
40+ - "host.docker.internal:1027"
41+ labels:
42+ motor_component: coordinator
43+ motor_metric_scope: instance
44+ cluster: example-namespace
45+ 
46+ - job_name: motor-coordinator-role-prefill
47+ metrics_path: /metrics?type=role&role=prefill
48+ static_configs:
49+ - targets:
50+ - "host.docker.internal:1027"
51+ labels:
52+ motor_component: coordinator
53+ motor_metric_scope: role
54+ role: prefill
55+ pd_role: prefill
56+ cluster: example-namespace
57+ 
58+ - job_name: motor-coordinator-role-decode
59+ metrics_path: /metrics?type=role&role=decode
60+ static_configs:
61+ - targets:
62+ - "host.docker.internal:1027"
63+ labels:
64+ motor_component: coordinator
65+ motor_metric_scope: role
66+ role: decode
67+ pd_role: decode
68+ cluster: example-namespace
69+ 
70+ # Engine management metrics(honor_labels 打开,保留应用侧标签)
71+ - job_name: motor-engine
72+ metrics_path: /metrics
73+ honor_labels: true
74+ static_configs:
75+ - targets:
76+ - "host.docker.internal:10001"
77+ labels:
78+ motor_component: engine
79+ pd_role: prefill
80+ role: prefill
81+ instance_id: p0
82+ cluster: example-namespace
83+ - targets:
84+ - "host.docker.internal:10001"
85+ labels:
86+ motor_component: engine
87+ pd_role: prefill
88+ role: prefill
89+ instance_id: p1
90+ cluster: example-namespace
91+ - targets:
92+ - "host.docker.internal:10001"
93+ labels:
94+ motor_component: engine
95+ pd_role: decode
96+ role: decode
97+ instance_id: d0
98+ cluster: example-namespace
99+ 
100+ # vLLM Profiling(来自 ms_service_metric 的 vllm_profiling_*)
101+ - job_name: vllm-profiling
102+ metrics_path: /metrics
103+ honor_labels: true
104+ static_configs:
105+ - targets:
106+ - "host.docker.internal:10001"
107+ labels:
108+ motor_component: vllm
109+ cluster: example-namespace
110+ 
111+ # Docker Compose 基础 exporter
112+ - job_name: node-exporter
113+ static_configs:
114+ - targets:
115+ - "node-exporter:9100"
116+ 
117+ - job_name: cadvisor
118+ static_configs:
119+ - targets:
120+ - "cadvisor:8080"
@@ -0,0 +1,27 @@
1+#!/usr/bin/env bash
2+# Export KEY=VALUE pairs from a dotenv-style file without using "source".
3+ 
4+load_dotenv() {
5+ local env_file=$1
6+ [[ -n "${env_file}" && -f "${env_file}" ]] || return 0
7+ local line key value
8+ while IFS= read -r line || [[ -n "${line}" ]]; do
9+ [[ "${line}" =~ ^[[:space:]]*# ]] && continue
10+ [[ "${line}" =~ ^[[:space:]]*$ ]] && continue
11+ line="${line#"${line%%[![:space:]]*}"}"
12+ line="${line%"${line##*[![:space:]]}"}"
13+ [[ "${line}" != *"="* ]] && continue
14+ key="${line%%=*}"
15+ value="${line#*=}"
16+ key="${key#"${key%%[![:space:]]*}"}"
17+ key="${key%"${key##*[![:space:]]}"}"
18+ value="${value#"${value%%[![:space:]]*}"}"
19+ value="${value%"${value##*[![:space:]]}"}"
20+ if [[ "${value}" == \"*\" ]]; then
21+ value="${value:1:${#value}-2}"
22+ elif [[ "${value}" == \'*\' ]]; then
23+ value="${value:1:${#value}-2}"
24+ fi
25+ export "${key}=${value}"
26+ done < "${env_file}"
27+}
@@ -0,0 +1,188 @@
1+#!/usr/bin/env bash
2+ 
3+set -euo pipefail
4+ 
5+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
6+STACK_DIR="$(cd -- "${SCRIPT_DIR}/.." &>/dev/null && pwd)"
7+. "${SCRIPT_DIR}/load-dotenv.sh"
8+RUN_DIR="${STACK_DIR}/generated"
9+ENV_FILE="${RUN_DIR}/discovered.env"
10+PID_FILE="${RUN_DIR}/k8s-port-forwards.pids"
11+META_FILE="${RUN_DIR}/k8s-port-forwards.meta"
12+LOG_DIR="${RUN_DIR}/logs"
13+BIND_HOST="${PORT_FORWARD_BIND_HOST:-0.0.0.0}"
14+ 
15+usage() {
16+ echo "Usage: $0 [--env-file <generated/discovered.env>]"
17+}
18+ 
19+while [[ $# -gt 0 ]]; do
20+ case "$1" in
21+ --env-file)
22+ [[ $# -lt 2 ]] && { echo "[port-forward] missing value for --env-file" >&2; exit 1; }
23+ ENV_FILE="$2"
24+ shift 2
25+ ;;
26+ -h|--help)
27+ usage
28+ exit 0
29+ ;;
30+ *)
31+ echo "[port-forward] unknown option: $1" >&2
32+ usage
33+ exit 1
34+ ;;
35+ esac
36+done
37+ 
38+[[ -f "${ENV_FILE}" ]] || { echo "[port-forward] env file not found: ${ENV_FILE}"; exit 0; }
39+mkdir -p "${RUN_DIR}" "${LOG_DIR}"
40+load_dotenv "${ENV_FILE}"
41+ 
42+PORT_FORWARD_COUNT="${PORT_FORWARD_COUNT:-0}"
43+if [[ "${PORT_FORWARD_COUNT}" -eq 0 ]]; then
44+ echo "[port-forward] no host TCP forwards requested."
45+ exit 0
46+fi
47+ 
48+desired_specs() {
49+ local idx var
50+ for ((idx = 0; idx < PORT_FORWARD_COUNT; idx++)); do
51+ var="PORT_FORWARD_${idx}"
52+ [[ -n "${!var:-}" ]] && echo "${!var}"
53+ done
54+}
55+ 
56+is_pid_running() {
57+ local pid="$1"
58+ [[ -n "${pid}" ]] && kill -0 "${pid}" >/dev/null 2>&1
59+}
60+ 
61+stop_existing() {
62+ [[ -f "${PID_FILE}" ]] || return 0
63+ while IFS= read -r pid; do
64+ [[ -z "${pid}" ]] && continue
65+ if is_pid_running "${pid}"; then
66+ echo "[port-forward] stopping pid=${pid}"
67+ kill "${pid}" || true
68+ fi
69+ done <"${PID_FILE}"
70+ rm -f "${PID_FILE}" "${META_FILE}"
71+}
72+ 
73+args_matches_listen_port() {
74+ local args="$1"
75+ local port="$2"
76+ [[ "${args}" =~ (^|[[:space:]])--listen-port[[:space:]]+${port}([^0-9]|$) ]] && return 0
77+ [[ "${args}" =~ --listen-port=${port}([^0-9]|$) ]] && return 0
78+ return 1
79+}
80+ 
81+cleanup_orphans_for_ports() {
82+ local ports=("$@")
83+ command -v pgrep >/dev/null 2>&1 || return 0
84+ local pid args port
85+ while IFS= read -r pid; do
86+ [[ -z "${pid}" || "${pid}" == "$$" ]] && continue
87+ args="$(ps -p "${pid}" -o args= 2>/dev/null || true)"
88+ [[ "${args}" == *"tcp-forward.py"* ]] || continue
89+ for port in "${ports[@]}"; do
90+ if args_matches_listen_port "${args}" "${port}"; then
91+ echo "[port-forward] cleaning orphan pid=${pid} listen_port=${port}"
92+ kill "${pid}" || true
93+ fi
94+ done
95+ done < <(pgrep -f "tcp-forward.py" || true)
96+}
97+ 
98+port_accepts_connections() {
99+ local port="$1"
100+ python3 - "${port}" <<'PY'
101+import socket
102+import sys
103+port = int(sys.argv[1])
104+sock = socket.socket()
105+sock.settimeout(1)
106+try:
107+ sock.connect(("127.0.0.1", port))
108+except OSError:
109+ sys.exit(1)
110+finally:
111+ sock.close()
112+PY
113+}
114+ 
115+start_one() {
116+ local spec="$1"
117+ local namespace pod_ip remote_port local_port pod_name
118+ IFS='|' read -r namespace pod_ip remote_port local_port pod_name <<<"${spec}"
119+ local log_file="${LOG_DIR}/tcp-forward-${local_port}.log"
120+ python3 "${SCRIPT_DIR}/tcp-forward.py" \
121+ --listen-host "${BIND_HOST}" \
122+ --listen-port "${local_port}" \
123+ --target-host "${pod_ip}" \
124+ --target-port "${remote_port}" >"${log_file}" 2>&1 &
125+ local pid=$!
126+ sleep 0.5
127+ if ! is_pid_running "${pid}" || ! port_accepts_connections "${local_port}"; then
128+ kill "${pid}" >/dev/null 2>&1 || true
129+ sleep 0.3
130+ python3 "${SCRIPT_DIR}/tcp-forward.py" \
131+ --listen-host "${BIND_HOST}" \
132+ --listen-port "${local_port}" \
133+ --target-host "${pod_ip}" \
134+ --target-port "${remote_port}" >>"${log_file}" 2>&1 &
135+ pid=$!
136+ sleep 0.5
137+ fi
138+ if ! is_pid_running "${pid}" || ! port_accepts_connections "${local_port}"; then
139+ echo "[port-forward] failed to start ${namespace}/${pod_name} ${pod_ip}:${remote_port} -> ${local_port}" >&2
140+ return 1
141+ fi
142+ echo "${pid}" >>"${PID_FILE}"
143+ echo "[port-forward] started ${namespace}/${pod_name} ${pod_ip}:${remote_port} -> ${BIND_HOST}:${local_port} pid=${pid}"
144+}
145+ 
146+DESIRED_FILE="$(mktemp)"
147+EXISTING_FILE="$(mktemp)"
148+trap 'rm -f "${DESIRED_FILE}" "${EXISTING_FILE}"' EXIT
149+desired_specs | sort >"${DESIRED_FILE}"
150+if [[ -f "${META_FILE}" ]]; then
151+ sort "${META_FILE}" >"${EXISTING_FILE}"
152+else
153+ : >"${EXISTING_FILE}"
154+fi
155+ 
156+all_alive=1
157+if [[ -f "${PID_FILE}" ]]; then
158+ while IFS= read -r pid; do
159+ [[ -z "${pid}" ]] && continue
160+ is_pid_running "${pid}" || all_alive=0
161+ done <"${PID_FILE}"
162+else
163+ all_alive=0
164+fi
165+ 
166+if cmp -s "${DESIRED_FILE}" "${EXISTING_FILE}" && [[ "${all_alive}" -eq 1 ]]; then
167+ echo "[port-forward] existing topology is current."
168+ exit 0
169+fi
170+ 
171+mapfile -t local_ports < <(desired_specs | awk -F'|' '{print $4}' | sort -u)
172+stop_existing
173+cleanup_orphans_for_ports "${local_ports[@]}"
174+: >"${PID_FILE}"
175+start_failed=0
176+while IFS= read -r spec; do
177+ [[ -z "${spec}" ]] && continue
178+ if ! start_one "${spec}"; then
179+ start_failed=1
180+ fi
181+done <"${DESIRED_FILE}"
182+ 
183+if [[ "${start_failed}" -eq 1 ]]; then
184+ rm -f "${META_FILE}"
185+ echo "[port-forward] one or more forwards failed; META not updated (will retry on next run)." >&2
186+ exit 1
187+fi
188+cp "${DESIRED_FILE}" "${META_FILE}"
@@ -0,0 +1,356 @@
1+#!/usr/bin/env bash
2+ 
3+set -euo pipefail
4+ 
5+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
6+STACK_DIR="$(cd -- "${SCRIPT_DIR}/.." &>/dev/null && pwd)"
7+cd "${STACK_DIR}"
8+. "${SCRIPT_DIR}/load-dotenv.sh"
9+ 
10+ENV_FILE="./generated/discovered.env"
11+PROMETHEUS_FILE="./generated/prometheus.yml"
12+ 
13+usage() {
14+ cat <<'EOF'
15+Usage: ./scripts/start-native.sh [options]
16+ 
17+Options:
18+ --env-file <path> Discovered env file (default: ./generated/discovered.env)
19+ --prometheus-file <path> Generated prometheus config (default: ./generated/prometheus.yml)
20+ -h, --help Show this help
21+EOF
22+}
23+ 
24+while [[ $# -gt 0 ]]; do
25+ case "$1" in
26+ --env-file)
27+ [[ $# -lt 2 ]] && { echo "[native] missing value for --env-file" >&2; exit 1; }
28+ ENV_FILE="$2"
29+ shift 2
30+ ;;
31+ --prometheus-file)
32+ [[ $# -lt 2 ]] && { echo "[native] missing value for --prometheus-file" >&2; exit 1; }
33+ PROMETHEUS_FILE="$2"
34+ shift 2
35+ ;;
36+ -h|--help)
37+ usage
38+ exit 0
39+ ;;
40+ *)
41+ echo "[native] unknown option: $1" >&2
42+ usage
43+ exit 1
44+ ;;
45+ esac
46+done
47+ 
48+if [[ -f "${STACK_DIR}/.env" ]]; then
49+ load_dotenv "${STACK_DIR}/.env"
50+elif [[ -f "${STACK_DIR}/.env.example" ]]; then
51+ load_dotenv "${STACK_DIR}/.env.example"
52+fi
53+load_dotenv "${ENV_FILE}"
54+ 
55+PROXY_SH="${PROXY_SH:-}"
56+if [[ -n "${PROXY_SH}" && -f "${PROXY_SH}" ]]; then
57+ load_dotenv "${PROXY_SH}"
58+ echo "[native] loaded proxy config: ${PROXY_SH}"
59+fi
60+ 
61+RUNTIME_DIR="${STACK_DIR}/.native-runtime"
62+BIN_DIR="${RUNTIME_DIR}/bin"
63+LOG_DIR="${RUNTIME_DIR}/logs"
64+RUN_DIR="${RUNTIME_DIR}/run"
65+DATA_DIR="${RUNTIME_DIR}/data"
66+GRAFANA_DIR="${RUNTIME_DIR}/grafana"
67+GRAFANA_PROVISIONING_DIR="${GRAFANA_DIR}/provisioning"
68+GRAFANA_DASHBOARD_DIR="${GRAFANA_DIR}/dashboards"
69+ 
70+mkdir -p "${BIN_DIR}" "${LOG_DIR}" "${RUN_DIR}" "${DATA_DIR}" \
71+ "${GRAFANA_PROVISIONING_DIR}/datasources" "${GRAFANA_PROVISIONING_DIR}/dashboards" "${GRAFANA_DASHBOARD_DIR}"
72+ 
73+GRAFANA_PORT="${GRAFANA_PORT:-3000}"
74+PROMETHEUS_PORT="${PROMETHEUS_PORT:-9090}"
75+TEMPO_QUERY_PORT="${TEMPO_QUERY_PORT:-3200}"
76+OTEL_GRPC_PORT="${OTEL_GRPC_PORT:-4317}"
77+OTEL_HTTP_PORT="${OTEL_HTTP_PORT:-4318}"
78+TEMPO_OTLP_GRPC_PORT="${TEMPO_OTLP_GRPC_PORT:-14317}"
79+TEMPO_OTLP_HTTP_PORT="${TEMPO_OTLP_HTTP_PORT:-14318}"
80+GF_SECURITY_ADMIN_USER="${GF_SECURITY_ADMIN_USER:-motor}"
81+GF_SECURITY_ADMIN_PASSWORD="${GF_SECURITY_ADMIN_PASSWORD:-motor}"
82+PROMETHEUS_VERSION="${PROMETHEUS_VERSION:-v2.55.1}"
83+TEMPO_VERSION="${TEMPO_VERSION:-2.6.1}"
84+OTEL_COLLECTOR_VERSION="${OTEL_COLLECTOR_VERSION:-0.115.1}"
85+GRAFANA_VERSION="${GRAFANA_VERSION:-11.3.0}"
86+ 
87+download_file() {
88+ local url="$1"
89+ local out_file="$2"
90+ if command -v curl >/dev/null 2>&1; then
91+ curl -fL --retry 3 --retry-delay 2 -o "${out_file}" "${url}"
92+ return
93+ fi
94+ if command -v wget >/dev/null 2>&1; then
95+ wget -O "${out_file}" "${url}"
96+ return
97+ fi
98+ echo "[native] neither curl nor wget is available" >&2
99+ exit 1
100+}
101+ 
102+install_prometheus() {
103+ local target="${BIN_DIR}/prometheus"
104+ [[ -x "${target}" ]] && return
105+ local ver="${PROMETHEUS_VERSION#v}"
106+ local archive="${RUNTIME_DIR}/prometheus-${ver}.tar.gz"
107+ local url="https://github.com/prometheus/prometheus/releases/download/${PROMETHEUS_VERSION}/prometheus-${ver}.linux-amd64.tar.gz"
108+ echo "[native] downloading Prometheus ${PROMETHEUS_VERSION}..."
109+ download_file "${url}" "${archive}"
110+ tar -xzf "${archive}" -C "${RUNTIME_DIR}"
111+ cp "${RUNTIME_DIR}/prometheus-${ver}.linux-amd64/prometheus" "${target}"
112+ chmod +x "${target}"
113+}
114+ 
115+install_tempo() {
116+ local target="${BIN_DIR}/tempo"
117+ [[ -x "${target}" ]] && return
118+ local ver="${TEMPO_VERSION#v}"
119+ local archive="${RUNTIME_DIR}/tempo-${ver}.tar.gz"
120+ local url="https://github.com/grafana/tempo/releases/download/v${ver}/tempo_${ver}_linux_amd64.tar.gz"
121+ echo "[native] downloading Tempo ${TEMPO_VERSION}..."
122+ download_file "${url}" "${archive}"
123+ tar -xzf "${archive}" -C "${RUNTIME_DIR}"
124+ cp "${RUNTIME_DIR}/tempo" "${target}"
125+ chmod +x "${target}"
126+}
127+ 
128+install_otel_collector() {
129+ local target="${BIN_DIR}/otelcol-contrib"
130+ [[ -x "${target}" ]] && return
131+ local ver="${OTEL_COLLECTOR_VERSION#v}"
132+ local archive="${RUNTIME_DIR}/otelcol-contrib-${ver}.tar.gz"
133+ local url="https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v${ver}/otelcol-contrib_${ver}_linux_amd64.tar.gz"
134+ echo "[native] downloading OTel Collector ${OTEL_COLLECTOR_VERSION}..."
135+ download_file "${url}" "${archive}"
136+ tar -xzf "${archive}" -C "${RUNTIME_DIR}"
137+ cp "${RUNTIME_DIR}/otelcol-contrib" "${target}"
138+ chmod +x "${target}"
139+}
140+ 
141+install_grafana() {
142+ local grafana_home="${RUNTIME_DIR}/grafana-v${GRAFANA_VERSION}"
143+ [[ -x "${grafana_home}/bin/grafana" ]] && return
144+ local archive="${RUNTIME_DIR}/grafana-${GRAFANA_VERSION}.tar.gz"
145+ local url="https://dl.grafana.com/oss/release/grafana-${GRAFANA_VERSION}.linux-amd64.tar.gz"
146+ echo "[native] downloading Grafana ${GRAFANA_VERSION}..."
147+ download_file "${url}" "${archive}"
148+ tar -xzf "${archive}" -C "${RUNTIME_DIR}"
149+}
150+ 
151+prepare_configs() {
152+ cat > "${RUNTIME_DIR}/tempo.yaml" <<EOF
153+server:
154+ http_listen_port: ${TEMPO_QUERY_PORT}
155+ grpc_listen_port: 9095
156+ 
157+distributor:
158+ receivers:
159+ otlp:
160+ protocols:
161+ grpc:
162+ endpoint: 0.0.0.0:${TEMPO_OTLP_GRPC_PORT}
163+ http:
164+ endpoint: 0.0.0.0:${TEMPO_OTLP_HTTP_PORT}
165+ 
166+ingester:
167+ trace_idle_period: 10s
168+ max_block_duration: 5m
169+ 
170+storage:
171+ trace:
172+ backend: local
173+ wal:
174+ path: ${DATA_DIR}/tempo/wal
175+ local:
176+ path: ${DATA_DIR}/tempo/traces
177+ 
178+compactor:
179+ compaction:
180+ block_retention: 72h
181+ 
182+usage_report:
183+ reporting_enabled: false
184+EOF
185+ 
186+ cat > "${RUNTIME_DIR}/otel-collector.yaml" <<EOF
187+receivers:
188+ otlp:
189+ protocols:
190+ grpc:
191+ endpoint: 0.0.0.0:${OTEL_GRPC_PORT}
192+ http:
193+ endpoint: 0.0.0.0:${OTEL_HTTP_PORT}
194+ 
195+processors:
196+ batch:
197+ send_batch_size: 1024
198+ timeout: 5s
199+ 
200+exporters:
201+ otlp/tempo:
202+ endpoint: localhost:${TEMPO_OTLP_GRPC_PORT}
203+ tls:
204+ insecure: true
205+ debug:
206+ verbosity: basic
207+ 
208+service:
209+ pipelines:
210+ traces:
211+ receivers: [otlp]
212+ processors: [batch]
213+ exporters: [otlp/tempo, debug]
214+ telemetry:
215+ logs:
216+ level: info
217+EOF
218+ 
219+ if [[ -f "${PROMETHEUS_FILE}" ]]; then
220+ cp "${PROMETHEUS_FILE}" "${RUNTIME_DIR}/prometheus.yml"
221+ else
222+ cp "./prometheus/prometheus.yml" "${RUNTIME_DIR}/prometheus.yml"
223+ fi
224+ 
225+ python3 - <<'PY'
226+from pathlib import Path
227+runtime_file = Path(".native-runtime/prometheus.yml")
228+text = runtime_file.read_text(encoding="utf-8")
229+text = text.replace("node-exporter:9100", "localhost:9100")
230+text = text.replace("cadvisor:8080", "localhost:8088")
231+runtime_file.write_text(text, encoding="utf-8")
232+PY
233+ 
234+ cat > "${GRAFANA_PROVISIONING_DIR}/datasources/datasources.yml" <<EOF
235+apiVersion: 1
236+datasources:
237+ - name: Prometheus
238+ type: prometheus
239+ uid: prometheus
240+ access: proxy
241+ url: http://localhost:${PROMETHEUS_PORT}
242+ isDefault: true
243+ jsonData:
244+ timeInterval: 5s
245+ httpMethod: POST
246+ 
247+ - name: Tempo
248+ type: tempo
249+ uid: tempo
250+ access: proxy
251+ url: http://localhost:${TEMPO_QUERY_PORT}
252+ jsonData:
253+ httpMethod: GET
254+ nodeGraph:
255+ enabled: true
256+ search:
257+ hide: false
258+EOF
259+ 
260+ cat > "${GRAFANA_PROVISIONING_DIR}/dashboards/dashboard-providers.yml" <<EOF
261+apiVersion: 1
262+providers:
263+ - name: motor-native
264+ orgId: 1
265+ folder: ""
266+ type: file
267+ disableDeletion: false
268+ editable: true
269+ updateIntervalSeconds: 30
270+ allowUiUpdates: true
271+ options:
272+ path: ${GRAFANA_DASHBOARD_DIR}
273+ foldersFromFilesStructure: false
274+EOF
275+ 
276+ cp "./grafana/dashboards/motor-all-metrics.json" "${GRAFANA_DASHBOARD_DIR}/"
277+ cp "./grafana/dashboards/motor-kv-cache.json" "${GRAFANA_DASHBOARD_DIR}/"
278+ cp "./grafana/dashboards/motor-vllm-profiling.json" "${GRAFANA_DASHBOARD_DIR}/"
279+}
280+ 
281+is_pid_running() {
282+ local pid="$1"
283+ [[ -n "${pid}" ]] && kill -0 "${pid}" >/dev/null 2>&1
284+}
285+ 
286+start_component() {
287+ local name="$1"
288+ shift
289+ local pid_file="${RUN_DIR}/${name}.pid"
290+ local log_file="${LOG_DIR}/${name}.log"
291+ if [[ -f "${pid_file}" ]]; then
292+ local old_pid
293+ old_pid="$(<"${pid_file}")"
294+ if is_pid_running "${old_pid}"; then
295+ echo "[native] ${name} already running (pid=${old_pid})"
296+ return 0
297+ fi
298+ fi
299+ nohup "$@" >"${log_file}" 2>&1 &
300+ local new_pid=$!
301+ echo "${new_pid}" > "${pid_file}"
302+ echo "[native] started ${name} (pid=${new_pid})"
303+}
304+ 
305+install_prometheus
306+install_tempo
307+install_otel_collector
308+install_grafana
309+prepare_configs
310+ 
311+mkdir -p "${DATA_DIR}/tempo" "${DATA_DIR}/prometheus" "${DATA_DIR}/grafana"
312+ 
313+start_component "tempo" \
314+ "${BIN_DIR}/tempo" \
315+ "-config.file=${RUNTIME_DIR}/tempo.yaml"
316+ 
317+start_component "otel-collector" \
318+ "${BIN_DIR}/otelcol-contrib" \
319+ "--config=${RUNTIME_DIR}/otel-collector.yaml"
320+ 
321+start_component "prometheus" \
322+ "${BIN_DIR}/prometheus" \
323+ "--config.file=${RUNTIME_DIR}/prometheus.yml" \
324+ "--storage.tsdb.path=${DATA_DIR}/prometheus" \
325+ "--storage.tsdb.retention.time=72h" \
326+ "--web.listen-address=:${PROMETHEUS_PORT}" \
327+ "--web.enable-lifecycle" \
328+ "--web.enable-remote-write-receiver" \
329+ "--enable-feature=utf8-names"
330+ 
331+GRAFANA_HOME="${RUNTIME_DIR}/grafana-v${GRAFANA_VERSION}"
332+start_component "grafana" \
333+ env \
334+ GF_SECURITY_ADMIN_USER="${GF_SECURITY_ADMIN_USER}" \
335+ GF_SECURITY_ADMIN_PASSWORD="${GF_SECURITY_ADMIN_PASSWORD}" \
336+ GF_USERS_ALLOW_SIGN_UP="false" \
337+ GF_LOG_LEVEL="warn" \
338+ GF_SERVER_HTTP_PORT="${GRAFANA_PORT}" \
339+ GF_PATHS_PROVISIONING="${GRAFANA_PROVISIONING_DIR}" \
340+ GF_PATHS_DATA="${DATA_DIR}/grafana" \
341+ "${GRAFANA_HOME}/bin/grafana" server --homepath "${GRAFANA_HOME}"
342+ 
343+cat <<EOF
344+ 
345+================================================================
346+pyMotor observability stack is running in native mode.
347+ 
348+ Grafana http://localhost:${GRAFANA_PORT} (user: ${GF_SECURITY_ADMIN_USER} / pass: ${GF_SECURITY_ADMIN_PASSWORD})
349+ Prometheus http://localhost:${PROMETHEUS_PORT}
350+ Tempo http://localhost:${TEMPO_QUERY_PORT}
351+ OTel OTLP localhost:${OTEL_GRPC_PORT} (gRPC) / ${OTEL_HTTP_PORT} (HTTP)
352+ 
353+Runtime dir: ${RUNTIME_DIR}
354+Logs dir: ${LOG_DIR}
355+================================================================
356+EOF
@@ -0,0 +1,38 @@
1+#!/usr/bin/env bash
2+ 
3+set -euo pipefail
4+ 
5+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
6+STACK_DIR="$(cd -- "${SCRIPT_DIR}/.." &>/dev/null && pwd)"
7+RUN_DIR="${STACK_DIR}/generated"
8+PID_FILE="${RUN_DIR}/k8s-port-forwards.pids"
9+META_FILE="${RUN_DIR}/k8s-port-forwards.meta"
10+ 
11+is_pid_running() {
12+ local pid="$1"
13+ [[ -n "${pid}" ]] && kill -0 "${pid}" >/dev/null 2>&1
14+}
15+ 
16+if [[ -f "${PID_FILE}" ]]; then
17+ while IFS= read -r pid; do
18+ [[ -z "${pid}" ]] && continue
19+ if is_pid_running "${pid}"; then
20+ echo "[port-forward-stop] stopping pid=${pid}"
21+ kill "${pid}" || true
22+ fi
23+ done <"${PID_FILE}"
24+fi
25+ 
26+if command -v pgrep >/dev/null 2>&1; then
27+ while IFS= read -r pid; do
28+ [[ -z "${pid}" || "${pid}" == "$$" ]] && continue
29+ args="$(ps -p "${pid}" -o args= 2>/dev/null || true)"
30+ [[ "${args}" == *"tcp-forward.py"* ]] || continue
31+ [[ "${args}" == *"${SCRIPT_DIR}/tcp-forward.py"* || "${args}" == *" tcp-forward.py"* ]] || continue
32+ echo "[port-forward-stop] stopping orphan pid=${pid}"
33+ kill "${pid}" || true
34+ done < <(pgrep -f "tcp-forward.py" || true)
35+fi
36+ 
37+rm -f "${PID_FILE}" "${META_FILE}"
38+echo "[port-forward-stop] done."
@@ -0,0 +1,77 @@
1+#!/usr/bin/env python3
2+# -*- coding: utf-8 -*-
3+ 
4+"""Small stdlib TCP forwarder for Docker-to-PodIP bridge targets."""
5+ 
6+from __future__ import annotations
7+ 
8+import argparse
9+import selectors
10+import socket
11+import sys
12+import threading
13+from typing import Tuple
14+ 
15+ 
16+def _pipe(left: socket.socket, right: socket.socket) -> None:
17+ selector = selectors.DefaultSelector()
18+ selector.register(left, selectors.EVENT_READ, right)
19+ selector.register(right, selectors.EVENT_READ, left)
20+ try:
21+ while True:
22+ for key, _ in selector.select():
23+ src = key.fileobj
24+ dst = key.data
25+ data = src.recv(65536)
26+ if not data:
27+ return
28+ dst.sendall(data)
29+ finally:
30+ selector.close()
31+ left.close()
32+ right.close()
33+ 
34+ 
35+def _handle(client: socket.socket, addr: Tuple[str, int], target_host: str, target_port: int) -> None:
36+ del addr
37+ try:
38+ upstream = socket.create_connection((target_host, target_port), timeout=5)
39+ except OSError:
40+ client.close()
41+ return
42+ _pipe(client, upstream)
43+ 
44+ 
45+def main() -> int:
46+ parser = argparse.ArgumentParser(description="Forward one local TCP port to one remote target.")
47+ parser.add_argument("--listen-host", default="0.0.0.0")
48+ parser.add_argument("--listen-port", type=int, required=True)
49+ parser.add_argument("--target-host", required=True)
50+ parser.add_argument("--target-port", type=int, required=True)
51+ args = parser.parse_args()
52+ 
53+ server = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
54+ server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
55+ server.bind((args.listen_host, args.listen_port))
56+ server.listen(128)
57+ print(
58+ f"[tcp-forward] {args.listen_host}:{args.listen_port} -> {args.target_host}:{args.target_port}",
59+ flush=True,
60+ )
61+ try:
62+ while True:
63+ client, addr = server.accept()
64+ thread = threading.Thread(
65+ target=_handle,
66+ args=(client, addr, args.target_host, args.target_port),
67+ daemon=True,
68+ )
69+ thread.start()
70+ except KeyboardInterrupt:
71+ return 0
72+ finally:
73+ server.close()
74+ 
75+ 
76+if __name__ == "__main__":
77+ sys.exit(main())
@@ -0,0 +1,81 @@
1+#!/usr/bin/env bash
2+# 验证 observability stack 的 Tracing 通路:OTel Collector (:4317) → Tempo (:3200)
3+#
4+# 用法(stack 已 ./start.sh 启动后):
5+# ./scripts/verify-tracing.sh
6+# OTEL_HOST=127.0.0.1 ./scripts/verify-tracing.sh
7+#
8+# 依赖:python3 + opentelemetry-exporter-otlp-proto-grpc
9+# pip install opentelemetry-exporter-otlp-proto-grpc opentelemetry-sdk
10+ 
11+set -euo pipefail
12+ 
13+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
14+STACK_DIR="$(cd "${SCRIPT_DIR}/.." && pwd)"
15+ 
16+OTEL_HOST="${OTEL_HOST:-127.0.0.1}"
17+OTEL_GRPC_PORT="${OTEL_GRPC_PORT:-4317}"
18+TEMPO_PORT="${TEMPO_QUERY_PORT:-3200}"
19+SERVICE_NAME="${SERVICE_NAME:-pymotor-tracing-verify}"
20+ 
21+if [[ -f "${STACK_DIR}/.env" ]]; then
22+ OTEL_GRPC_PORT="$(grep -E '^OTEL_GRPC_PORT=' "${STACK_DIR}/.env" 2>/dev/null | tail -n1 | cut -d= -f2 || true)"
23+ OTEL_GRPC_PORT="${OTEL_GRPC_PORT:-4317}"
24+ TEMPO_PORT="$(grep -E '^TEMPO_QUERY_PORT=' "${STACK_DIR}/.env" 2>/dev/null | tail -n1 | cut -d= -f2 || true)"
25+ TEMPO_PORT="${TEMPO_PORT:-3200}"
26+fi
27+ 
28+echo "[verify-tracing] sending test span to ${OTEL_HOST}:${OTEL_GRPC_PORT} (service=${SERVICE_NAME})"
29+ 
30+if ! python3 -c "from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter" 2>/dev/null; then
31+ echo "[verify-tracing] installing opentelemetry packages..." >&2
32+ pip install -q opentelemetry-exporter-otlp-proto-grpc opentelemetry-sdk
33+fi
34+ 
35+python3 - "${OTEL_HOST}" "${OTEL_GRPC_PORT}" "${SERVICE_NAME}" <<'PY'
36+import sys
37+import time
38+import uuid
39+ 
40+from opentelemetry import trace
41+from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
42+from opentelemetry.sdk.resources import Resource
43+from opentelemetry.sdk.trace import TracerProvider
44+from opentelemetry.sdk.trace.export import BatchSpanProcessor
45+ 
46+host, port, service_name = sys.argv[1:4]
47+endpoint = f"{host}:{port}"
48+ 
49+resource = Resource.create({"service.name": service_name})
50+provider = TracerProvider(resource=resource)
51+provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint=endpoint, insecure=True)))
52+trace.set_tracer_provider(provider)
53+ 
54+tracer = trace.get_tracer("pymotor.observability.verify")
55+with tracer.start_as_current_span("verify-tracing-span") as span:
56+ span.set_attribute("verify", True)
57+ span.set_attribute("stack", "pymotor-observability")
58+ time.sleep(0.05)
59+ 
60+provider.force_flush(timeout_millis=5000)
61+provider.shutdown()
62+print(f"[verify-tracing] span exported to grpc://{endpoint}")
63+PY
64+ 
65+echo "[verify-tracing] waiting for Tempo ingest..."
66+sleep 3
67+ 
68+SEARCH_URL="http://${OTEL_HOST}:${TEMPO_PORT}/api/search?limit=20"
69+echo "[verify-tracing] querying Tempo: ${SEARCH_URL}"
70+RESP="$(curl -sf "${SEARCH_URL}" || true)"
71+ 
72+if echo "${RESP}" | grep -q "${SERVICE_NAME}"; then
73+ echo "[verify-tracing] OK — trace visible in Tempo (service.name=${SERVICE_NAME})"
74+ echo "[verify-tracing] open Grafana → Explore → Tempo, search service: ${SERVICE_NAME}"
75+ exit 0
76+fi
77+ 
78+echo "[verify-tracing] WARN — span sent but not yet found in Tempo search response." >&2
79+echo "[verify-tracing] response snippet: $(echo "${RESP}" | head -c 200)" >&2
80+echo "[verify-tracing] check: docker compose logs otel-collector tempo | tail -50" >&2
81+exit 1
@@ -0,0 +1,8 @@
1+#!/usr/bin/env bash
2+ 
3+set -euo pipefail
4+ 
5+SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" &>/dev/null && pwd)"
6+cd "${SCRIPT_DIR}"
7+ 
8+exec ./launch.sh "$@"