Files
og/docs/WORKER-TELEMETRY.md
2026-09-11 06:14:38 +02:00

128 lines
4.8 KiB
Markdown

# Remote worker telemetry
Ollama Fair Gateway can consume optional out-of-band resource telemetry through `workers[].telemetry_url`. Checkpoint 20 ships a small `worker-telemetry` agent so remote Ollama hosts can report their own RAM/GPU state instead of accidentally reporting the gateway host through `local_system_stats`.
The agent is **not** part of the inference request path. Gateway workers fetch it during the normal health/telemetry refresh cycle.
## Why this matters for remote workers
`local_system_stats: true` runs the memory collector inside the gateway process. It is correct only when the Ollama worker and gateway share the same host. For a worker such as `http://10.2.10.48:11434`, it otherwise reports the gateway machine's RAM as if it belonged to `10.2.10.48`.
For remote workers, use:
```json
{
"name": "M75q - Gen5 - 1048",
"url": "http://10.2.10.48:11434",
"local_system_stats": false,
"telemetry_url": "http://10.2.10.48:11500/telemetry"
}
```
and run the agent on that Ollama host.
## Validate locally first
Host memory only:
```sh
./ollama-gateway-worker-telemetry-linux-amd64 -once
```
AMD GPU through the Linux amdgpu sysfs interface:
```sh
./ollama-gateway-worker-telemetry-linux-amd64 -once -amd-sysfs
```
If auto-detection picks the wrong GPU, specify the device explicitly:
```sh
./ollama-gateway-worker-telemetry-linux-amd64 \
-once \
-amd-sysfs \
-amd-device /sys/class/drm/card1/device
```
NVIDIA:
```sh
./ollama-gateway-worker-telemetry-linux-amd64 -once -nvidia-smi
```
The JSON schema matches the existing gateway `telemetry_url` contract:
```json
{
"memory_used_bytes": 123,
"memory_total_bytes": 456,
"vram_used_bytes": 789,
"vram_total_bytes": 1024,
"gpu_utilization_percent": 42,
"gpu_temperature_c": 63.5,
"gpu_power_watts": 88,
"source": "host-memory+amdgpu-sysfs",
"updated_at": "2026-09-08T18:00:00Z"
}
```
A field may be absent when the operating system/driver does not expose it. The agent still returns the remaining usable metrics and puts collector failures in the `error` field.
## Run as a service
The listener defaults to loopback for safety. To expose it to the gateway, bind a worker interface and restrict clients by CIDR. Prefer an exact gateway host address when possible:
```sh
./ollama-gateway-worker-telemetry-linux-amd64 \
-listen 0.0.0.0:11500 \
-allow-cidrs 10.2.19.1/32 \
-amd-sysfs
```
Example systemd unit:
```ini
[Unit]
Description=Ollama Gateway Worker Telemetry
After=network-online.target
[Service]
ExecStart=/usr/local/bin/ollama-gateway-worker-telemetry-linux-amd64 -listen 0.0.0.0:11500 -allow-cidrs 10.2.19.1/32 -amd-sysfs
Restart=on-failure
User=nobody
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
[Install]
WantedBy=multi-user.target
```
Adjust the CIDR to the source IP the worker actually sees from the gateway. Do not use `0.0.0.0/0` merely to make the endpoint reachable.
## AMD telemetry details
On Linux the agent auto-detects the first `/sys/class/drm/card*/device` whose PCI vendor is `0x1002`. It reads standard amdgpu attributes when available:
- `mem_info_vram_total`
- `mem_info_vram_used`
- `gpu_busy_percent`
- `hwmon/*/temp1_input`
- `hwmon/*/power1_average`
No ROCm library, `rocm-smi`, NVML or CGO is required for the AMD path. Integrated/shared-memory GPUs may expose different or incomplete VRAM semantics; keep the statically configured `memory_capacity_bytes`/`vram_capacity_bytes` values as hard capacity hints and treat runtime telemetry as routing pressure information.
## Security
The agent intentionally has no application credential store. It relies on a narrow listener/firewall/CIDR allowlist so it remains dependency-free and does not introduce a second secret lifecycle. Resource telemetry can still reveal infrastructure details, so expose it only on a trusted management network, VPN or host firewall. For untrusted network segments, put the endpoint behind a TLS-authenticated reverse proxy rather than exposing it directly.
`workers[].telemetry_url` remains administrator-controlled and is an SSRF-capable URL. Point it only to trusted worker exporters.
## Freshness and caching
Checkpoint 21 validates `updated_at` when the external exporter supplies it. The shipped agent always sends the field and also returns `Cache-Control: no-store`.
A sample is rejected when it is older than `max(30s, 6 x workers[].health_interval)` or more than 30 seconds in the future. Rejected values are not merged into the routing telemetry; the reason is exposed through the worker telemetry error instead. This prevents a caching proxy or badly skewed exporter clock from silently presenting old resource pressure as current data.
For compatibility, third-party exporters that omit `updated_at` are still accepted. Their freshness cannot be verified, so new integrations should always provide an RFC3339 timestamp.