-
This commit is contained in:
@@ -0,0 +1,78 @@
|
||||
# Alerts and signed webhooks
|
||||
|
||||
The alert manager evaluates bounded operational conditions and persists firing/resolved history. It never includes prompt, response, tool, or conversation content in webhook payloads.
|
||||
|
||||
## Conditions
|
||||
|
||||
Supported conditions include:
|
||||
|
||||
- worker unhealthy for `worker_down_for`;
|
||||
- circuit breaker open;
|
||||
- queue depth threshold;
|
||||
- oldest queue wait threshold;
|
||||
- VRAM pressure;
|
||||
- repeated OOM indication from the worker/circuit error;
|
||||
- local gateway storage size;
|
||||
- actor or tenant quota remaining percentage.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"alerts": {
|
||||
"enabled": true,
|
||||
"evaluation_interval": "15s",
|
||||
"cooldown": "5m",
|
||||
"history_limit": 500,
|
||||
"webhook_timeout": "5s",
|
||||
"webhook_max_concurrent": 4,
|
||||
"webhook_queue": 1024,
|
||||
"webhook_retry_attempts": 3,
|
||||
"webhook_retry_backoff": "500ms",
|
||||
"thresholds": {
|
||||
"worker_down_for": "30s",
|
||||
"circuit_open": true,
|
||||
"queue_depth": 100,
|
||||
"queue_wait": "30s",
|
||||
"vram_percent": 95,
|
||||
"storage_bytes": 0,
|
||||
"quota_remaining_percent": 10,
|
||||
"oom": true
|
||||
},
|
||||
"webhooks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Webhook delivery
|
||||
|
||||
A webhook payload has a stable event ID and is delivered asynchronously from a bounded queue. Transient transport failures, HTTP 429, and HTTP 5xx are retried with bounded exponential backoff. Other HTTP 4xx responses are treated as permanent failures.
|
||||
|
||||
Headers:
|
||||
|
||||
- `X-Ollama-Gateway-Event-ID`
|
||||
- `X-Ollama-Gateway-Timestamp`
|
||||
- `X-Ollama-Gateway-Delivery-Attempt`
|
||||
- `X-Ollama-Gateway-Signature` when a secret is configured
|
||||
|
||||
Signature format:
|
||||
|
||||
```text
|
||||
sha256=HMAC_SHA256(secret, timestamp + "." + raw_json_body)
|
||||
```
|
||||
|
||||
Consumers should deduplicate by event ID because retries intentionally reuse the same event ID.
|
||||
|
||||
## Hot-path isolation
|
||||
|
||||
Quota observations happen during request admission, but disk persistence and webhook I/O are not performed synchronously on that request path. State persistence is coalesced asynchronously and webhook delivery uses a bounded queue.
|
||||
|
||||
## Secrets
|
||||
|
||||
Webhook secrets are redacted from the Admin JSON configuration view. OpenTelemetry header values are redacted there as well. Saving the redacted configuration restores the existing secret values instead of persisting the literal `<redacted>` placeholder.
|
||||
|
||||
## Metrics
|
||||
|
||||
- `ollama_gateway_queue_oldest_wait_seconds`
|
||||
- `ollama_gateway_alerts_active`
|
||||
- `ollama_gateway_alerts_last_evaluate_timestamp_seconds`
|
||||
@@ -0,0 +1,77 @@
|
||||
# Architecture
|
||||
|
||||
## Request path
|
||||
|
||||
```text
|
||||
HTTP request
|
||||
-> authentication / identity
|
||||
-> compute request parsing + cost estimate
|
||||
-> in-memory quota reservation
|
||||
-> hierarchical in-memory WFQ
|
||||
-> model-aware worker selection
|
||||
-> streaming reverse proxy
|
||||
-> actual usage reconciliation
|
||||
-> in-memory usage aggregation
|
||||
-> optional asynchronous JSONL journal
|
||||
```
|
||||
|
||||
Management and blob endpoints do not enter the compute scheduler and keep their request body streaming end-to-end.
|
||||
|
||||
## Single authoritative engine
|
||||
|
||||
The gateway deliberately has no external coordination backend. All mutable control state belongs to one process. This means there is exactly one authoritative view of queue order, quota buckets and worker slots and therefore no distributed lock, lease renewal, consensus or fail-open state split.
|
||||
|
||||
The hot-path structures are Go-native:
|
||||
|
||||
- scheduler maps and heaps protected by one scheduler mutex;
|
||||
- worker load using atomic counters;
|
||||
- quota buckets protected by a small mutex;
|
||||
- runtime policy and session maps behind RWMutex/Mutex;
|
||||
- live request snapshots behind the lifecycle tracker;
|
||||
- usage summaries updated in memory and journaled asynchronously.
|
||||
|
||||
## Fair scheduler
|
||||
|
||||
`internal/scheduler` implements two-level hierarchical weighted fair queueing.
|
||||
|
||||
A tenant has a root service score. Inside that tenant each queued request has an actor virtual-finish score derived from request cost and actor weight. When capacity becomes free, the scheduler selects the tenant with the lowest root score and then its request with the lowest actor finish score.
|
||||
|
||||
This yields two properties:
|
||||
|
||||
- tenant fairness is not defeated by creating many actors;
|
||||
- actors within one tenant still share service fairly.
|
||||
|
||||
The scheduler is event-driven. A worker slot release signals the scheduler immediately; there is no polling backend.
|
||||
|
||||
## Quotas
|
||||
|
||||
`internal/quota` is an in-memory hierarchical token-bucket ledger whose bucket balances are periodically snapshotted for restart continuity. A request reserves estimated compute credits from actor and tenant buckets. Final usage reconciles the reservation. Token buckets refill lazily from monotonic wall-clock deltas.
|
||||
|
||||
## Worker routing
|
||||
|
||||
`internal/worker` maintains health and model residency snapshots for each Ollama endpoint. Local `active` counters are atomic and are the hard worker slot limiter. Candidate scoring combines active-slot pressure with a strong affinity bonus when the requested model is already loaded.
|
||||
|
||||
There are no worker leases or renewal goroutines.
|
||||
|
||||
## Live flow and infrastructure map
|
||||
|
||||
`internal/liveflow` tracks bounded prompt-free lifecycle metadata. `internal/infrastructure` turns that local state into the UI topology snapshot. The browser receives it over SSE and renders the animated pulse map with Canvas.
|
||||
|
||||
The infrastructure view is explicitly local to the process. It shows one gateway node plus all configured Ollama workers and models.
|
||||
|
||||
## Runtime policy store and sessions
|
||||
|
||||
Tenant policy overrides are backed by the local persistent policy store. Browser OIDC sessions remain in memory and disappear on restart by design.
|
||||
|
||||
## Usage
|
||||
|
||||
`internal/usage` maintains actor, tenant and global summaries in memory. The optional file journal runs asynchronously on a buffered channel and never participates in admission or request dispatch.
|
||||
|
||||
## Scaling model
|
||||
|
||||
Scale inference by adding Ollama workers to the one gateway. Running multiple independent gateway replicas would create independent fairness domains. If strict global fairness is required, requests must pass through the same gateway engine.
|
||||
|
||||
|
||||
## Optional content-bearing conversation state
|
||||
|
||||
The Responses conversation layer is separate from browser OIDC sessions. Browser sessions remain volatile. When explicitly enabled, Responses conversation contexts are tenant/actor scoped, encrypted at rest, retention bounded, and used only to expand `previous_response_id` before normal admission/routing. The ordinary proxy path does not capture response bodies; bounded capture is enabled only for store-eligible `/v1/responses` requests.
|
||||
@@ -0,0 +1,140 @@
|
||||
# Durable batch jobs
|
||||
|
||||
Durable batch jobs are the P3.1 background-execution layer. They are **disabled by default** because each submitted job intentionally persists its request body and, when available, its response body.
|
||||
|
||||
## Configuration
|
||||
|
||||
```json
|
||||
"batch_jobs": {
|
||||
"enabled": false,
|
||||
"retention": "168h",
|
||||
"max_jobs": 1000,
|
||||
"max_concurrent": 1,
|
||||
"max_input_bytes": 16777216
|
||||
}
|
||||
```
|
||||
|
||||
The service class `batch` must also exist under `service_classes.classes`. `batch_jobs.max_input_bytes` must not exceed `server.max_body_bytes`.
|
||||
|
||||
Storage paths are bootstrap-only:
|
||||
|
||||
```json
|
||||
"storage": {
|
||||
"batch_jobs_file": "batch-jobs.json",
|
||||
"batch_jobs_dir": "batch"
|
||||
}
|
||||
```
|
||||
|
||||
## Client API
|
||||
|
||||
A batch job wraps one normal configured compute `POST` endpoint. The request body is the JSON body that the gateway will replay later.
|
||||
|
||||
```http
|
||||
POST /gateway/v1/batches
|
||||
Authorization: Bearer <credential>
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"path": "/api/chat",
|
||||
"body": {
|
||||
"model": "qwen3:8b",
|
||||
"messages": [{"role":"user","content":"Summarize the report"}],
|
||||
"stream": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
A successful submission returns `202 Accepted`, a `Location: /gateway/v1/batches/<id>` header and the durable job metadata.
|
||||
|
||||
Owner-scoped endpoints:
|
||||
|
||||
```text
|
||||
GET /gateway/v1/batches
|
||||
GET /gateway/v1/batches/<id>
|
||||
GET /gateway/v1/batches/<id>/output
|
||||
POST /gateway/v1/batches/<id>/pause
|
||||
POST /gateway/v1/batches/<id>/resume
|
||||
POST /gateway/v1/batches/<id>/cancel
|
||||
```
|
||||
|
||||
A caller can only see or control jobs created by the same authenticated tenant and scheduler actor. A cross-identity lookup is returned as not found.
|
||||
|
||||
## Admin UI and API
|
||||
|
||||
Administrators have a dedicated **Batch Jobs** page with all durable jobs, state, attempts, identity metadata, model/path, pause/resume/cancel controls and authenticated output download.
|
||||
|
||||
Admin endpoints require `gateway:admin`:
|
||||
|
||||
```text
|
||||
GET /gateway/ui-api/batches
|
||||
GET /gateway/ui-api/batches/<id>
|
||||
GET /gateway/ui-api/batches/<id>/output
|
||||
POST /gateway/ui-api/batches/<id>/pause
|
||||
POST /gateway/ui-api/batches/<id>/resume
|
||||
POST /gateway/ui-api/batches/<id>/cancel
|
||||
```
|
||||
|
||||
## Execution semantics
|
||||
|
||||
Every attempt is replayed into the same gateway compute path as an interactive request, with the original identity metadata and a forced service class of `batch`. Therefore the attempt still passes through:
|
||||
|
||||
- tenant model ACLs and the submitted key-specific ACL snapshot;
|
||||
- alias resolution, capability preflight, placement, maintenance state and circuit eligibility;
|
||||
- quota reservation/reconciliation;
|
||||
- weighted fair scheduling and the configured `batch` class concurrency ceiling;
|
||||
- adaptive worker routing and normal safe pre-stream retry behavior;
|
||||
- usage metering, request IDs, alerts and OpenTelemetry.
|
||||
|
||||
The submission credential itself is never stored. The durable identity snapshot contains tenant/subject/application/scopes and the key-specific model ACL needed to execute an already accepted job. Current tenant policy, worker state, placement, quota state and routing configuration are evaluated again at execution time. Revoking the submitting API key does not retroactively cancel an already accepted durable job; cancel the job explicitly if that is required operationally.
|
||||
|
||||
## States
|
||||
|
||||
```text
|
||||
queued -> running -> completed
|
||||
| | \
|
||||
| | -> failed
|
||||
| -> pausing -> paused -> queued
|
||||
-> paused
|
||||
-> cancelled
|
||||
running/pausing -> cancelling -> cancelled
|
||||
```
|
||||
|
||||
`attempts` increments when an execution attempt starts. A failed gateway/backend response is terminal; the durable batch manager itself does not blindly retry application failures.
|
||||
|
||||
## Restart and shutdown behavior
|
||||
|
||||
The metadata snapshot is written atomically. Input/output payloads live in separate spool files referenced from metadata.
|
||||
|
||||
On startup:
|
||||
|
||||
- `running` becomes `queued` and is eligible for retry;
|
||||
- `pausing` becomes `paused`;
|
||||
- `cancelling` becomes `cancelled`.
|
||||
|
||||
During graceful shutdown the root context cancels active attempts, and the process waits for those attempts to persist their restart-safe `queued` state before final shutdown accounting is closed. A hard kill can still leave metadata in `running`; startup recovery converts that state to `queued`.
|
||||
|
||||
This gives durable batches **at-least-once execution across an interruption boundary**, not exactly-once execution. If a worker completed side effects but the gateway did not durably commit the batch result before shutdown/crash, that job can run again after restart. Use idempotent batch workloads or application-level idempotency keys when duplicate execution would be harmful.
|
||||
|
||||
## Storage, privacy and retention
|
||||
|
||||
Layout:
|
||||
|
||||
```text
|
||||
<data_dir>/batch-jobs.json
|
||||
<data_dir>/batch/input/<batch-id>.json
|
||||
<data_dir>/batch/output/<batch-id>.response
|
||||
```
|
||||
|
||||
Metadata and spool payload files are created with mode `0600`; spool directories use `0750`. Input/output references are stored in `batch-jobs.json`, not the potentially large payloads themselves.
|
||||
|
||||
**Batch payloads are content-bearing and are not encrypted by the gateway.** This differs from the optional encrypted conversation store. Keep `batch_jobs.enabled=false` unless durable content persistence is intended, protect `storage.data_dir` with filesystem/disk encryption and access controls, and treat gateway backups as sensitive.
|
||||
|
||||
Terminal jobs and their input/output files are removed after `batch_jobs.retention`. `max_jobs` caps retained metadata/jobs; `max_input_bytes` caps one submitted request body.
|
||||
|
||||
## Accounting
|
||||
|
||||
Each execution attempt receives its own normal `X-Request-ID`. The durable job stores the latest execution request ID and HTTP status. The attempt appears in the ordinary usage journal with `service_class: "batch"`; prompt/completion tokens and compute credits therefore use the same accounting and rollup path as other inference traffic.
|
||||
|
||||
## Current scope
|
||||
|
||||
P3.1 deliberately implements one durable compute request per job. It is not an OpenAI Batch API clone and does not yet ingest multi-request JSONL files. Global batch coordination across multiple gateway replicas remains intentionally deferred outside the active roadmap; the current design stays single-process and durable.
|
||||
@@ -0,0 +1,69 @@
|
||||
# Context / `num_ctx` hardening
|
||||
|
||||
Checkpoint 27 treats context size as a runtime resource rather than only a model capability.
|
||||
|
||||
## Effective context
|
||||
|
||||
For requests that cannot set native Ollama `options.num_ctx` (OpenAI `/v1/*`, Responses and Anthropic-compatible routes), each worker gets an effective context window from the strongest available evidence:
|
||||
|
||||
1. loaded `/api/ps` `context_length`;
|
||||
2. Modelfile `num_ctx` returned by `/api/show` `parameters`;
|
||||
3. `workers[].default_context_tokens`;
|
||||
4. `model_capabilities.context.default_worker_tokens`.
|
||||
|
||||
The result is clamped by the theoretical model maximum and `workers[].context_limits`. The request is routed only to workers that can hold the estimated input + output budget. If no worker can hold it and `context_guard` is `reject`, admission fails before queueing or inference.
|
||||
|
||||
## Native `options.num_ctx`
|
||||
|
||||
`options.num_ctx` is accepted only on native Ollama `/api/*` requests. It must be a positive JSON integer. The gateway checks it against:
|
||||
|
||||
- `model_capabilities.context.max_requested_tokens` (default 32768; `-1` removes the gateway cap);
|
||||
- the model maximum discovered from `/api/show`;
|
||||
- a matching `workers[].context_limits` entry.
|
||||
|
||||
A model already loaded with a smaller context may still receive a larger explicit `num_ctx` because Ollama can resize/reload it, but that worker loses the loaded-model routing bonus for that request.
|
||||
|
||||
## Conservative request estimate
|
||||
|
||||
The context guard now includes:
|
||||
|
||||
- `prompt`, messages and `system`;
|
||||
- Responses `instructions` and `input`;
|
||||
- native generate `suffix`;
|
||||
- tool schema JSON;
|
||||
- `max_tokens`, `max_completion_tokens`, `max_output_tokens` and native `num_predict`;
|
||||
- a configurable reserve per image (`vision_reserve_tokens_per_image`);
|
||||
- a configurable input safety margin (`estimation_margin_percent`).
|
||||
|
||||
The estimator intentionally remains provider-independent and is not an exact tokenizer. The margin and image reserve compensate for that uncertainty.
|
||||
|
||||
## Recommended configuration
|
||||
|
||||
```json
|
||||
"model_capabilities": {
|
||||
"mode": "enforce",
|
||||
"cache_ttl": "10m",
|
||||
"context_guard": "reject",
|
||||
"context": {
|
||||
"max_requested_tokens": 32768,
|
||||
"default_worker_tokens": 4096,
|
||||
"estimation_margin_percent": 15,
|
||||
"vision_reserve_tokens_per_image": 2048
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For a worker with a known different Ollama default or a stricter memory envelope:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "worker-a",
|
||||
"default_context_tokens": 4096,
|
||||
"context_limits": {
|
||||
"large-*": 16384,
|
||||
"*": 32768
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Use a Modelfile `num_ctx` when OpenAI-compatible clients need a larger stable context, because those APIs do not provide native per-request `options.num_ctx` semantics.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Optional encrypted Responses conversations
|
||||
|
||||
`conversations` implements the P2.3 stateful layer for clients that use the OpenAI Responses API with `previous_response_id`.
|
||||
|
||||
The feature is **disabled by default**. When disabled, the gateway does not parse, capture, persist, or reinterpret response content for conversation purposes; `/v1/responses` keeps its normal passthrough behavior.
|
||||
|
||||
## Configuration
|
||||
|
||||
```json
|
||||
"conversations": {
|
||||
"enabled": false,
|
||||
"encryption_key": "${GATEWAY_CONVERSATION_KEY}",
|
||||
"retention": "24h",
|
||||
"max_entries": 1000,
|
||||
"max_content_bytes": 2097152
|
||||
},
|
||||
"storage": {
|
||||
"conversations_file": "conversations.enc.json"
|
||||
}
|
||||
```
|
||||
|
||||
When enabled:
|
||||
|
||||
- `encryption_key` must contain at least 32 characters. Use an environment variable or another deployment secret source; do not commit it.
|
||||
- `retention` is independent from usage-journal retention because this store contains prompt/output content.
|
||||
- `max_entries` bounds the number of stored response contexts; oldest entries are evicted first.
|
||||
- `max_content_bytes` bounds one flattened conversation context and also bounds the response capture used to construct it.
|
||||
|
||||
The configured key is SHA-256-derived into an AES-256 key. The state snapshot is encrypted with AES-256-GCM and written mode `0600` through the same fsync + atomic-replace mechanism used by the other local state files. The on-disk JSON envelope contains only version/algorithm metadata, nonce, ciphertext, and update time; conversation content is not plaintext on disk.
|
||||
|
||||
Changing or losing the encryption key makes the existing store unreadable. Startup fails closed if an enabled store cannot be decrypted.
|
||||
|
||||
## `previous_response_id` behavior
|
||||
|
||||
For an enabled store, a successful `/v1/responses` request is stored when the request does not set `"store": false`.
|
||||
|
||||
For a later request containing `previous_response_id`:
|
||||
|
||||
1. the ID is looked up only inside the authenticated tenant + actor boundary;
|
||||
2. the prior flattened input/output item context is loaded;
|
||||
3. the current `input` is appended;
|
||||
4. `previous_response_id` is removed before forwarding upstream;
|
||||
5. the resulting complete `input` is sent through the normal ACL, preflight, quota, queue, placement, routing and proxy path.
|
||||
|
||||
A missing ID and an ID owned by another tenant/actor are intentionally indistinguishable to the caller. Both fail as an invalid previous response reference, preventing cross-identity enumeration.
|
||||
|
||||
String `input` values are normalized to a user message item when history must be flattened. Existing array/object inputs are retained as items. Previous `instructions` are not copied into the stored conversation context; only input/output items are chained.
|
||||
|
||||
## Streaming
|
||||
|
||||
The proxy normally does not retain response bodies. Conversation capture is activated only for `/v1/responses` when conversations are enabled and the request is eligible for storage.
|
||||
|
||||
For streaming Responses, the gateway reconstructs persisted output from `response.completed` (or completed output-item events when necessary). Client streaming remains unchanged. If the capture exceeds `max_content_bytes`, the inference response still succeeds but that response is not stored for future chaining.
|
||||
|
||||
## Retention and deletion
|
||||
|
||||
Expired entries are pruned on reads/writes and by a background cleanup loop. The cleanup interval is derived from the configured retention and capped at 15 minutes, so content does not depend on future traffic to expire.
|
||||
|
||||
`store: false` is the per-request opt-out. It still allows a request to consume a valid prior response context, but the new response is not persisted as the next link.
|
||||
|
||||
The store is intentionally not exposed as a content browser in the Admin UI. The persistence page reports only operational metadata such as whether it is enabled and the number of retained entries.
|
||||
|
||||
## Backups and threat model
|
||||
|
||||
The Admin storage backup includes the encrypted conversation state file when present. Treat backups as sensitive anyway: gateway configuration backups can contain deployment secrets, and possession of both the conversation ciphertext and its encryption key allows decryption.
|
||||
|
||||
This feature protects content at rest from casual filesystem disclosure. It does not protect content from a compromised running gateway process, from an administrator who possesses the configured key, or from the Ollama worker that receives the expanded prompt context.
|
||||
@@ -0,0 +1,117 @@
|
||||
# Deployment hardening and preflight
|
||||
|
||||
Checkpoint 19 adds deployment checks intended to catch configuration/mount mistakes before production traffic reaches the gateway. These checks do not participate in the inference hot path.
|
||||
|
||||
## Effective configuration preflight
|
||||
|
||||
Run the gateway binary with `-check-config` before starting the service:
|
||||
|
||||
```sh
|
||||
./ollama-gateway -config ./gateway-config.json -check-config
|
||||
```
|
||||
|
||||
The command validates the same strict bootstrap schema as normal startup, loads the persistent UI override from `storage.data_dir` when present, enforces the bootstrap-only storage rule, verifies that the effective state directory is writable, and prints a secret-free JSON summary. A non-zero exit code means the gateway should not be started with that configuration.
|
||||
|
||||
The summary includes:
|
||||
|
||||
- bootstrap config path,
|
||||
- whether a persistent override was loaded,
|
||||
- persistent override path,
|
||||
- effective worker count,
|
||||
- effective/absolute state directory,
|
||||
- storage writability,
|
||||
- non-fatal deployment warnings.
|
||||
|
||||
Warnings currently cover common production foot-guns such as `local_system_stats=true` on a remote worker URL, an unknown `native.control_worker`, public metrics, a relative state directory and UI authentication/cookie combinations that deserve review.
|
||||
|
||||
For Compose deployments, run the check in the exact image/user/mount context that will be used in production:
|
||||
|
||||
```sh
|
||||
GATEWAY_CONFIG=./gateway-config.json docker compose run --rm gateway \
|
||||
-config /etc/ollama-gateway/config.json -check-config
|
||||
```
|
||||
|
||||
This is preferable to validating only on the Docker host because it also proves that the container user can read the bootstrap file and write the mounted state volume.
|
||||
|
||||
## Built-in HTTP probe
|
||||
|
||||
The scratch image contains no shell, curl or wget. The gateway binary therefore has a minimal HTTP probe mode:
|
||||
|
||||
```sh
|
||||
/ollama-gateway -probe http://127.0.0.1:8080/healthz -probe-timeout 2s
|
||||
```
|
||||
|
||||
It exits 0 for any 2xx response and non-zero for connection errors, timeouts or non-2xx responses. Probe mode does not load gateway configuration or state.
|
||||
|
||||
The image/Compose liveness healthcheck uses `/healthz`. Load balancers and rollout automation should additionally require `/readyz`; readiness checks scheduler, workers and persistent runtime stores and can return 503 while the process itself remains healthy.
|
||||
|
||||
## Compose entrypoint rule
|
||||
|
||||
The image already defines:
|
||||
|
||||
```text
|
||||
ENTRYPOINT ["/ollama-gateway"]
|
||||
```
|
||||
|
||||
Therefore Compose `command` must contain only arguments:
|
||||
|
||||
```yaml
|
||||
command: ["-config", "/etc/ollama-gateway/config.json"]
|
||||
```
|
||||
|
||||
Do **not** repeat `/ollama-gateway` in `command`. Repeating the executable turns it into the first positional argument passed to Go's flag parser and can prevent following flags from being interpreted as intended.
|
||||
|
||||
The supplied Compose file supports selecting a bootstrap configuration without editing the file:
|
||||
|
||||
```sh
|
||||
GATEWAY_CONFIG=./gateway-config.json docker compose up -d --build
|
||||
```
|
||||
|
||||
If `GATEWAY_CONFIG` is omitted, the development `config.example.json` is mounted. Do not mistake that default example for a production configuration; `-check-config` reports the effective worker count and makes that error easier to detect before rollout.
|
||||
|
||||
## Recommended staged update
|
||||
|
||||
1. Back up the bootstrap config and `storage.data_dir`.
|
||||
2. Build/pull the candidate image.
|
||||
3. Run `-check-config` in the candidate container with the production mounts.
|
||||
4. Stop the old gateway gracefully.
|
||||
5. Start the candidate and wait for the Docker liveness check.
|
||||
6. Require `/readyz` before restoring traffic.
|
||||
7. Verify authenticated `/gateway/ui-api/session`, one non-streaming request and one streaming request if used.
|
||||
8. Keep the previous image/binary and state backup available until the observation window is complete.
|
||||
|
||||
## Checkpoint 23: reverse-proxy boundary and container sandbox
|
||||
|
||||
The supplied Compose file now publishes the gateway on host loopback by default:
|
||||
|
||||
```text
|
||||
127.0.0.1:9080 -> container :8080
|
||||
```
|
||||
|
||||
This is the preferred topology when the TLS/reverse proxy runs on the same Docker host. It prevents LAN clients from bypassing the proxy and reaching the gateway's published port directly. Override `GATEWAY_PUBLISH_ADDRESS` only when the real proxy is on another host. In that case bind to the specific host interface where possible and enforce a host firewall/ACL that permits only the real proxy source address.
|
||||
|
||||
The Compose runtime also uses a read-only root filesystem, drops all Linux capabilities and enables `no-new-privileges`. `/data` remains the only persistent writable application state mount and `/tmp` is an ephemeral tmpfs.
|
||||
|
||||
For the deployment observed during the checkpoint-22 rollout, browser traffic arrived at the container through Docker with TCP peer `172.30.3.1`. The hardened production configuration therefore replaces broad RFC1918 `trusted_proxies` ranges with the exact `172.30.3.1/32` peer. `ip_bypass_use_forwarded_ip` remains false. If the network path changes, re-observe the gateway's `remote_addr` and trust only the actual proxy/NAT peer that is expected to supply forwarded headers.
|
||||
|
||||
A host-side preflight catches the common rollout mistakes before Compose starts production traffic:
|
||||
|
||||
```sh
|
||||
cp .env.production.example .env
|
||||
GATEWAY_CONFIG=./gateway-config.json ./scripts/production-preflight.sh
|
||||
```
|
||||
|
||||
It fails if the production config points to `config.example.json`, if the gateway is published outside loopback without explicit acknowledgement, if Compose rendering fails, or if the gateway's own `-check-config` rejects the effective configuration/state mounts.
|
||||
|
||||
For a remote reverse proxy, use an explicit interface and firewall policy, for example:
|
||||
|
||||
```sh
|
||||
GATEWAY_PUBLISH_ADDRESS=10.2.10.20 \
|
||||
ALLOW_NON_LOOPBACK_BIND=1 \
|
||||
GATEWAY_CONFIG=./gateway-config.json \
|
||||
./scripts/production-preflight.sh
|
||||
```
|
||||
|
||||
The explicit opt-in is not a substitute for a firewall. The published port should be reachable only from the reverse proxy.
|
||||
|
||||
For production, prefer the stricter `docker-compose.production.yml`. Unlike the development Compose file, it requires `GATEWAY_CONFIG` to be set and never falls back to `config.example.json`.
|
||||
@@ -0,0 +1,131 @@
|
||||
> **Status: archived diagnostic path.** Multi-gateway HA was removed from the active roadmap in checkpoint 28 because current single-process performance does not justify distributed coordination. Keep this tooling for future re-evaluation if requirements change.
|
||||
|
||||
# P3.2 HA readiness and single-process limit characterization
|
||||
|
||||
P3.2 is deliberately gated: do not add distributed consensus, global queue state or leader election until the single-process gateway is shown to be the limiting availability/capacity component rather than Ollama/GPU inference.
|
||||
|
||||
This repository includes a small evidence toolchain for that decision:
|
||||
|
||||
- `cmd/mock-ollama`: deterministic Ollama/OpenAI-compatible mock backend that removes model compute from the measurement;
|
||||
- `cmd/bench`: concurrent OpenAI request driver with latency, TTFB, status, bytes, token throughput and JSON output;
|
||||
- `scripts/ha-readiness.sh`: repeatable concurrency sweep that writes one JSON result per level and captures Prometheus snapshots before/after each level;
|
||||
- `cmd/ha-snapshot`: optional host/process endpoint snapshotter for gateway RSS/CPU/threads/FDs and host CPU/memory/load evidence;
|
||||
- `cmd/ha-sampler`: sustained resource sampler for each load level, including sampled peak RSS/CPU/threads/FDs and host-memory/load pressure;
|
||||
- `cmd/ha-report`: reconciles client benchmark counts with gateway Prometheus deltas, attaches endpoint and sustained resource evidence, and writes machine-readable JSON plus a Markdown evidence summary.
|
||||
|
||||
## 1. Build/run the mock backend
|
||||
|
||||
```bash
|
||||
go run ./cmd/mock-ollama \
|
||||
-listen 127.0.0.1:11435 \
|
||||
-model qwen3:8b \
|
||||
-delay 0 \
|
||||
-response-bytes 128
|
||||
```
|
||||
|
||||
Point one gateway worker at `http://127.0.0.1:11435`. The mock implements the worker inventory/metadata endpoints plus native and OpenAI chat endpoints required by the gateway.
|
||||
|
||||
Useful mock controls:
|
||||
|
||||
```text
|
||||
-delay 5ms fixed pre-response latency
|
||||
-stream-delay 10ms delay between streaming chunks
|
||||
-response-bytes N generated payload size
|
||||
-fail-every N deterministic HTTP 503 every Nth inference request
|
||||
```
|
||||
|
||||
The zero-delay mode approximates gateway/control-plane overhead. Delay/stream-delay modes exercise many concurrent in-flight connections without requiring a GPU.
|
||||
|
||||
## 2. Run a concurrency sweep
|
||||
|
||||
For an authenticated gateway:
|
||||
|
||||
```bash
|
||||
export GATEWAY_BENCH_API_KEY='...'
|
||||
export GATEWAY_PID="$(pgrep -n ollama-gateway)"
|
||||
BASE_URL=http://127.0.0.1:8080 \
|
||||
MODEL=qwen3:8b \
|
||||
REQUESTS=1000 \
|
||||
WARMUP=50 \
|
||||
CONCURRENCIES='1 4 16 32 64 128' \
|
||||
./scripts/ha-readiness.sh
|
||||
```
|
||||
|
||||
Each level prints a human summary and writes `concurrency-<N>.json` plus `metrics-before-c<N>.prom` and `metrics-after-c<N>.prom`. When `GATEWAY_PID` is set, the sweep also writes `resources-before-c<N>.json`, `resources-after-c<N>.json`, and `resources-samples-c<N>.json`. After the sweep, `cmd/ha-report` writes `report.json` and `report.md`. The report verifies that the gateway request, queue and service-observation deltas equal **warmup + measured requests** for every level and records endpoint-resource and sustained-resource evidence completeness separately.
|
||||
|
||||
The sweep authenticates `/metrics` with `GATEWAY_BENCH_API_KEY` when that environment variable is set; the key is used only as an HTTP header and is never written to the result directory. Set `METRICS_URL` when metrics are exposed at a separate admin endpoint. Set `CAPTURE_METRICS=false` only for an intentionally client-only run; such a run does not produce reconciled HA evidence.
|
||||
|
||||
Endpoint resource capture defaults to `auto`: it is enabled when `GATEWAY_PID` is present and otherwise skipped. Set `CAPTURE_RESOURCES=true` to require it explicitly, or `CAPTURE_RESOURCES=false` to suppress it. Sustained capture separately defaults to `auto` through `CAPTURE_SUSTAINED_RESOURCES`; set it to `false` when an external profiler already supplies the sustained evidence. `GATEWAY_PID` must identify the actual gateway process, not a shell wrapper. `RESOURCE_SAMPLE_INTERVAL` defaults to `250ms`, and `RESOURCE_SAMPLE_MAX_DURATION` defaults to `15m` as a safety bound for a stuck benchmark.
|
||||
|
||||
The `resources-before/after` files are **endpoint snapshots immediately before and after a load level**. The `resources-samples` file is collected throughout the load level. On Linux, sampled process CPU uses deltas between `/proc/<pid>/stat` and aggregate `/proc/stat`, scaled to all logical CPUs, so a multi-core process can legitimately exceed 100%. The sampler also records sampled peak RSS, threads, open FDs and load plus minimum available host memory. On other platforms the tool falls back to repeated `ps` sampling and leaves unavailable fields absent.
|
||||
|
||||
Each benchmark JSON contains:
|
||||
|
||||
- successful/error counts and HTTP status distribution;
|
||||
- requests/second;
|
||||
- latency p50/p95/p99/max;
|
||||
- TTFB p50/p95/p99/max;
|
||||
- bytes received and bytes/second;
|
||||
- prompt/completion tokens and completion tokens/second when non-streaming usage is available;
|
||||
- stream/keep-alive/service-class settings.
|
||||
|
||||
To isolate connection setup cost, run an additional pass with `cmd/bench -disable-keepalive`. For long-lived response pressure, run with `STREAM=true` and configure a non-zero mock `-stream-delay`.
|
||||
|
||||
## 3. Measure three distinct regimes
|
||||
|
||||
### A. Gateway ceiling
|
||||
|
||||
Mock Ollama with `-delay 0`, small response body, non-streaming. Increase concurrency until throughput flattens or latency/error rate rises sharply. This measures HTTP/auth/admission/scheduler/routing/proxy/accounting overhead rather than LLM inference.
|
||||
|
||||
### B. Connection/stream pressure
|
||||
|
||||
Mock Ollama with streaming enabled and `-stream-delay` long enough to keep many requests open. This tests goroutine/socket/buffer behavior and cancellation/shutdown paths.
|
||||
|
||||
### C. Real inference
|
||||
|
||||
Repeat against the actual Ollama workers. This gives the practical service curve and reveals whether GPU/VRAM/model residency saturates far earlier than the gateway process.
|
||||
|
||||
## 4. Record gateway-side signals
|
||||
|
||||
The default sweep captures `/metrics` automatically. `cmd/ha-report` consumes the full `ollama_gateway_*` Prometheus names and derives per-level deltas. Important signals include:
|
||||
|
||||
- scheduler queued/running and queue wait;
|
||||
- service-class queued/running;
|
||||
- worker active/model-active slots;
|
||||
- request/error/retry/circuit counters;
|
||||
- usage-journal drop/storage gauges;
|
||||
- process CPU/RSS/thread/file-descriptor and host load/memory endpoint snapshots from `cmd/ha-snapshot`;
|
||||
- sampled peak process CPU/RSS/thread/file-descriptor and host pressure from `cmd/ha-sampler`, plus socket/runtime metrics from sustained host/container monitoring when deeper evidence is required.
|
||||
|
||||
The benchmark numbers are meaningful only together with CPU, memory and worker saturation. A throughput plateau while GPU/worker slots are saturated is not evidence that HA is needed for gateway capacity.
|
||||
|
||||
## 5. HA decision gate
|
||||
|
||||
Proceed from readiness work to a real P3.2 coordinator/cluster design only when at least one of these is demonstrated:
|
||||
|
||||
1. **Capacity:** gateway CPU/network/connection capacity saturates materially before the Ollama worker pool and cannot be resolved by tuning one process.
|
||||
2. **Availability:** the required recovery objective cannot tolerate one authoritative gateway process, even with supervisor restart and durable local state.
|
||||
3. **Operational topology:** multiple gateway instances are mandatory across failure domains, while strict global fairness/quota/batch semantics must still be preserved.
|
||||
|
||||
Do **not** deploy independent replicas behind a generic load balancer and call that P3.2. Independent replicas create separate fair queues, quota buckets, transient job registries and batch dispatchers. They improve process redundancy only by weakening global correctness.
|
||||
|
||||
### Evidence report gate states
|
||||
|
||||
- `incomplete`: one or more metric snapshots are missing, or gateway request/queue/service counters do not reconcile with the expected warmup + measured request count. Fix the measurement before drawing conclusions.
|
||||
- `not-proven`: client and gateway counters reconcile, but the evidence still does not demonstrate one of the capacity/availability/topology criteria above. `resource_evidence_complete` states whether every level has before/after snapshots; `sustained_resource_evidence_complete` states whether every level has a valid sustained sampler trace. This is the normal result of a mock-backend smoke test.
|
||||
|
||||
The tool intentionally does **not** auto-promote the project into clustered HA based only on a throughput plateau. Sustained sampling closes the short-peak gap in endpoint snapshots, but it still must be interpreted with socket/runtime metrics and, for real inference runs, worker/GPU saturation.
|
||||
|
||||
## 6. P3.2 design constraints once the gate is met
|
||||
|
||||
The HA design must preserve the current hot-path properties:
|
||||
|
||||
- one authoritative global order for fair admission;
|
||||
- globally consistent quota and durable batch state;
|
||||
- fencing/epochs so a stale leader cannot dispatch duplicate authoritative work;
|
||||
- membership and leader failover without making every streamed token depend on consensus;
|
||||
- local worker proxying remains direct after a request is admitted/assigned;
|
||||
- content-bearing batch/conversation storage retains its current explicit privacy boundaries;
|
||||
- no mandatory Redis dependency is introduced merely as a shortcut.
|
||||
|
||||
A likely architecture is a small consensus-backed coordinator/control plane plus local proxy workers, not replicated copies of today’s independent in-memory scheduler.
|
||||
@@ -0,0 +1,198 @@
|
||||
# Verbesserungen und Prioritäten (September 2026)
|
||||
|
||||
Dieses Dokument beschreibt die nach einem Abgleich mit Ollama, OpenWebUI und etablierten LLM-Gateways priorisierten Verbesserungen. Ziel bleibt ein einzelner, sehr schneller Go-Prozess ohne Redis oder Datenbank im Inference-Hot-Path.
|
||||
|
||||
## Priorität A – in diesem Release umgesetzt
|
||||
|
||||
### 1. Capability-aware Model Preflight
|
||||
|
||||
Ollama liefert über `POST /api/show` Modell-Capabilities und Modellinformationen. Das Gateway cached diese Metadaten pro Worker/Modell und erkennt aktuell:
|
||||
|
||||
- `completion`
|
||||
- `tools`
|
||||
- `vision`
|
||||
- `thinking`
|
||||
- `embedding`
|
||||
|
||||
Ein Request wird vor Admission gegen seine benötigten Capabilities geprüft. Das verhindert z. B., dass ein OpenWebUI-Tool-Request erst tief im Ollama-Backend mit `does not support tools` scheitert.
|
||||
|
||||
Konfiguration:
|
||||
|
||||
```json
|
||||
"model_capabilities": {
|
||||
"mode": "enforce",
|
||||
"cache_ttl": "10m",
|
||||
"context_guard": "reject",
|
||||
"context": {
|
||||
"max_requested_tokens": 32768,
|
||||
"default_worker_tokens": 4096,
|
||||
"estimation_margin_percent": 15,
|
||||
"vision_reserve_tokens_per_image": 2048
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`mode`:
|
||||
|
||||
- `enforce`: bekannte Capability-Konflikte mit HTTP 400 ablehnen.
|
||||
- `observe`: nur Warnheader/Logs setzen.
|
||||
- `off`: Metadaten-Preflight deaktivieren.
|
||||
|
||||
Wenn `/api/show` temporär nicht erreichbar ist, lässt das Gateway den Request bewusst durch. Verfügbarkeit hat bei unbekannten Metadaten Vorrang; nur *bekannt inkompatible* Requests werden geblockt.
|
||||
|
||||
### 2. Context Guard
|
||||
|
||||
Der Gateway trennt das **theoretische Modellmaximum** von dem Kontext, den ein konkreter Worker tatsächlich bereitstellt. Für OpenAI-/Responses-/Anthropic-kompatible Requests ohne natives `num_ctx` gilt als effektive Quelle in dieser Reihenfolge:
|
||||
|
||||
1. `context_length` des bereits geladenen Modells aus `/api/ps`;
|
||||
2. `num_ctx` aus den Modelfile-Parametern von `/api/show`;
|
||||
3. `workers[].default_context_tokens`;
|
||||
4. `model_capabilities.context.default_worker_tokens`.
|
||||
|
||||
Anschließend greifen `workers[].context_limits` und der Gateway-Cap. Ein natives `options.num_ctx` wird nur auf `/api/*` akzeptiert, muss ein positiver Integer sein und wird gegen Gateway-Cap, Modellmaximum und Worker-Cap geprüft. Ein kleinerer aktuell geladener Kontext darf für ein explizit größeres `num_ctx` neu allokiert werden, erhält dann aber keinen Loaded-Routingbonus.
|
||||
|
||||
Die Admission-Schätzung zählt neben Prompt/Messages/Tool-Schemas jetzt auch `instructions`, `suffix`, `max_output_tokens`, eine konfigurierbare Vision-Reserve pro Bild und eine Sicherheitsmarge. Das ist bewusst eine konservative Schätzung und kein Tokenizer-Ersatz.
|
||||
|
||||
### 3. Per-model Concurrency
|
||||
|
||||
Ein 24-GB-Worker sollte kleine 8B-Modelle anders behandeln können als 27B-Modelle. Jeder Worker kann deshalb zusätzlich zu `max_concurrent` Modellgrenzen definieren:
|
||||
|
||||
```json
|
||||
"model_concurrency": {
|
||||
"qwen3:8b": 2,
|
||||
"gemma4:*": 1,
|
||||
"*": 1
|
||||
}
|
||||
```
|
||||
|
||||
Matching-Reihenfolge:
|
||||
|
||||
1. exakter Modellname;
|
||||
2. längster passender Prefix mit abschließendem `*`;
|
||||
3. `*`;
|
||||
4. `max_concurrent`.
|
||||
|
||||
Der Slot wird im selben in-memory Worker-State geführt und bei Cancel/Fehler/Completion wieder freigegeben.
|
||||
|
||||
### 4. Adaptive Worker Routing
|
||||
|
||||
Worker-Routing berücksichtigt jetzt zusätzlich zu Health und aktiven Slots:
|
||||
|
||||
- bereits geladenes Modell;
|
||||
- auf dem Worker installiertes Modell;
|
||||
- per-model Slot-Auslastung;
|
||||
- gelernte Output-Tokenrate je Worker/Modell (EWMA), persistent über Neustarts;
|
||||
- tatsächlichen VRAM-Druck, wenn Telemetrie verfügbar ist;
|
||||
- GPU-Auslastung;
|
||||
- eine harte Zusatzstrafe für das Laden eines neuen Modells bei sehr hohem VRAM-Druck.
|
||||
|
||||
Die Gewichte sind konfigurierbar:
|
||||
|
||||
```json
|
||||
"routing": {
|
||||
"loaded_bonus": 60,
|
||||
"installed_bonus": 30,
|
||||
"throughput_bonus": 20,
|
||||
"vram_pressure_penalty": 35,
|
||||
"gpu_utilization_penalty": 10,
|
||||
"avoid_vram_percent": 97
|
||||
}
|
||||
```
|
||||
|
||||
Durchsatzwerte werden ausschließlich aus echten abgeschlossenen Inferenzrequests gelernt und unter `storage.worker_performance_file` persistiert. Nach einem Neustart startet das Routing deshalb mit den zuletzt bekannten EWMA-Werten und lernt sie anschließend weiter.
|
||||
|
||||
### 5. NVIDIA-Telemetrie ohne zusätzlichen Agenten
|
||||
|
||||
Auf NVIDIA-Hosts kann ein Worker direkt `nvidia-smi` verwenden:
|
||||
|
||||
```json
|
||||
"nvidia_smi": true,
|
||||
"nvidia_gpu": "0"
|
||||
```
|
||||
|
||||
Erfasst werden:
|
||||
|
||||
- GPU utilization
|
||||
- VRAM used / total
|
||||
- Temperatur
|
||||
- Power Draw
|
||||
|
||||
Die Abfrage läuft außerhalb des Request-Hot-Paths im normalen Worker-Health-Zyklus. Das Gateway bleibt `CGO_ENABLED=0` und benötigt keine NVML-Bibliothek. Ein vorhandenes `telemetry_url` kann weiterhin verwendet werden und externe Werte überschreiben.
|
||||
|
||||
### 6. Operations UI und Metrics
|
||||
|
||||
Die Modellansicht zeigt jetzt erkannte Capabilities und Context Length. Worker-Karten zeigen – soweit vorhanden – GPU, VRAM, Temperatur, Leistung und gelernte tok/s.
|
||||
|
||||
Prometheus exportiert zusätzlich:
|
||||
|
||||
```text
|
||||
ollama_gateway_worker_memory_used_bytes
|
||||
ollama_gateway_worker_memory_total_bytes
|
||||
ollama_gateway_worker_vram_used_bytes
|
||||
ollama_gateway_worker_vram_total_bytes
|
||||
ollama_gateway_worker_gpu_utilization_percent
|
||||
ollama_gateway_worker_gpu_temperature_celsius
|
||||
ollama_gateway_worker_gpu_power_watts
|
||||
ollama_gateway_worker_model_active
|
||||
ollama_gateway_worker_model_prompt_tokens_per_second
|
||||
ollama_gateway_worker_model_output_tokens_per_second
|
||||
ollama_gateway_worker_model_performance_samples
|
||||
```
|
||||
|
||||
Modelle sind eine begrenzte operative Dimension. Tenant-/User-/Request-IDs werden weiterhin nicht als Prometheus-Labels ausgegeben.
|
||||
|
||||
### 7. Unbegrenzte Credits verständlicher im UI
|
||||
|
||||
`0` bedeutet weiterhin unbegrenzt. Die Policy-Oberfläche zeigt dafür jetzt `∞` und bietet explizite Checkboxen für unbegrenzte Actor- bzw. Tenant-Credits.
|
||||
|
||||
## RTX-4090-Startprofil
|
||||
|
||||
`config.rtx4090.example.json` enthält einen Startpunkt für eine RTX 4090 mit 24 GiB VRAM und beispielhaft 64 GiB System-RAM. Den RAM-Wert bitte an den tatsächlichen Host anpassen. Das Profil funktioniert unter Linux und Windows; beide Plattformen erfassen Host-RAM nativ, und `nvidia-smi` liefert die GPU-Telemetrie, sofern es im `PATH` verfügbar ist.
|
||||
|
||||
Für 24 GiB VRAM ist der wichtigste Unterschied zum 256-GB-Unified-Memory-Mac die Modell-spezifische Concurrency. Große ~17-GB-Modelle starten im Beispiel bei 1, kleine Modelle dürfen 2 parallele Slots erhalten.
|
||||
|
||||
## Priorität B – als nächste Ausbaustufe empfohlen
|
||||
|
||||
### Upstream Circuit Breaker + begrenzte Retries
|
||||
|
||||
Retries sind bei LLM-Streaming gefährlich, weil eine bereits begonnene Generierung nicht transparent wiederholt werden darf. Sinnvoll wäre deshalb:
|
||||
|
||||
- Retry nur vor dem ersten Response-Byte und nur bei eindeutig transienten Transportfehlern;
|
||||
- Circuit Breaker pro Worker;
|
||||
- Exponential Backoff bei Health-/Load-Fehlern;
|
||||
- kein blindes Request Hedging, da es GPU-Arbeit dupliziert.
|
||||
|
||||
### API-Key Model ACLs
|
||||
|
||||
Pro API-Key/Tenant sollten erlaubte/verbotene Modell-Patterns möglich sein, z. B. `allowed_models: ["qwen3:*", "embeddinggemma:*"]`. Das ist besonders für teure oder unzensierte Modelle sinnvoll.
|
||||
|
||||
### Capability Routing / expliziter Fallback
|
||||
|
||||
Ein optionaler, explizit konfigurierter Fallback könnte Tool-Requests auf ein toolfähiges Modell umleiten. Default sollte `reject` bleiben, da ein stiller Modellwechsel semantisch überraschend ist.
|
||||
|
||||
### Persistentes Audit-Log als optionales Modul
|
||||
|
||||
Control-Plane-, Accounting- und Routing-Zustand ist inzwischen lokal persistent, während aktive Scheduling-/Socket-Zustände in-memory bleiben. Für Compliance wäre zusätzlich ein optionaler asynchroner Audit-Sink sinnvoll, der nur Admin-Aktionen und Metadaten schreibt und niemals Prompts/Antworten.
|
||||
|
||||
### OpenTelemetry
|
||||
|
||||
Prometheus deckt Aggregationen ab. Für verteilte Client-/Gateway-/Ollama-Latenz wäre optionales OTel Tracing sinnvoll, weiterhin ohne Prompt-Inhalte.
|
||||
|
||||
## Referenzen
|
||||
|
||||
- Ollama Show model details: https://docs.ollama.com/api-reference/show-model-details
|
||||
- Ollama Tool calling: https://docs.ollama.com/capabilities/tool-calling
|
||||
- Ollama Thinking: https://docs.ollama.com/capabilities/thinking
|
||||
- Ollama Vision: https://docs.ollama.com/capabilities/vision
|
||||
- Ollama Embeddings: https://docs.ollama.com/capabilities/embeddings
|
||||
- OpenWebUI Ollama connection/context notes: https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-ollama/
|
||||
- NVIDIA `nvidia-smi` selective query: https://docs.nvidia.com/deploy/nvidia-smi/
|
||||
|
||||
|
||||
## Checkpoint 27 — Context / `num_ctx` Hardening
|
||||
|
||||
Der Context Guard unterscheidet jetzt zwischen dem theoretischen Modellmaximum (`/api/show model_info.*.context_length`) und dem tatsächlich nutzbaren Kontext eines Workers. Für Requests ohne explizites natives `options.num_ctx` gilt in dieser Reihenfolge: geladenes `/api/ps context_length`, Modelfile-`num_ctx`, `workers[].default_context_tokens`, dann der globale `model_capabilities.context.default_worker_tokens`. `workers[].context_limits` kann Modelle pro Worker zusätzlich begrenzen.
|
||||
|
||||
Native `options.num_ctx` muss ein positiver Integer sein. Es wird gegen Gateway-Cap, Modellmaximum und Worker-Cap geprüft. Ein bereits kleiner geladenes Modell bleibt für einen explizit größeren `num_ctx` grundsätzlich routbar, verliert aber den Loaded-Bonus, weil Ollama den KV-Cache neu allozieren muss.
|
||||
|
||||
Die Preflight-Schätzung berücksichtigt nun außerdem `max_output_tokens`, `instructions`, `suffix`, einen konfigurierbaren Vision-Reservewert pro Bild sowie eine Sicherheitsmarge. Die Modellansicht zeigt Modell-Maximum, Modelfile-Kontext und geladenen Kontext getrennt.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Jobs and cancellation
|
||||
|
||||
The gateway distinguishes between **model lifecycle operations**, **transient inference jobs**, and **durable batch jobs**.
|
||||
|
||||
## Model Stop
|
||||
|
||||
The Models page uses **Stop** to unload a resident Ollama model. Ollama does not expose `/api/stop`; the gateway sends:
|
||||
|
||||
```http
|
||||
POST /api/generate
|
||||
Content-Type: application/json
|
||||
|
||||
{"model":"<model>","keep_alive":0,"stream":false}
|
||||
```
|
||||
|
||||
This asks Ollama to unload the model immediately. It does not cancel a specific request.
|
||||
|
||||
## Inference jobs
|
||||
|
||||
Every compute request receives an `X-Request-ID` and is registered as an in-memory job while it is queued, routing, running, or streaming.
|
||||
|
||||
Admin endpoints:
|
||||
|
||||
```text
|
||||
GET /gateway/ui-api/jobs
|
||||
POST /gateway/ui-api/jobs/<request-id>/cancel
|
||||
```
|
||||
|
||||
The web UI exposes these under **Jobs** and also places an **Abbrechen** button next to active requests in Live Flow.
|
||||
|
||||
### Cancellation semantics
|
||||
|
||||
- **queued**: remove the ticket from the WFQ heap immediately and release the credit reservation;
|
||||
- **waiting for worker**: cancel worker acquisition and release the scheduler lease;
|
||||
- **running/streaming**: cancel the upstream HTTP context so the Ollama request terminates, then release worker/scheduler slots;
|
||||
- reconcile any usage already observable from the response;
|
||||
- store terminal live state `cancelled` with internal status `499`.
|
||||
|
||||
For a streaming response whose upstream HTTP headers were already forwarded, the downstream client may still have HTTP status `200`; cancellation is observable as an early stream termination. Internal gateway telemetry records the request as `499`.
|
||||
|
||||
Transient inference-job state is process-local and disappears on gateway restart, consistent with the in-memory hot-path design.
|
||||
|
||||
## Durable batch jobs
|
||||
|
||||
P3.1 adds a separate durable job type under `/gateway/v1/batches`. Batch definitions, state, input references and output references survive restart; an active execution attempt itself is cancelled on shutdown and converted back to a restart-safe queued state.
|
||||
|
||||
The Admin UI exposes durable work on the dedicated **Batch Jobs** page rather than mixing it into the transient **Jobs** page. See `docs/BATCH-JOBS.md` for API, state-machine, persistence, privacy, retention and accounting semantics.
|
||||
@@ -0,0 +1,65 @@
|
||||
# Migration to the in-memory engine
|
||||
|
||||
This release removes the external coordination backend entirely.
|
||||
|
||||
## Removed configuration
|
||||
|
||||
Delete these old top-level/settings from existing configuration files:
|
||||
|
||||
```text
|
||||
redis
|
||||
scheduler.distributed
|
||||
scheduler.lease_ttl
|
||||
scheduler.poll_interval
|
||||
workers[].request_lease_ttl
|
||||
usage.redis_aggregates
|
||||
cluster
|
||||
```
|
||||
|
||||
Replace the old local/cluster visualization settings with:
|
||||
|
||||
```json
|
||||
"infrastructure": {
|
||||
"node_name": "mac-studio-gateway",
|
||||
"refresh_interval": "250ms",
|
||||
"max_requests": 256
|
||||
}
|
||||
```
|
||||
|
||||
The configuration parser uses `DisallowUnknownFields`, so obsolete settings fail fast instead of being silently ignored.
|
||||
|
||||
## Runtime behavior
|
||||
|
||||
The following state now exists only inside one gateway process:
|
||||
|
||||
- weighted fair queue state;
|
||||
- running concurrency;
|
||||
- per-worker slots;
|
||||
- actor and tenant quota buckets;
|
||||
- runtime policy overrides;
|
||||
- browser sessions;
|
||||
- usage summaries;
|
||||
- live/infrastructure telemetry.
|
||||
|
||||
Active scheduling state still resets on restart. Durable control-plane state (UI-created API-key hashes, policy overrides, quota balances, metrics snapshots and usage history) is now stored locally under `storage.data_dir`; the inference scheduler itself remains in memory.
|
||||
|
||||
## Multiple Ollama workers
|
||||
|
||||
No change is required. One gateway may still route to many Ollama workers and maintains one fair queue across all of them.
|
||||
|
||||
## Multiple gateway processes
|
||||
|
||||
Do not place multiple independent gateway processes behind a generic load balancer if you require strict global fairness. Each process is now a separate fairness/quota/session domain.
|
||||
|
||||
For the intended deployment, use one authoritative gateway and scale the Ollama worker pool behind it.
|
||||
|
||||
## API change
|
||||
|
||||
The former cluster visualization endpoints are now local infrastructure endpoints:
|
||||
|
||||
```text
|
||||
GET /gateway/ui-api/infrastructure
|
||||
GET /gateway/ui-api/infrastructure/stream
|
||||
```
|
||||
|
||||
The web UI has been updated accordingly.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Model Placement
|
||||
|
||||
Model Placement is the gateway's hard model-to-worker routing policy. It is evaluated **before** adaptive worker scoring, so GPU load, VRAM pressure, model affinity or learned throughput can never override a placement prohibition.
|
||||
|
||||
## Why it is separate from tenant policies
|
||||
|
||||
Tenant/Fairness policies answer:
|
||||
|
||||
> How much compute may a tenant or actor consume?
|
||||
|
||||
Model Placement answers:
|
||||
|
||||
> On which Ollama workers may a model execute?
|
||||
|
||||
Keeping those concerns separate makes both the configuration and the UI predictable.
|
||||
|
||||
## UI workflow
|
||||
|
||||
Open **Admin -> Model Placement**.
|
||||
|
||||
The matrix has one row per discovered model and one column per worker:
|
||||
|
||||
- `✓ installed` — allowed and installed on this worker.
|
||||
- `○ allowed` — policy allows it, but the model is not currently installed there.
|
||||
- `● loaded` — allowed, installed and currently resident according to `/api/ps`.
|
||||
- `⛔ blocked` — the effective placement rule denies the model.
|
||||
- `↳` — the cell is resolved by an exact model rule.
|
||||
|
||||
Click a cell to create an exact allow/deny exception. Clicking an exact UI exception again removes that exact rule so the broader prefix/default rule becomes effective again.
|
||||
|
||||
The worker cards open a ruleset editor with:
|
||||
|
||||
- **Allow all**: allow by default and optionally deny selected models/prefixes.
|
||||
- **Whitelist**: deny by default and only allow selected models/prefixes.
|
||||
- **Only installed** preset: builds a whitelist from the worker's current `/api/tags` inventory.
|
||||
- **Block all** preset: an empty whitelist.
|
||||
- **Reset to config default**: deletes the persistent UI override and immediately restores the worker's bootstrap/persistent-config baseline.
|
||||
|
||||
All UI changes are written to `storage.model_placement_file` (default `model-placement.json`) using the gateway's atomic state writer and take effect for new requests immediately.
|
||||
|
||||
## Pattern semantics
|
||||
|
||||
Patterns support:
|
||||
|
||||
```text
|
||||
qwen3:8b exact model
|
||||
gemma4:* prefix wildcard
|
||||
orcarouter/* prefix wildcard
|
||||
* all models
|
||||
```
|
||||
|
||||
Wildcards are only supported as a single trailing `*`.
|
||||
|
||||
Resolution uses specificity:
|
||||
|
||||
1. exact model rule;
|
||||
2. longest matching prefix rule;
|
||||
3. worker mode (`allow_all` or `whitelist`).
|
||||
|
||||
At equal specificity, deny wins.
|
||||
|
||||
This means the following is valid and useful:
|
||||
|
||||
```json
|
||||
{
|
||||
"mode": "allow_all",
|
||||
"allowed_models": ["gemma4:latest"],
|
||||
"denied_models": ["gemma4:*"]
|
||||
}
|
||||
```
|
||||
|
||||
`gemma4:latest` is allowed because the exact rule is more specific, while other `gemma4:*` variants are blocked.
|
||||
|
||||
## Example: model A on node 1, model B on node 1 + 2
|
||||
|
||||
```json
|
||||
"workers": [
|
||||
{
|
||||
"name": "node-1",
|
||||
"url": "http://10.10.11.10:11434",
|
||||
"max_concurrent": 4,
|
||||
"model_placement": {
|
||||
"mode": "whitelist",
|
||||
"allowed_models": ["model-a:*", "model-b:*"],
|
||||
"denied_models": []
|
||||
}
|
||||
},
|
||||
{
|
||||
"name": "node-2",
|
||||
"url": "http://10.10.11.20:11434",
|
||||
"max_concurrent": 2,
|
||||
"model_placement": {
|
||||
"mode": "whitelist",
|
||||
"allowed_models": ["model-b:*"],
|
||||
"denied_models": []
|
||||
}
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
The adaptive router then chooses between node 1 and node 2 for `model-b:*`, but `model-a:*` can only run on node 1.
|
||||
|
||||
## Inventory behavior
|
||||
|
||||
Each worker's `/api/tags` inventory is tracked separately from `/api/ps` loaded state.
|
||||
|
||||
Routing behavior is:
|
||||
|
||||
1. remove unhealthy workers;
|
||||
2. remove workers blocked by Model Placement;
|
||||
3. if inventory is known, require the model to be installed;
|
||||
4. enforce global and per-model concurrency;
|
||||
5. score the remaining workers by loaded affinity, load, learned tok/s, GPU utilization and VRAM pressure.
|
||||
|
||||
If at least one eligible worker has a known installed copy, only those installed workers are considered. If all eligible inventories are known and none contains the model, the gateway returns a descriptive `model ... is not installed on any eligible worker` error instead of sending the request to an arbitrary node.
|
||||
|
||||
A worker whose inventory has never been successfully retrieved is treated as **inventory unknown**, not as "model absent". That provides a controlled fail-open path for transient `/api/tags` discovery failures while still respecting hard placement rules.
|
||||
|
||||
## Client model discovery
|
||||
|
||||
`/api/tags` and `/v1/models` are gateway-native aggregated endpoints. Models that are installed only on placement-blocked workers are excluded from client discovery. This keeps OpenWebUI and other clients aligned with the gateway's actual routing policy.
|
||||
|
||||
The admin model inventory remains unfiltered so operators can still see and manage models that are installed but deliberately blocked from inference.
|
||||
|
||||
## Persistence precedence
|
||||
|
||||
For a worker:
|
||||
|
||||
```text
|
||||
persistent UI placement override
|
||||
↓ (if absent)
|
||||
config.json / persistent config baseline
|
||||
↓
|
||||
effective routing rule
|
||||
```
|
||||
|
||||
Deleting the UI override restores the baseline immediately. The placement file is included in gateway ZIP backups and is shown on **Admin -> Persistenz**.
|
||||
@@ -0,0 +1,134 @@
|
||||
# OpenWebUI integration
|
||||
|
||||
## Recommended connection
|
||||
|
||||
Use OpenWebUI's Ollama provider connection with the gateway base URL (no `/api` suffix) and a dedicated gateway API key. You can either configure a durable key in `auth.api_keys` or create a persistent key under **Admin -> Sicherheit -> API Keys**.
|
||||
|
||||
Gateway configuration:
|
||||
|
||||
```json
|
||||
"auth": {
|
||||
"oidc": { "enabled": false },
|
||||
"api_keys": [
|
||||
{
|
||||
"name": "openwebui",
|
||||
"key": "${OPENWEBUI_GATEWAY_KEY}",
|
||||
"tenant": "interactive",
|
||||
"subject": "openwebui",
|
||||
"application": "openwebui",
|
||||
"scopes": []
|
||||
}
|
||||
],
|
||||
"ip_bypass": [],
|
||||
"trusted_proxies": []
|
||||
}
|
||||
```
|
||||
|
||||
### Create the key in the web interface
|
||||
|
||||
Open **Admin -> Sicherheit -> API Keys -> API-Key erstellen** and use, for example:
|
||||
|
||||
```text
|
||||
Name: openwebui
|
||||
Tenant: interactive
|
||||
Application: openwebui
|
||||
Scopes: (empty)
|
||||
```
|
||||
|
||||
Copy the generated `ofg_...` secret immediately and paste it into OpenWebUI. The gateway does not retain the plaintext value. The plaintext is shown once; only its SHA-256 hash and metadata are persisted, so the same OpenWebUI credential remains valid across gateway restarts.
|
||||
|
||||
OpenWebUI connection:
|
||||
|
||||
```text
|
||||
URL: http://host.docker.internal:8080
|
||||
API Key: <same value as OPENWEBUI_GATEWAY_KEY>
|
||||
```
|
||||
|
||||
OpenWebUI's backend appends native Ollama paths itself. Do not configure the URL as `...:8080/api`.
|
||||
|
||||
## Diagnostics
|
||||
|
||||
From a machine that can reach the gateway:
|
||||
|
||||
```bash
|
||||
curl -i -H "Authorization: Bearer $OPENWEBUI_GATEWAY_KEY" http://GATEWAY:8080/api/version
|
||||
curl -i -H "Authorization: Bearer $OPENWEBUI_GATEWAY_KEY" http://GATEWAY:8080/api/tags
|
||||
```
|
||||
|
||||
Expected status for both is HTTP 200. `/api/tags` must contain a `models` array and every returned model is normalized to contain both `name` and `model`.
|
||||
|
||||
OpenAI-compatible discovery is also available:
|
||||
|
||||
```bash
|
||||
curl -i -H "Authorization: Bearer $OPENWEBUI_GATEWAY_KEY" http://GATEWAY:8080/v1/models
|
||||
```
|
||||
|
||||
## Docker networking
|
||||
|
||||
The default example IP bypass only trusts loopback. An OpenWebUI Docker container normally reaches the gateway from a Docker/host network address, so loopback bypass does not apply. Prefer a dedicated API key over widening the bypass CIDR.
|
||||
|
||||
On Docker Desktop for macOS, `host.docker.internal` normally resolves to the host. On Linux, use an address/service name reachable from the OpenWebUI backend or configure Docker's host-gateway mapping.
|
||||
|
||||
## Gateway compatibility behavior
|
||||
|
||||
The gateway owns model discovery instead of forwarding it to a single control worker:
|
||||
|
||||
- `GET /api/tags`: parallel inventory across all workers, deduplicated by model ID.
|
||||
- `GET /api/ps`: aggregated loaded-model view.
|
||||
- `GET /v1/models`: OpenAI model list generated from the same aggregate inventory.
|
||||
- `POST /api/show`: routed to a worker that owns the requested installed model when known.
|
||||
- Compute requests are constrained to workers known to own the model; a loaded copy receives additional affinity preference.
|
||||
|
||||
If one worker fails during discovery but another returns models, the gateway returns the available model set and sets `X-Gateway-Partial-Errors`. If no worker can provide tags, discovery returns HTTP 503 rather than an empty, misleading list.
|
||||
|
||||
## OpenWebUI with "Authentication: None" and IP bypass
|
||||
|
||||
If the OpenWebUI connection is configured with **Authentication: None**, the gateway must authenticate the OpenWebUI backend through `auth.ip_bypass`. The bypass is evaluated against the **source IP observed by the gateway**, not against the gateway URL.
|
||||
|
||||
For example, if the gateway log shows:
|
||||
|
||||
```text
|
||||
authentication rejected client_ip=10.10.11.42 ... path=/api/version ...
|
||||
```
|
||||
|
||||
add only that OpenWebUI host address when possible:
|
||||
|
||||
```json
|
||||
"ip_bypass": [
|
||||
{
|
||||
"cidrs": ["10.10.11.42/32"],
|
||||
"tenant": "interactive",
|
||||
"subject": "openwebui",
|
||||
"application": "openwebui",
|
||||
"scopes": []
|
||||
},
|
||||
{
|
||||
"cidrs": ["127.0.0.1/32", "::1/128"],
|
||||
"tenant": "local",
|
||||
"subject": "localhost",
|
||||
"application": "local-tools",
|
||||
"scopes": ["gateway:admin"]
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Then OpenWebUI may remain configured with:
|
||||
|
||||
```text
|
||||
URL: http://10.10.11.123:8080
|
||||
Authentication: None
|
||||
```
|
||||
|
||||
Do **not** blindly add `10.10.11.123/32` just because that is the gateway URL. That is the destination address. Use the `client_ip` value printed by the gateway when OpenWebUI performs `/api/version` or `/api/tags`.
|
||||
|
||||
For Docker, the observed source may be a Docker bridge/VM address rather than the LAN address of the host. A dedicated static API key is more stable than allowing a broad Docker subnet.
|
||||
|
||||
### Authentication diagnostics
|
||||
|
||||
Authentication failures now log the resolved client address and include it in the `X-Gateway-Client-IP` response header. Native Ollama endpoints also return Ollama-compatible errors such as:
|
||||
|
||||
```json
|
||||
{"error":"authentication required"}
|
||||
```
|
||||
|
||||
rather than an OpenAI-shaped nested object. This prevents OpenWebUI from rendering the gateway error as `[object Object]`.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Persistent local state
|
||||
|
||||
The gateway deliberately separates the inference hot path from durable state. Fair queue ordering, active worker slots, live requests and browser OIDC sessions remain in memory. Restart-worthy control-plane/accounting data is stored under `storage.data_dir` using only the Go standard library.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Contents | Write pattern |
|
||||
|---|---|---|
|
||||
| `gateway-config.json` | configuration override saved from Admin -> Konfiguration | atomic replace |
|
||||
| `api-keys.json` | UI-created key metadata + SHA-256 hashes | atomic replace |
|
||||
| `policies.json` | tenant policy overrides | atomic replace |
|
||||
| `metrics.json` | Prometheus counters and histogram state | periodic atomic snapshot + shutdown |
|
||||
| `quota.json` | actor/tenant credit bucket balances | periodic atomic snapshot + shutdown |
|
||||
| `worker-performance.json` | learned per-worker/model prompt/output tok/s EWMA | periodic atomic snapshot + shutdown |
|
||||
| `model-placement.json` | live per-worker model placement overrides from the admin UI | atomic replace |
|
||||
| `conversations.enc.json` | optional encrypted Responses conversation contexts | synchronous encrypted atomic replace + retention cleanup |
|
||||
| `batch-jobs.json` | optional durable batch definitions, identity metadata, state and spool references | synchronous atomic replace on state transitions |
|
||||
| `batch/input/*.json` | durable batch request payloads | create + fsync + rename, mode 0600 |
|
||||
| `batch/output/*.response` | durable batch response payloads | temp file + fsync + rename, mode 0600 |
|
||||
| `usage/usage-YYYY-MM-DD.jsonl` | request-level accounting inside detail retention | append-only, buffered |
|
||||
| `usage/rollups/daily/rollup-daily-YYYY-MM-DD.json` | daily dimensional usage aggregates | atomic replace during compaction |
|
||||
| `usage/rollups/monthly/rollup-monthly-YYYY-MM.json` | long-term monthly dimensional aggregates | idempotent atomic replace |
|
||||
|
||||
By default, no prompt text or generated response content is written by these stores. There are two explicit opt-in exceptions: `conversations.enabled=true` stores Responses context encrypted at rest, while `batch_jobs.enabled=true` allows submitted batch request/response content to be written to the local spool in plaintext files protected by filesystem permissions. See `docs/CONVERSATIONS.md` and `docs/BATCH-JOBS.md`.
|
||||
|
||||
## Bootstrap and persistent configuration
|
||||
|
||||
The file passed with `-config` is the bootstrap configuration. It determines `storage.data_dir` and the state filenames. If `<data_dir>/gateway-config.json` exists and validates, it becomes the effective configuration for the process.
|
||||
|
||||
The web JSON editor saves a complete validated override but never exposes API-key secrets, the UI session secret or the browser OIDC client secret. Redacted secret values are preserved from the currently effective configuration. API keys themselves should be managed through the dedicated Security page.
|
||||
|
||||
The `storage` section is bootstrap-only. This prevents the persistent configuration from moving the file that is needed to locate itself. To move the state directory, stop the gateway, update the bootstrap config and move/copy the state directory explicitly.
|
||||
|
||||
## Crash and shutdown semantics
|
||||
|
||||
Atomic state files are written through a temporary file, `fsync`, and rename. API keys, tenant policies, model-placement overrides and durable batch state transitions are saved synchronously when correctness requires it. Metrics, quota and learned worker-performance snapshots are refreshed every `storage.flush_interval` and once again during graceful shutdown. A hard process/host crash can therefore lose at most the most recent snapshot interval for those snapshot files; a batch attempt left as `running` is repaired to `queued` during startup recovery.
|
||||
|
||||
Usage events are appended asynchronously and flushed according to `usage.flush_interval`. On normal shutdown the queue is drained and the current journal is flushed/synced. Before startup replay, retention compaction runs once so expired request rows are never needlessly loaded back into memory. Startup then reconstructs all-time counters from monthly rollups, daily rollups and the remaining detail journals. Only detail-journal events populate the recent-request UI.
|
||||
|
||||
|
||||
## Tiered usage retention and aggregation
|
||||
|
||||
The default retention policy is:
|
||||
|
||||
```json
|
||||
"retention": {
|
||||
"detail_days": 30,
|
||||
"daily_days": 400,
|
||||
"monthly_months": 0,
|
||||
"compaction_interval": "6h"
|
||||
}
|
||||
```
|
||||
|
||||
- **Detail**: raw request metadata for the most recent 30 calendar days.
|
||||
- **Daily**: older request rows are replaced by one aggregate file per day until day 400.
|
||||
- **Monthly**: older daily files are folded into monthly files. `monthly_months: 0` means keep them forever.
|
||||
|
||||
Each rollup stores independent aggregate dimensions for global traffic, tenants, actors, applications, models and workers. Stored measures include requests/errors, prompt/completion/cached tokens, credits, queue/service duration, prompt/eval nanoseconds and transferred bytes. Prompt/output token throughput can therefore be derived after request details have expired.
|
||||
|
||||
Monthly files keep their source day keyed internally. This is deliberate: if the process crashes after the monthly file is synced but before the daily source file is deleted, retrying compaction replaces that day contribution instead of adding it twice.
|
||||
|
||||
Compaction only touches closed historical journal days. The active day's writer is never rewritten. The background interval can be configured, and admins can trigger the same compaction with `POST /gateway/ui-api/storage/compact` or **Admin -> Persistenz -> Jetzt kompaktieren**.
|
||||
|
||||
Historical aggregate queries are exposed as:
|
||||
|
||||
```text
|
||||
GET /gateway/ui-api/usage/rollups?granularity=daily&dimension=global&limit=120
|
||||
GET /gateway/ui-api/usage/rollups?granularity=daily&dimension=tenant&name=team-a
|
||||
GET /gateway/ui-api/usage/rollups?granularity=daily&dimension=actor&name=team-a%00alice
|
||||
GET /gateway/ui-api/usage/rollups?granularity=monthly&dimension=model&name=qwen3:8b
|
||||
```
|
||||
|
||||
The Prometheus endpoint exports gauges for raw/daily/monthly file counts and byte footprints, plus the timestamp of the last compaction and the number of bytes reclaimed by that run. The cumulative Prometheus counter snapshot itself remains fixed-size and does not need retention compaction.
|
||||
|
||||
## Intentionally non-persistent state
|
||||
|
||||
- fair-queue heap and virtual clocks;
|
||||
- running/streaming transient inference jobs and cancellation handles;
|
||||
- active batch attempt contexts/cancellation handles (the durable batch definition and state remain persistent);
|
||||
- active worker/model slots;
|
||||
- live-flow/infrastructure animation state;
|
||||
- browser OIDC sessions;
|
||||
- transient model pull operation state.
|
||||
|
||||
These objects are tied to sockets, requests or process-local execution and cannot safely be resumed after a restart.
|
||||
|
||||
## Backup
|
||||
|
||||
For a consistent offline backup, stop the gateway and copy `storage.data_dir`. For normal filesystem snapshots, the atomic JSON files and append-only usage journals are safe to copy while the process is running, though the newest buffered usage records may not yet be present on disk.
|
||||
|
||||
## Admin storage controls
|
||||
|
||||
**Admin -> Persistenz** shows every durable state file plus detail/daily/monthly usage and batch-spool footprint. **Jetzt flushen** forces metrics, quota, worker performance, encrypted conversation retention state, durable batch metadata, and buffered usage to disk. **Jetzt kompaktieren** applies usage/conversation/batch retention immediately. **Backup herunterladen** flushes first and streams a ZIP containing durable state, batch spool content when present, remaining detail journals and all rollups. Treat the backup as sensitive because configuration can contain deployment secrets and enabled batch jobs can contain prompt/response content.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Production update checklist
|
||||
|
||||
This release is intended to be safe to stage as an update of an existing single-process Ollama Fair Gateway. The sustained HA-readiness work is isolated from the inference hot path: `cmd/ha-sampler`, the readiness shell wrapper and report tooling do not participate in normal request handling.
|
||||
|
||||
## Scope of this checkpoint
|
||||
|
||||
Relative to checkpoint 16, the gateway runtime behavior and persisted-state schemas are unchanged. The release adds/refines only HA-readiness tooling and documentation. The packaged gateway binaries should therefore remain byte-identical to checkpoint 16 when built with the same Go toolchain and flags; the release verification records this explicitly.
|
||||
|
||||
If the production system is older than checkpoint 16, treat this as a normal application upgrade because earlier checkpoints added runtime alias/tenant ACL controls, conversations and durable batch state. The configuration loader supplies defaults for omitted optional sections, but a backup is still required before replacing an older binary.
|
||||
|
||||
## Before updating
|
||||
|
||||
1. Record the currently deployed binary checksum and keep the old binary available for rollback.
|
||||
2. Back up the bootstrap configuration and `storage.data_dir`. For the most consistent backup, stop the gateway before copying the data directory; see `docs/PERSISTENCE.md`.
|
||||
3. Validate the intended effective configuration with `ollama-gateway -config <bootstrap.json> -check-config`. For containers, run this in the candidate image with the production config/state mounts so file permissions are tested too; see `docs/DEPLOYMENT-HARDENING.md`.
|
||||
4. Confirm there is enough free space for the existing usage/batch/conversation retention settings.
|
||||
5. Do not run the HA readiness sweep against production traffic unless the additional synthetic inference load is acceptable.
|
||||
|
||||
## Update procedure
|
||||
|
||||
1. Stop the supervised gateway gracefully with SIGTERM and wait for the `gateway stopped` log line.
|
||||
2. Replace only the gateway executable appropriate for the host architecture. Verify its SHA-256 value against `dist/SHA256SUMS`.
|
||||
3. Keep the existing bootstrap config and data directory in place.
|
||||
4. Start the gateway under the same supervisor/service account.
|
||||
5. Require both `/healthz` and `/readyz` to succeed before restoring external traffic.
|
||||
6. Send at least one representative non-streaming and, if used in production, one streaming request through the normal authenticated client path.
|
||||
7. Verify `/metrics`, the admin UI/API used operationally, and recent logs for persistence/configuration errors.
|
||||
|
||||
## Rollback
|
||||
|
||||
If startup/readiness or representative inference fails, stop the new process and restore the previous executable. When upgrading from checkpoint 16 specifically, no state-schema rollback is required by this checkpoint because the runtime/persistence code is unchanged. When upgrading from an older build, retain the pre-upgrade data-directory backup until the update has been observed successfully under normal load.
|
||||
|
||||
## HA readiness tooling in production
|
||||
|
||||
`GATEWAY_PID` must be the actual gateway PID. With it set, `scripts/ha-readiness.sh` captures before/after snapshots and starts `cmd/ha-sampler` during each load level. The sampler is read-only with respect to the gateway process; on Linux it reads `/proc` and uses `ps`, and it writes evidence only to the selected `OUT_DIR`. It does not write secrets or gateway state.
|
||||
|
||||
## Checkpoint 18 Admin UI hotfix
|
||||
|
||||
If upgrading from checkpoint 17, no configuration or persistent-state migration is required. Rebuild/redeploy the gateway executable or container because the Admin UI assets are embedded in the binary. Checkpoint 18 restores the Model Placement UI implementation and adds a regression test for missing render-dispatch functions.
|
||||
|
||||
|
||||
## Checkpoint 19 deployment hardening
|
||||
|
||||
Checkpoint 19 keeps P3.2/HA gated and instead hardens the single-node production update path. The shipped Compose command no longer repeats the image entrypoint, the scratch image has a built-in liveness healthcheck via the gateway's `-probe` mode, and `-check-config` validates the effective bootstrap+persistent configuration plus state-directory writability before startup. No persisted-state schema migration is introduced by this checkpoint.
|
||||
|
||||
## Checkpoint 20 remote worker telemetry
|
||||
|
||||
Checkpoint 20 adds an optional out-of-band worker telemetry agent and hardens the existing `workers[].telemetry_url` merge path. It does **not** change any persistent-state schema and the agent is not part of inference request execution. Existing deployments can upgrade without enabling the agent; behavior remains unchanged until `telemetry_url` is configured.
|
||||
|
||||
For remote workers, prefer `local_system_stats: false` and point `telemetry_url` at the agent running on the actual Ollama host. Roll this out in two stages: first deploy/verify each agent with `-once` and its `/telemetry` endpoint, then change the gateway worker configuration. Keep the static `memory_capacity_bytes` and `vram_capacity_bytes` values as capacity hints even when runtime telemetry is enabled.
|
||||
|
||||
If telemetry becomes unavailable, the gateway continues operating; the worker telemetry snapshot records the error and routing falls back to the remaining configured/observed signals. Do not expose the agent on an untrusted network. Restrict `-allow-cidrs` to the gateway source address or protect the endpoint with a TLS-authenticated reverse proxy.
|
||||
|
||||
## Checkpoint 21 telemetry freshness hardening
|
||||
|
||||
Checkpoint 21 does not change persistent-state schemas or inference semantics. It validates timestamps from `telemetry_url` exporters, rejects stale/future samples before they affect routing pressure, preserves pre-existing local collector warnings when external telemetry fails, and makes the shipped worker agent explicitly non-cacheable. Existing exporters without `updated_at` remain compatible.
|
||||
|
||||
|
||||
## Checkpoint 22 IP-bypass hardening
|
||||
|
||||
Checkpoint 22 changes only the authentication source used by credential-free `auth.ip_bypass`: the TCP peer is now authoritative by default. Forwarded client-IP resolution remains available for logs and usage attribution. Deployments that intentionally relied on an original `X-Forwarded-For` client address to satisfy IP bypass must explicitly set `auth.ip_bypass_use_forwarded_ip=true` and should restrict `auth.trusted_proxies` to exact proxy addresses plus network ACLs. API-key and OIDC authentication are unaffected. No persistent-state schema migration is introduced.
|
||||
|
||||
## Checkpoint 23 reverse-proxy/deployment boundary hardening
|
||||
|
||||
Checkpoint 23 does not change inference, scheduler, quota or persistent-state schemas. The supplied Compose topology now defaults to host-loopback publishing (`127.0.0.1:9080`), a read-only root filesystem, `no-new-privileges` and all Linux capabilities dropped. Deployments whose reverse proxy is on another host must explicitly opt into a non-loopback bind and should firewall the published port to that proxy only.
|
||||
|
||||
Before rollout, run `scripts/production-preflight.sh` with the production bootstrap config. The script refuses the development `config.example.json` and non-loopback publishing unless explicitly acknowledged, then runs the gateway's secret-redacted `-check-config` inside the candidate container with the production mounts.
|
||||
@@ -0,0 +1,64 @@
|
||||
# Public Status Dashboard
|
||||
|
||||
Checkpoint 28 adds an optional unauthenticated **read-only** dashboard for users who should be able to observe current gateway capacity without receiving admin access.
|
||||
|
||||
## Enable
|
||||
|
||||
```json
|
||||
"public_dashboard": {
|
||||
"enabled": true,
|
||||
"path": "/status",
|
||||
"title": "Ollama Gateway Status",
|
||||
"subtitle": "Live-Auslastung und Infrastruktur",
|
||||
"refresh_interval": "2s",
|
||||
"max_live_requests": 64,
|
||||
"show_worker_names": false,
|
||||
"show_model_names": true,
|
||||
"show_resource_metrics": true,
|
||||
"worker_display_names": {
|
||||
"internal-worker-name": "GPU Node A"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The dashboard is then available at `/status/` without a gateway credential. It is independent of the authenticated `/admin/` UI and also works when the admin UI itself is disabled.
|
||||
|
||||
## Privacy boundary
|
||||
|
||||
The public API does **not** reuse or serialize the admin snapshots. It builds a new response from an explicit allow-list. The public response excludes:
|
||||
|
||||
- tenant, actor, subject and application identifiers;
|
||||
- API keys, scopes and authentication metadata;
|
||||
- worker URLs, IP labels and arbitrary worker labels;
|
||||
- prompts, responses and request bodies;
|
||||
- persistent storage paths and configuration;
|
||||
- policies, quota buckets, aliases and ACLs;
|
||||
- internal error strings, circuit error details and telemetry errors;
|
||||
- cost/credit estimates.
|
||||
|
||||
Request IDs are replaced by short SHA-256-derived opaque IDs. Worker names are anonymized by default and can be replaced with explicit public display names. Model names are also anonymized unless `show_model_names` is explicitly enabled.
|
||||
|
||||
## Public data
|
||||
|
||||
The dashboard exposes only operational information intended for status display:
|
||||
|
||||
- current queue/running/streaming counts;
|
||||
- gateway uptime;
|
||||
- healthy/total worker count;
|
||||
- number of loaded models;
|
||||
- worker slot usage and health state;
|
||||
- loaded model display names when enabled;
|
||||
- numeric RAM/VRAM/GPU/temperature/power values when resource metrics are enabled;
|
||||
- anonymized recent/live request state, queue/service timing and token counts.
|
||||
|
||||
The snapshot endpoint is `GET <path>/api/snapshot`. It is cacheable for one second to reduce repeated work behind a reverse proxy. The browser refresh interval is configurable but cannot be set below one second.
|
||||
|
||||
## Reverse proxy
|
||||
|
||||
Checkpoint 23+ binds the production Compose service to loopback by default. Expose `/status/` through the same trusted reverse proxy used for the public service. Do not expose the gateway container port directly merely to publish the dashboard.
|
||||
|
||||
If the reverse proxy applies authentication globally, carve out only the configured public dashboard path. Keep `/admin/`, `/gateway/`, `/api/`, `/v1/` and `/metrics` under their existing gateway authentication rules.
|
||||
|
||||
## Security headers
|
||||
|
||||
The public UI sends a restrictive CSP, denies framing, disables referrer leakage and disables browser permissions such as camera, microphone, geolocation, payments and USB.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Checkpoint 17 — production update verification
|
||||
|
||||
Validation date: 2026-09-08
|
||||
Validation toolchain: Go 1.23.2 on Linux/amd64
|
||||
|
||||
## Release scope
|
||||
|
||||
Checkpoint 17 completes sustained HA-readiness resource sampling. It adds `cmd/ha-sampler`, shared read-only resource collection under `internal/haresource`, sweep integration, report integration, tests and operator documentation.
|
||||
|
||||
The normal gateway request, scheduler, routing and persistence code is unchanged relative to checkpoint 16. Rebuilding with the same Go toolchain and release flags produced byte-identical gateway executables for all four packaged gateway targets.
|
||||
|
||||
## Required verification results
|
||||
|
||||
The following checks passed on the final source tree:
|
||||
|
||||
- `go test ./...`
|
||||
- `go test ./... -count=3` during the release cycle
|
||||
- `go vet ./...`
|
||||
- `go test -race ./...`
|
||||
- `gofmt -l .` returned no files
|
||||
- `sh -n scripts/ha-readiness.sh scripts/ollama-env.sh`
|
||||
- all shipped JSON example configurations parse with `jq`
|
||||
- all files in `dist/SHA256SUMS` verify successfully
|
||||
|
||||
## Runtime smoke validation
|
||||
|
||||
The packaged Linux/amd64 gateway was tested against the deterministic mock Ollama backend.
|
||||
|
||||
1. **Persistent restart smoke:** two complete start/readiness/request/shutdown cycles reused the same `storage.data_dir`. Each cycle served a non-streaming and a streaming OpenAI-compatible request, logged `gateway stopped`, and exited with status 0.
|
||||
2. **Authenticated readiness smoke:** private `/metrics` plus a Bearer API key completed the readiness sweep. Client counts reconciled with gateway request/queue/service counters, endpoint resource snapshots were complete, and sustained resource traces were complete.
|
||||
3. **Secret-leak check:** the Bearer API key used by the authenticated sweep was not present in any generated readiness result file.
|
||||
4. **Final sustained-sampler E2E:** concurrency 1 and 4 both completed with zero request errors, exact counter reconciliation, and `resource_samples.stop_reason == "stop-file"`. A sampler that exits because of `max-duration`, process exit, or interruption is deliberately not accepted as complete sustained evidence.
|
||||
|
||||
## Gateway binary identity vs checkpoint 16
|
||||
|
||||
| Target | Checkpoint 17 SHA-256 | Compared with checkpoint 16 |
|
||||
|---|---|---|
|
||||
| Linux amd64 | `01efc1b236faa9b7b912e6bee6af8266f6a24722af419a6dd43b6f294c955f47` | byte-identical |
|
||||
| Linux arm64 | `a364f03e67fe377f140936c2c0cdd2b480166e2c5d1fe519878d730268e74248` | byte-identical |
|
||||
| macOS arm64 | `8f9185f561752656902ef8cf44f5f1c8f65b660f40fb5caf790c2d23773c14d8` | byte-identical |
|
||||
| Windows amd64 | `5a5bbbf962f1a9fea88a76fc0ee1e28669170496800e9ea39d9ed11caf14b053` | byte-identical |
|
||||
|
||||
This identity is the strongest update-safety property of this checkpoint for systems already on checkpoint 16: replacing the gateway executable with the checkpoint 17 gateway executable does not change the executable bytes.
|
||||
|
||||
## Operator note
|
||||
|
||||
For systems older than checkpoint 16, use `docs/PRODUCTION-UPDATE.md` and take a configuration/data-directory backup first. Earlier checkpoints contain real gateway runtime and persistence features, so the byte-identity statement above does not apply to an upgrade directly from an older production build.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Checkpoint 18 — Admin UI placement hotfix
|
||||
|
||||
Validation date: 2026-09-08
|
||||
|
||||
## Release scope
|
||||
|
||||
Checkpoint 18 is a production hotfix on top of checkpoint 17. It restores the complete Model Placement UI helper block that was accidentally removed during the runtime model-alias UI work introduced after checkpoint 13.
|
||||
|
||||
The regression manifested after successful admin authentication as a browser error such as:
|
||||
|
||||
`renderPlacement is not defined`
|
||||
|
||||
The missing block also contained the generic `lines()` helper used by alias, placement and tenant model-access forms, so leaving only `renderPlacement` patched would have exposed follow-on UI failures when saving those forms.
|
||||
|
||||
No gateway request, scheduler, routing, quota, persistence, batch, conversation, HA-sampling or worker-runtime behavior was intentionally changed by this hotfix. The gateway executables do change because the Admin UI JavaScript is embedded into the binary.
|
||||
|
||||
## Fix
|
||||
|
||||
Restored these UI helpers:
|
||||
|
||||
- `placementRule`
|
||||
- `placementRows`
|
||||
- `placementCell`
|
||||
- `placementSourceLabel`
|
||||
- `renderPlacement`
|
||||
- `placementCellHTML`
|
||||
- `placementWorkerCard`
|
||||
- `loadPlacement`
|
||||
- `openPlacementDialog`
|
||||
- `lines`
|
||||
- `installedForWorker`
|
||||
- `savePlacementRule`
|
||||
- `placementExactAction`
|
||||
|
||||
The existing placement matrix, worker-rule editor, exact allow/deny overrides, presets and reset controls now have their implementation functions again.
|
||||
|
||||
## Regression protection
|
||||
|
||||
`internal/webui/webui_test.go` now verifies that every `renderX()` function called by the main `render()` dispatcher is actually defined in the embedded Admin UI bundle. A second test verifies the complete Placement UI helper set.
|
||||
|
||||
This specifically prevents the original `renderPlacement is not defined` class of regression from passing `go test ./...` again.
|
||||
|
||||
## Required verification results
|
||||
|
||||
The final source tree passed:
|
||||
|
||||
- `go test ./...`
|
||||
- `go vet ./...`
|
||||
- `go test -race ./...`
|
||||
- `node --check internal/webui/assets/app.js`
|
||||
- release cross-builds for Linux amd64, Linux arm64, macOS arm64 and Windows amd64
|
||||
- `dist/SHA256SUMS` verification
|
||||
|
||||
## Update guidance
|
||||
|
||||
For Docker/Compose deployments, rebuild or replace the gateway image/binary. Copying only the source `app.js` beside an already-built binary does not fix the problem because `internal/webui/assets/app.js` is embedded at build time.
|
||||
|
||||
The production configuration and persistent `/data` volume do not require a migration from checkpoint 17 to checkpoint 18.
|
||||
|
||||
## Gateway binary SHA-256
|
||||
|
||||
| Target | SHA-256 |
|
||||
|---|---|
|
||||
| Linux amd64 | `550f481313e5869d884280dfc3248d70b6eabf51b4b3bb44083c36f622065cf2` |
|
||||
| Linux arm64 | `27f4ab0b7a1dc9a2379e7e99471e177f9da9c715445b05b089d028011b8101aa` |
|
||||
| macOS arm64 | `d72dd623844050b1b7396ae364e49a30ab76d2bee21a2ced22dad308c0e839bb` |
|
||||
| Windows amd64 | `dedf06c3d42fdc8c2fe3de19cd9c00fcdb6eb6fab91442df3b18ac5e8a416072` |
|
||||
@@ -0,0 +1,58 @@
|
||||
# Checkpoint 19 — deployment preflight and container hardening
|
||||
|
||||
Validation date: 2026-09-08
|
||||
|
||||
## Scope
|
||||
|
||||
Checkpoint 19 follows the Checkpoint 18 Admin UI hotfix. It does not implement P3.2/HA and does not change any persistent-state schema. Its purpose is to reduce single-node production update risk after real deployment feedback exposed configuration mount/entrypoint and file-permission mistakes that were otherwise only visible after startup.
|
||||
|
||||
The inference request/scheduler/quota/proxy path is unchanged by this checkpoint. Runtime changes are limited to startup configuration loading being factored through the same helper used by the new preflight, plus two explicit CLI-only modes (`-check-config` and `-probe`).
|
||||
|
||||
## Changes
|
||||
|
||||
- `-check-config`: strict effective configuration validation plus persistent override loading and state-directory write probe.
|
||||
- Secret-free JSON preflight summary with worker count and deployment warnings.
|
||||
- `-probe`: minimal HTTP probe implemented inside the gateway binary for scratch/container health checks.
|
||||
- Dockerfile/Compose liveness healthcheck against public `/healthz`.
|
||||
- Corrected Compose command: arguments only, no duplicate `/ollama-gateway` after the image `ENTRYPOINT`.
|
||||
- `GATEWAY_CONFIG` Compose variable for selecting the bootstrap config explicitly.
|
||||
- Regression tests for persistent-secret restoration, preflight output, HTTP probing and Compose entrypoint/healthcheck rules.
|
||||
- Production deployment documentation.
|
||||
|
||||
## Verification completed
|
||||
|
||||
The final source tree passed:
|
||||
|
||||
- `go test ./... -count=1`
|
||||
- `go vet ./...`
|
||||
- Race detector over all packages, split as `go test -race ./cmd/...` and `go test -race ./internal/...` so each command completed with an observable exit status. (The single monolithic invocation exceeded the execution harness timeout; no package was omitted from the split run.)
|
||||
- `node --check internal/webui/assets/app.js`
|
||||
- Compose YAML parse/assertions for the corrected command, `GATEWAY_CONFIG` mount and healthcheck. The Docker CLI itself was not available in the verification environment, so `docker compose config` was not claimed.
|
||||
- Real Linux-amd64 binary `-check-config` run against the migrated production configuration: `status=ok`, exactly 2 effective workers, writable state directory, expected remote `local_system_stats` warnings, and no API-key secret in the preflight output.
|
||||
- Real gateway + deterministic mock Ollama smoke: `/healthz` probe, `/readyz` probe, authenticated `/gateway/ui-api/session`, OpenAI-compatible inference, SIGTERM, and gateway exit code 0.
|
||||
- Cross-builds for Linux amd64, Linux arm64, macOS arm64 and Windows amd64.
|
||||
- `dist/SHA256SUMS` verification for every packaged binary/helper.
|
||||
|
||||
## Gateway binary SHA-256
|
||||
|
||||
| Target | SHA-256 |
|
||||
|---|---|
|
||||
| Linux amd64 | `e30a7b01f8a96f6396c3582ab208bf8e4e330c6a882b724998448dc66dfa7cf7` |
|
||||
| Linux arm64 | `675a72c054c21ba383c36891b9bdeb1b93cbd4690abedc22b74ae02c8bc081e4` |
|
||||
| macOS arm64 | `586e489cf00c54aa5906a63f4ceada6ff22c90943a80d31fb1c864b557dde559` |
|
||||
| Windows amd64 | `1d24976651a358bc80031f36eb2f59cca9233dfa89bf3189bb1bb4b6e9de96fe` |
|
||||
|
||||
The HA helper binaries were not modified by this checkpoint and their existing checksums remain listed and verified in `dist/SHA256SUMS`.
|
||||
|
||||
## Update guidance
|
||||
|
||||
No configuration-schema or persistent-state migration is required from Checkpoint 18. Container deployments should rebuild/recreate the image because the Dockerfile healthcheck and gateway CLI modes are part of this release.
|
||||
|
||||
Before starting production traffic, run the candidate container in preflight mode with the exact config/state mounts:
|
||||
|
||||
```sh
|
||||
GATEWAY_CONFIG=./gateway-config.json docker compose run --rm gateway \
|
||||
-config /etc/ollama-gateway/config.json -check-config
|
||||
```
|
||||
|
||||
Then start the service, wait for liveness, require `/readyz`, authenticate to the admin UI/API and send representative inference traffic. See `docs/DEPLOYMENT-HARDENING.md` and `docs/PRODUCTION-UPDATE.md`.
|
||||
@@ -0,0 +1,69 @@
|
||||
# Checkpoint 20 — remote worker telemetry
|
||||
|
||||
Date: 2026-09-08
|
||||
|
||||
## Scope
|
||||
|
||||
Checkpoint 20 extends the single-node production hardening path with an optional `worker-telemetry` agent for remote Ollama hosts. It also fixes the existing external telemetry merge so omitted JSON fields no longer overwrite previously collected values with zero, while explicit zero values remain meaningful.
|
||||
|
||||
No persistent-state schema changes are introduced. The inference, quota and scheduler request paths are unchanged.
|
||||
|
||||
## New artifacts
|
||||
|
||||
- `cmd/worker-telemetry`
|
||||
- `dist/ollama-gateway-worker-telemetry-linux-amd64`
|
||||
- `dist/ollama-gateway-worker-telemetry-linux-arm64`
|
||||
- `docs/WORKER-TELEMETRY.md`
|
||||
- Linux AMDGPU sysfs collector in `internal/hoststats`
|
||||
|
||||
## Verification completed
|
||||
|
||||
The final source tree passed:
|
||||
|
||||
- `go test ./...`
|
||||
- `go vet ./...`
|
||||
- JavaScript syntax validation with `node --check internal/webui/assets/app.js`
|
||||
- race tests across all command packages and all tested internal packages, including `internal/worker` and `internal/hoststats`
|
||||
- SHA-256 verification for every file listed in `dist/SHA256SUMS`
|
||||
- cross-builds for gateway Linux amd64/arm64, macOS arm64 and Windows amd64
|
||||
- cross-builds for worker telemetry Linux amd64/arm64
|
||||
|
||||
## Remote telemetry E2E
|
||||
|
||||
A Linux-amd64 release gateway was started with:
|
||||
|
||||
- one mock Ollama worker,
|
||||
- `local_system_stats: false`,
|
||||
- `telemetry_url` pointing at the release worker-telemetry agent,
|
||||
- a temporary simulated AMDGPU sysfs tree.
|
||||
|
||||
The gateway reported the worker healthy and exposed telemetry source:
|
||||
|
||||
`telemetry-url:host-memory+amdgpu-sysfs`
|
||||
|
||||
The simulated GPU values were preserved end-to-end:
|
||||
|
||||
- VRAM used: 4,294,967,296 bytes
|
||||
- VRAM total: 17,179,869,184 bytes
|
||||
- GPU utilization: 73%
|
||||
- GPU temperature: 62.5 C
|
||||
- GPU power: 88 W
|
||||
|
||||
Both the gateway and telemetry agent exited cleanly with status 0 after SIGTERM.
|
||||
|
||||
## Merge regression coverage
|
||||
|
||||
`externalTelemetry` uses pointer fields. This explicitly distinguishes an omitted JSON property from a supplied zero. Tests verify that omitted memory/GPU fields preserve existing telemetry while an explicit zero for VRAM usage or GPU utilization is applied.
|
||||
|
||||
## Production rollout
|
||||
|
||||
Enabling the agent is optional and can be staged independently from the gateway update. For a remote worker:
|
||||
|
||||
```json
|
||||
{
|
||||
"local_system_stats": false,
|
||||
"telemetry_url": "http://WORKER-IP:11500/telemetry"
|
||||
}
|
||||
```
|
||||
|
||||
Validate the agent locally with `-once` first and restrict network access using `-allow-cidrs` plus the worker host firewall. See `docs/WORKER-TELEMETRY.md`.
|
||||
@@ -0,0 +1,40 @@
|
||||
# Checkpoint 21 — telemetry freshness hardening
|
||||
|
||||
Date: 2026-09-08
|
||||
|
||||
## Scope
|
||||
|
||||
Checkpoint 21 hardens the optional remote worker telemetry path introduced in checkpoint 20. No persistent-state schema changes are introduced and the inference/quota/scheduler hot paths are unchanged.
|
||||
|
||||
Changes:
|
||||
|
||||
- parse optional external `updated_at`,
|
||||
- reject samples older than `max(30s, 6 x health_interval)`,
|
||||
- reject timestamps more than 30 seconds in the future,
|
||||
- keep backward compatibility for exporters without timestamps,
|
||||
- append external HTTP/decode/freshness errors instead of overwriting an existing collector warning,
|
||||
- send `Cache-Control: no-store` and `X-Content-Type-Options: nosniff` from the shipped telemetry agent.
|
||||
|
||||
## Safety behavior
|
||||
|
||||
A rejected external sample is not merged into `ResourceTelemetry`. Static worker capacity hints, Ollama health/inventory and all other routing inputs remain available. The worker remains usable; the telemetry error is surfaced for operators.
|
||||
|
||||
## Verification
|
||||
|
||||
Release verification includes unit tests for fresh/stale/future/missing timestamps, merge semantics, full repository tests/vet, race tests, release cross-builds, dist checksums and a real gateway + mock Ollama + telemetry-agent smoke path.
|
||||
|
||||
## E2E freshness evidence
|
||||
|
||||
Fresh-agent path:
|
||||
|
||||
- agent response included `Cache-Control: no-store`,
|
||||
- source was `telemetry-url:host-memory+amdgpu-sysfs`,
|
||||
- simulated VRAM/GPU/temperature/power values were merged correctly.
|
||||
|
||||
Stale-exporter path:
|
||||
|
||||
- exporter was made ready before gateway startup,
|
||||
- payload timestamp was intentionally more than one hour old,
|
||||
- gateway reported `telemetry sample is stale: ... exceeds 30s`,
|
||||
- stale VRAM/GPU pressure values were not merged,
|
||||
- gateway still shut down cleanly with exit status 0.
|
||||
@@ -0,0 +1,39 @@
|
||||
# Checkpoint 22 — IP-bypass / trusted-proxy hardening
|
||||
|
||||
Date: 2026-09-08
|
||||
|
||||
## Security issue addressed
|
||||
|
||||
Before checkpoint 22, `Authenticator.Authenticate` evaluated `auth.ip_bypass` against the fully resolved client IP. If the immediate peer was in `auth.trusted_proxies`, that address could originate from `X-Forwarded-For`. A deployment that trusted a broad reachable subnet could therefore allow a direct client in that subnet to present a bypass address such as `127.0.0.1`.
|
||||
|
||||
## New default
|
||||
|
||||
- `ClientIP()` still resolves trusted forwarding chains for observability.
|
||||
- `auth.ip_bypass` is evaluated against the direct TCP peer by default.
|
||||
- `Identity.ClientIP` remains the resolved client address for usage/logging.
|
||||
- Legacy forwarded-IP bypass is available only with explicit `auth.ip_bypass_use_forwarded_ip=true`.
|
||||
- Config preflight warns when forwarded bypass is enabled and when `auth.trusted_proxies` contains broad non-loopback CIDRs.
|
||||
|
||||
API-key and OIDC authentication semantics are unchanged. No persistent-state schema change is introduced.
|
||||
|
||||
## Regression coverage
|
||||
|
||||
Tests prove that:
|
||||
|
||||
1. a trusted peer with `X-Forwarded-For: 127.0.0.1` cannot satisfy loopback IP bypass by default,
|
||||
2. forwarded client-IP resolution still reports `127.0.0.1` for observability in that test,
|
||||
3. explicit compatibility mode restores forwarded-IP bypass,
|
||||
4. a direct loopback peer still satisfies the normal loopback bypass,
|
||||
5. preflight warns on compatibility mode and a broad `10.0.0.0/8` trusted-proxy range while not flagging loopback-only proxy ranges.
|
||||
|
||||
## Built-binary spoof E2E
|
||||
|
||||
A release Linux-amd64 gateway was started with a trusted loopback proxy peer and an IP-bypass rule for `192.0.2.123/32`.
|
||||
|
||||
With the new default (`ip_bypass_use_forwarded_ip=false`), a request from the trusted TCP peer carrying `X-Forwarded-For: 192.0.2.123` returned **401 Unauthorized**.
|
||||
|
||||
The same test with the explicit compatibility flag set to `true` returned **200 OK** with an `ip-bypass` admin identity and resolved `client_ip=192.0.2.123`. This proves both the secure default and the intentional compatibility escape hatch in the release binary.
|
||||
|
||||
## Production-config preflight
|
||||
|
||||
The user's two-worker configuration validates successfully with the checkpoint-22 candidate. Preflight reports broad trusted-proxy warnings for `10.0.0.0/8`, `172.16.0.0/12` and `192.168.0.0/16`; these no longer permit forwarded headers to satisfy IP bypass under the default, but should still be narrowed to improve client-IP attribution integrity.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Release verification: checkpoint 23
|
||||
|
||||
Checkpoint 23 hardens the Docker/reverse-proxy deployment boundary. It does not change the inference hot path or persisted-state schemas.
|
||||
|
||||
## Scope verification
|
||||
|
||||
`cmd/` and `internal/` were byte-for-byte/diff checked against checkpoint 22 after the deployment changes. No runtime Go source changed. Existing release binaries therefore remain unchanged and all entries in `dist/SHA256SUMS` still verify.
|
||||
|
||||
## Deployment changes
|
||||
|
||||
- Development Compose defaults to `127.0.0.1:9080 -> :8080` instead of publishing on all host interfaces.
|
||||
- `docker-compose.production.yml` requires `GATEWAY_CONFIG`; there is no production fallback to `config.example.json`.
|
||||
- Container root filesystem is read-only; `/data` is the persistent writable volume and `/tmp` is tmpfs.
|
||||
- Runtime explicitly uses UID/GID `65532:65532`, drops all Linux capabilities and enables `no-new-privileges`.
|
||||
- JSON-file logging is bounded to five 10 MiB files in the production Compose example.
|
||||
- `scripts/production-preflight.sh` rejects `config.example.json`, rejects non-loopback publishing unless explicitly acknowledged, renders Compose, and invokes the gateway's secret-redacted `-check-config` in the candidate container.
|
||||
- `scripts/production-preflight_test.sh` regression-tests loopback success, example-config rejection, wildcard-bind rejection and explicit remote-proxy opt-in.
|
||||
|
||||
## Production config validation
|
||||
|
||||
The hardened two-worker production configuration was validated with the release binary:
|
||||
|
||||
- status: `ok`
|
||||
- workers: `2`
|
||||
- data directory: `/data`
|
||||
- state directory writable: `true`
|
||||
- broad trusted-proxy warnings: none
|
||||
|
||||
The configuration narrows `auth.trusted_proxies` to loopback plus the observed Docker peer `172.30.3.1/32`; `auth.ip_bypass_use_forwarded_ip` remains false.
|
||||
|
||||
## Test results
|
||||
|
||||
- `go test ./...`: PASS
|
||||
- `go vet ./...`: PASS
|
||||
- `go test -race ./cmd/...`: PASS
|
||||
- `go test -race ./internal/...`: PASS
|
||||
- `sh -n scripts/production-preflight.sh`: PASS
|
||||
- `sh -n scripts/production-preflight_test.sh`: PASS
|
||||
- `scripts/production-preflight_test.sh`: PASS
|
||||
- Compose YAML parse check: PASS
|
||||
- `dist/SHA256SUMS`: all 10 packaged binaries PASS
|
||||
|
||||
## Release-binary smoke test
|
||||
|
||||
The shipped Linux-amd64 gateway binary was started against an isolated deterministic Ollama mock and temporary state directory.
|
||||
|
||||
- `/healthz`: 200 / `status=ok`
|
||||
- `/readyz`: 200 / `status=ready`
|
||||
- authenticated `/gateway/ui-api/session`: admin identity returned
|
||||
- authenticated `/v1/chat/completions`: successful response (`chatcmpl-mock`)
|
||||
- SIGTERM gateway shutdown: exit code 0 and `gateway stopped` logged
|
||||
|
||||
The mock process is test infrastructure and is not part of the production artifact.
|
||||
|
||||
## Deployment recommendation
|
||||
|
||||
For the observed topology, keep `GATEWAY_PUBLISH_ADDRESS=127.0.0.1` so only a reverse proxy on the Docker host can reach port 9080. If the reverse proxy is remote, bind to one specific host interface, set `ALLOW_NON_LOOPBACK_BIND=1` only for the preflight, and firewall the published port to the exact reverse-proxy source IP.
|
||||
@@ -0,0 +1,53 @@
|
||||
# Release verification — Checkpoint 24
|
||||
|
||||
Checkpoint 24 is an admin-WebUI-only visual refresh based on the production-hardened Checkpoint 23 runtime. It changes the embedded HTML/CSS/JavaScript presentation layer and the four gateway binaries that embed those assets. Scheduler, quota, routing, authentication, persistence, telemetry and inference semantics are unchanged.
|
||||
|
||||
## Scope
|
||||
|
||||
- Professional light control-plane theme inspired by dense technical/mission-database UIs.
|
||||
- Offline-safe font stack: preferred local IBM Plex/JetBrains/Cascadia fonts with system fallbacks; no external font or CDN dependency.
|
||||
- Explicit light color scheme for browser-native controls.
|
||||
- Light sidebar, panels, KPI cards, tables, forms, dialogs, callouts, code/config editor and login screen.
|
||||
- Light Live Flow and Infrastructure canvases, including canvas-drawn labels/nodes/routes rather than CSS-only restyling.
|
||||
- Accessibility-oriented color tokens. Primary text, muted text, accent text, primary buttons and semantic status colors meet approximately WCAG AA 4.5:1 contrast against white for normal text.
|
||||
- Regression tests assert that the light theme, light metadata and light canvas colors remain embedded.
|
||||
|
||||
## Verification performed
|
||||
|
||||
```sh
|
||||
node --check internal/webui/assets/app.js
|
||||
python -c 'import tinycss2; ... parse internal/webui/assets/app.css ...'
|
||||
go test ./...
|
||||
go vet ./...
|
||||
go test -race ./...
|
||||
```
|
||||
|
||||
All commands completed successfully.
|
||||
|
||||
Four gateway release targets were rebuilt because WebUI assets are embedded in the gateway binary:
|
||||
|
||||
- linux/amd64
|
||||
- linux/arm64
|
||||
- darwin/arm64
|
||||
- windows/amd64
|
||||
|
||||
`dist/SHA256SUMS` was regenerated and `sha256sum -c dist/SHA256SUMS` succeeded for all shipped binaries.
|
||||
|
||||
## Release-binary smoke
|
||||
|
||||
The rebuilt linux/amd64 binary was started with an isolated writable data directory and deterministic mock Ollama backend. The following checks passed:
|
||||
|
||||
- `-check-config`
|
||||
- `/readyz`
|
||||
- embedded `/admin/`
|
||||
- embedded `/admin/app.css` contains the Checkpoint 24 light-theme marker
|
||||
- embedded `/admin/app.js` passes `node --check`
|
||||
- `/gateway/ui-api/session` via direct loopback admin IP bypass returns an admin identity
|
||||
- `/v1/chat/completions` returns a successful mock response
|
||||
- SIGTERM shutdown exits with code 0
|
||||
|
||||
A headless Chromium screenshot was attempted, but the container's Chromium/DBus environment did not terminate reliably. This was not counted as a passed test. Asset parsing, embedding and runtime serving were verified independently as described above.
|
||||
|
||||
## Deployment
|
||||
|
||||
No configuration or persistent-state migration is required from Checkpoint 23. Rebuild/redeploy the gateway image because the WebUI is compiled into the binary. After update, hard-refresh the browser once to discard any previously cached `app.css` or `app.js`.
|
||||
@@ -0,0 +1,15 @@
|
||||
# Checkpoint 26 release verification
|
||||
|
||||
Scope: embedded Admin WebUI only (dialog semantics + regression tests).
|
||||
|
||||
Verified:
|
||||
|
||||
- `go test ./...`
|
||||
- `go vet ./...`
|
||||
- `node --check internal/webui/assets/app.js`
|
||||
- race tests across all packages (full run completed in two segments due execution timeout)
|
||||
- rebuilt Gateway binaries for Linux amd64/arm64, macOS arm64, Windows amd64
|
||||
- `dist/SHA256SUMS` verified
|
||||
- real Linux-amd64 release-binary smoke confirmed embedded policy/pull close controls and clean SIGTERM shutdown
|
||||
|
||||
Regression covered: required HTML form controls must not prevent modal cancellation/close.
|
||||
@@ -0,0 +1,84 @@
|
||||
# Checkpoint 27 release verification
|
||||
|
||||
Checkpoint 27 hardens effective context-window handling and `num_ctx` admission/routing.
|
||||
|
||||
## Scope
|
||||
|
||||
Runtime changes are intentionally limited to context estimation, model metadata, context-aware worker eligibility/routing, the policy simulator, configuration validation, and the corresponding Admin UI context columns. Scheduler fairness, authentication, persistence formats, batch execution, conversations, and quota accounting are otherwise unchanged.
|
||||
|
||||
## Effective context model
|
||||
|
||||
The gateway now keeps these values distinct:
|
||||
|
||||
1. **Model maximum** from Ollama `/api/show` model metadata.
|
||||
2. **Configured context** from a Modelfile `PARAMETER num_ctx`, when present.
|
||||
3. **Loaded context** from Ollama `/api/ps context_length`, when the model is resident.
|
||||
4. **Worker default** from `workers[].default_context_tokens` or the gateway context default.
|
||||
5. **Administrative cap** from `workers[].context_limits` and `model_capabilities.context.max_requested_tokens`.
|
||||
|
||||
For OpenAI-compatible requests, the effective usable context is derived from the actually loaded/configured/default context and administrative limits; the theoretical model maximum is no longer treated as the runtime context by itself.
|
||||
|
||||
For native Ollama `/api/*` requests, a valid explicit `options.num_ctx` can request a resize/reload, but is still bounded by the model maximum, gateway cap, and worker context limit. A worker whose currently loaded context is smaller remains eligible only as a resize candidate and does not receive a misleading loaded-context routing advantage.
|
||||
|
||||
## Request estimation fixes
|
||||
|
||||
Context admission now includes:
|
||||
|
||||
- Responses API `max_output_tokens`;
|
||||
- Responses API `instructions`;
|
||||
- native `/api/generate` `suffix`;
|
||||
- configurable estimation margin;
|
||||
- configurable per-image vision reserve;
|
||||
- the largest positive output-token budget rather than silently preferring a smaller field.
|
||||
|
||||
`options.num_ctx` is strictly validated as a positive integer. It is accepted only on native Ollama `/api/*` endpoints. OpenAI-compatible requests cannot inject it as a false context override.
|
||||
|
||||
## Validation performed
|
||||
|
||||
The release candidate passed:
|
||||
|
||||
- `go test ./...`
|
||||
- `go vet ./...`
|
||||
- `node --check internal/webui/assets/app.js`
|
||||
- `go test -race ./...`
|
||||
- strict `-check-config` against the supplied two-M75q production configuration
|
||||
- Linux amd64, Linux arm64, macOS arm64, and Windows amd64 gateway cross-builds
|
||||
- `dist/SHA256SUMS` verification
|
||||
- a real Linux-amd64 release-binary smoke test against the deterministic mock Ollama backend:
|
||||
- `/readyz` ready
|
||||
- Admin API-key session HTTP 200
|
||||
- `/v1/chat/completions` HTTP 200
|
||||
- native string-valued `options.num_ctx` HTTP 400 before backend execution
|
||||
- OpenAI `options.num_ctx` spoof attempt HTTP 400
|
||||
- SIGTERM shutdown exit code 0
|
||||
|
||||
The final ZIP is additionally unpacked and tested again before release.
|
||||
|
||||
## Recommended production settings for the current two-M75q deployment
|
||||
|
||||
The supplied production config uses conservative starting bounds for the ~16 GiB AMD worker class:
|
||||
|
||||
```json
|
||||
"model_capabilities": {
|
||||
"mode": "enforce",
|
||||
"cache_ttl": "10m",
|
||||
"context_guard": "reject",
|
||||
"context": {
|
||||
"max_requested_tokens": 32768,
|
||||
"default_worker_tokens": 4096,
|
||||
"estimation_margin_percent": 15,
|
||||
"vision_reserve_tokens_per_image": 2048
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
and per worker:
|
||||
|
||||
```json
|
||||
"default_context_tokens": 4096,
|
||||
"context_limits": {
|
||||
"*": 16384
|
||||
}
|
||||
```
|
||||
|
||||
`16384` is an operational safety cap, not a claim that every model/workload will fit efficiently at 16k. Increase it only after measuring actual VRAM pressure, offload behavior, latency, and throughput on the target worker/model combination.
|
||||
@@ -0,0 +1,61 @@
|
||||
# Release verification — Checkpoint 28 public status dashboard
|
||||
|
||||
Checkpoint 28 replaces the active P3.2 HA roadmap item with an optional public, unauthenticated, read-only operational dashboard. Existing HA-readiness tools remain packaged but no cluster/coordinator behavior is introduced.
|
||||
|
||||
## Runtime scope
|
||||
|
||||
Changed runtime areas:
|
||||
|
||||
- strict configuration schema: new `public_dashboard` block;
|
||||
- unauthenticated routing for the configured public dashboard path only;
|
||||
- a dedicated sanitized public snapshot builder;
|
||||
- embedded `internal/publicui` assets.
|
||||
|
||||
Unchanged semantics:
|
||||
|
||||
- inference/proxy protocol handling;
|
||||
- scheduler/fairness/quota algorithms;
|
||||
- worker selection and model placement;
|
||||
- durable state schemas and storage files;
|
||||
- admin authentication and admin UI authorization;
|
||||
- batch/conversation persistence.
|
||||
|
||||
No state migration is required.
|
||||
|
||||
## Privacy verification
|
||||
|
||||
The public API is not derived by JSON-marshalling admin/infrastructure objects. Tests require an explicit allow-list and fail if known sensitive fixture values appear in the public response. Verified absent fields/data include tenant, actor, application, internal request ID, worker URL, labels, internal error strings, telemetry errors and estimated credits.
|
||||
|
||||
Worker names and model names default to anonymized display names. Worker aliases can be explicitly supplied in configuration. Resource metrics contain only numeric operational values.
|
||||
|
||||
## Automated verification
|
||||
|
||||
- `go test ./...` — PASS
|
||||
- `go vet ./...` — PASS
|
||||
- `go test -race ./...` — PASS
|
||||
- `node --check internal/publicui/assets/app.js` — PASS
|
||||
- `node --check internal/webui/assets/app.js` — PASS
|
||||
- all four gateway release targets rebuilt — PASS
|
||||
- every entry in `dist/SHA256SUMS` — PASS
|
||||
- `config.example.json` strict parse test — PASS
|
||||
- production-derived checkpoint-28 config through `-check-config` — PASS (2 workers, `/data` writable)
|
||||
|
||||
## Release-binary smoke
|
||||
|
||||
The packaged Linux amd64 gateway was started with a deterministic mock Ollama backend and an isolated state directory.
|
||||
|
||||
Verified:
|
||||
|
||||
- `/readyz` — 200
|
||||
- `/status/` without credentials — 200
|
||||
- `/status/api/snapshot` without credentials — 200
|
||||
- `/gateway/ui-api/session` without credentials — 401
|
||||
- `/gateway/ui-api/session` with admin API key — 200
|
||||
- `/v1/chat/completions` with API key — 200
|
||||
- post-request public snapshot contains the public worker alias and opaque `REQ-*` request ID
|
||||
- post-request public snapshot does not contain private gateway/worker names, auth key, tenant/actor/application fields or worker URL
|
||||
- SIGTERM shutdown — gateway exit code 0
|
||||
|
||||
## Deployment
|
||||
|
||||
Enable the dashboard explicitly in configuration. A recommended production-derived example is supplied separately as `gateway-config.checkpoint28.public-dashboard.json`; it publishes model names but maps the two internal M75q worker names to `GPU Node A` and `GPU Node B`.
|
||||
@@ -0,0 +1,27 @@
|
||||
# Reliability and worker maintenance
|
||||
|
||||
## Circuit breaker
|
||||
|
||||
`reliability.enabled` activates a per-worker circuit breaker. Consecutive transport/backend failures increment the worker failure counter. At `failure_threshold` the circuit enters `open` for `open_duration`. After that period the next eligible request acts as the half-open probe; success closes the circuit, failure reopens it.
|
||||
|
||||
```json
|
||||
"reliability": {
|
||||
"enabled": true,
|
||||
"failure_threshold": 3,
|
||||
"open_duration": "30s",
|
||||
"retry_attempts": 2,
|
||||
"retry_backoff": "50ms"
|
||||
}
|
||||
```
|
||||
|
||||
Retries are deliberately conservative: only transport failures before any upstream response is committed to the client are retried. A backend HTTP 5xx contributes to the circuit breaker but is not replayed once its headers/body have begun.
|
||||
|
||||
## Worker maintenance
|
||||
|
||||
The admin Worker page exposes:
|
||||
|
||||
- `active`: accepts new inference jobs;
|
||||
- `draining`: no new inference jobs, existing jobs continue;
|
||||
- `disabled`: no new inference jobs until explicitly enabled.
|
||||
|
||||
Drain/disable state is persisted in `<data_dir>/worker-state.json`. Circuit state is intentionally transient and can be manually reset from the Worker page.
|
||||
@@ -0,0 +1,80 @@
|
||||
# Usage retention and long-term aggregation
|
||||
|
||||
The gateway stores request-level usage as append-only daily JSONL journals and compacts old data into bounded rollups.
|
||||
|
||||
## Default policy
|
||||
|
||||
```json
|
||||
"usage": {
|
||||
"journal_dir": "./data/usage",
|
||||
"buffer": 16384,
|
||||
"flush_interval": "1s",
|
||||
"retention": {
|
||||
"detail_days": 30,
|
||||
"daily_days": 400,
|
||||
"monthly_months": 0,
|
||||
"compaction_interval": "6h"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `detail_days`: keep per-request records for this many calendar days.
|
||||
- `daily_days`: retain one daily rollup until this age. Must be >= `detail_days`.
|
||||
- `monthly_months`: retain monthly rollups for this many months. `0` means forever.
|
||||
- `compaction_interval`: background compaction cadence; minimum one minute.
|
||||
|
||||
## Storage layout
|
||||
|
||||
```text
|
||||
data/usage/
|
||||
├── usage-2026-09-06.jsonl
|
||||
└── rollups/
|
||||
├── daily/
|
||||
│ └── rollup-daily-2026-07-23.json
|
||||
└── monthly/
|
||||
└── rollup-monthly-2025-08.json
|
||||
```
|
||||
|
||||
Each daily/monthly rollup retains aggregate dimensions for global traffic, tenant, actor, application, model, and worker.
|
||||
|
||||
Measures retained after request details expire include request/error counts, prompt/completion/cached tokens, credits, queue/service time, prompt/eval nanoseconds, bytes in/out, and last-request timestamp. Prompt and output tokens/second remain derivable from aggregate evaluation durations.
|
||||
|
||||
## Crash/idempotency behavior
|
||||
|
||||
Raw -> daily compaction writes and fsyncs the new daily file before deleting the source journal. Repeating this after a crash replaces the same daily period rather than summing it twice.
|
||||
|
||||
Daily -> monthly files retain a map of their source days. If a process dies after the monthly file is committed but before the source daily file is deleted, replaying compaction replaces that day entry. This makes the monthly fold idempotent.
|
||||
|
||||
## Startup and all-time accounting
|
||||
|
||||
Before replaying request journals, startup applies retention once. It then reconstructs all-time usage from retained monthly rollups, retained daily rollups, and the remaining request-level journals. Only remaining request-level journals populate the recent-request table.
|
||||
|
||||
If `monthly_months > 0`, data older than that window is deleted completely and removed from reconstructed all-time usage.
|
||||
|
||||
## Admin/API
|
||||
|
||||
```text
|
||||
POST /gateway/ui-api/storage/compact
|
||||
GET /gateway/ui-api/usage/rollups?granularity=daily&dimension=global&limit=120
|
||||
GET /gateway/ui-api/usage/rollups?granularity=daily&dimension=tenant&name=team-a
|
||||
GET /gateway/ui-api/usage/rollups?granularity=monthly&dimension=model&name=qwen3:8b
|
||||
```
|
||||
|
||||
Allowed dimensions: `global`, `tenant`, `actor`, `application`, `model`, `worker`.
|
||||
|
||||
The Admin UI exposes the same controls under **Persistenz** and historical aggregates under **Nutzung**.
|
||||
|
||||
## Prometheus
|
||||
|
||||
```text
|
||||
ollama_gateway_usage_raw_files
|
||||
ollama_gateway_usage_daily_rollup_files
|
||||
ollama_gateway_usage_monthly_rollup_files
|
||||
ollama_gateway_usage_raw_bytes
|
||||
ollama_gateway_usage_daily_rollup_bytes
|
||||
ollama_gateway_usage_monthly_rollup_bytes
|
||||
ollama_gateway_usage_last_compaction_timestamp_seconds
|
||||
ollama_gateway_usage_last_reclaimed_bytes
|
||||
```
|
||||
|
||||
No tenant/user labels are added to `/metrics`; dimensional long-term analysis belongs to the rollup API/UI, avoiding unbounded Prometheus cardinality.
|
||||
@@ -0,0 +1,195 @@
|
||||
# Ollama Fair Gateway — P0–P3 implementation roadmap
|
||||
|
||||
This roadmap turns the gateway from a local fair scheduler into a production-oriented LLM control plane while preserving the project's core properties: one Go binary, no mandatory external state service, streaming-first proxying, and no prompt/response content persistence by default.
|
||||
|
||||
## Delivery principles
|
||||
|
||||
1. **Never block the inference hot path on durable storage.** Control-plane writes may be synchronous when correctness requires it; usage/metrics remain buffered/snapshotted.
|
||||
2. **Never retry after client-visible output starts.** A retry is permitted only before response headers/body are committed.
|
||||
3. **Hard policies precede adaptive scoring.** Model ACL → alias resolution → capability preflight → placement → maintenance/circuit eligibility → adaptive routing.
|
||||
4. **Every runtime control must be observable and persistent when operationally meaningful.** Drain/disable state, API keys, policies, placement, aliases and durable batch definitions survive restart; active sockets, transient inference jobs and in-flight batch attempt contexts do not.
|
||||
5. **Protocol compatibility remains explicit.** Native Ollama, OpenAI-compatible and later Anthropic-compatible errors/streaming semantics are handled independently.
|
||||
|
||||
---
|
||||
|
||||
## P0 — Production safety and policy foundation
|
||||
|
||||
### P0.1 Circuit breaker + safe pre-stream retry
|
||||
|
||||
**Goal:** stop repeatedly routing to unhealthy/OOM/transport-failing workers.
|
||||
|
||||
- Worker circuit states: `closed`, `open`, `half_open`.
|
||||
- Configurable consecutive-failure threshold and open duration.
|
||||
- Transport failures and backend 5xx contribute to the circuit.
|
||||
- Only transport failures that happen before response commitment are retried.
|
||||
- Retry excludes workers that already failed the current request.
|
||||
- Admin UI shows circuit state, last circuit error and manual reset.
|
||||
- Prometheus counters planned for opens/retries/failures.
|
||||
|
||||
**Acceptance:** a connection-reset worker opens its circuit and a request succeeds on another eligible worker without duplicate client-visible output.
|
||||
|
||||
### P0.2 Virtual models / aliases
|
||||
|
||||
**Goal:** clients use stable names such as `fast`, `coding`, `vision` rather than physical Ollama tags.
|
||||
|
||||
- Ordered fallback list of real models.
|
||||
- Optional required capability set.
|
||||
- Alias participates in `/api/tags` and `/v1/models` discovery.
|
||||
- Gateway rewrites the outbound model while returning diagnostic headers:
|
||||
- `X-Gateway-Model-Alias`
|
||||
- `X-Gateway-Resolved-Model`
|
||||
- Alias target must still satisfy worker placement and inventory.
|
||||
|
||||
**Acceptance:** OpenWebUI can select a virtual model and the backend receives the resolved physical model.
|
||||
|
||||
### P0.3 Model ACLs
|
||||
|
||||
**Goal:** answer “who may use which model?” independently from placement (“where may it run?”).
|
||||
|
||||
- Tenant baseline rules.
|
||||
- API-key-specific allow/deny rules.
|
||||
- Exact and trailing `*` patterns; most specific match wins, deny wins ties.
|
||||
- ACL is enforced for direct model requests and model discovery.
|
||||
- API-key create UI exposes allow/deny lists.
|
||||
|
||||
**Acceptance:** a key allowed only for `coding` cannot discover or directly call a denied physical model.
|
||||
|
||||
### P0.4 Worker maintenance / drain
|
||||
|
||||
**Goal:** take workers out of rotation without killing existing streams.
|
||||
|
||||
- `active`: accepts new work.
|
||||
- `draining`: no new work; active jobs finish.
|
||||
- `disabled`: no new work until explicitly re-enabled.
|
||||
- State persisted in `worker-state.json`.
|
||||
- UI buttons: Drain, Disable, Activate; circuit reset beside them.
|
||||
|
||||
**Acceptance:** a draining worker's active job completes while all new jobs route elsewhere.
|
||||
|
||||
---
|
||||
|
||||
## P1 — Client reach, QoS and observability
|
||||
|
||||
### P1.1 Anthropic `/v1/messages`
|
||||
|
||||
- Native Anthropic-compatible request/stream passthrough to Ollama.
|
||||
- Token/usage extraction and protocol-correct errors.
|
||||
- Tool, vision and thinking capability preflight.
|
||||
- Same ACL/quota/placement/scheduler pipeline as Ollama/OpenAI.
|
||||
|
||||
### P1.2 Service classes / priority
|
||||
|
||||
- Classes: `interactive`, `system`, `background`, `batch`.
|
||||
- Weighted scheduling without starvation.
|
||||
- Per-class queue wait and concurrency ceilings.
|
||||
- API-key/default mapping and optional request header override with scope.
|
||||
|
||||
### P1.3 Auto-tuning and benchmark profiles
|
||||
|
||||
- Measure TTFT, prompt tok/s, output tok/s, throughput and VRAM by model/concurrency.
|
||||
- Suggest `max_concurrent` / per-model concurrency.
|
||||
- Optional “apply recommendation” workflow with audit entry.
|
||||
- Never auto-change production settings unless explicitly enabled.
|
||||
|
||||
### P1.4 OpenTelemetry
|
||||
|
||||
- OTLP traces/metrics, content capture off by default.
|
||||
- Spans: auth, admission, queue, route, upstream, first-byte/stream.
|
||||
- Correlate with `X-Request-ID`.
|
||||
|
||||
---
|
||||
|
||||
## P2 — Capacity automation and operator workflows
|
||||
|
||||
### P2.1 Warm/preload policies — ✅ implemented
|
||||
|
||||
- `hot`, `warm`, `cold` model classes.
|
||||
- Explicit/preferred workers.
|
||||
- Idle unload using Ollama `keep_alive: 0`.
|
||||
- Optional pre-warm at gateway start/worker recovery.
|
||||
- Memory-pressure-aware eviction suggestions.
|
||||
|
||||
### P2.2 Alerts and webhooks — ✅ implemented
|
||||
|
||||
- Conditions: worker down, circuit open, repeated OOM, queue depth/wait, storage growth, quota near exhaustion.
|
||||
- Cooldown/deduplication.
|
||||
- Generic signed webhook first; provider-specific integrations later.
|
||||
|
||||
### P2.3 Optional stateful conversations — ✅ implemented
|
||||
|
||||
- Opt-in conversation store for clients that need `previous_response_id` semantics.
|
||||
- Separate encryption/retention policy because this stores content.
|
||||
- Disabled by default to preserve current privacy posture.
|
||||
|
||||
### P2.4 Policy simulator — ✅ implemented early in P0/P1
|
||||
|
||||
- Simulate identity/model/capabilities/context without executing inference.
|
||||
- Explain ACL, alias, placement, inventory, circuit, drain, concurrency and final routing score.
|
||||
- UI provides a step-by-step decision trace.
|
||||
|
||||
---
|
||||
|
||||
## P3 — Batch and public operations visibility
|
||||
|
||||
### P3.1 Batch jobs — ✅ implemented
|
||||
|
||||
- Durable job definitions and status.
|
||||
- Separate background scheduling class.
|
||||
- Pause/resume/cancel.
|
||||
- Input/output references rather than embedding large payloads in control state.
|
||||
- Retention and accounting integrated with existing usage rollups.
|
||||
|
||||
### P3.2 Public status dashboard — ✅ implemented
|
||||
|
||||
- Separate unauthenticated, strictly read-only status surface.
|
||||
- Current queue/load, worker capacity, infrastructure map and anonymized live flow.
|
||||
- Explicit public-data allow-list; no tenant/actor/application identities or admin/control-plane state.
|
||||
- Worker/model-name privacy controls and public worker aliases.
|
||||
- HA/gateway clustering is removed from the active roadmap because observed single-process performance does not justify the distributed coordination cost. The HA-readiness tooling remains available if that operational assumption changes later.
|
||||
|
||||
---
|
||||
|
||||
## Planned implementation sequence
|
||||
|
||||
1. **Release A — P0 reliability foundation**: circuit breaker, safe retry, drain/disable persistence, UI controls.
|
||||
2. **Release B — P0 policy surface**: model aliases and model ACLs with full runtime CRUD UI and persistence.
|
||||
3. **Release C — P1 protocol/QoS**: Anthropic Messages + service classes.
|
||||
4. **Release D — P1 observability/tuning**: OpenTelemetry + benchmark recommendations.
|
||||
5. **Release E — P2 automation**: warm model manager, alerts, policy simulator.
|
||||
6. **Release F — P2 stateful optional layer**: conversations with explicit content-retention controls.
|
||||
7. **Release G — P3 batch**.
|
||||
8. **Release H — P3 public status dashboard**; multi-gateway HA is deferred outside the active roadmap.
|
||||
|
||||
## Current implementation status
|
||||
|
||||
- P0.1: **implemented** (circuit breaker, safe pre-stream retry and reliability metrics).
|
||||
- P0.2: **implemented** (aliases, discovery, capability-qualified resolution and persistent runtime CRUD from the admin API/UI with immediate publication for new requests).
|
||||
- P0.3: **implemented** (tenant/API-key ACL enforcement, API-key ACL CRUD and persistent runtime Tenant Model Access CRUD from the admin API/UI).
|
||||
- P0.4: **implemented** (persistent drain/disable + UI controls).
|
||||
- P1.1: **implemented** (Anthropic Messages routing, metering and protocol-aware handling).
|
||||
- P1.2: **implemented** (weighted service classes with queue/concurrency controls and scoped override header).
|
||||
- P1.3: **implemented** (explicit benchmark profiles and apply-recommendation workflow).
|
||||
- P1.4: **implemented** (OTLP/HTTP tracing with content capture off by default).
|
||||
- P2.1: **implemented** (warm/preload policies).
|
||||
- P2.2: **implemented** (alerts + signed webhooks).
|
||||
- P2.3: **implemented** (optional AES-256-GCM encrypted Responses conversation store with retention, `store:false`, identity scoping and `previous_response_id` expansion).
|
||||
- P2.4: **implemented** (policy simulator).
|
||||
- P3.1: **implemented** (durable metadata + separate content spool, batch service-class execution through the normal gateway pipeline, pause/resume/cancel, restart recovery, retention, owner API, admin UI/API and accounting integration).
|
||||
- P3.2: **implemented** as the public status dashboard in checkpoint 28. Multi-gateway HA is intentionally removed from the active roadmap after observed single-process performance showed no current capacity justification. The existing HA-readiness tools remain as diagnostic evidence tooling, not as a commitment to add distributed coordination.
|
||||
|
||||
### Checkpoint 23 note
|
||||
|
||||
At checkpoint 23, HA remained gated. Checkpoint 23 further reduces single-node production risk by hardening the Docker/reverse-proxy boundary: loopback-only publishing by default, an explicit production Compose file with no example-config fallback, read-only container rootfs/capability dropping/no-new-privileges, and a host preflight that rejects accidental public binding. This does not constitute HA implementation and does not change inference or persistent-state semantics.
|
||||
|
||||
### Checkpoint 24 note
|
||||
|
||||
At checkpoint 24, HA remained gated. Checkpoint 24 does not alter runtime scheduling, routing, persistence or HA semantics; it replaces the dark admin presentation with a professional light, data-dense control-plane theme. The embedded Live Flow and Infrastructure canvases were updated together with CSS so the UI remains visually coherent. No config/state migration is required.
|
||||
|
||||
### Checkpoint 27 note
|
||||
|
||||
At checkpoint 27, HA remained gated. Checkpoint 27 hardens context-window admission and routing before any HA work: the gateway no longer equates the theoretical `/api/show` model maximum with the context actually available on a worker. Effective context now derives from loaded `/api/ps context_length`, Modelfile `num_ctx`, explicit worker defaults and per-worker/model caps. Native `options.num_ctx` is validated/capped, while OpenAI/Responses estimation now covers `max_output_tokens`, `instructions`, `suffix`, vision reserves and an admission margin. Context-suitable workers are selected before inference, and the policy simulator/model UI expose the effective context evidence.
|
||||
|
||||
|
||||
### Checkpoint 28 note
|
||||
|
||||
Checkpoint 28 closes the active P3 roadmap with an optional public read-only status dashboard. `/status/` exposes current load, queue, worker capacity, an infrastructure map and anonymized live request state through a dedicated sanitized API. It never reuses the admin snapshot schema. HA remains available only as a future re-evaluation if measured availability or capacity requirements change.
|
||||
@@ -0,0 +1,76 @@
|
||||
# Security notes
|
||||
|
||||
## OIDC
|
||||
|
||||
The gateway validates bearer JWTs against OIDC discovery and JWKS. It rejects `alg=none`, restricts accepted signature algorithms, validates issuer, audience, expiry and `nbf`, and refreshes JWKS when a `kid` is unknown. API bearer-token validation is also the source of truth for the optional browser Authorization Code + PKCE login flow.
|
||||
|
||||
## Trusted-IP bypass
|
||||
|
||||
IP bypass is evaluated before bearer authentication. Starting with checkpoint 22, the **TCP peer address** is the authentication source for `auth.ip_bypass` by default, even when `X-Forwarded-For` is accepted for logging and usage attribution. This prevents a trusted-proxy chain from turning a spoofed forwarded address into a credential-free identity.
|
||||
|
||||
Set `auth.ip_bypass_use_forwarded_ip=true` only when a deployment intentionally needs an original client address behind a reverse proxy to satisfy IP bypass. That mode requires a tightly restricted `auth.trusted_proxies` list plus network ACLs that prevent clients from reaching the gateway directly. Prefer narrow `/32` or `/128` bypass/proxy entries. Never configure a broad Internet-wide bypass for an exposed gateway.
|
||||
|
||||
## API keys
|
||||
|
||||
API keys are represented in the authenticator by their SHA-256 digest; plaintext runtime secrets are not retained. Clients may send either `X-API-Key` or `Authorization: Bearer ...`.
|
||||
|
||||
There are two sources:
|
||||
|
||||
- **Configuration keys** come from `auth.api_keys`. Put their secrets in environment variables and reference them with `${NAME}` in JSON. They survive restarts and are read-only in the web UI.
|
||||
- **Runtime keys** are created by an administrator in the web UI or through `POST /gateway/ui-api/api-keys`. A cryptographically random `ofg_...` secret is returned once, then only the digest and metadata are persisted in the local key store. UI-created keys can be revoked immediately and survive process restarts until revoked.
|
||||
|
||||
Only identities with `gateway:admin` may list, create or revoke API keys. Cookie-authenticated mutations are covered by the same CSRF checks as the rest of the control plane. Key secrets are never included in list responses, the redacted configuration endpoint or admin logs.
|
||||
|
||||
## Ollama management
|
||||
|
||||
With `native.management_requires_admin=true`, model-management operations require `gateway:admin`. Do not expose Ollama port 11434 directly to untrusted clients; expose only the gateway.
|
||||
|
||||
## Web control plane
|
||||
|
||||
Browser OIDC uses Authorization Code + PKCE. The transient verifier/state is HMAC-protected with `ui.session_secret`. After exchange, the browser receives only an opaque session identifier; the OIDC access token remains in the gateway's in-memory session store.
|
||||
|
||||
A process restart invalidates all browser sessions. Cookie-authenticated state-changing requests require a same-origin CSRF token. The embedded UI uses a restrictive Content Security Policy and has no CDN, analytics or third-party JavaScript dependency.
|
||||
|
||||
Policy overrides are persisted locally and take effect immediately for new requests.
|
||||
|
||||
## In-memory security boundary
|
||||
|
||||
Removing external state storage reduces the network attack surface: there is no coordination service, state-store credential or message-bus endpoint to protect. Sensitive runtime material such as OIDC session tokens and quota state exists only in process memory.
|
||||
|
||||
A crash or restart clears browser sessions and active transient inference jobs, but API keys, policy overrides, quota state, usage history, metric counters and durable batch definitions are restored from local durable state.
|
||||
|
||||
## Optional content-bearing conversations
|
||||
|
||||
The default gateway privacy posture remains content-free: prompts and generated responses are not persisted unless `conversations.enabled=true` is explicitly configured. The optional Responses conversation store is isolated from usage retention and uses AES-256-GCM encryption at rest, a dedicated retention window, an entry cap, and a per-context byte cap.
|
||||
|
||||
`previous_response_id` lookup is bound to the authenticated tenant and scheduler actor. A response ID from another identity is reported exactly like a nonexistent ID to avoid enumeration. `store:false` prevents persistence of the new response. The Admin UI exposes store status but not decrypted conversation content. See `docs/CONVERSATIONS.md`.
|
||||
|
||||
The conversation encryption key is a deployment secret. Losing or changing it makes the existing ciphertext unreadable; exposing it together with the state file defeats at-rest confidentiality. Backups should therefore be protected as secrets.
|
||||
|
||||
## Optional content-bearing batch jobs
|
||||
|
||||
Durable batch jobs are also disabled by default. Enabling them does not persist content by itself, but every submitted job stores its request body and later its response body under `storage.batch_jobs_dir`. These spool payloads are **not encrypted by the gateway**; they rely on mode-0600 files, a mode-0750 spool directory and the deployment filesystem/disk security boundary. Use filesystem or full-disk encryption if at-rest content encryption is required.
|
||||
|
||||
Batch metadata stores an authenticated identity/authorization snapshot but never the bearer token or API-key secret. Owner-facing `/gateway/v1/batches` endpoints are tenant+actor scoped; cross-identity IDs are hidden as not found. Global batch inspection/control and output download are restricted to the admin UI/API and therefore require `gateway:admin`.
|
||||
|
||||
An accepted job can execute after the original credential has been revoked because the credential secret is intentionally not retained. Current tenant model policy, worker placement/maintenance/circuit state, quotas and scheduling are still evaluated when each attempt runs. Operationally, revoke the credential **and cancel its accepted durable jobs** when retroactive cancellation is required.
|
||||
|
||||
Batch input/output files and backups can contain prompts, tool arguments, retrieved data and model output. Configure a short `batch_jobs.retention`, protect `storage.data_dir`, and treat backup archives as content-bearing secrets. See `docs/BATCH-JOBS.md`.
|
||||
|
||||
## Live telemetry
|
||||
|
||||
Live telemetry contains tenant/actor identifiers, request IDs, model names, worker names, timings, token counts and credit estimates, but never prompt text or generated model output. UI telemetry endpoints require an authenticated administrator.
|
||||
|
||||
`workers[].telemetry_url` is administrator-controlled and fetched by the gateway. Point it only to a trusted exporter because it is an SSRF-capable configured URL, just like an Ollama worker URL. The optional shipped worker-telemetry agent should be bound to a trusted management interface and restricted with `-allow-cidrs`; it intentionally does not create a separate application-secret store.
|
||||
|
||||
## Model access control
|
||||
|
||||
Model access control is independent from worker placement. Tenant defaults are configured under `model_access`; API keys can carry a narrower or different allow/deny list. The gateway applies the identity ACL before scheduling and filters `/api/tags` and `/v1/models` accordingly.
|
||||
|
||||
Patterns are exact model/alias names, `*`, or one trailing wildcard such as `qwen3:*`. The most specific matching rule wins; deny wins a specificity tie. API-key ACL metadata is persisted together with the key hash and is editable at creation time in **Admin -> Sicherheit -> API Keys**.
|
||||
|
||||
## Reverse-proxy network boundary
|
||||
|
||||
Prefer publishing the gateway only on loopback when the reverse proxy is local to the Docker host. A proxy should be the only network component able to reach the gateway's published port. Avoid trusting whole RFC1918 ranges merely because the deployment uses private addressing; list only concrete proxy/NAT peer addresses in `auth.trusted_proxies`.
|
||||
|
||||
Checkpoint 23's Compose defaults additionally run with a read-only root filesystem, all Linux capabilities dropped and `no-new-privileges`. Persistent writes are constrained to `/data`; `/tmp` is ephemeral. These controls reduce the impact of a process compromise but do not replace API-key/OIDC authentication or host firewalling.
|
||||
@@ -0,0 +1,65 @@
|
||||
# Warm Model Management
|
||||
|
||||
The warm-model manager controls Ollama model residency without putting model-management I/O on the inference hot path.
|
||||
|
||||
## Classes
|
||||
|
||||
- `hot`: keep the model resident on the requested number of eligible workers. A hot model is proactively preloaded.
|
||||
- `warm`: optionally preload the model and unload it after `idle_timeout`.
|
||||
- `cold`: do not preload and unload aggressively after `idle_timeout`.
|
||||
|
||||
Rules support exact names and one trailing `*` wildcard. The most specific matching rule wins.
|
||||
|
||||
```json
|
||||
{
|
||||
"warm_models": {
|
||||
"enabled": true,
|
||||
"reconcile_interval": "30s",
|
||||
"operation_timeout": "2m",
|
||||
"policies": {
|
||||
"qwen3:8b": {
|
||||
"class": "hot",
|
||||
"workers": ["rtx-4090"],
|
||||
"replicas": 1,
|
||||
"preload": true,
|
||||
"idle_timeout": "30m"
|
||||
},
|
||||
"gemma4:*": {
|
||||
"class": "warm",
|
||||
"replicas": 1,
|
||||
"preload": false,
|
||||
"idle_timeout": "20m"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Safety rules
|
||||
|
||||
Warm-model operations are subordinate to normal routing policy:
|
||||
|
||||
1. worker must be healthy and in `active` maintenance mode;
|
||||
2. preload must be allowed by Model Placement;
|
||||
3. when worker inventory is known, preload requires the model to be installed;
|
||||
4. workers with an open/half-open circuit are not used for proactive preload;
|
||||
5. a preload/unload reserves a per-worker/per-model maintenance slot before calling Ollama;
|
||||
6. inference acquisition for the same worker/model is blocked during that short operation;
|
||||
7. unload is refused while the model has active inference jobs.
|
||||
|
||||
The maintenance reservation closes the race where an idle unload could otherwise begin at the same moment a new request is routed to the model.
|
||||
|
||||
## Persistence
|
||||
|
||||
Runtime rules edited in **Admin → Warm Models** are stored in `warm-models.json`. The static config remains the baseline and can be restored with **Auf Config zurücksetzen**.
|
||||
|
||||
Action history is retained with the same state file. Last-use timestamps are deliberately runtime state; after restart, currently loaded models receive a fresh grace period before idle eviction.
|
||||
|
||||
## VRAM pressure
|
||||
|
||||
The manager does not automatically evict arbitrary models under memory pressure. It exposes eviction suggestions for inactive, non-hot models. This avoids surprise unloads and lets the operator decide whether to act.
|
||||
|
||||
## Metrics
|
||||
|
||||
- `ollama_gateway_warm_actions_running`
|
||||
- `ollama_gateway_warm_eviction_suggestions`
|
||||
@@ -0,0 +1,127 @@
|
||||
# Remote worker telemetry
|
||||
|
||||
Ollama Fair Gateway can consume optional out-of-band resource telemetry through `workers[].telemetry_url`. Checkpoint 20 ships a small `worker-telemetry` agent so remote Ollama hosts can report their own RAM/GPU state instead of accidentally reporting the gateway host through `local_system_stats`.
|
||||
|
||||
The agent is **not** part of the inference request path. Gateway workers fetch it during the normal health/telemetry refresh cycle.
|
||||
|
||||
## Why this matters for remote workers
|
||||
|
||||
`local_system_stats: true` runs the memory collector inside the gateway process. It is correct only when the Ollama worker and gateway share the same host. For a worker such as `http://10.2.10.48:11434`, it otherwise reports the gateway machine's RAM as if it belonged to `10.2.10.48`.
|
||||
|
||||
For remote workers, use:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "M75q - Gen5 - 1048",
|
||||
"url": "http://10.2.10.48:11434",
|
||||
"local_system_stats": false,
|
||||
"telemetry_url": "http://10.2.10.48:11500/telemetry"
|
||||
}
|
||||
```
|
||||
|
||||
and run the agent on that Ollama host.
|
||||
|
||||
## Validate locally first
|
||||
|
||||
Host memory only:
|
||||
|
||||
```sh
|
||||
./ollama-gateway-worker-telemetry-linux-amd64 -once
|
||||
```
|
||||
|
||||
AMD GPU through the Linux amdgpu sysfs interface:
|
||||
|
||||
```sh
|
||||
./ollama-gateway-worker-telemetry-linux-amd64 -once -amd-sysfs
|
||||
```
|
||||
|
||||
If auto-detection picks the wrong GPU, specify the device explicitly:
|
||||
|
||||
```sh
|
||||
./ollama-gateway-worker-telemetry-linux-amd64 \
|
||||
-once \
|
||||
-amd-sysfs \
|
||||
-amd-device /sys/class/drm/card1/device
|
||||
```
|
||||
|
||||
NVIDIA:
|
||||
|
||||
```sh
|
||||
./ollama-gateway-worker-telemetry-linux-amd64 -once -nvidia-smi
|
||||
```
|
||||
|
||||
The JSON schema matches the existing gateway `telemetry_url` contract:
|
||||
|
||||
```json
|
||||
{
|
||||
"memory_used_bytes": 123,
|
||||
"memory_total_bytes": 456,
|
||||
"vram_used_bytes": 789,
|
||||
"vram_total_bytes": 1024,
|
||||
"gpu_utilization_percent": 42,
|
||||
"gpu_temperature_c": 63.5,
|
||||
"gpu_power_watts": 88,
|
||||
"source": "host-memory+amdgpu-sysfs",
|
||||
"updated_at": "2026-09-08T18:00:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
A field may be absent when the operating system/driver does not expose it. The agent still returns the remaining usable metrics and puts collector failures in the `error` field.
|
||||
|
||||
## Run as a service
|
||||
|
||||
The listener defaults to loopback for safety. To expose it to the gateway, bind a worker interface and restrict clients by CIDR. Prefer an exact gateway host address when possible:
|
||||
|
||||
```sh
|
||||
./ollama-gateway-worker-telemetry-linux-amd64 \
|
||||
-listen 0.0.0.0:11500 \
|
||||
-allow-cidrs 10.2.19.1/32 \
|
||||
-amd-sysfs
|
||||
```
|
||||
|
||||
Example systemd unit:
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=Ollama Gateway Worker Telemetry
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
ExecStart=/usr/local/bin/ollama-gateway-worker-telemetry-linux-amd64 -listen 0.0.0.0:11500 -allow-cidrs 10.2.19.1/32 -amd-sysfs
|
||||
Restart=on-failure
|
||||
User=nobody
|
||||
NoNewPrivileges=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=true
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
Adjust the CIDR to the source IP the worker actually sees from the gateway. Do not use `0.0.0.0/0` merely to make the endpoint reachable.
|
||||
|
||||
## AMD telemetry details
|
||||
|
||||
On Linux the agent auto-detects the first `/sys/class/drm/card*/device` whose PCI vendor is `0x1002`. It reads standard amdgpu attributes when available:
|
||||
|
||||
- `mem_info_vram_total`
|
||||
- `mem_info_vram_used`
|
||||
- `gpu_busy_percent`
|
||||
- `hwmon/*/temp1_input`
|
||||
- `hwmon/*/power1_average`
|
||||
|
||||
No ROCm library, `rocm-smi`, NVML or CGO is required for the AMD path. Integrated/shared-memory GPUs may expose different or incomplete VRAM semantics; keep the statically configured `memory_capacity_bytes`/`vram_capacity_bytes` values as hard capacity hints and treat runtime telemetry as routing pressure information.
|
||||
|
||||
## Security
|
||||
|
||||
The agent intentionally has no application credential store. It relies on a narrow listener/firewall/CIDR allowlist so it remains dependency-free and does not introduce a second secret lifecycle. Resource telemetry can still reveal infrastructure details, so expose it only on a trusted management network, VPN or host firewall. For untrusted network segments, put the endpoint behind a TLS-authenticated reverse proxy rather than exposing it directly.
|
||||
|
||||
`workers[].telemetry_url` remains administrator-controlled and is an SSRF-capable URL. Point it only to trusted worker exporters.
|
||||
|
||||
## Freshness and caching
|
||||
|
||||
Checkpoint 21 validates `updated_at` when the external exporter supplies it. The shipped agent always sends the field and also returns `Cache-Control: no-store`.
|
||||
|
||||
A sample is rejected when it is older than `max(30s, 6 x workers[].health_interval)` or more than 30 seconds in the future. Rejected values are not merged into the routing telemetry; the reason is exposed through the worker telemetry error instead. This prevents a caching proxy or badly skewed exporter clock from silently presenting old resource pressure as current data.
|
||||
|
||||
For compatibility, third-party exporters that omit `updated_at` are still accepted. Their freshness cannot be verified, so new integrations should always provide an RFC3339 timestamp.
|
||||
Reference in New Issue
Block a user