Review point on #7239: the two setup checks in joinClient reported through
t.Fatalf while every other check in the suite goes through require. Same
outcome, one less shape to read.
The WaitProxyPeer check stays behind an if rather than being passed straight
to require.NoError: the message interpolates the proxy container's whole log,
and require evaluates its arguments before it knows the assertion passed. As
written the fetch happens only on the failure it exists to explain.
ResolveProxyIP exists to wake the lazy proxy peer, and it retried only
curl exit 6 — DNS. The wake-up attempt that arrives before WireGuard has
brought the tunnel up fails with exit 7 instead, and that returned
immediately:
no HTTP response from vast-azalea.netbird.local: exit status 7
(curl: (7) Failed to connect ... after 0 ms)
So the one function whose job is to tolerate a not-yet-ready endpoint
failed on the readiness state it was written for, one second after the
client container reported ready. Retry both exit codes within the same
window; anything else would still be failing when the window closed and
still fails immediately.
Raise the access-log ingest window to 60s for the same reason. The proxy
streams each entry with a 10s send timeout of its own, so 30s left
barely three attempts of headroom before a test that had already got its
200 was failed for a row still in flight.
The first live run answered the question the mock could not. OpenAI and
Anthropic both filter correctly against real catalogues — Anthropic's
dated claude-haiku-4-5-20251001 survives a record registering the
undated id, and OpenAI's listing comes back as the single model the
guardrail permits. Vertex is refused by the proxy, as intended.
Bedrock is the one that was wrong, and wrong about something worth
recording: GET /inference-profiles reaches AWS and AWS answers
<UnknownOperationException/>. ListInferenceProfiles is a control-plane
operation on bedrock.<region>.amazonaws.com; a provider record carries a
single upstream and it must be the runtime host for InvokeModel to work,
so no Bedrock record can serve a listing as the model stands. The mock
serves that path on the same listener as everything else, which is
exactly why this went unnoticed.
Replace the routed/filtered pair with an explicit outcome, since the
three cases are different contracts rather than degrees of success, and
tell apart 'the proxy refused' from 'the vendor refused' by whether the
body names a middleware — no upstream error body does. The two
non-listing outcomes now issue a single request instead of retrying for
the full window waiting on a status that is never coming, which is where
92 of the failing run's 136 seconds went.
The mock upstream advertises ids we chose, so a listing narrowing to the
ones we authorised is arithmetic we controlled both sides of. It cannot
show the filter surviving a real catalogue: ids we never enumerated,
dated builds whose suffix the vendor picks, surfaces that answer a
listing request with something that is not a listing.
Cover the four surfaces against their real endpoints, each gated on its
own credential so a partial key set still yields partial coverage:
- OpenAI enumerates two real models and the policy permits one, so
both bounds are observable at once against a catalogue of dozens.
- Anthropic returns dated build ids while the record registers the
undated one, which exercises date-normalisation on ids the vendor
chose. This is also the surface Claude Code actually calls.
- Bedrock lists inference profiles rather than models; the request is
routed but not model-bounded, since filtering keys on /v1/models.
- Vertex serves no listing at all, so discovery must be refused rather
than rewritten onto an upstream that would 404 it.
One proxy serves every case, with a group, policy and client per
provider: a model-less request matches exactly one route, so two
providers authorised for the same caller would leave one untested.
Every response is logged before anything is asserted on it. A live
catalogue is the one input the suite does not control, so a failure has
to arrive carrying the response that caused it.
The discovery isolation test drove a single client in the main group.
VLLMUnlistedModel's absence from that client's listing was consistent
with two different worlds: the listing being scoped to the caller's
policy, or the other team's policy never having applied at all. The
test passed either way, so it did not prove what it claimed.
Mint a setup key per group and join a second client on the other
group, reusing the running proxy. The other client must see its own
model before the main client's listing is asserted, and must not see
the main group's model — isolation is checked in both directions.
The listing was narrowed by the provider record's enumerated models, which is
the right bound only while one policy reaches a provider. Where two teams
share a provider under different allowlists, every caller was offered the
union: each model outside their own policy is a request the guardrail refuses
a moment later, which is the empty-or-wrong picker this endpoint exists to
avoid, moved one level up. A gateway record enumerating nothing was worse
still — it offered the upstream's entire catalogue however narrow the policy.
The synthesiser already knows which policies authorise a provider and which
groups each binds, so the router can answer this at request time where it
knows the caller's groups. Each route now carries one rule per authorising
policy — its source groups and the models it permits — and the listing is
bounded to the union across the rules matching the caller, intersected with
what the provider serves.
This is deliberately finer than the guardrail's own per-provider allowlist,
which stays as it is: that list is a fail-closed backstop and cannot tell who
is asking, so discovery is now narrower than the backstop rather than wider.
A policy setting no allowlist lifts the restriction for the groups it binds,
so nil and empty model lists stay distinct end to end — collapsing them would
let a listing that should offer nothing fall open to everything.
A provider update added the proxy mapping and then rebuilt the middleware
chain. Between the two, the route was live with no chain behind it, and a
request that landed there was served straight through — a successful
inference that was neither routed by policy nor metered.
Rebuild first. The worst a request in the remaining window meets is the new
chain in front of the previous target, which is still counted.
The loop bounded when a new attempt could start, not how long one could
run. The chat container is capped at 90 seconds of its own and the row
lookup at another 20, so an attempt begun just inside the 180-second window
could report a repricing failure nearly two minutes after that window
closed. Run every call in the loop under a context that expires with the
deadline.
Three open review findings.
The discovery filter treated every slash in a listed id as a gateway prefix
and matched the tail against the policy. A self-hosted id carries slashes of
its own, and an upstream may scope ids per tenant, so "tenant-b/claude-sonnet-5"
matched a permitted "claude-sonnet-5" and reached the picker — a model the
policy never named, and one the guardrail denies on sight, since enforcement
compares the id as written. Strip only the namespaces a gateway is known to
prepend, taken from the first slash rather than the last.
The e2e retry loops slept between attempts without watching the context, so a
cancelled run kept retrying calls that fail instantly and spent its remaining
window sleeping between them. They now stop when the context is done.
The streamed provider's setup key outlived its test: deleting the group does
not delete the key that auto-joins it.
The failure message read repriced.InputCostUsd, and repriced is the zero
value on every path that reaches that line — so a run that gave up always
reported "last input_cost_usd=$0.000000", which reads as a row priced at
zero rather than as a row still at the old rate, or as no row at all.
Keep the last cost read and report that, saying so plainly when no row was
ever read.
The poll interval was not bounded by the window: a page that came back
without the row just before the deadline still slept a full two seconds
before the loop noticed it was out of budget, so the caller waited longer
than the window it asked for to be told nothing arrived.
Cap the wait at whatever is left of the window.
Two review findings on the e2e suite.
lookupAccessLogBySession polled under the caller's context, so one stalled
request could hold the loop open well past the 30s ingest window it exists to
enforce — and the caller would read that delay as a missing row rather than a
slow one. Each poll now expires with the window; the parent's cancellation
still applies, since the request context derives from it.
The streaming test asserted only that the total cost was positive. Input and
output are positive on their own, so a cache-read bucket that was parsed and
then never billed would have passed. Assert the sum of the three buckets: the
gap a dropped cache read leaves is 7e-6, well outside the delta.
Six behaviours had unit coverage only, either because they arrived from code
review after the end-to-end tests were written or because no request in the
suite had the shape that reaches them.
Streaming is the important one. Input tokens exist only in a stream's opening
message_start event, and reading a stream with the wrong vendor's parser
misses it — the metering bug this endpoint's protocol work fixed. Nothing in
the suite sent stream: true, so the branch never ran. The mock now serves an
SSE surface on a second listener, reporting counts that differ from its
buffered ones so a passing assertion can only mean the stream accumulator ran,
and one case drives it through a record typed for the wrong surface.
The rest need no new harness capability: the per-model lookup against the
allowlist, the read-method gate on the non-inference paths, dated ids reaching
an undated registration while a pinned build refuses a different one, the
Bedrock inference-profile lookup reaching its upstream rather than a policy
denial, and a custom dated id keeping its own price.
Sub-agent ids stay uncovered: the parser lifts them onto request metadata but
nothing persists them, so there is no queryable surface to assert against
until that half lands. Covered here only to the extent that sending the
headers leaves the request served and metered.
TestPriceChangeUpdatesRecordedCost drives requests in a loop until one is
priced at the new rate, because the price push and the proxy's chain rebuild
are async. The loop could not actually retry: it looked the row up through
findAccessLogBySession, which fails the test outright when no row lands
within 30s, so the first post-update request that produced no row ended the
run instead of yielding to the next attempt.
That is the observed failure — the nightly run has been red on this test
roughly half the time, always with "session id ...-reprice-b-... must be
recorded in an access-log row" after ~41s: container setup, one request, one
30s wait, dead.
A missing row there is expected rather than exceptional. The provider update
rebuilds the middleware chain, and a request served mid-rebuild can complete
without a resolved provider: 200 to the caller, nothing to attribute, so no
row is ever written for it. Split the polling helper into a non-fatal lookup
and keep the fail-fast wrapper for callers whose row must exist, then treat a
miss in the loop like any other not-yet-repriced iteration. Only the outer
deadline is fatal.
Shorten the per-attempt wait to 20s and raise the overall deadline to 180s so
several attempts fit where before the budget allowed barely one.
The routing and parser-selection fixes touch every provider surface, and
the unit tests only prove each side of a seam in isolation. Two suites
close that:
The provider matrix drives one request per wire shape over a single tunnel,
with a record per catalog surface behind it, and asserts the surface each
request was metered under together with the token counts that surface's own
usage block carries. A response read by the wrong provider's parser meters
zero, so a regression fails instead of passing on a coincidental non-zero.
It also covers the Bedrock and Vertex token-counting paths, the warm-up
probe, and the vendor error envelope on a refusal.
The discovery suite covers the configuration that broke: an account with a
model allowlist, where the listing carries no model and the gate failed
closed. It asserts the listing is served, that it is bounded to the
authorised model, and that inference outside the allowlist is still
refused, so the exemption cannot be read as a way around the gate.
The mock upstream grows the Anthropic, Bedrock and token-counting shapes so
one container stands in for every surface, and the client gains GET and
arbitrary-POST helpers for the endpoints that carry no chat body.
e2e/harness documents itself as feature-agnostic, but three details
assumed the caller lives in this repo, so the terraform provider's
acceptance suite would otherwise carry a second harness for the same
product.
repoRoot took the first module root above the working directory as the
Docker build context, which from another module is the caller's own
root, with no combined/Dockerfile.multistage in it. It now requires that
ancestor to be this module, and otherwise asks the go tool for the
source: for a dependent, the extracted directory of the version it pins,
so the server matches the client library it was compiled against. That
lookup uses -mod=readonly, since automatic vendor mode otherwise reports
an empty Dir.
Geolocation was disabled unconditionally. Agent-network ingest does not
use it, but location-based posture checks need the database, and a rule
management cannot evaluate fails rather than passing.
StartClient pinned one network alias and set no hostname, so a second
agent could not start and a peer's name was arbitrary. Management
records that hostname, making it the peer's name in the API.
The client entrypoint is copied with an explicit mode: git tracks it
100755, but the module cache extracts 0444, so a dependent's build
produced a container exiting with "permission denied".
Adds CombinedOption, WithGeolocation, WithServerEnv, ClientOption and
WithClientName.
Store the per-account gateway endpoint as {domain, proxy_address} with a
global unique index on the full hostname; dedicated = (domain ==
proxy_address). Bootstrap becomes an explicit POST carrying exactly one
of proxy_address (server allocates an adjective-noun label beneath it)
or endpoint (claimed verbatim, address-first); provider create loses its
bootstrap side effect. PUT is a full replace with every field required —
the immutable identity fields must be echoed unchanged and a mismatch is
rejected with 422. A guarded DELETE releases the endpoint: refused with
412 while providers exist or a proxy is actively serving the endpoint
hostname (matched case-insensitively); re-creating bootstraps fresh. A
self-addressed pin excludes its address from the account's cluster allow
list, and the live mapping update path now addresses the serving proxy
from the synthesized service. Existing rows are migrated on all three
store engines.
## Describe your changes
Work on the Terraform provider (terraform-provider-netbird #177–#183)
surfaced places where the agent-network API broke its own contracts or
deviated from the conventions the rest of the management API follows,
forcing client-side workarounds.
Settings reads now follow the settings-endpoint convention: GET always
answers with a JSON object. Before bootstrap it returns the defaults
with an empty cluster/subdomain/endpoint (previously 200 with a JSON
`null` body, while the spec said 404). The settings PUT can bootstrap
the account by carrying a `cluster` — previously the row could only come
into existence through the first provider create, and a settings-first
setup was impossible; a differing cluster on a bootstrapped account is
rejected instead of silently ignored. PUT remains full-state.
The provider PUT schema promised omit-preserves semantics for several
operator-editable fields that the handler never delivered (it builds the
row from the request, like every other update handler). The schema
wording now matches the shipped full-state behavior; only the api_key
(secret) and session keys stay preserved by the manager. Identity
headers are always present in provider responses so an explicitly
cleared value round-trips as an empty string.
The Go REST client gains the full agent-network surface (catalog,
providers, policies, guardrails, budget rules, settings), including a
shim translating the legacy 200+`null` settings body from older servers
into an `IsNotFound` error.
Note for reviewers: the dashboard special-cased the `null` settings
body; it needs a small follow-up for the new defaults response (in
progress).
## Describe your changes
The model-allowlist guardrail was merged into one account-wide union and
enforced flat on every request, ignoring which policy/group/provider
authorised it. With multiple policies — especially a mix of guardrailed
and un-guardrailed ones — this caused:
- **false-allow**: a model allowlisted for one group/provider leaked to
any caller; and
- **false-deny**: an un-guardrailed policy (intended unrestricted) was
blocked by another policy's allowlist.
Enforcement is now policy/group-aware, mirroring `llm_limit_check`:
- **Management (`SelectPolicyForRequest`) is authoritative.** It uses
the request model (already carried in
`CheckLLMPolicyLimitsRequest.model`, previously ignored) to keep only
applicable policies whose guardrails permit the model; no
allowlist-enabled guardrail = unrestricted. Denies
`llm_policy.model_blocked` when policies govern the (provider, groups)
but none permits the model.
- **Proxy `llm_guardrail` becomes a per-provider fail-closed backstop.**
The synthesiser emits an allowlist only for providers every authorising
policy restricts; the middleware keys off the resolved provider id and
keeps unknown-model fail-closed.
## Describe your changes
Use http get to trigger lazy connections as e2e will have single proxy
and status checks would fail without proper lazy wake-up
## Describe your changes
Supersedes #6767 (same change, moved to an unprefixed branch; review
feedback from there is addressed here).
With lazy connections enabled, a peer is not dialed until on-demand
traffic arrives, so the first request to a peer resolved by name (e.g.
the agent-network reverse-proxy) races — or loses to — the WireGuard
handshake. Activation was previously reactive only: a data-path packet
or an inbound signal.
This adds a proactive, DNS-time trigger. When the local resolver answers
for an overlay name, it now warms the lazy connection to the peer(s) the
answer points at and waits briefly for one to connect before returning
the response — so by the time the client sends its first packet the
tunnel is already up.
- `dns/local`: new `PeerActivator` capability + `SetPeerActivator`
setter (mirrors the existing `PeerConnectivity`/`SetPeerConnectivity`
injection). `ServeDNS` warms on the pre-filter answer, so activating a
lazily-idle peer also lets it survive the disconnected-peer filter.
Warm-up is scoped to match-only (non-authoritative) zones — the
synthesized private-service zones and user-created zones — so plain
peer-name lookups in the account's authoritative peer zone never wake
idle peers. No-op when no activator is wired (lazy off) or the answer
carries no peer IPs. Budget is `NB_DNS_LAZY_WARMUP_TIMEOUT` (default 2s,
parsed once at construction, invalid values logged); on timeout the
answer is returned anyway (never SERVFAIL).
- The resolver-facing interfaces (`PeerActivator`, `PeerConnectivity`)
take `netip.Addr` instead of string IPs; record addresses are extracted
as `netip.Addr` (v4-mapped forms unmapped) and converted to string only
at the `peer.Status` boundary.
- `SetPeerActivator` is part of the `dns.Server` interface (no-op on the
mock), so the engine wires it without a type assertion.
- `client/internal`: a small engine-side adapter (`dnsPeerActivator`)
resolves answer IPs to peers via `Status.PeerStateByIP`, activates them
through `ConnMgr.ActivatePeer` (HA fan-out included), and polls
`PeerStateByIP` until one is connected. The activation dial is tied to
the engine's long-lived context so a handshake that outlasts the
per-query wait still completes in the background. `ConnMgr.ActivatePeer`
is safe for concurrent use (the lazy manager pointer is guarded by a
dedicated RWMutex and the manager is internally synchronized), so the
DNS path never contends with network-map processing on `syncMsgMux`.
Scope is overlay-only for free: the trigger lives in the local resolver,
which only answers for NetBird-managed names; public/upstream DNS is
unaffected. Already-connected peers short-circuit, so steady-state DNS
latency is unchanged.
## Issue ticket number and link
N/A — follow-up to the lazy-connection rollout; fixes the agent-network
cold-start observed in the e2e (proxy peer stuck disconnected until
traffic).
### Checklist
- [x] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [x] Created tests that fail without the change (if possible)
- [x] This change does **not** modify the public API, gRPC protocols,
functionality behavior, CLI / service flags, or introduce a new feature
— **OR** I have discussed it with the NetBird team beforehand (link the
issue / Slack thread in the description). See
[CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#discuss-changes-with-the-netbird-team-first).
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
Internal client behavior; the only knob is the optional
`NB_DNS_LAZY_WARMUP_TIMEOUT` tuning env var with a safe default.
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
https://github.com/netbirdio/docs/pull/__
## Tests
- `dns/local`: warm-up invokes the activator with the answer's peer
address in match-only zones; authoritative-zone answers never trigger
warm-up; no-activator path unchanged; no-answer queries don't invoke the
activator; `NB_DNS_LAZY_WARMUP_TIMEOUT` parsing
(valid/invalid/non-positive); `extractRecordAddr` unmaps v4-mapped
record data.
- `client/internal`: `dnsPeerActivator` skips
connected/unknown/conn-less peers with no wait, returns as soon as a
pending peer connects, and releases the DNS response at the budget when
the peer stays idle; `ConnMgr.ActivatePeer` races the manager lifecycle
cleanly under `-race`.
- Agent-network e2e green on this change (with lazy connections
enabled): https://github.com/netbirdio/netbird/actions/runs/29891467544
## Describe your changes
* [proxy] enforce model allowlist for URL-routed providers
(Bedrock/Vertex) by @mlsmaycon in
https://github.com/netbirdio/netbird/pull/6764
* [management] Remove proxy peer stale deduplication logic by @mlsmaycon
in https://github.com/netbirdio/netbird/pull/6768
## Issue ticket number and link
## Stack
<!-- branch-stack -->
### Checklist
- [ ] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [ ] Created tests that fail without the change (if possible)
- [ ] This change does **not** modify the public API, gRPC protocols,
functionality behavior, CLI / service flags, or introduce a new feature
— **OR** I have discussed it with the NetBird team beforehand (link the
issue / Slack thread in the description). See
[CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#discuss-changes-with-the-netbird-team-first).
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
https://github.com/netbirdio/docs/pull/__
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added model-allowlist guardrails for path-routed providers, including
Bedrock and Vertex.
- Added Bedrock request support for chat interactions.
- Added guardrail management capabilities.
- **Bug Fixes**
- Requests with missing or blank model identifiers are now denied when a
model allowlist is configured, improving fail-closed protection.
- Corrected provider-specific request handling and session tracking for
Bedrock interactions.
- **Tests**
- Expanded coverage for allowlist enforcement and provider routing
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Theodor Midtlien <theodor@midtlien.com>
Co-authored-by: blaugrau90 <61945343+blaugrau90@users.noreply.github.com>
Co-authored-by: Viktor Liu <17948409+lixmal@users.noreply.github.com>
* [management,proxy] Add per-provider skip_tls_verification for agent-network
Let agent-network providers opt into skipping upstream TLS verification for
self-hosted / internal gateways behind a private or self-signed cert.
- provider: add SkipTLSVerification (persisted via AutoMigrate) with
request/response mapping (nil on update preserves, explicit false clears).
- openapi: skip_tls_verification on the provider request + response; types
regenerated.
- synthesizer: carry the flag into the llm_router route config so it reaches
the proxy.
- proxy: llm_router sets it on the UpstreamRewrite mutation, and the reverse
proxy applies roundtrip.WithSkipTLSVerify per selected route when forwarding
upstream (the router dials per provider, so a per-target flag alone wouldn't
cover it).
- tests: synthesizer route config carries the flag, router rewrite propagates
it, and the request/response round-trip incl. update semantics.
* [e2e] Validate per-provider skip_tls_verification end to end
Add a self-signed HTTPS upstream (nginx) to the harness and a test that
provisions two providers on that same upstream — one with
skip_tls_verification=true, one false — behind one proxy + client. The
skip=true provider's chat reaches the upstream (200); the skip=false
provider's fails the TLS handshake (5xx). Same upstream, opposite outcome,
which proves the flag is honoured per provider (a single target-level flag
could not, since all of an account's providers share one synthesised
target).
* [e2e] WaitProxyPeer: require >=1 connected peer, not exact 1/1
Each proxy container registers a fresh WireGuard key and its peer is not
removed on teardown, so proxy peers from earlier tests linger in the
account as disconnected. WaitProxyPeer matched the exact string
"1/1 Connected", which failed once a second proxy-using test ran in the
same package (status "1/2"). Parse the "Peers count: X/Y Connected" line
and wait for X>=1 instead: only the live proxy can be connected, and the
caller's subsequent chat is the real end-to-end assertion. Fixes the CI
failure of TestProviderSkipTLSVerification (runs after TestProvidersMatrix).
* [management,proxy] Agent network: per-account LLM gateway (policy, metering, multi-provider) (#6555)
* [agent-network] Shared proto, OpenAPI schema, and generated types
* [agent-network] Management: store, manager, synthesizer, policy engine, provider catalog, HTTP/gRPC API
Adds the account-scoped agent-network module: provider/policy/budget CRUD and
store, the reverse-proxy service synthesizer, policy selection + limit
enforcement, the provider catalog (incl. Vertex AI and AWS Bedrock entries),
and the management HTTP + proxy gRPC surfaces.
* [management] Fix agent-network proxy-peer fan-out on affected-peer recompute
The affected-peers resolver loaded only persisted reverse-proxy services, but
agent-network services are synthesized on demand and never persisted. As a
result the embedded proxy peer was never folded into the affected set when a
client's group changed, so the proxy received no network-map update for a newly
authorised client and rejected its handshake until a full resync (restart).
loadProxyServices now merges the synthesized agent-network services (injected
via a registration hook to avoid an import cycle), so proxy peers learn newly
authorised clients immediately.
* [proxy] Reverse-proxy middleware framework, chain, and request plumbing
The per-target middleware chain (slots, dispatcher, mutation gate, metadata
merger), body capture, access-log terminal sink, and the proxy wiring that
builds + runs chains for synthesized agent-network services.
* [proxy] LLM parsers, pricing, and builtin middlewares (OpenAI, Anthropic, Vertex AI, AWS Bedrock)
Request/response parsers and SSE/event-stream metering, the embedded pricing
table, and the builtin middleware set: request parser, router, policy
limit-check/record, cost meter, guardrail, identity inject, response parser.
Includes the path-routed providers — Google Vertex AI (keyfile:: service-account
OAuth minting) and AWS Bedrock (bearer auth, invoke/converse/streaming, optional
/bedrock prefix) — plus the Models allowlist and unmeterable-publisher deny.
* [proxy] IPv6 in-place apply and TCP accept-loop hardening on netstack listeners
* [agent-network] End-to-end test suite, module docs, and deployment preset
* [agent-network] Fix codespell typos and exclude false positives
- labelgen word pool: vermillion -> vermilion, racoon -> raccoon.
- codespell ignore list: add flate (Go compress/flate package), recordin
(a test-local identifier), and unparseable (a valid alternative spelling used
consistently across identifiers + a metadata-value constant).
* [management] Set LastSeen on injected proxy peer in realstack test (MySQL strict-mode)
The injected embedded proxy peer had a PeerStatus with a zero LastSeen, which
serializes to '0000-00-00' and is rejected by MySQL in strict mode (SQLite
tolerates it). Set LastSeen to a valid time so SaveAccount succeeds on both
engines.
* [agent-network] Remove e2e shell-script suite from this branch
The end-to-end shell scripts under scripts/e2e/ are maintained in a separate
testing suite and are not part of this change set.
* [agent-network] Polish module docs: remove internal review scaffolding, fix links, verify diagrams
Strip PR-review framing, commit references, absolute paths, and stale internal
references from the agent-network module docs; fix broken relative links; verify
all diagrams against the current architecture. Remove the internal AI-reviewer
prompt file.
* [management] Refine session expiration handling to support 3-state encoding for SSO deadlines
* [agent-network] Relocate agentnetwork package to internals/modules
Move management/server/agentnetwork (and its catalog/, labelgen/, types/
subpackages) to management/internals/modules/agentnetwork, alongside the
reverse-proxy module, and rewrite all importers. Pure relocation: package names,
the synthesizer + affectedpeers registration hook, and store access (shared
store.Store) are unchanged, so no import cycle is introduced (affectedpeers
still depends only on the agentnetwork/types leaf).
* [agent-network] Co-locate HTTP handlers in the module (RegisterEndpoints)
Move the agent-network HTTP handlers from server/http/handlers/agentnetwork into
the module at internals/modules/agentnetwork/handlers (package handlers) and
rename the entrypoint AddEndpoints -> RegisterEndpoints, matching the
reverse-proxy module convention. Wiring in http/handler.go updated accordingly.
* Update getting started to point to rc when agent network enabled
* Add a reference to a commercial license
* Fix docs localhost link
* Fix docs localhost link
* Add private services domain note
* [management] Add agent-network telemetry metrics (#6561)
Surface agent-network adoption and usage in the self-hosted metrics
worker: distinct accounts, providers, policies, budget rules, accounts
with log collection enabled, and aggregated input/output tokens plus
cost.
Tokens and cost are summed from agent_network_request_usage (the
always-written per-request ledger) so the figures are accurate
regardless of the log-collection toggle and carry no double-counting.
All values come from a handful of indexed aggregate queries run only on
the worker's periodic tick.
Adds store.AgentNetworkMetrics with GetAgentNetworkMetrics on the Store
interface, the SqlStore implementation, and a zero-valued FileStore stub.
* Update NetBird server and proxy image versions to 0.74.0-rc.2
* [management,proxy] Reduce agent-network cognitive complexity (#6566)
Address the SonarCloud quality-gate findings in new agent-network code
by extracting focused helpers. No behavior change.
- synthesizer.go: split buildIdentityInjectConfigJSON into per-shape
rule builders; extract mergeGuardrail from mergeGuardrails to cut
nesting depth.
- llm_identity_inject: extract injectionEmitsAnything validation
predicate from New.
- llm_response_parser/streaming.go: extract applyOpenAIStreamUsage and
applyAnthropicStreamUsage (via a named anthropicStreamUsage type) and
simplify the OpenAI scanner loop.
- reverseproxy.go: decompose ServeHTTP into serveRouteError,
buildTargetContext, serveDirect, serveWithChain, captureRequestForChain,
serveDeny, newResponseWriter, observeResponse, and forwardUpstream,
preserving the defer ordering so response observation still reads the
captured writer before it is released.
* [management] Move agent-network access-log ingest into the agentnetwork module (#6568)
The agent-network access-log ingest path (metaKey wire contract, flatten,
usage derivation, and the dual-write of the usage ledger + settings-gated
full row) lived in the reverseproxy accesslogs manager, even though the
agentnetwork module already owns the rest of that domain — types, read
(ListAccessLogs / GetUsageOverview), the budget-counter writes, and
retention cleanup.
Move it next to the rest: a stateless agentnetwork.IngestAccessLog(ctx,
store, entry) that the reverseproxy SaveAccessLog delegates to when the
entry is agent-network. Removes the agentNetworkTypes import from the
reverseproxy manager. No behavior change; the write/read table separation
is unchanged.
Adds real-store coverage for the disable->enable log-collection toggle
(usage ledger always written, full row gated) plus the metadata parse and
group-dedup helpers, which previously had no dedicated tests.
* Add session view support in the access log
* [management,proxy] Container-based agent-network e2e harness (#6577)
* [e2e] Add container-based agent-network e2e harness (Pillar 1)
Introduce a self-contained, OIDC-free e2e harness that stands up NetBird
in containers, so suites no longer depend on the hand-maintained Tilt
stack or a real IdP.
- harness brings up the combined server (management + signal + relay +
STUN + embedded IdP) in a single container built from
combined/Dockerfile.multistage, and mints an admin PAT through the
unauthenticated /api/setup bootstrap (NB_SETUP_PAT_ENABLED). API access
goes through the existing shared/management/client/rest typed client.
- the image is built via the docker CLI (BuildKit) so the Dockerfile's
cache mounts are honored; testcontainers then runs the tagged image.
- everything is behind the `e2e` build tag so normal builds and unit
tests never pull in testcontainers.
Adds BuildKit cache mounts to combined/Dockerfile.multistage so source
changes recompile incrementally rather than from scratch.
Pillar 1 proven by TestCombinedBootstrap: server builds, boots, mints a
PAT, and the PAT authenticates a real management API call.
* [e2e] Add management-side agent-network scenarios (Pillar 2)
Port the API-driven agent-network scenarios from the bash suites to Go,
sharing one combined server per package run (TestMain) with each test
owning its resource cleanup. Drives the /api/agent-network/* endpoints
through the shared REST client's NewRequest primitive with the generated
api types.
Scenarios:
- provider lifecycle (create/get/list/delete + 404 after delete)
- provider validation (missing api_key, unknown catalog id → 4xx)
- settings collection-toggle round-trip with cluster/subdomain immutability
- policy window floor (reject <60s enabled limit, accept at 60s)
- consumption read endpoint returns an array
All deterministic and dependency-free (dummy provider keys; no upstream
calls), so they run headless in CI.
* [e2e] Add live chat-through-proxy scenario (Pillar 3)
Stand up the full agent-network data path in containers and drive a real
chat-completion through the gateway:
- harness: a shared docker network (combined server reachable by alias),
a proxy container built from the published reverse-proxy image
(NB_PROXY_PRIVATE, NB_PROXY_ALLOW_INSECURE, NB_RELAY_TRANSPORT=ws to match
the combined server's WS-multiplexed relay) with a generated self-signed
wildcard cert, and a netbird client container that joins via a setup key.
- the combined image, proxy image, and client image default to the
published rc.2 releases (overridable via NB_E2E_*_IMAGE; a bare local tag
is built from source instead). Geolocation download is disabled so the
server starts without external fetches.
- one shared domain is used for the management exposed address, the proxy
domain, and the agent-network cluster; the proxy token is minted via the
server CLI (global) to match the manual install.
TestChatCompletionThroughProxy provisions provider+policy+group+setup key,
runs proxy+client, drives an OpenAI chat-completion through the tunnel, and
asserts a 200 plus the ingested access-log row. Requires OPENAI_TOKEN
(skips otherwise). The provider must be created with enabled=true explicitly
— the create default is false despite the API doc.
* [e2e] Run the live chat scenario across a provider matrix
Replace the single-provider chat test with a data-driven matrix that runs
the same scenario through every provider whose credentials are present in
the environment (keys/URLs sourced from ~/.llm-keys locally, Actions
secrets in CI):
- OpenAI (chat), Anthropic (messages), Vercel, OpenRouter, Cloudflare
(OpenAI-compatible gateways), and Bedrock (path-routed, bearer, via the
messages shape) — covering both wire shapes and the gateway routing.
- all providers are created enabled with a unique model string so the
proxy's connect-time snapshot carries them all and model->provider
routing is unambiguous (provider toggles after connect don't reconcile
to a connected proxy).
- the client supports both wire shapes (/v1/chat/completions and
/v1/messages); Cloudflare gets the openai provider segment appended to
its gateway URL.
Each provider must return 200 through the tunnel and produce an ingested
access-log row. Vertex is intentionally excluded from the uniform matrix:
it needs a bespoke rawPredict request shape rather than the shared
chat/messages path, so it warrants a dedicated scenario.
* [ci] Add manual workflow for the agent-network e2e suite
The e2e suite (build tag `e2e`) stands up the combined server + proxy +
client in Docker and drives live chat-completions, so it is slow and needs
provider credentials. Gate it out of normal CI (it already is, via the
build tag) and run it on demand via workflow_dispatch. Provider scenarios
skip when their secret is unset, so it degrades gracefully.
* [e2e] Add Vertex to the provider matrix; run e2e on ubuntu-latest
Vertex (Anthropic-on-Vertex) doesn't share the chat/messages wire shapes:
the model travels in a rawPredict path and the proxy mints the service
account's OAuth token. Add a Vertex client method that posts
/v1/projects/<project>/locations/<region>/publishers/anthropic/models/<model>:rawPredict
with the Vertex anthropic_version body, and wire it into the matrix as a
path-routed provider (created without a models array). It is keyed off
GOOGLE_VERTEX_SA_BASE64 + GOOGLE_VERTEX_PROJECT (region defaults to
"global", model to a pinned claude snapshot, both overridable).
Also bump the e2e workflow runner to ubuntu-latest and add the Vertex
secrets.
* Add docker/docker and docker/go-connections as direct dependencies in go.mod
* [ci] Trigger agent-network e2e workflow on push to main and pull requests
* [e2e] Fix proxy cert permission denied on Linux CI runners
The proxy bind-mounts a temp dir of self-signed certs. MkdirTemp creates
it 0700 and the key was 0600, which Docker Desktop on macOS ignores but a
non-root proxy container on Linux runners cannot traverse/read, so the
cert watcher failed with "open /certs/tls.crt: permission denied" and the
container exited. Widen the cert dir to 0755 and write the throwaway key
0644 so the proxy uid can read the bind-mounted material.
* [e2e] Build images from source by default instead of pulling rc.2
The agent-network code under test lives in this branch, so the e2e should
exercise it rather than a frozen published release. Flip the harness
default: combined/proxy/client are now built from their in-repo
Dockerfiles (combined/Dockerfile.multistage, proxy/Dockerfile.multistage,
e2e/harness/Dockerfile.client) under local tags. Pulling a published image
stays available by setting NB_E2E_*_IMAGE to a registry reference.
Builds now go through buildx --load so the Dockerfile cache mounts are
honored and the result is loaded for testcontainers. The CI workflow adds
a container-driver builder and a local layer cache (NB_E2E_BUILDX_CACHE)
persisted via actions/cache, which caches the base/apt/dep-download layers
across runs. The Go compile still re-runs each time, as BuildKit mount
caches cannot be exported to the GitHub cache.
* [e2e] Cover real providers in lifecycle + assert real consumption metering
- TestProviderLifecycle now runs per available real provider (create → get →
list → delete → 404) instead of a single dummy provider, exercising each
catalog's create and field round-trip. Create is offline, so it stays fast
and burns no provider quota; falls back to a synthetic OpenAI provider when
no keys are set.
- TestProvidersMatrix attaches a token limit (high caps, 60s window) to its
policy, which switches on usage metering, and asserts consumption rows are
recorded with positive token counts after the live traffic. Consumption is
account-scoped (keyed by source group / user and window, not per provider),
so the assertion is aggregate.
- TestProviderValidation gains invalid-upstream and blank-name cases. Create
validation is uniform across catalogs (no per-provider required-field rules),
so per-provider rejection cases would be redundant.
* [e2e] Assert session id propagates per provider
Each matrix request now sends a unique session id as the universal
x-session-id header and asserts it round-trips into that provider's
access-log row. This guards the session-grouping contract end to end for
every provider (header extraction runs in llm_request_parser ahead of the
parser-specific body extraction, so it is provider-agnostic).
* [e2e] Drop accidentally committed sync-phases dashboard
netbird-sync-phases.json was swept into the Pillar 1 commit by a broad
git add; it belongs to the unrelated sync-phases metrics work, not this
e2e harness. Remove it from the branch so the PR diff is scoped to the
e2e changes.
* [e2e] Revert accidentally committed sync-phase ingest spec
The netbird_sync_phase measurement spec in metrics ingest was swept into
the Pillar 1 commit; it belongs to the unrelated sync-phases metrics work,
not this e2e harness. Its emission side never landed here, so the spec was
orphaned anyway. Restore ingest/main.go to its origin/main state.
* Fix golint issues
* Fix sonar
* Add access log session test
* Fix access log tests
---------
Co-authored-by: braginini <bangvalo@gmail.com>
Co-authored-by: Zoltan Papp <zoltan.pmail@gmail.com>