mirror of
https://github.com/netbirdio/netbird.git
synced 2026-08-25 09:01:29 +02:00
[proxy,management] Conform the Agent Network endpoint to the LLM gateway protocol Reviewed the proxy against Claude Code's published gateway contract. The transport layer already held up; fourteen gaps sat one layer up, in the model catalog and in the non-inference endpoints clients call. Two of them cost money. The catalog carried no claude-opus-5 or claude-sonnet-5, so an operator could not authorise the models coding agents default to — those requests denied as not-routable, or priced at zero where a catch-all carried them. And gateway records pin ParserID "openai" while the same record serves /v1/messages, so Anthropic responses were read with the OpenAI parser, which never looks at message_start where input tokens live: input metered as roughly zero on every stream and cost was skipped entirely. The rest fix requests refused for structural rather than policy reasons: model discovery denied for every account with a model allowlist, token counting denied on Bedrock and mis-parsed on Vertex, startup probes refused and written into the access log at every session start, and denials rendered in a shape no LLM client parses. Two changes are additive by design — the deny body keeps every field it had and adds the vendor's error object alongside, and body-level identity injection is now gated on the request's dialect so it stops sending OpenAI-shape fields into Anthropic bodies that reject them. The end-to-end work turned up one more: the discovery filter treated any slash in a model id as a gateway prefix, which would have dropped every self-hosted "Qwen/..." model from the picker.
30 lines
1.2 KiB
Go
30 lines
1.2 KiB
Go
package llm
|
|
|
|
import (
|
|
sharedllm "github.com/netbirdio/netbird/shared/llm"
|
|
)
|
|
|
|
// NormalizeBedrockModel strips an ARN wrapper, a cross-region inference-profile
|
|
// prefix, and the version/throughput suffix from a Bedrock model id so it
|
|
// matches the catalog/pricing key. Thin delegate to the shared implementation
|
|
// (shared/llm), which management also uses at synthesis time so both sides of
|
|
// the pricing / routing contract normalize identically.
|
|
func NormalizeBedrockModel(modelID string) string {
|
|
return sharedllm.NormalizeBedrockModel(modelID)
|
|
}
|
|
|
|
// NormalizeAnthropicModel strips the trailing "-YYYYMMDD" release-date suffix
|
|
// from an Anthropic model id so a dated id a client pins matches the undated
|
|
// one the operator registered. Thin delegate to shared/llm for the same
|
|
// contract reason as the two below.
|
|
func NormalizeAnthropicModel(modelID string) string {
|
|
return sharedllm.NormalizeAnthropicModel(modelID)
|
|
}
|
|
|
|
// NormalizeVertexModel strips the "@version" suffix from a Vertex AI model id
|
|
// so it matches the catalog/pricing key. Thin delegate to shared/llm, kept
|
|
// beside NormalizeBedrockModel for the same contract reason.
|
|
func NormalizeVertexModel(modelID string) string {
|
|
return sharedllm.NormalizeVertexModel(modelID)
|
|
}
|