Three open review findings.
The discovery filter treated every slash in a listed id as a gateway prefix
and matched the tail against the policy. A self-hosted id carries slashes of
its own, and an upstream may scope ids per tenant, so "tenant-b/claude-sonnet-5"
matched a permitted "claude-sonnet-5" and reached the picker — a model the
policy never named, and one the guardrail denies on sight, since enforcement
compares the id as written. Strip only the namespaces a gateway is known to
prepend, taken from the first slash rather than the last.
The e2e retry loops slept between attempts without watching the context, so a
cancelled run kept retrying calls that fail instantly and spent its remaining
window sleeping between them. They now stop when the context is done.
The streamed provider's setup key outlived its test: deleting the group does
not delete the key that auto-joins it.
Two review findings on the e2e suite.
lookupAccessLogBySession polled under the caller's context, so one stalled
request could hold the loop open well past the 30s ingest window it exists to
enforce — and the caller would read that delay as a missing row rather than a
slow one. Each poll now expires with the window; the parent's cancellation
still applies, since the request context derives from it.
The streaming test asserted only that the total cost was positive. Input and
output are positive on their own, so a cache-read bucket that was parsed and
then never billed would have passed. Assert the sum of the three buckets: the
gap a dropped cache read leaves is 7e-6, well outside the delta.
Six behaviours had unit coverage only, either because they arrived from code
review after the end-to-end tests were written or because no request in the
suite had the shape that reaches them.
Streaming is the important one. Input tokens exist only in a stream's opening
message_start event, and reading a stream with the wrong vendor's parser
misses it — the metering bug this endpoint's protocol work fixed. Nothing in
the suite sent stream: true, so the branch never ran. The mock now serves an
SSE surface on a second listener, reporting counts that differ from its
buffered ones so a passing assertion can only mean the stream accumulator ran,
and one case drives it through a record typed for the wrong surface.
The rest need no new harness capability: the per-model lookup against the
allowlist, the read-method gate on the non-inference paths, dated ids reaching
an undated registration while a pinned build refuses a different one, the
Bedrock inference-profile lookup reaching its upstream rather than a policy
denial, and a custom dated id keeping its own price.
Sub-agent ids stay uncovered: the parser lifts them onto request metadata but
nothing persists them, so there is no queryable surface to assert against
until that half lands. Covered here only to the extent that sending the
headers leaves the request served and metered.