Files
glpi-neural-brain/VECTOR-GRAPH-AGENT-OFFLOAD-V3.md
jbergner 440423c5b6
release-tag / release-image (push) Successful in 2m43s
RC-3
2026-08-09 11:29:13 +02:00

7.6 KiB
Raw Permalink Blame History

Vector Graph v3: Orphan Pass, Vector-guided THINKING and Agent CPU Offload

This revision turns the mathematical semantic_neighbor layer into the default candidate infrastructure for expensive graph reasoning while keeping every strong semantic relation under Brain control.

Goals

  1. Connect a conservative first pass with an optional second pass for remaining Knowledge orphans.
  2. Let AI-THINK evaluate promising mathematical neighbours instead of searching the whole vector space again.
  3. Prevent generic Security/Hardening vocabulary from steering article web research toward the wrong entity.
  4. Move the CPU-heavy vector calculation to integrated Source Agents when requested, without Ollama/chat/embedding inference on the Agent.
  5. Stop analysis/audit events from timing out behind the Store's single primary SQLite connection.

Mathematical passes

Primary pass

mutual-knn-local-scaling-v1 remains unchanged in interpretation:

  • deterministic sparse random-projection LSH;
  • bounded exact Cosine shortlist;
  • local scaling;
  • reciprocal k-NN preferred;
  • output relation: semantic_neighbor, origin vector-math.

Optional orphan second pass

After the primary graph is built, the Brain determines which production Knowledge nodes would still be direct Knowledge/evidence orphans when old vector-math edges are ignored. Nodes already touched by the new primary result are removed from that focus set.

orphan-knn-local-scaling-v1 then searches only those focus nodes against the full vector corpus. It is intentionally one-sided and conservative; it does not try to force every node into the graph.

BRAIN_VECTOR_GRAPH_ORPHAN_PASS=true
BRAIN_VECTOR_GRAPH_ORPHAN_NEIGHBORS=2
BRAIN_VECTOR_GRAPH_ORPHAN_CANDIDATES=256
BRAIN_VECTOR_GRAPH_ORPHAN_MIN_SIMILARITY=0.80
BRAIN_VECTOR_GRAPH_ORPHAN_MIN_AFFINITY=0.30

The second pass is disabled by default until its false-positive rate has been reviewed on the target corpus.

Vector-guided THINKING

With:

BRAIN_THINKING_VECTOR_GUIDED=true

AI-THINK first scans existing vector-math/semantic_neighbor edges and chooses the strongest pair that has not already received an ai-inference decision. Reciprocal mathematical links are preferred.

Only if no unreviewed vector candidate exists does THINKING fall back to the previous embedding candidate search.

This separates responsibilities:

  • semantic_neighbor: cheap mathematical candidate relation;
  • same_topic, related_to, depends_on, etc.: expensive interpreted relation.

The candidate source is written to analysis metadata as candidate_source=vector_graph or embedding_search.

Agent CPU offload

The integrated BRAIN_MODE=agent runtime can now advertise the capability:

vector_graph

The Agent does not need a source polling task to calculate these jobs.

Brain configuration

BRAIN_VECTOR_GRAPH_ENABLED=true
BRAIN_VECTOR_GRAPH_AGENT_OFFLOAD=true
BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=false
BRAIN_VECTOR_GRAPH_AGENT_WAIT=2m

AGENT_REQUIRED=false is recommended initially. If no compatible Agent is online, if the job times out, or if the graph changes while the job is in flight, the Brain falls back to the same local CPU implementation.

Set BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=true only when local CPU fallback is intentionally forbidden.

Agent configuration

BRAIN_MODE=agent
BRAIN_AGENT_BRAIN_URL=http://brain:8090
BRAIN_AGENT_ID=cpu-agent-01
BRAIN_AGENT_TOKEN=brain_agent_...
BRAIN_AGENT_COMPUTE_ENABLED=true
BRAIN_AGENT_COMPUTE_POLL_INTERVAL=5s
BRAIN_AGENT_COMPUTE_MAX_BYTES=134217728

The Source Agent still initializes no graph, Ollama, SearXNG, AI-THINK or article pipeline.

Wire protocol

The Brain owns the vector index and submits an in-memory pull job. An Agent claims it through the authenticated Agent API.

The request is streamed as a compact binary payload:

  • fixed protocol magic;
  • JSON job/config header;
  • node ID;
  • vector dimension;
  • raw little-endian float32 vector values.

This avoids expanding a roughly 62 MiB 21k×768 float32 matrix into much larger decimal JSON.

The Agent returns only the mathematical result (links, optional positions, statistics). The Brain validates:

  • compute kind;
  • graph version;
  • every source/target ID against the submitted set;
  • no self-links;
  • finite 0..1 similarity/affinity/confidence values.

The Brain then applies the result through the same graph code as a local build. The Agent cannot choose an arbitrary edge type: the Brain persists only semantic_neighbor with origin vector-math.

If layout is enabled, 3D vector positions are also calculated on the Agent; no model inference is involved.

No-model guarantee

A vector_graph compute job calls neither Chat nor Embed. It consumes embeddings that already exist in the Brain. Missing/new embeddings are still the Brain learning pipeline's responsibility.

Therefore:

  • edge calculation: CPU-only arithmetic;
  • orphan second pass: CPU-only arithmetic;
  • layout: CPU-only arithmetic;
  • Agent compute: CPU-only arithmetic;
  • creation of a missing embedding: still embedding-model inference, outside the compute job.

Research topic guard

Web research now has a deterministic primary-entity/topic gate before source quality can rescue a result.

Generic terms such as security, hardening, support, testing, documentation, forensics and template verbs are removed from the primary topic anchors. German compounds are supported (browser matches Webbrowser).

Examples covered by regression tests:

  • Browser Security → BSI Webbrowser: allowed;
  • Browser Security → Proxmox Server Hardening: rejected;
  • Rate Limit Testing → Rate Limit source: allowed;
  • Rate Limit Testing → generic Web Security Testing: rejected.

A failed topic guard caps relevance and clears gap coverage even if an LLM assessment or high-quality domain would otherwise rank the source highly.

Audit persistence

The append-only analysis writer now uses a dedicated SQLite connection to the same WAL database. The primary graph connection deliberately remains MaxOpenConns(1), but audit telemetry no longer waits for that same Go connection slot.

The audit connection uses a 60-second SQLite busy timeout and longer writer deadlines. This addresses the observed context deadline exceeded persistence drops without weakening graph transaction ownership.

Real-corpus validation

Against the bundled production graph snapshot (21,289 production Knowledge vectors, 768 dimensions):

  • primary pass: 16,940 semantic_neighbor links, 16,007 reciprocal;
  • orphan focus after the primary pass: 5,218 nodes;
  • orphan second pass: 3,038 additional links touching 2,200 focused orphan nodes;
  • full Agent wire payload: 66,568,747 bytes (63.48 MiB);
  • binary encode/decode in the supplied environment: ~0.38 s / ~0.43 s;
  • primary + orphan compute after wire decode: ~8.1 s CPU wall time;
  • no Chat or Embed call occurred inside the compute job.

The orphan pass kept min_similarity=0.80; the additional coverage therefore comes from a larger focused candidate search rather than globally weakening the similarity floor.

Brain:

BRAIN_VECTOR_GRAPH_ENABLED=true
BRAIN_VECTOR_GRAPH_ORPHAN_PASS=true
BRAIN_VECTOR_GRAPH_AGENT_OFFLOAD=true
BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=false
BRAIN_THINKING_VECTOR_GUIDED=true
BRAIN_VECTOR_GRAPH_LAYOUT=false

Agent:

BRAIN_AGENT_COMPUTE_ENABLED=true
BRAIN_AGENT_COMPUTE_POLL_INTERVAL=5s
BRAIN_AGENT_COMPUTE_MAX_BYTES=134217728

Keep layout disabled for the first run. Verify vector.graph.agent.completed, orphan count, second-pass edge examples and Agent status/capabilities before enabling layout or making Agent offload mandatory.