7.6 KiB
Vector Graph v3: Orphan Pass, Vector-guided THINKING and Agent CPU Offload
This revision turns the mathematical semantic_neighbor layer into the default candidate infrastructure for expensive graph reasoning while keeping every strong semantic relation under Brain control.
Goals
- Connect a conservative first pass with an optional second pass for remaining Knowledge orphans.
- Let AI-THINK evaluate promising mathematical neighbours instead of searching the whole vector space again.
- Prevent generic Security/Hardening vocabulary from steering article web research toward the wrong entity.
- Move the CPU-heavy vector calculation to integrated Source Agents when requested, without Ollama/chat/embedding inference on the Agent.
- Stop analysis/audit events from timing out behind the Store's single primary SQLite connection.
Mathematical passes
Primary pass
mutual-knn-local-scaling-v1 remains unchanged in interpretation:
- deterministic sparse random-projection LSH;
- bounded exact Cosine shortlist;
- local scaling;
- reciprocal k-NN preferred;
- output relation:
semantic_neighbor, originvector-math.
Optional orphan second pass
After the primary graph is built, the Brain determines which production Knowledge nodes would still be direct Knowledge/evidence orphans when old vector-math edges are ignored. Nodes already touched by the new primary result are removed from that focus set.
orphan-knn-local-scaling-v1 then searches only those focus nodes against the full vector corpus. It is intentionally one-sided and conservative; it does not try to force every node into the graph.
BRAIN_VECTOR_GRAPH_ORPHAN_PASS=true
BRAIN_VECTOR_GRAPH_ORPHAN_NEIGHBORS=2
BRAIN_VECTOR_GRAPH_ORPHAN_CANDIDATES=256
BRAIN_VECTOR_GRAPH_ORPHAN_MIN_SIMILARITY=0.80
BRAIN_VECTOR_GRAPH_ORPHAN_MIN_AFFINITY=0.30
The second pass is disabled by default until its false-positive rate has been reviewed on the target corpus.
Vector-guided THINKING
With:
BRAIN_THINKING_VECTOR_GUIDED=true
AI-THINK first scans existing vector-math/semantic_neighbor edges and chooses the strongest pair that has not already received an ai-inference decision. Reciprocal mathematical links are preferred.
Only if no unreviewed vector candidate exists does THINKING fall back to the previous embedding candidate search.
This separates responsibilities:
semantic_neighbor: cheap mathematical candidate relation;same_topic,related_to,depends_on, etc.: expensive interpreted relation.
The candidate source is written to analysis metadata as candidate_source=vector_graph or embedding_search.
Agent CPU offload
The integrated BRAIN_MODE=agent runtime can now advertise the capability:
vector_graph
The Agent does not need a source polling task to calculate these jobs.
Brain configuration
BRAIN_VECTOR_GRAPH_ENABLED=true
BRAIN_VECTOR_GRAPH_AGENT_OFFLOAD=true
BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=false
BRAIN_VECTOR_GRAPH_AGENT_WAIT=2m
AGENT_REQUIRED=false is recommended initially. If no compatible Agent is online, if the job times out, or if the graph changes while the job is in flight, the Brain falls back to the same local CPU implementation.
Set BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=true only when local CPU fallback is intentionally forbidden.
Agent configuration
BRAIN_MODE=agent
BRAIN_AGENT_BRAIN_URL=http://brain:8090
BRAIN_AGENT_ID=cpu-agent-01
BRAIN_AGENT_TOKEN=brain_agent_...
BRAIN_AGENT_COMPUTE_ENABLED=true
BRAIN_AGENT_COMPUTE_POLL_INTERVAL=5s
BRAIN_AGENT_COMPUTE_MAX_BYTES=134217728
The Source Agent still initializes no graph, Ollama, SearXNG, AI-THINK or article pipeline.
Wire protocol
The Brain owns the vector index and submits an in-memory pull job. An Agent claims it through the authenticated Agent API.
The request is streamed as a compact binary payload:
- fixed protocol magic;
- JSON job/config header;
- node ID;
- vector dimension;
- raw little-endian float32 vector values.
This avoids expanding a roughly 62 MiB 21k×768 float32 matrix into much larger decimal JSON.
The Agent returns only the mathematical result (links, optional positions, statistics). The Brain validates:
- compute kind;
- graph version;
- every source/target ID against the submitted set;
- no self-links;
- finite 0..1 similarity/affinity/confidence values.
The Brain then applies the result through the same graph code as a local build. The Agent cannot choose an arbitrary edge type: the Brain persists only semantic_neighbor with origin vector-math.
If layout is enabled, 3D vector positions are also calculated on the Agent; no model inference is involved.
No-model guarantee
A vector_graph compute job calls neither Chat nor Embed. It consumes embeddings that already exist in the Brain. Missing/new embeddings are still the Brain learning pipeline's responsibility.
Therefore:
- edge calculation: CPU-only arithmetic;
- orphan second pass: CPU-only arithmetic;
- layout: CPU-only arithmetic;
- Agent compute: CPU-only arithmetic;
- creation of a missing embedding: still embedding-model inference, outside the compute job.
Research topic guard
Web research now has a deterministic primary-entity/topic gate before source quality can rescue a result.
Generic terms such as security, hardening, support, testing, documentation, forensics and template verbs are removed from the primary topic anchors. German compounds are supported (browser matches Webbrowser).
Examples covered by regression tests:
- Browser Security → BSI Webbrowser: allowed;
- Browser Security → Proxmox Server Hardening: rejected;
- Rate Limit Testing → Rate Limit source: allowed;
- Rate Limit Testing → generic Web Security Testing: rejected.
A failed topic guard caps relevance and clears gap coverage even if an LLM assessment or high-quality domain would otherwise rank the source highly.
Audit persistence
The append-only analysis writer now uses a dedicated SQLite connection to the same WAL database. The primary graph connection deliberately remains MaxOpenConns(1), but audit telemetry no longer waits for that same Go connection slot.
The audit connection uses a 60-second SQLite busy timeout and longer writer deadlines. This addresses the observed context deadline exceeded persistence drops without weakening graph transaction ownership.
Real-corpus validation
Against the bundled production graph snapshot (21,289 production Knowledge vectors, 768 dimensions):
- primary pass: 16,940
semantic_neighborlinks, 16,007 reciprocal; - orphan focus after the primary pass: 5,218 nodes;
- orphan second pass: 3,038 additional links touching 2,200 focused orphan nodes;
- full Agent wire payload: 66,568,747 bytes (63.48 MiB);
- binary encode/decode in the supplied environment: ~0.38 s / ~0.43 s;
- primary + orphan compute after wire decode: ~8.1 s CPU wall time;
- no Chat or Embed call occurred inside the compute job.
The orphan pass kept min_similarity=0.80; the additional coverage therefore comes from a larger focused candidate search rather than globally weakening the similarity floor.
Recommended first distributed test
Brain:
BRAIN_VECTOR_GRAPH_ENABLED=true
BRAIN_VECTOR_GRAPH_ORPHAN_PASS=true
BRAIN_VECTOR_GRAPH_AGENT_OFFLOAD=true
BRAIN_VECTOR_GRAPH_AGENT_REQUIRED=false
BRAIN_THINKING_VECTOR_GUIDED=true
BRAIN_VECTOR_GRAPH_LAYOUT=false
Agent:
BRAIN_AGENT_COMPUTE_ENABLED=true
BRAIN_AGENT_COMPUTE_POLL_INTERVAL=5s
BRAIN_AGENT_COMPUTE_MAX_BYTES=134217728
Keep layout disabled for the first run. Verify vector.graph.agent.completed, orphan count, second-pass edge examples and Agent status/capabilities before enabling layout or making Agent offload mandatory.