update
mega-ci / static-release-gates (push) Failing after 11s
release-tag / release-image (push) Successful in 6m46s
mega-ci / go-quality (services/control) (push) Successful in 10m7s
mega-ci / go-quality (platform/neuroforge) (push) Successful in 10m20s
mega-ci / go-quality (services/agent) (push) Successful in 11m1s
mega-ci / go-quality (services/knowledge) (push) Successful in 11m13s
mega-ci / docker-build (push) Has been skipped
mega-ci / static-release-gates (push) Failing after 11s
release-tag / release-image (push) Successful in 6m46s
mega-ci / go-quality (services/control) (push) Successful in 10m7s
mega-ci / go-quality (platform/neuroforge) (push) Successful in 10m20s
mega-ci / go-quality (services/agent) (push) Successful in 11m1s
mega-ci / go-quality (services/knowledge) (push) Successful in 11m13s
mega-ci / docker-build (push) Has been skipped
This commit is contained in:
@@ -1,4 +1,17 @@
|
||||
# Changelog
|
||||
- v1.6.1 hotfix: streaming authoritative checkpoint I/O avoids full raw/encoded `state.json` copies during load/save.
|
||||
- v1.6.1 hotfix: corrupt `state.json`/`secrets.json` fail closed instead of being silently ignored.
|
||||
- v1.6.1 hotfix: completed `vector.relink` blobs are compacted during WAL replay as well as checkpoint migration.
|
||||
|
||||
## v0.8.3
|
||||
|
||||
- Crash-/Recovery-Hardening für große Knowledge-Korpora und Graph-Backfills.
|
||||
- Streaming-Migration entfernt historische `vector.relink`-Payloads vor dem vollständigen Checkpoint-Unmarshal.
|
||||
- Erfolgreiche Relink-Jobs verwerfen Target-/Candidate-Vektoren sofort nach Master-Apply.
|
||||
- Queue-Payload-Budget schützt den Master vor ungebremstem transientem Job-State.
|
||||
- HNSW-Deltas kopieren nur veränderte Nodes statt bei jedem Checkpoint den kompletten Index.
|
||||
- Checkpoint-Reihenfolge index-first verhindert `state.json`-Vorlauf nach Crash in der Indexpersistenz.
|
||||
- Bootstrap-HTTP und Startup-Phasenlogs machen lange Recovery sichtbar.
|
||||
|
||||
## v0.8.2
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# NeuroForge v0.8.2 – Production Guide
|
||||
# NeuroForge v0.8.3 – Production Guide
|
||||
|
||||
## 1. Sicherheitsgrenze
|
||||
|
||||
@@ -146,3 +146,7 @@ Research-Trace-Events sind **Observability/Audit**, nicht autoritative Knowledge
|
||||
|
||||
Die `preview`-Felder im Research-Trace sind absichtlich gekürzt und enthalten keine vollständigen Dokumente. Für vollständige Inhalte den Source-/Memory-Inspector verwenden.
|
||||
|
||||
|
||||
### v1.6.1 Recovery / large graph backfill
|
||||
|
||||
For large graph backfills keep completed relink payloads compact and bound pending payload bytes. The master streams `state.json` on load/save and fails closed on corrupt authoritative state/secrets. Never delete or replace a corrupt checkpoint automatically; retain the volume for forensic recovery/restore. During startup with explicit `-listen`, `/livez` and the bootstrap `/admin` page expose the current recovery phase while `/readyz` remains 503.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# NeuroForge v0.8.2
|
||||
# NeuroForge v0.8.3
|
||||
|
||||
NeuroForge ist eine persistente, assoziativ lernende KI-Schicht in Go. Ollama und optional OpenAI liefern Inferenz/Embeddings; NeuroForge besitzt den dauerhaften Wissenszustand: Vektoren, HNSW/Disk-PQ-Recall, Synapsen, Rewards, Provenance, Konflikte, Konsolidierung, Goals und Learning Cycles.
|
||||
|
||||
**v0.8.2 erweitert den Production-/Explainability-Stand um einen source-grounded Lernpfad, Dokument-/Text-Ingestion, SearXNG-Research und ein vollständig neu gestaltetes CSS/Vanilla-JS-Admin-UI mit responsive Knowledge-Graph und Level-of-Detail (LOD).** Wissen soll nicht nur gespeichert, sondern als Kette `Quelle → Evidence → Recall → Learning` nachvollziehbar sein.
|
||||
**v0.8.3 übernimmt den v0.8.2-Funktionsumfang und härtet große Bulk-/Graph-Workloads gegen Speicher- und Recovery-Spitzen. v0.8.2 erweitert den Production-/Explainability-Stand um einen source-grounded Lernpfad, Dokument-/Text-Ingestion, SearXNG-Research und ein vollständig neu gestaltetes CSS/Vanilla-JS-Admin-UI mit responsive Knowledge-Graph und Level-of-Detail (LOD).** Wissen soll nicht nur gespeichert, sondern als Kette `Quelle → Evidence → Recall → Learning` nachvollziehbar sein.
|
||||
|
||||
|
||||
## Neu in v0.8.2: Live Research pro Goal
|
||||
|
||||
@@ -1 +1 @@
|
||||
0.8.2
|
||||
0.8.3
|
||||
|
||||
@@ -6,12 +6,14 @@ import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"log"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
@@ -101,6 +103,59 @@ func maxIntMain(a, b int) int {
|
||||
return b
|
||||
}
|
||||
|
||||
type bootstrapHandler struct {
|
||||
mu sync.RWMutex
|
||||
phase string
|
||||
handler http.Handler
|
||||
}
|
||||
|
||||
func newBootstrapHandler() *bootstrapHandler {
|
||||
return &bootstrapHandler{phase: "process.start"}
|
||||
}
|
||||
|
||||
func (b *bootstrapHandler) SetPhase(phase string) {
|
||||
b.mu.Lock()
|
||||
b.phase = phase
|
||||
b.mu.Unlock()
|
||||
}
|
||||
|
||||
func (b *bootstrapHandler) SetHandler(h http.Handler) {
|
||||
b.mu.Lock()
|
||||
b.handler = h
|
||||
b.phase = "ready"
|
||||
b.mu.Unlock()
|
||||
}
|
||||
|
||||
func (b *bootstrapHandler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
|
||||
b.mu.RLock()
|
||||
h := b.handler
|
||||
phase := b.phase
|
||||
b.mu.RUnlock()
|
||||
if h != nil {
|
||||
h.ServeHTTP(w, r)
|
||||
return
|
||||
}
|
||||
w.Header().Set("Cache-Control", "no-store")
|
||||
switch r.URL.Path {
|
||||
case "/livez", "/healthz":
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
w.WriteHeader(http.StatusOK)
|
||||
_, _ = fmt.Fprintf(w, `{"ok":true,"status":"starting","phase":%q}`, phase)
|
||||
case "/readyz":
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
_, _ = fmt.Fprintf(w, `{"ok":false,"status":"starting","phase":%q}`, phase)
|
||||
case "/", "/admin":
|
||||
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
_, _ = fmt.Fprintf(w, `<!doctype html><html><head><meta charset="utf-8"><title>NeuroForge starting</title><meta http-equiv="refresh" content="5"></head><body><h1>NeuroForge startet</h1><p>Recovery-/Startphase: <code>%s</code></p><p>Die autoritativen Daten werden geladen. Diese Seite aktualisiert sich automatisch.</p></body></html>`, phase)
|
||||
default:
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
_, _ = fmt.Fprintf(w, `{"error":"neuroforge is starting","phase":%q}`, phase)
|
||||
}
|
||||
}
|
||||
|
||||
func main() {
|
||||
if err := run(); err != nil {
|
||||
log.Printf("fatal: %v", err)
|
||||
@@ -117,7 +172,55 @@ func run() (retErr error) {
|
||||
return err
|
||||
}
|
||||
|
||||
s, err := store.New(*data)
|
||||
var boot *bootstrapHandler
|
||||
var srv *http.Server
|
||||
var errCh chan error
|
||||
// Container deployments pass an explicit -listen address. Bind it before
|
||||
// opening the potentially large store so liveness and a minimal startup UI
|
||||
// remain reachable during WAL/index recovery instead of looking like a dead
|
||||
// container with no logs.
|
||||
if strings.TrimSpace(*listen) != "" {
|
||||
boot = newBootstrapHandler()
|
||||
d := core.DefaultConfig()
|
||||
h := d.HTTP
|
||||
srv = &http.Server{
|
||||
Addr: *listen,
|
||||
Handler: boot,
|
||||
ReadHeaderTimeout: time.Duration(h.ReadHeaderTimeoutSeconds) * time.Second,
|
||||
ReadTimeout: time.Duration(h.ReadTimeoutSeconds) * time.Second,
|
||||
WriteTimeout: time.Duration(h.WriteTimeoutSeconds) * time.Second,
|
||||
IdleTimeout: time.Duration(h.IdleTimeoutSeconds) * time.Second,
|
||||
MaxHeaderBytes: h.MaxHeaderBytes,
|
||||
}
|
||||
ln, err := net.Listen("tcp", *listen)
|
||||
if err != nil {
|
||||
return fmt.Errorf("bootstrap listen %s: %w", *listen, err)
|
||||
}
|
||||
errCh = make(chan error, 1)
|
||||
go func() {
|
||||
err := srv.Serve(ln)
|
||||
if errors.Is(err, http.ErrServerClosed) {
|
||||
err = nil
|
||||
}
|
||||
errCh <- err
|
||||
}()
|
||||
log.Printf("NeuroForge bootstrap listener active on %s", *listen)
|
||||
defer func() {
|
||||
if retErr != nil && srv != nil {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
|
||||
defer cancel()
|
||||
_ = srv.Shutdown(ctx)
|
||||
}
|
||||
}()
|
||||
}
|
||||
reportStartup := func(phase string) {
|
||||
log.Printf("startup phase: %s", phase)
|
||||
if boot != nil {
|
||||
boot.SetPhase(phase)
|
||||
}
|
||||
}
|
||||
|
||||
s, err := store.NewWithProgress(*data, reportStartup)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -258,7 +361,7 @@ func run() (retErr error) {
|
||||
// and environment-owned only when set, preserving admin-managed config otherwise.
|
||||
workerEnv := []string{
|
||||
"NEUROFORGE_WORKER_LEASE_SECONDS", "NEUROFORGE_WORKER_HEARTBEAT_SECONDS", "NEUROFORGE_WORKER_STALE_AFTER_SECONDS",
|
||||
"NEUROFORGE_WORKER_DEFAULT_MAX_ATTEMPTS", "NEUROFORGE_WORKER_RETRY_BACKOFF_SECONDS", "NEUROFORGE_WORKER_MAX_QUEUED_JOBS",
|
||||
"NEUROFORGE_WORKER_DEFAULT_MAX_ATTEMPTS", "NEUROFORGE_WORKER_RETRY_BACKOFF_SECONDS", "NEUROFORGE_WORKER_MAX_QUEUED_JOBS", "NEUROFORGE_WORKER_MAX_QUEUED_PAYLOAD_MB",
|
||||
"NEUROFORGE_WORKER_JOB_RETENTION_HOURS", "NEUROFORGE_WORKER_MAX_TERMINAL_JOBS",
|
||||
"NEUROFORGE_WORKER_MASTER_APPLY_MAX_ATTEMPTS", "NEUROFORGE_WORKER_MASTER_APPLY_BACKOFF_SECONDS",
|
||||
"NEUROFORGE_GRAPH_BACKFILL_ENABLED", "NEUROFORGE_GRAPH_BACKFILL_INTERVAL_SECONDS", "NEUROFORGE_GRAPH_BACKFILL_BATCH_SIZE",
|
||||
@@ -294,6 +397,9 @@ func run() (retErr error) {
|
||||
if v, ok := envInt("NEUROFORGE_WORKER_MAX_QUEUED_JOBS"); ok {
|
||||
cfg.Worker.MaxQueuedJobs = v
|
||||
}
|
||||
if v, ok := envInt("NEUROFORGE_WORKER_MAX_QUEUED_PAYLOAD_MB"); ok {
|
||||
cfg.Worker.MaxQueuedPayloadMB = v
|
||||
}
|
||||
if v, ok := envInt("NEUROFORGE_WORKER_JOB_RETENTION_HOURS"); ok {
|
||||
cfg.Worker.JobRetentionHours = v
|
||||
}
|
||||
@@ -453,30 +559,33 @@ func run() (retErr error) {
|
||||
}
|
||||
|
||||
h := cfg.HTTP
|
||||
srv := &http.Server{
|
||||
Addr: addr,
|
||||
Handler: api.Handler(),
|
||||
ReadHeaderTimeout: time.Duration(h.ReadHeaderTimeoutSeconds) * time.Second,
|
||||
ReadTimeout: time.Duration(h.ReadTimeoutSeconds) * time.Second,
|
||||
WriteTimeout: time.Duration(h.WriteTimeoutSeconds) * time.Second,
|
||||
IdleTimeout: time.Duration(h.IdleTimeoutSeconds) * time.Second,
|
||||
MaxHeaderBytes: h.MaxHeaderBytes,
|
||||
if srv == nil {
|
||||
srv = &http.Server{
|
||||
Addr: addr,
|
||||
Handler: api.Handler(),
|
||||
ReadHeaderTimeout: time.Duration(h.ReadHeaderTimeoutSeconds) * time.Second,
|
||||
ReadTimeout: time.Duration(h.ReadTimeoutSeconds) * time.Second,
|
||||
WriteTimeout: time.Duration(h.WriteTimeoutSeconds) * time.Second,
|
||||
IdleTimeout: time.Duration(h.IdleTimeoutSeconds) * time.Second,
|
||||
MaxHeaderBytes: h.MaxHeaderBytes,
|
||||
}
|
||||
errCh = make(chan error, 1)
|
||||
go func() {
|
||||
err := srv.ListenAndServe()
|
||||
if errors.Is(err, http.ErrServerClosed) {
|
||||
err = nil
|
||||
}
|
||||
errCh <- err
|
||||
}()
|
||||
} else {
|
||||
boot.SetHandler(api.Handler())
|
||||
}
|
||||
log.Printf("NeuroForge v0.8.2 listening on %s", addr)
|
||||
log.Printf("NeuroForge v0.8.3 listening on %s", addr)
|
||||
log.Printf("Admin dashboard: /admin · readiness: /readyz · metrics: /metrics")
|
||||
if os.Getenv("NEUROFORGE_ADMIN_TOKEN") == "" {
|
||||
log.Printf("Admin token is intentionally not printed; read it locally from %s or set NEUROFORGE_ADMIN_TOKEN", filepath.Join(*data, "secrets.json"))
|
||||
}
|
||||
|
||||
errCh := make(chan error, 1)
|
||||
go func() {
|
||||
err := srv.ListenAndServe()
|
||||
if errors.Is(err, http.ErrServerClosed) {
|
||||
err = nil
|
||||
}
|
||||
errCh <- err
|
||||
}()
|
||||
|
||||
var serveErr error
|
||||
select {
|
||||
case serveErr = <-errCh:
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestBootstrapHandlerExposesLivenessAndStartupUI(t *testing.T) {
|
||||
b := newBootstrapHandler()
|
||||
b.SetPhase("hnsw.rebuild")
|
||||
|
||||
rr := httptest.NewRecorder()
|
||||
b.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/livez", nil))
|
||||
if rr.Code != http.StatusOK || !strings.Contains(rr.Body.String(), "hnsw.rebuild") {
|
||||
t.Fatalf("livez=%d %s", rr.Code, rr.Body.String())
|
||||
}
|
||||
|
||||
rr = httptest.NewRecorder()
|
||||
b.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/readyz", nil))
|
||||
if rr.Code != http.StatusServiceUnavailable {
|
||||
t.Fatalf("readyz=%d", rr.Code)
|
||||
}
|
||||
|
||||
rr = httptest.NewRecorder()
|
||||
b.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/admin", nil))
|
||||
if rr.Code != http.StatusServiceUnavailable || !strings.Contains(rr.Body.String(), "hnsw.rebuild") {
|
||||
t.Fatalf("admin=%d %s", rr.Code, rr.Body.String())
|
||||
}
|
||||
|
||||
b.SetHandler(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { w.WriteHeader(http.StatusNoContent) }))
|
||||
rr = httptest.NewRecorder()
|
||||
b.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/admin", nil))
|
||||
if rr.Code != http.StatusNoContent {
|
||||
t.Fatalf("delegated=%d", rr.Code)
|
||||
}
|
||||
}
|
||||
@@ -398,6 +398,7 @@ type Config struct {
|
||||
DefaultMaxAttempts int `json:"default_max_attempts"`
|
||||
RetryBackoffSeconds int `json:"retry_backoff_seconds"`
|
||||
MaxQueuedJobs int `json:"max_queued_jobs"`
|
||||
MaxQueuedPayloadMB int `json:"max_queued_payload_mb"`
|
||||
MasterApplyMaxAttempts int `json:"master_apply_max_attempts"`
|
||||
MasterApplyBackoffSeconds int `json:"master_apply_backoff_seconds"`
|
||||
JobRetentionHours int `json:"job_retention_hours"`
|
||||
@@ -921,14 +922,15 @@ func DefaultConfig() Config {
|
||||
c.Worker.DefaultMaxAttempts = 3
|
||||
c.Worker.RetryBackoffSeconds = 15
|
||||
c.Worker.MaxQueuedJobs = 5000
|
||||
c.Worker.MaxQueuedPayloadMB = 128
|
||||
c.Worker.MasterApplyMaxAttempts = 5
|
||||
c.Worker.MasterApplyBackoffSeconds = 5
|
||||
c.Worker.JobRetentionHours = 168
|
||||
c.Worker.MaxTerminalJobs = 20000
|
||||
c.Worker.JobRetentionHours = 24
|
||||
c.Worker.MaxTerminalJobs = 2000
|
||||
c.Worker.GraphBackfillEnabled = true
|
||||
c.Worker.GraphBackfillIntervalS = 10
|
||||
c.Worker.GraphBackfillBatchSize = 64
|
||||
c.Worker.GraphBackfillMaxQueued = 512
|
||||
c.Worker.GraphBackfillBatchSize = 16
|
||||
c.Worker.GraphBackfillMaxQueued = 64
|
||||
c.Worker.GraphBackfillMinDegree = 3
|
||||
c.Worker.GraphCandidateMultiplier = 6
|
||||
c.Worker.GraphRetryAfterMinutes = 360
|
||||
|
||||
@@ -59,7 +59,7 @@ func (s *Server) routes() {
|
||||
s.mux.HandleFunc("GET /healthz", s.livez)
|
||||
s.mux.HandleFunc("GET /livez", s.livez)
|
||||
s.mux.HandleFunc("GET /readyz", s.readyz)
|
||||
s.mux.HandleFunc("GET /version", func(w http.ResponseWriter, r *http.Request) { s.json(w, 200, map[string]any{"version": "0.8.2"}) })
|
||||
s.mux.HandleFunc("GET /version", func(w http.ResponseWriter, r *http.Request) { s.json(w, 200, map[string]any{"version": "0.8.3"}) })
|
||||
s.mux.Handle("POST /api/v1/chat", s.appAuth(http.HandlerFunc(s.chat)))
|
||||
s.mux.Handle("POST /api/v1/learn", s.appAuth(http.HandlerFunc(s.learn)))
|
||||
s.mux.Handle("POST /api/v1/search", s.appAuth(http.HandlerFunc(s.search)))
|
||||
@@ -882,7 +882,7 @@ func (s *Server) requestLimits(next http.Handler) http.Handler {
|
||||
}
|
||||
|
||||
func (s *Server) livez(w http.ResponseWriter, r *http.Request) {
|
||||
s.json(w, http.StatusOK, map[string]any{"ok": true, "status": "alive", "time": time.Now().UTC(), "version": "0.8.2"})
|
||||
s.json(w, http.StatusOK, map[string]any{"ok": true, "status": "alive", "time": time.Now().UTC(), "version": "0.8.3"})
|
||||
}
|
||||
|
||||
func configuredModelAvailable(models map[string]bool, configured string) bool {
|
||||
|
||||
@@ -276,6 +276,9 @@ func (s *Server) metricsEndpoint(w http.ResponseWriter, r *http.Request) {
|
||||
for _, status := range statuses {
|
||||
promSample(&b, "neuroforge_jobs", orch.JobsByStatus[status], "status", status)
|
||||
}
|
||||
promHeader(&b, "neuroforge_job_payload_bytes", "Durable orchestrator payload/result bytes retained in memory and checkpoints.", "gauge")
|
||||
promSample(&b, "neuroforge_job_payload_bytes", orch.PendingPayloadBytes, "state", "pending")
|
||||
promSample(&b, "neuroforge_job_payload_bytes", orch.TerminalPayloadBytes, "state", "terminal")
|
||||
promHeader(&b, "neuroforge_jobs_by_resource", "Current durable orchestrator jobs by resource class.", "gauge")
|
||||
for resource, n := range orch.JobsByResource {
|
||||
promSample(&b, "neuroforge_jobs_by_resource", n, "resource", resource)
|
||||
|
||||
@@ -169,7 +169,7 @@ func FetchResource(ctx context.Context, cfg FetchConfig, rawURL string) (Resourc
|
||||
req.Header.Set("Accept", "text/html,application/xhtml+xml,application/pdf,application/vnd.openxmlformats-officedocument.wordprocessingml.document,text/plain,text/markdown,text/csv,application/json;q=0.9,*/*;q=0.2")
|
||||
ua := strings.TrimSpace(cfg.UserAgent)
|
||||
if ua == "" {
|
||||
ua = "NeuroForge/0.8.2 research bot"
|
||||
ua = "NeuroForge/0.8.3 research bot"
|
||||
}
|
||||
req.Header.Set("User-Agent", ua)
|
||||
resp, err := client.Do(req)
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"neuroforge/internal/core"
|
||||
)
|
||||
|
||||
func readCheckpointRevision(t *testing.T, dir string) uint64 {
|
||||
t.Helper()
|
||||
b, err := os.ReadFile(filepath.Join(dir, "state.json"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var st core.PersistedState
|
||||
if err := json.Unmarshal(b, &st); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return st.Revision
|
||||
}
|
||||
|
||||
func TestCheckpointDoesNotAdvanceStateBeforeIndexSnapshotSucceeds(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
s, err := New(dir)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
cfg := s.Config()
|
||||
cfg.Storage.CheckpointEvery = 1
|
||||
if err := s.UpdateConfig(cfg); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
before := readCheckpointRevision(t, dir)
|
||||
|
||||
idxDir := filepath.Join(dir, "hnsw-index")
|
||||
if err := os.RemoveAll(idxDir); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(idxDir, []byte("block directory creation"), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
m := &core.Memory{ID: "m1", Kind: core.MemorySemantic, MemoryType: core.MemorySemantic, Text: "x", Vector: []float32{1, 0, 0}, VectorDim: 3, Status: core.MemoryActive}
|
||||
if err := s.AddMemory(m); err == nil {
|
||||
t.Fatal("expected checkpoint/index failure")
|
||||
}
|
||||
after := readCheckpointRevision(t, dir)
|
||||
if after != before {
|
||||
t.Fatalf("state checkpoint advanced despite index failure: before=%d after=%d", before, after)
|
||||
}
|
||||
|
||||
_ = os.Remove(idxDir)
|
||||
_ = s.Close()
|
||||
}
|
||||
@@ -211,41 +211,28 @@ func (s *Store) writeSegmentedIndexSnapshotLocked() error {
|
||||
return nil
|
||||
}
|
||||
|
||||
current := s.currentSnapshotsLocked()
|
||||
currentShadow := make(map[int]indexSnapshotShadow, len(s.indexes))
|
||||
delta := indexDeltaBundle{Revision: s.state.Revision, Dimensions: map[string]indexDimensionDelta{}}
|
||||
dims := map[int]bool{}
|
||||
for dim := range current {
|
||||
for dim := range s.indexes {
|
||||
dims[dim] = true
|
||||
}
|
||||
for dim := range s.indexShadow {
|
||||
dims[dim] = true
|
||||
}
|
||||
for dim := range dims {
|
||||
cur, curOK := current[dim]
|
||||
idx, curOK := s.indexes[dim]
|
||||
prev, prevOK := s.indexShadow[dim]
|
||||
key := strconv.Itoa(dim)
|
||||
if !curOK {
|
||||
if !curOK || idx == nil {
|
||||
delta.Dimensions[key] = indexDimensionDelta{DeletedDimension: true}
|
||||
continue
|
||||
}
|
||||
d := indexDimensionDelta{Config: cur.Config, EntryID: cur.EntryID, MaxLevel: cur.MaxLevel}
|
||||
curIDs := map[string]bool{}
|
||||
for _, n := range cur.Nodes {
|
||||
curIDs[n.ID] = true
|
||||
h := hashSnapshotNode(n)
|
||||
if ph, ok := prev.Nodes[n.ID]; !prevOK || !ok || ph != h {
|
||||
d.Upserts = append(d.Upserts, n)
|
||||
}
|
||||
}
|
||||
if prevOK {
|
||||
for id := range prev.Nodes {
|
||||
if !curIDs[id] {
|
||||
d.Deletes = append(d.Deletes, id)
|
||||
}
|
||||
}
|
||||
}
|
||||
sort.Strings(d.Deletes)
|
||||
if !prevOK || len(d.Upserts) > 0 || len(d.Deletes) > 0 || prev.EntryID != cur.EntryID || prev.MaxLevel != cur.MaxLevel || prev.Config != cur.Config {
|
||||
vprev := vector.HNSWShadow{Config: prev.Config, EntryID: prev.EntryID, MaxLevel: prev.MaxLevel, Nodes: prev.Nodes}
|
||||
cur, upserts, deletes := idx.Delta(vprev)
|
||||
currentShadow[dim] = indexSnapshotShadow{Config: cur.Config, EntryID: cur.EntryID, MaxLevel: cur.MaxLevel, Nodes: cur.Nodes}
|
||||
d := indexDimensionDelta{Config: cur.Config, EntryID: cur.EntryID, MaxLevel: cur.MaxLevel, Upserts: upserts, Deletes: deletes}
|
||||
if !prevOK || len(upserts) > 0 || len(deletes) > 0 || prev.EntryID != cur.EntryID || prev.MaxLevel != cur.MaxLevel || prev.Config != cur.Config {
|
||||
delta.Dimensions[key] = d
|
||||
}
|
||||
}
|
||||
@@ -258,7 +245,7 @@ func (s *Store) writeSegmentedIndexSnapshotLocked() error {
|
||||
if err := writeAtomic(manifestPath, 0600, &manifest); err != nil {
|
||||
return err
|
||||
}
|
||||
s.indexShadow = buildIndexShadow(current)
|
||||
s.indexShadow = currentShadow
|
||||
s.indexSnapshotRevision = s.state.Revision
|
||||
s.indexDeltaCount = len(manifest.Deltas)
|
||||
return nil
|
||||
|
||||
@@ -43,18 +43,20 @@ type WorkerHeartbeat struct {
|
||||
}
|
||||
|
||||
type OrchestratorStatus struct {
|
||||
Workers []core.WorkerState `json:"workers"`
|
||||
JobsByStatus map[string]int `json:"jobs_by_status"`
|
||||
JobsByType map[string]int `json:"jobs_by_type"`
|
||||
JobsByResource map[string]int `json:"jobs_by_resource"`
|
||||
OldestQueued time.Time `json:"oldest_queued,omitempty"`
|
||||
Queued int `json:"queued"`
|
||||
Claimed int `json:"claimed"`
|
||||
Retrying int `json:"retrying"`
|
||||
Applying int `json:"applying"`
|
||||
Blocked int `json:"blocked"`
|
||||
Failed int `json:"failed"`
|
||||
Done int `json:"done"`
|
||||
Workers []core.WorkerState `json:"workers"`
|
||||
JobsByStatus map[string]int `json:"jobs_by_status"`
|
||||
JobsByType map[string]int `json:"jobs_by_type"`
|
||||
JobsByResource map[string]int `json:"jobs_by_resource"`
|
||||
OldestQueued time.Time `json:"oldest_queued,omitempty"`
|
||||
Queued int `json:"queued"`
|
||||
Claimed int `json:"claimed"`
|
||||
Retrying int `json:"retrying"`
|
||||
Applying int `json:"applying"`
|
||||
Blocked int `json:"blocked"`
|
||||
Failed int `json:"failed"`
|
||||
Done int `json:"done"`
|
||||
PendingPayloadBytes int64 `json:"pending_payload_bytes"`
|
||||
TerminalPayloadBytes int64 `json:"terminal_payload_bytes"`
|
||||
}
|
||||
|
||||
func normalizeCapabilities(in []string) []string {
|
||||
@@ -104,15 +106,22 @@ func (s *Store) EnqueueJobSpec(spec JobSpec) (*core.Job, error) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
cfg := s.state.Config.Worker
|
||||
if cfg.MaxQueuedJobs > 0 {
|
||||
pending := 0
|
||||
for _, j := range s.state.Jobs {
|
||||
if j != nil && (j.Status == "queued" || j.Status == "claimed" || j.Status == "retry_wait" || j.Status == "blocked" || j.Status == "apply_wait") {
|
||||
pending++
|
||||
}
|
||||
pending := 0
|
||||
var pendingPayloadBytes int64
|
||||
for _, j := range s.state.Jobs {
|
||||
if j == nil || (j.Status != "queued" && j.Status != "claimed" && j.Status != "retry_wait" && j.Status != "blocked" && j.Status != "apply_wait") {
|
||||
continue
|
||||
}
|
||||
if pending >= cfg.MaxQueuedJobs {
|
||||
return nil, fmt.Errorf("orchestrator queue full: %d >= %d", pending, cfg.MaxQueuedJobs)
|
||||
pending++
|
||||
pendingPayloadBytes += int64(len(j.Payload) + len(j.Result))
|
||||
}
|
||||
if cfg.MaxQueuedJobs > 0 && pending >= cfg.MaxQueuedJobs {
|
||||
return nil, fmt.Errorf("orchestrator queue full: %d >= %d", pending, cfg.MaxQueuedJobs)
|
||||
}
|
||||
if cfg.MaxQueuedPayloadMB > 0 {
|
||||
limit := int64(cfg.MaxQueuedPayloadMB) << 20
|
||||
if pendingPayloadBytes+int64(len(b)) > limit {
|
||||
return nil, fmt.Errorf("orchestrator queued payload budget exceeded: %d + %d > %d bytes", pendingPayloadBytes, len(b), limit)
|
||||
}
|
||||
}
|
||||
if key := strings.TrimSpace(spec.IdempotencyKey); key != "" {
|
||||
@@ -603,6 +612,13 @@ func (s *Store) OrchestratorStatus() OrchestratorStatus {
|
||||
if j == nil {
|
||||
continue
|
||||
}
|
||||
payloadBytes := int64(len(j.Payload) + len(j.Result))
|
||||
switch j.Status {
|
||||
case "queued", "claimed", "retry_wait", "blocked", "apply_wait":
|
||||
out.PendingPayloadBytes += payloadBytes
|
||||
case "done", "failed", "canceled":
|
||||
out.TerminalPayloadBytes += payloadBytes
|
||||
}
|
||||
out.JobsByStatus[j.Status]++
|
||||
out.JobsByType[j.Type]++
|
||||
resource := j.ResourceClass
|
||||
@@ -701,6 +717,15 @@ func (s *Store) FinishMasterApply(id string, applyErr error) (*core.Job, error)
|
||||
j.ApplyError = ""
|
||||
j.ApplyNextAttemptAt = time.Time{}
|
||||
j.FinishedAt = now
|
||||
// vector.relink payloads contain the target plus a bounded set of full
|
||||
// candidate vectors. Once the authoritative master apply succeeded these
|
||||
// blobs have no retry value and retaining thousands of them can consume
|
||||
// gigabytes during a large graph backfill. Keep the durable audit metadata
|
||||
// (type/idempotency/status/timestamps) but release the transient vectors.
|
||||
if j.Type == "vector.relink" {
|
||||
j.Payload = nil
|
||||
j.Result = nil
|
||||
}
|
||||
} else {
|
||||
j.ApplyError = strings.TrimSpace(applyErr.Error())
|
||||
maxAttempts := j.MaxApplyAttempts
|
||||
@@ -732,6 +757,41 @@ func (s *Store) FinishMasterApply(id string, applyErr error) (*core.Job, error)
|
||||
return &cp, s.commitLocked("job.upsert", cp)
|
||||
}
|
||||
|
||||
// compactCompletedRelinkJobsLocked removes transient vector blobs from jobs
|
||||
// that already reached their durable terminal state. It is also used on boot to
|
||||
// migrate v1.6.0 checkpoints that may contain thousands of completed backfill
|
||||
// payloads. Caller must hold s.mu.
|
||||
func (s *Store) compactCompletedRelinkJobsLocked() int {
|
||||
n := 0
|
||||
for _, j := range s.state.Jobs {
|
||||
if j == nil || j.Type != "vector.relink" || j.Status != "done" {
|
||||
continue
|
||||
}
|
||||
if len(j.Payload) == 0 && len(j.Result) == 0 {
|
||||
continue
|
||||
}
|
||||
j.Payload = nil
|
||||
j.Result = nil
|
||||
n++
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
func (s *Store) CompactCompletedRelinkJobs() (int, error) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
n := s.compactCompletedRelinkJobsLocked()
|
||||
if n == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
// One checkpoint is substantially cheaper than one WAL event per historical
|
||||
// job and atomically rewrites state.json without the obsolete vector blobs.
|
||||
if err := s.checkpointLocked(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
return n, nil
|
||||
}
|
||||
|
||||
func (s *Store) CancelJob(id, reason string) error {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
@@ -292,3 +292,52 @@ func TestApplyWaitJobSurvivesRestart(t *testing.T) {
|
||||
t.Fatalf("recovered master apply queue=%+v", pending)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOrchestratorQueuedPayloadBudget(t *testing.T) {
|
||||
s, err := New(t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer s.Close()
|
||||
s.mu.Lock()
|
||||
s.state.Config.Worker.MaxQueuedPayloadMB = 1
|
||||
s.mu.Unlock()
|
||||
blob := strings.Repeat("x", 700<<10)
|
||||
if _, err := s.EnqueueJobSpec(JobSpec{Type: "large", Payload: map[string]any{"blob": blob}, ResourceClass: "cpu"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := s.EnqueueJobSpec(JobSpec{Type: "large2", Payload: map[string]any{"blob": blob}, ResourceClass: "cpu"}); err == nil || !strings.Contains(err.Error(), "payload budget") {
|
||||
t.Fatalf("expected payload budget rejection, err=%v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCompletedRelinkDropsTransientVectorBlobs(t *testing.T) {
|
||||
s, err := New(t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer s.Close()
|
||||
j, err := s.EnqueueJobSpec(JobSpec{Type: "vector.relink", Payload: map[string]any{"target": []float32{1, 2, 3}}, ResourceClass: "cpu", RequiredCapabilities: []string{"cpu", "vector.relink"}, RequiresMasterApply: true})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
w := WorkerHeartbeat{ID: "cpu", ResourceClass: "cpu", Capabilities: []string{"cpu", "vector.relink"}, MaxConcurrency: 1}
|
||||
claimed, err := s.ClaimJobForWorker(w, time.Minute)
|
||||
if err != nil || claimed == nil || claimed.ID != j.ID {
|
||||
t.Fatalf("claim=%+v err=%v", claimed, err)
|
||||
}
|
||||
waiting, err := s.CompleteJobLease(j.ID, w.ID, claimed.LeaseToken, json.RawMessage(`{"target_id":"x","neighbors":[]}`), "")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if waiting.Status != "apply_wait" {
|
||||
t.Fatalf("status=%s", waiting.Status)
|
||||
}
|
||||
done, err := s.FinishMasterApply(j.ID, nil)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if done.Status != "done" || len(done.Payload) != 0 || len(done.Result) != 0 {
|
||||
t.Fatalf("done job retained transient blobs: %+v", done)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,120 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"neuroforge/internal/core"
|
||||
)
|
||||
|
||||
func TestNewFailsClosedOnCorruptAuthoritativeState(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
if err := os.WriteFile(filepath.Join(dir, v161JobCompactionMarker), []byte("compacted=0\n"), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "state.json"), []byte(`{"config":`), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := New(dir); err == nil || !strings.Contains(err.Error(), "load authoritative state checkpoint") {
|
||||
t.Fatalf("New error=%v, want authoritative state corruption failure", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewFailsClosedOnCorruptAuthoritativeSecrets(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
st := core.PersistedState{
|
||||
Config: core.DefaultConfig(),
|
||||
Memories: map[string]*core.Memory{},
|
||||
Synapses: map[string]*core.Synapse{},
|
||||
Jobs: map[string]*core.Job{},
|
||||
Goals: map[string]*core.Goal{},
|
||||
Sources: map[string]*core.KnowledgeSource{},
|
||||
ResearchRuns: map[string]*core.ResearchRun{},
|
||||
}
|
||||
if err := writeAtomic(filepath.Join(dir, "state.json"), 0600, &st); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "secrets.json"), []byte(`{"admin_token":"unterminated}`), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := New(dir); err == nil || !strings.Contains(err.Error(), "load authoritative secrets") {
|
||||
t.Fatalf("New error=%v, want authoritative secrets corruption failure", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadJSONRejectsTrailingJSONValue(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "state.json")
|
||||
if err := os.WriteFile(path, []byte(`{"revision":1}{"revision":2}`), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
s := &Store{}
|
||||
var st core.PersistedState
|
||||
if err := s.loadJSON(path, &st); err == nil || !strings.Contains(err.Error(), "unexpected trailing JSON token") {
|
||||
t.Fatalf("loadJSON error=%v, want trailing JSON rejection", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReplayWALCompactsCompletedRelinkPayloadImmediately(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
walDir := filepath.Join(dir, "wal")
|
||||
if err := os.MkdirAll(walDir, 0700); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
job := core.Job{
|
||||
ID: "job_done_relink",
|
||||
Type: "vector.relink",
|
||||
Status: "done",
|
||||
Payload: json.RawMessage(`{"target":{"vector":[1,2,3]},"candidates":[{"vector":[4,5,6]}]}`),
|
||||
Result: json.RawMessage(`{"edges":[{"a":"a","b":"b","weight":0.9}]}`),
|
||||
CreatedAt: time.Now().UTC(),
|
||||
UpdatedAt: time.Now().UTC(),
|
||||
}
|
||||
data, err := json.Marshal(job)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ev := walEvent{Revision: 1, Time: time.Now().UTC(), Type: "job.upsert", Data: data}
|
||||
f, err := os.OpenFile(filepath.Join(walDir, "wal-active.jsonl"), os.O_CREATE|os.O_WRONLY|os.O_TRUNC, 0600)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
bw := bufio.NewWriter(f)
|
||||
if err := json.NewEncoder(bw).Encode(ev); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := bw.Flush(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := f.Close(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
s, err := New(dir)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer s.Close()
|
||||
got, ok := s.Job(job.ID)
|
||||
if !ok {
|
||||
t.Fatal("replayed job missing")
|
||||
}
|
||||
if len(got.Payload) != 0 || len(got.Result) != 0 {
|
||||
t.Fatalf("replayed completed relink retained vector blobs: payload=%d result=%d", len(got.Payload), len(got.Result))
|
||||
}
|
||||
}
|
||||
|
||||
func TestFreshDirectoryDoesNotConsumeV161CompactionMarker(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
if n, err := compactLegacyTerminalRelinkCheckpoint(dir); err != nil || n != 0 {
|
||||
t.Fatalf("compact fresh dir = %d, %v", n, err)
|
||||
}
|
||||
if _, err := os.Stat(filepath.Join(dir, v161JobCompactionMarker)); !os.IsNotExist(err) {
|
||||
t.Fatalf("fresh directory unexpectedly created migration marker: %v", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,277 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
|
||||
"neuroforge/internal/core"
|
||||
)
|
||||
|
||||
const v161JobCompactionMarker = ".migration-v1.6.1-terminal-relink-compaction"
|
||||
|
||||
// compactLegacyTerminalRelinkCheckpoint is a one-time, streaming migration for
|
||||
// v1.6.0 checkpoints. That release retained full target/candidate vectors in
|
||||
// completed vector.relink jobs. A large graph backfill can therefore make
|
||||
// state.json hundreds of MB or larger and cause OOM during the next startup.
|
||||
//
|
||||
// The migration deliberately runs before state.json is unmarshaled. It rewrites
|
||||
// JSON token-by-token and only materializes one Job at a time, so peak memory is
|
||||
// bounded by the largest single job instead of the whole checkpoint.
|
||||
func compactLegacyTerminalRelinkCheckpoint(dir string) (int, error) {
|
||||
marker := filepath.Join(dir, v161JobCompactionMarker)
|
||||
if _, err := os.Stat(marker); err == nil {
|
||||
return 0, nil
|
||||
}
|
||||
path := filepath.Join(dir, "state.json")
|
||||
in, err := os.Open(path)
|
||||
if err != nil {
|
||||
if errors.Is(err, os.ErrNotExist) {
|
||||
// Fresh data directory. Do not create the migration marker yet: an
|
||||
// operator may restore a v1.6.0 checkpoint into this directory before
|
||||
// the next boot, and that restored checkpoint must still be compacted.
|
||||
return 0, nil
|
||||
}
|
||||
return 0, err
|
||||
}
|
||||
defer in.Close()
|
||||
|
||||
tmp := path + ".v161-compact.tmp"
|
||||
out, err := os.OpenFile(tmp, os.O_CREATE|os.O_TRUNC|os.O_WRONLY, 0600)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
ok := false
|
||||
defer func() {
|
||||
_ = out.Close()
|
||||
if !ok {
|
||||
_ = os.Remove(tmp)
|
||||
}
|
||||
}()
|
||||
|
||||
dec := json.NewDecoder(bufio.NewReaderSize(in, 1<<20))
|
||||
dec.UseNumber()
|
||||
bw := bufio.NewWriterSize(out, 1<<20)
|
||||
compacted, err := rewriteCheckpointObject(dec, bw)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err := bw.Flush(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err := out.Sync(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err := out.Close(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if compacted > 0 {
|
||||
if err := os.Rename(tmp, path); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
} else {
|
||||
_ = os.Remove(tmp)
|
||||
}
|
||||
if err := os.WriteFile(marker, []byte(fmt.Sprintf("compacted=%d\n", compacted)), 0600); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
ok = true
|
||||
return compacted, nil
|
||||
}
|
||||
|
||||
func rewriteCheckpointObject(dec *json.Decoder, w *bufio.Writer) (int, error) {
|
||||
tok, err := dec.Token()
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if d, ok := tok.(json.Delim); !ok || d != '{' {
|
||||
return 0, errors.New("state checkpoint must be a JSON object")
|
||||
}
|
||||
if err := w.WriteByte('{'); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
first := true
|
||||
compacted := 0
|
||||
for dec.More() {
|
||||
kt, err := dec.Token()
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
key, ok := kt.(string)
|
||||
if !ok {
|
||||
return 0, errors.New("state checkpoint object key is not a string")
|
||||
}
|
||||
if !first {
|
||||
if err := w.WriteByte(','); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
first = false
|
||||
kb, _ := json.Marshal(key)
|
||||
if _, err := w.Write(kb); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err := w.WriteByte(':'); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if key == "jobs" {
|
||||
n, err := rewriteJobsObject(dec, w)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
compacted += n
|
||||
continue
|
||||
}
|
||||
if err := copyJSONValue(dec, w); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
if _, err := dec.Token(); err != nil { // closing }
|
||||
return 0, err
|
||||
}
|
||||
if err := w.WriteByte('}'); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if tok, err := dec.Token(); err != io.EOF {
|
||||
if err == nil {
|
||||
return 0, fmt.Errorf("unexpected trailing JSON token %v", tok)
|
||||
}
|
||||
return 0, err
|
||||
}
|
||||
return compacted, nil
|
||||
}
|
||||
|
||||
func rewriteJobsObject(dec *json.Decoder, w *bufio.Writer) (int, error) {
|
||||
tok, err := dec.Token()
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if tok == nil {
|
||||
_, err = w.WriteString("null")
|
||||
return 0, err
|
||||
}
|
||||
if d, ok := tok.(json.Delim); !ok || d != '{' {
|
||||
return 0, errors.New("jobs must be a JSON object")
|
||||
}
|
||||
if err := w.WriteByte('{'); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
first := true
|
||||
compacted := 0
|
||||
for dec.More() {
|
||||
kt, err := dec.Token()
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
key := kt.(string)
|
||||
var job core.Job
|
||||
if err := dec.Decode(&job); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if job.Type == "vector.relink" && job.Status == "done" {
|
||||
if len(job.Payload) > 0 || len(job.Result) > 0 {
|
||||
compacted++
|
||||
}
|
||||
job.Payload = nil
|
||||
job.Result = nil
|
||||
}
|
||||
if !first {
|
||||
if err := w.WriteByte(','); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
first = false
|
||||
kb, _ := json.Marshal(key)
|
||||
jb, err := json.Marshal(job)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if _, err := w.Write(kb); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err := w.WriteByte(':'); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if _, err := w.Write(jb); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
if _, err := dec.Token(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
return compacted, w.WriteByte('}')
|
||||
}
|
||||
|
||||
func copyJSONValue(dec *json.Decoder, w *bufio.Writer) error {
|
||||
tok, err := dec.Token()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if d, ok := tok.(json.Delim); ok {
|
||||
switch d {
|
||||
case '{':
|
||||
if err := w.WriteByte('{'); err != nil {
|
||||
return err
|
||||
}
|
||||
first := true
|
||||
for dec.More() {
|
||||
kt, err := dec.Token()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if !first {
|
||||
if err := w.WriteByte(','); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
first = false
|
||||
kb, _ := json.Marshal(kt.(string))
|
||||
if _, err := w.Write(kb); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := w.WriteByte(':'); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := copyJSONValue(dec, w); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
if _, err := dec.Token(); err != nil {
|
||||
return err
|
||||
}
|
||||
return w.WriteByte('}')
|
||||
case '[':
|
||||
if err := w.WriteByte('['); err != nil {
|
||||
return err
|
||||
}
|
||||
first := true
|
||||
for dec.More() {
|
||||
if !first {
|
||||
if err := w.WriteByte(','); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
first = false
|
||||
if err := copyJSONValue(dec, w); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
if _, err := dec.Token(); err != nil {
|
||||
return err
|
||||
}
|
||||
return w.WriteByte(']')
|
||||
default:
|
||||
return fmt.Errorf("unexpected JSON delimiter %q", d)
|
||||
}
|
||||
}
|
||||
b, err := json.Marshal(tok)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
_, err = w.Write(b)
|
||||
return err
|
||||
}
|
||||
@@ -0,0 +1,56 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"neuroforge/internal/core"
|
||||
)
|
||||
|
||||
func TestCompactLegacyTerminalRelinkCheckpoint(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
st := core.PersistedState{Config: core.DefaultConfig(), Jobs: map[string]*core.Job{}, Synapses: map[string]*core.Synapse{}, Goals: map[string]*core.Goal{}}
|
||||
st.Jobs["done"] = &core.Job{ID: "done", Type: "vector.relink", Status: "done", Payload: json.RawMessage(`{"target":[1,2,3]}`), Result: json.RawMessage(`{"ok":true}`)}
|
||||
st.Jobs["queued"] = &core.Job{ID: "queued", Type: "vector.relink", Status: "queued", Payload: json.RawMessage(`{"target":[4,5,6]}`)}
|
||||
st.Jobs["other"] = &core.Job{ID: "other", Type: "model.chat", Status: "done", Payload: json.RawMessage(`{"prompt":"keep"}`), Result: json.RawMessage(`{"text":"keep"}`)}
|
||||
b, err := json.Marshal(st)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "state.json"), b, 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
n, err := compactLegacyTerminalRelinkCheckpoint(dir)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if n != 1 {
|
||||
t.Fatalf("compacted=%d want 1", n)
|
||||
}
|
||||
|
||||
var got core.PersistedState
|
||||
bb, err := os.ReadFile(filepath.Join(dir, "state.json"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := json.Unmarshal(bb, &got); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
cleared := func(b json.RawMessage) bool { return len(b) == 0 || string(b) == "null" }
|
||||
if !cleared(got.Jobs["done"].Payload) || !cleared(got.Jobs["done"].Result) {
|
||||
t.Fatalf("done relink blobs not cleared: payload=%q result=%q", got.Jobs["done"].Payload, got.Jobs["done"].Result)
|
||||
}
|
||||
if len(got.Jobs["queued"].Payload) == 0 {
|
||||
t.Fatalf("queued relink payload must be retained")
|
||||
}
|
||||
if len(got.Jobs["other"].Payload) == 0 || len(got.Jobs["other"].Result) == 0 {
|
||||
t.Fatalf("unrelated job blobs must be retained")
|
||||
}
|
||||
|
||||
if n2, err := compactLegacyTerminalRelinkCheckpoint(dir); err != nil || n2 != 0 {
|
||||
t.Fatalf("second migration = %d, %v", n2, err)
|
||||
}
|
||||
}
|
||||
@@ -1,11 +1,13 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"crypto/rand"
|
||||
"encoding/hex"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"math"
|
||||
"net/url"
|
||||
"os"
|
||||
@@ -49,7 +51,19 @@ type Store struct {
|
||||
synapseAdj map[string]map[string]*core.Synapse
|
||||
}
|
||||
|
||||
type OpenProgressFunc func(phase string)
|
||||
|
||||
func New(dir string) (*Store, error) {
|
||||
return NewWithProgress(dir, nil)
|
||||
}
|
||||
|
||||
func NewWithProgress(dir string, report OpenProgressFunc) (*Store, error) {
|
||||
progress := func(phase string) {
|
||||
if report != nil {
|
||||
report(phase)
|
||||
}
|
||||
}
|
||||
progress("store.prepare")
|
||||
if dir == "" {
|
||||
dir = "./data"
|
||||
}
|
||||
@@ -58,8 +72,18 @@ func New(dir string) (*Store, error) {
|
||||
}
|
||||
s := &Store{dir: dir, indexes: map[int]*vector.HNSW{}, diskIndexes: map[int]*vector.PQIndex{}, provenanceSourceIDs: map[string]map[string]struct{}{}, workers: map[string]core.WorkerState{}, synapseAdj: map[string]map[string]*core.Synapse{}}
|
||||
s.state = core.PersistedState{Config: core.DefaultConfig(), Memories: map[string]*core.Memory{}, Synapses: map[string]*core.Synapse{}, Jobs: map[string]*core.Job{}, Goals: map[string]*core.Goal{}, Sources: map[string]*core.KnowledgeSource{}, ResearchRuns: map[string]*core.ResearchRun{}}
|
||||
_ = s.loadJSON(filepath.Join(dir, "state.json"), &s.state)
|
||||
_ = s.loadJSON(filepath.Join(dir, "secrets.json"), &s.secrets)
|
||||
progress("checkpoint.precompact")
|
||||
if _, err := compactLegacyTerminalRelinkCheckpoint(dir); err != nil {
|
||||
return nil, fmt.Errorf("compact legacy terminal relink jobs: %w", err)
|
||||
}
|
||||
progress("checkpoint.load")
|
||||
if err := s.loadJSON(filepath.Join(dir, "state.json"), &s.state); err != nil && !errors.Is(err, os.ErrNotExist) {
|
||||
return nil, fmt.Errorf("load authoritative state checkpoint: %w", err)
|
||||
}
|
||||
progress("secrets.load")
|
||||
if err := s.loadJSON(filepath.Join(dir, "secrets.json"), &s.secrets); err != nil && !errors.Is(err, os.ErrNotExist) {
|
||||
return nil, fmt.Errorf("load authoritative secrets: %w", err)
|
||||
}
|
||||
if s.state.Memories == nil {
|
||||
s.state.Memories = map[string]*core.Memory{}
|
||||
}
|
||||
@@ -84,6 +108,7 @@ func New(dir string) (*Store, error) {
|
||||
// v0.6 binary vector sidecar: rebuildable acceleration data used by the
|
||||
// disk ANN builder. Corruption must never prevent the authoritative memory
|
||||
// store from opening; move a bad cache aside and recreate it empty.
|
||||
progress("vector-journal.open")
|
||||
vjPath := filepath.Join(dir, "vector-journal.nfv")
|
||||
vj, vjErr := openVectorJournal(vjPath, vectorJournalOptionsFromConfig(s.state.Config))
|
||||
if vjErr != nil {
|
||||
@@ -95,6 +120,7 @@ func New(dir string) (*Store, error) {
|
||||
}
|
||||
|
||||
if s.state.Config.Storage.Segments.Enabled {
|
||||
progress("memory-segments.scan")
|
||||
seg, err := openSegmentStore(filepath.Join(dir, "memory-segments"), s.state.Config.Storage.Segments.MaxSegmentBytes, s.state.Config.Storage.Segments.MmapSealed)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("open memory segments: %w", err)
|
||||
@@ -110,9 +136,13 @@ func New(dir string) (*Store, error) {
|
||||
s.state.Memories = meta
|
||||
}
|
||||
}
|
||||
progress("wal.replay")
|
||||
if err := s.replayWAL(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
progress("jobs.compact")
|
||||
_ = s.compactCompletedRelinkJobsLocked()
|
||||
progress("graph.restore")
|
||||
s.rebuildSynapseAdjLocked()
|
||||
applyNewDefaults(&s.state.Config)
|
||||
if s.state.Cluster.Term < s.state.Config.Cluster.Term {
|
||||
@@ -198,19 +228,26 @@ func New(dir string) (*Store, error) {
|
||||
if s.secrets.ClusterToken == "" {
|
||||
s.secrets.ClusterToken = randomID(24)
|
||||
}
|
||||
progress("disk-ann.load")
|
||||
_ = s.loadDiskANNLocked()
|
||||
progress("hnsw.snapshot.load")
|
||||
if !s.loadIndexSnapshotLocked() {
|
||||
progress("hnsw.rebuild")
|
||||
s.rebuildIndexesLocked()
|
||||
}
|
||||
progress("secrets.persist")
|
||||
if err := s.persistSecretsLocked(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
progress("checkpoint.write")
|
||||
if err := s.checkpointLocked(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if s.state.Config.Storage.Tiering.Enabled && s.segments != nil {
|
||||
progress("tiering.initialize")
|
||||
s.tierMemoryBodiesLocked(time.Now().UTC())
|
||||
}
|
||||
progress("store.ready")
|
||||
return s, nil
|
||||
}
|
||||
|
||||
@@ -524,6 +561,9 @@ func applyNewDefaults(c *core.Config) {
|
||||
if c.Worker.MaxQueuedJobs == 0 {
|
||||
c.Worker.MaxQueuedJobs = d.Worker.MaxQueuedJobs
|
||||
}
|
||||
if c.Worker.MaxQueuedPayloadMB == 0 {
|
||||
c.Worker.MaxQueuedPayloadMB = d.Worker.MaxQueuedPayloadMB
|
||||
}
|
||||
if c.Worker.MasterApplyMaxAttempts == 0 {
|
||||
c.Worker.MasterApplyMaxAttempts = d.Worker.MasterApplyMaxAttempts
|
||||
}
|
||||
@@ -613,23 +653,80 @@ func inferMemoryType(kind string) string {
|
||||
}
|
||||
|
||||
func (s *Store) loadJSON(path string, v any) error {
|
||||
b, err := os.ReadFile(path)
|
||||
f, err := os.Open(path)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return json.Unmarshal(b, v)
|
||||
defer f.Close()
|
||||
|
||||
// Decode directly from the file instead of ReadFile+Unmarshal. Large
|
||||
// checkpoints can contain millions of graph edges/jobs; keeping a second
|
||||
// raw []byte copy of state.json during boot needlessly doubles peak memory.
|
||||
dec := json.NewDecoder(bufio.NewReaderSize(f, 1<<20))
|
||||
if err := dec.Decode(v); err != nil {
|
||||
return fmt.Errorf("decode %s: %w", filepath.Base(path), err)
|
||||
}
|
||||
// Reject a second JSON value/trailing non-whitespace. Silently accepting a
|
||||
// partially corrupt authoritative checkpoint can create split-brain state.
|
||||
if tok, err := dec.Token(); err != io.EOF {
|
||||
if err == nil {
|
||||
return fmt.Errorf("decode %s: unexpected trailing JSON token %v", filepath.Base(path), tok)
|
||||
}
|
||||
return fmt.Errorf("decode %s trailing data: %w", filepath.Base(path), err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func writeAtomic(path string, perm os.FileMode, v any) error {
|
||||
b, err := json.MarshalIndent(v, "", " ")
|
||||
func writeAtomic(path string, perm os.FileMode, v any) (retErr error) {
|
||||
// Stream JSON directly into the temporary file instead of MarshalIndent+
|
||||
// WriteFile. state.json and index manifests can be large; allocating the
|
||||
// complete encoded checkpoint as another in-memory byte slice is avoidable.
|
||||
tmp := path + ".tmp"
|
||||
_ = os.Remove(tmp)
|
||||
f, err := os.OpenFile(tmp, os.O_CREATE|os.O_EXCL|os.O_WRONLY, perm)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
tmp := path + ".tmp"
|
||||
if err := os.WriteFile(tmp, b, perm); err != nil {
|
||||
closed := false
|
||||
defer func() {
|
||||
if !closed {
|
||||
_ = f.Close()
|
||||
}
|
||||
if retErr != nil {
|
||||
_ = os.Remove(tmp)
|
||||
}
|
||||
}()
|
||||
|
||||
bw := bufio.NewWriterSize(f, 1<<20)
|
||||
enc := json.NewEncoder(bw)
|
||||
enc.SetIndent("", " ")
|
||||
if err := enc.Encode(v); err != nil {
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmp, path)
|
||||
if err := bw.Flush(); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := f.Sync(); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := f.Close(); err != nil {
|
||||
return err
|
||||
}
|
||||
closed = true
|
||||
if err := os.Rename(tmp, path); err != nil {
|
||||
return err
|
||||
}
|
||||
// Persist the directory entry as well. This matters for secrets/state after
|
||||
// host power loss and is cheap compared with the checkpoint itself.
|
||||
d, err := os.Open(filepath.Dir(path))
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := d.Sync(); err != nil {
|
||||
_ = d.Close()
|
||||
return err
|
||||
}
|
||||
return d.Close()
|
||||
}
|
||||
|
||||
func (s *Store) persistLocked() error {
|
||||
@@ -1609,7 +1706,7 @@ func (s *Store) validateConfigLocked(c core.Config) error {
|
||||
if c.Worker.LeaseSeconds < 10 || c.Worker.LeaseSeconds > 3600 || c.Worker.HeartbeatSeconds < 2 || c.Worker.HeartbeatSeconds >= c.Worker.LeaseSeconds || c.Worker.StaleAfterSeconds < c.Worker.HeartbeatSeconds || c.Worker.StaleAfterSeconds > 7200 {
|
||||
return errors.New("invalid worker lease/heartbeat/stale timing")
|
||||
}
|
||||
if c.Worker.DefaultMaxAttempts < 1 || c.Worker.DefaultMaxAttempts > 20 || c.Worker.RetryBackoffSeconds < 1 || c.Worker.RetryBackoffSeconds > 3600 || c.Worker.MaxQueuedJobs < 16 || c.Worker.MaxQueuedJobs > 1000000 || c.Worker.MasterApplyMaxAttempts < 1 || c.Worker.MasterApplyMaxAttempts > 20 || c.Worker.MasterApplyBackoffSeconds < 1 || c.Worker.MasterApplyBackoffSeconds > 3600 || c.Worker.JobRetentionHours < 1 || c.Worker.JobRetentionHours > 8760 || c.Worker.MaxTerminalJobs < 100 || c.Worker.MaxTerminalJobs > 1000000 {
|
||||
if c.Worker.DefaultMaxAttempts < 1 || c.Worker.DefaultMaxAttempts > 20 || c.Worker.RetryBackoffSeconds < 1 || c.Worker.RetryBackoffSeconds > 3600 || c.Worker.MaxQueuedJobs < 16 || c.Worker.MaxQueuedJobs > 1000000 || c.Worker.MaxQueuedPayloadMB < 16 || c.Worker.MaxQueuedPayloadMB > 65536 || c.Worker.MasterApplyMaxAttempts < 1 || c.Worker.MasterApplyMaxAttempts > 20 || c.Worker.MasterApplyBackoffSeconds < 1 || c.Worker.MasterApplyBackoffSeconds > 3600 || c.Worker.JobRetentionHours < 1 || c.Worker.JobRetentionHours > 8760 || c.Worker.MaxTerminalJobs < 100 || c.Worker.MaxTerminalJobs > 1000000 {
|
||||
return errors.New("invalid worker retry/queue/retention configuration")
|
||||
}
|
||||
if c.Worker.GraphBackfillIntervalS < 2 || c.Worker.GraphBackfillIntervalS > 3600 || c.Worker.GraphBackfillBatchSize < 1 || c.Worker.GraphBackfillBatchSize > 4096 || c.Worker.GraphBackfillMaxQueued < 1 || c.Worker.GraphBackfillMaxQueued > c.Worker.MaxQueuedJobs || c.Worker.GraphBackfillMinDegree < 1 || c.Worker.GraphBackfillMinDegree > 100 || c.Worker.GraphCandidateMultiplier < 2 || c.Worker.GraphCandidateMultiplier > 64 || c.Worker.GraphRetryAfterMinutes < 1 || c.Worker.GraphRetryAfterMinutes > 43200 {
|
||||
|
||||
@@ -232,6 +232,14 @@ func (s *Store) applyWALEvent(ev walEvent) error {
|
||||
if err := json.Unmarshal(ev.Data, &x); err != nil {
|
||||
return err
|
||||
}
|
||||
// v1.6.0 could leave a large WAL containing completed vector.relink
|
||||
// jobs with full target/candidate vectors. Compact each terminal event as
|
||||
// it is replayed so recovery memory remains bounded by one WAL record
|
||||
// instead of accumulating every historical vector payload in state.
|
||||
if x.Type == "vector.relink" && x.Status == "done" {
|
||||
x.Payload = nil
|
||||
x.Result = nil
|
||||
}
|
||||
s.state.Jobs[x.ID] = &x
|
||||
case "job.delete":
|
||||
var ids []string
|
||||
@@ -332,14 +340,20 @@ func (s *Store) checkpointLocked() error {
|
||||
// Keep state.json O(non-memory-state) instead of O(memory-count).
|
||||
checkpoint.Memories = nil
|
||||
}
|
||||
if err := writeAtomic(filepath.Join(s.dir, "state.json"), 0600, &checkpoint); err != nil {
|
||||
return err
|
||||
}
|
||||
// Persist acceleration state before the authoritative non-memory checkpoint.
|
||||
// If the process dies after the index snapshot but before state.json, the WAL
|
||||
// remains intact; boot replays it to the same revision and can immediately use
|
||||
// the already-written index. The previous order could advance state.json first,
|
||||
// then die during a large HNSW snapshot and force a full synchronous rebuild on
|
||||
// every restart.
|
||||
if s.state.Config.Storage.IndexSnapshot && s.state.Config.Brain.Index.Enabled {
|
||||
if err := s.writeIndexSnapshotLocked(); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
if err := writeAtomic(filepath.Join(s.dir, "state.json"), 0600, &checkpoint); err != nil {
|
||||
return err
|
||||
}
|
||||
// The checkpoint and (when enabled) memory segments now cover every WAL
|
||||
// event through state.Revision. Prune those already-checkpointed log files
|
||||
// only after all checkpoint artifacts succeeded, otherwise long bulk
|
||||
|
||||
@@ -679,6 +679,50 @@ type HNSWShadow struct {
|
||||
Nodes map[string][32]byte
|
||||
}
|
||||
|
||||
func (h *HNSW) nodeFingerprintLocked(n *hnswNode) [32]byte {
|
||||
hash := sha256.New()
|
||||
writeHashString(hash, n.ID)
|
||||
writeHashU32(hash, uint32(n.Level))
|
||||
writeHashU32(hash, uint32(len(n.Vector)))
|
||||
var b [4]byte
|
||||
for _, x := range n.Vector {
|
||||
binary.LittleEndian.PutUint32(b[:], math.Float32bits(x))
|
||||
_, _ = hash.Write(b[:])
|
||||
}
|
||||
for level := 0; level <= n.Level; level++ {
|
||||
writeHashU32(hash, uint32(level))
|
||||
edges := n.Neighbors[level]
|
||||
writeHashU32(hash, uint32(len(edges)))
|
||||
for _, edge := range edges {
|
||||
idx := int(edge.idx)
|
||||
if idx >= 0 && idx < len(h.nodes) {
|
||||
writeHashString(hash, h.nodes[idx].ID)
|
||||
}
|
||||
}
|
||||
}
|
||||
var sum [32]byte
|
||||
copy(sum[:], hash.Sum(nil))
|
||||
return sum
|
||||
}
|
||||
|
||||
func (h *HNSW) snapshotNodeLocked(n *hnswNode) HNSWSnapshotNode {
|
||||
cn := HNSWSnapshotNode{ID: n.ID, Vector: append([]float32(nil), n.Vector...), Level: n.Level, Neighbors: map[int][]string{}}
|
||||
for level, ids := range n.Neighbors {
|
||||
if len(ids) == 0 {
|
||||
continue
|
||||
}
|
||||
refs := make([]string, 0, len(ids))
|
||||
for _, edge := range ids {
|
||||
idx := int(edge.idx)
|
||||
if idx >= 0 && idx < len(h.nodes) {
|
||||
refs = append(refs, h.nodes[idx].ID)
|
||||
}
|
||||
}
|
||||
cn.Neighbors[level] = refs
|
||||
}
|
||||
return cn
|
||||
}
|
||||
|
||||
func (h *HNSW) Shadow() HNSWShadow {
|
||||
h.mu.RLock()
|
||||
defer h.mu.RUnlock()
|
||||
@@ -688,33 +732,41 @@ func (h *HNSW) Shadow() HNSWShadow {
|
||||
}
|
||||
out := HNSWShadow{Config: h.cfg, EntryID: entryID, MaxLevel: h.maxLevel, Nodes: make(map[string][32]byte, len(h.nodes))}
|
||||
for _, n := range h.nodes {
|
||||
hash := sha256.New()
|
||||
writeHashString(hash, n.ID)
|
||||
writeHashU32(hash, uint32(n.Level))
|
||||
writeHashU32(hash, uint32(len(n.Vector)))
|
||||
var b [4]byte
|
||||
for _, x := range n.Vector {
|
||||
binary.LittleEndian.PutUint32(b[:], math.Float32bits(x))
|
||||
_, _ = hash.Write(b[:])
|
||||
}
|
||||
for level := 0; level <= n.Level; level++ {
|
||||
writeHashU32(hash, uint32(level))
|
||||
edges := n.Neighbors[level]
|
||||
writeHashU32(hash, uint32(len(edges)))
|
||||
for _, edge := range edges {
|
||||
idx := int(edge.idx)
|
||||
if idx >= 0 && idx < len(h.nodes) {
|
||||
writeHashString(hash, h.nodes[idx].ID)
|
||||
}
|
||||
}
|
||||
}
|
||||
var sum [32]byte
|
||||
copy(sum[:], hash.Sum(nil))
|
||||
out.Nodes[n.ID] = sum
|
||||
out.Nodes[n.ID] = h.nodeFingerprintLocked(n)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// Delta returns only nodes whose vector/graph representation changed since prev,
|
||||
// plus a compact hash shadow for the current graph. Unlike Snapshot(), this does
|
||||
// not deep-copy every vector on each checkpoint, which keeps bulk-ingest memory
|
||||
// bounded as the HNSW grows.
|
||||
func (h *HNSW) Delta(prev HNSWShadow) (HNSWShadow, []HNSWSnapshotNode, []string) {
|
||||
h.mu.RLock()
|
||||
defer h.mu.RUnlock()
|
||||
entryID := ""
|
||||
if h.entry >= 0 && h.entry < len(h.nodes) {
|
||||
entryID = h.nodes[h.entry].ID
|
||||
}
|
||||
cur := HNSWShadow{Config: h.cfg, EntryID: entryID, MaxLevel: h.maxLevel, Nodes: make(map[string][32]byte, len(h.nodes))}
|
||||
upserts := make([]HNSWSnapshotNode, 0)
|
||||
for _, n := range h.nodes {
|
||||
fp := h.nodeFingerprintLocked(n)
|
||||
cur.Nodes[n.ID] = fp
|
||||
if old, ok := prev.Nodes[n.ID]; !ok || old != fp {
|
||||
upserts = append(upserts, h.snapshotNodeLocked(n))
|
||||
}
|
||||
}
|
||||
deletes := make([]string, 0)
|
||||
for id := range prev.Nodes {
|
||||
if _, ok := cur.Nodes[id]; !ok {
|
||||
deletes = append(deletes, id)
|
||||
}
|
||||
}
|
||||
sort.Strings(deletes)
|
||||
return cur, upserts, deletes
|
||||
}
|
||||
|
||||
func writeHashString(w io.Writer, s string) {
|
||||
writeHashU32(w, uint32(len(s)))
|
||||
_, _ = io.WriteString(w, s)
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
package vector
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestHNSWDeltaDoesNotSnapshotUnchangedVectors(t *testing.T) {
|
||||
h := NewHNSW(HNSWConfig{M: 4, EfConstruction: 16, EfSearch: 8})
|
||||
h.Add("a", []float32{1, 0, 0})
|
||||
h.Add("b", []float32{0.9, 0.1, 0})
|
||||
base := h.Shadow()
|
||||
cur, upserts, deletes := h.Delta(base)
|
||||
if len(upserts) != 0 || len(deletes) != 0 {
|
||||
t.Fatalf("unchanged delta upserts=%d deletes=%d", len(upserts), len(deletes))
|
||||
}
|
||||
if len(cur.Nodes) != 2 {
|
||||
t.Fatalf("shadow nodes=%d", len(cur.Nodes))
|
||||
}
|
||||
|
||||
h.Add("c", []float32{0.8, 0.2, 0})
|
||||
cur2, upserts, deletes := h.Delta(base)
|
||||
if len(upserts) == 0 {
|
||||
t.Fatalf("expected changed/new nodes")
|
||||
}
|
||||
if len(deletes) != 0 {
|
||||
t.Fatalf("unexpected deletes: %v", deletes)
|
||||
}
|
||||
if len(cur2.Nodes) != 3 {
|
||||
t.Fatalf("shadow nodes=%d", len(cur2.Nodes))
|
||||
}
|
||||
foundC := false
|
||||
for _, n := range upserts {
|
||||
if n.ID == "c" {
|
||||
foundC = true
|
||||
}
|
||||
}
|
||||
if !foundC {
|
||||
t.Fatalf("new node c missing from delta")
|
||||
}
|
||||
}
|
||||
@@ -1,7 +1,7 @@
|
||||
openapi: 3.1.0
|
||||
info:
|
||||
title: NeuroForge API
|
||||
version: 0.8.2
|
||||
version: 0.8.3
|
||||
description: REST API for NeuroForge associative learning, explainable vector recall,
|
||||
provenance, document/text ingestion, SearXNG-backed autonomous research, source-grounded
|
||||
goal cycles, responsive knowledge graph, model routing, cost controls, Prometheus
|
||||
|
||||
Reference in New Issue
Block a user