Every engineering experiment with its hypothesis, command and verdict (53 total), documented in the open. The interactive lab is at the lab page.
-
Memory Contamination Live-Arm: live LLM false-acceptance across 14 models
2026-08-14 · mscodebase-intelligence · verdict: confirmed.
The fail-closed look was anchor bias, not model paranoia: with token-string evidence qwen3.6/3.7 accepted 2-5/25 true claims (recall 0.08-0.20); with a real file fragment (V4 arm) recall jumps to 0.88 at FA 0.02-0.04. CORRECTED: the earlier 'every remaining false-accept is prese…
-
Memory Contamination proxy-control: retraction -88%, verify-on-read adoption to 0.0
2026-08-11 · mscodebase-intelligence · verdict: confirmed.
The proxy always decides (unknown=0) and false_accept=0 by construction - the headline numbers are a property of the heuristic, not of LLM behavior; the live arm (exp-1) measures the real gap.
-
Manifest anchoring (pkg: anchors, ADR-0005): closed-world manifest kills 7 false REFUTED
2026-08-14 · mscodebase-intelligence · verdict: confirmed.
A closed-world manifest is a stronger evidence source than bare tokens; the fix was verified by 1-V-REP (0 false REFUTED of TRUE facts).
-
Context aggregator vs multi-tool: 1 call instead of 4-5, -78% calls at recall parity
2026-08-08 · mscodebase-intelligence · verdict: confirmed.
The B-scheme (intent filter) is the optimum: same recall as multi-tool at 1 request and 275 tokens; the earlier recall gap was noise, not signal.
-
Multi-RAG ablation: FTS5-only beats the full pipeline by recall; H1 refuted
2026-08-11 · mscodebase-intelligence · verdict: refuted.
Recall comes from keyword tiers (FTS5/BM25), precision is bought by the reranker, and the vector tier is the weakest on symbol tasks. Production bug found: hybrid_search_async cache-hit silently skips the dense tier - cache isolation per arm is mandatory in ablations (first run…
-
Shadow-canary fail-open: 5/5 attacks passed before the fix, 0 after
2026-08-11 · mscodebase-intelligence · verdict: confirmed.
Relative metrics are attackable: absence of signal (empty canary) must fail closed, and a quality floor must be absolute, not relative.
-
Concurrency vs semantic correctness: '0 errors' != correct data under race
2026-08-11 · mscodebase-intelligence · verdict: confirmed.
VC and VOR are complementary layers: consistency (VC) catches lost writes, semantics (VOR) catches lies. A verdict that flips when one scenario variable changes (A6) is an illustration, not a law.
-
Mutation testing for the reranker grader: mutation score 8% -> 100%
2026-08-14 · mscodebase-intelligence · verdict: confirmed.
Validate values, not just types - NaN/Infinity otherwise silently earn the maximum score.
-
Root-cause prediction audit: top-1 accuracy 0.13 on 31 incidents
2026-07-22 · mscodebase-intelligence · verdict: partial.
Root-cause prediction is far from solved: the audit dataset and gold standard are the baseline, and the engine's current top-1 accuracy is 13%.
-
PID-lock self-healing: orphan lock wait 30s -> 120ms
2026-08-08 · mscodebase-intelligence · verdict: confirmed.
Never block boot on a lock whose holder may be a zombie: classify the holder and self-heal instead of failing closed for 30 seconds.
-
Late enrichment: imports metadata covers 0.0% of search chunks
2026-08-08 · mscodebase-intelligence · verdict: partial.
Index-time enrichment produced nothing for search chunks: imports coverage is 0.0, so the late path is kept behind a flag until the hypothesis is re-probed.
-
Benchmark 2.0 scaffold: 12 repository-reasoning tasks at levels L3-L5
2026-08-08 · mscodebase-intelligence · verdict: partial.
A reasoning-quality benchmark needs L3-L5 tasks with evidence and checks; short keyword queries are a separate failure mode (CoREB) that must be measured independently.
-
Server unavailable during reindex: sync update_all blocks the event loop (P0 root cause)
2026-08-13 · mscodebase-intelligence · verdict: confirmed.
552s indexing + ~220s update_all = 771s unavailability, in line with the log. Indexing-in-executor (H1) refuted; fast-fail search while is_reindexing exists; the blocker was the sync update_all (same class as BS-11: run_full_diagnostic had already been moved to asyncio.to_thread…
-
Vacuous-test scan: 1133/1143 tests are provable (hypothesis refuted in the good way)
2026-08-11 · mscodebase-intelligence · verdict: refuted.
The suite is almost fully provable - 1133/1143 (99.7%) contain a failing construct; the 3 'vacuous' are smoke tests that can still fail via exception propagation from helpers (discrimination weaker, not zero). MSCodeBase is NOT in the fintech state (7/40 proven) - which is preci…
-
ln.strip() bug class replayed: broken assert-extractor gives 3/8 false passes
2026-08-11 · mscodebase-intelligence · verdict: confirmed.
Bug class replayed: valid Python that never executes, exit 0, 'verified:true' for a wrong answer. The form count differs (3/8 vs fintech's 5/8) but the essence is identical - signature/exit-code verify the PROCESS, not the SEMANTICS ('ask for the output, not the exit code').
-
Population blind spot: '0 eligible' is indistinguishable from '0 collected'
2026-08-11 · mscodebase-intelligence · verdict: confirmed.
(a) and (b) share the same failure signal (passed=0, same warning class) - and (a)'s message claims 'empty/garbage chunks' when raw results were 0, a false explanation. '0 rows with 0 eligible' (healthy idle) is indistinguishable from '0 rows with N eligible' (broken collector).…
-
verify_clean_state.sh falsifiability: the drift-gate is structurally unable to fail
2026-08-11 · mscodebase-intelligence · verdict: partial.
(B) confirmed: the gate prints PASSED for a suite with zero asserts - semantic blindness. (A) refuted in reverse: the drift-gate is STRUCTURALLY unable to fire - its grep -iE '^pkg==' requires the pin at line start, but pins live in a TOML array (' "lancedb==0.34.0",') so PIN…
-
Evidence Ladder E1-E3: evidence form (anchor -> file fragment -> graph) vs verification quality
2026-08-15 · mscodebase-intelligence · verdict: partial.
Evidence form is a per-model knob, not a global answer: the file fragment is the strongest recall driver for all models; the graph closes the trap failure mode only for evidence-honest models (and on corrected labels the trap gap itself was a label issue). Fail-open models (glm…
-
Evidence Ladder E3b+E4: file+graph hybrid is NOT additive; git-provenance distinguishes existed-then vs exists-now
2026-08-15 · mscodebase-intelligence · verdict: refuted.
Hybridity is NOT additive: with a file fragment present, token presence dominates and the graph adds nothing (acc 0.900 < file 0.940) - VOR design must pick ONE evidence format: fragment for recall OR graph for trap-precision. Git-provenance is a cheap, powerful temporal signal…
-
ONNX embedder down: off-by-one project paths (not the model or ports)
2026-08-03 · mscodebase-intelligence · verdict: confirmed.
Root cause is paths, not model/ports: (1) onnx_client looked for '...src/src/core/embedder/onnx_server.py' (duplicated src), (2) onnx_server looked for the model in '.../src/.codebase_models/...' (not the root). The off-by-one was copied between files (onnx_client <- onnx_server…
-
Evidence Ladder E5: extended trap category (N=20, subject-validated labels) - present-trap is mass, graph really helps deepseek (75%->40%)
2026-08-16 · mscodebase-intelligence · verdict: confirmed.
Present-trap is NOT a mislabel issue: on honest subject-validated labels file_content FA is 2-15/20 (10-75%) - the earlier 'residual trap hole' was distorted by N=1. Graph evidence REALLY closes present-trap for deepseek (75%->40% FA); for glm it does not (14/20); qwen already h…
-
Pinned-rerun on the corrected dataset (fp e6ce7b90): canonical series numbers, routing band removed
2026-08-16 · mscodebase-intelligence · verdict: confirmed.
Routing band removed but per-model conclusions unchanged: file_content is the best recall arm for qwen (0.88), graph for glm (0.84, FA trap 1), hybrid for deepseek (0.92); temporal present-trap is universal (NOW 12/12, 9/12, 12/12), past solved by formulation (40/40). glm stays…
-
LSP live probe: call hierarchy + semantic tokens confirmed on pyright; type hierarchy / moniker NOT advertised
2026-08-19 · mscodebase-intelligence · verdict: partial.
Spec-present != server-implemented. On the pyright family the graph really gets call hierarchy (compiler-accurate cross-file CALLS edges) and semantic tokens (exact spans/kinds); type hierarchy (3.17) and moniker (3.16) are NOT advertised by pyright -> dormant until a server sup…
-
M1: real tool telemetry — 5/62 MCP tools ever used; the working-set boundary is single-digit
2026-08-25 · mscodebase-intelligence · verdict: partial.
Only 5 of 62 tools ever recorded a metric; dead-tool metrics are NOT measurable with current telemetry — a per-call counter in the error_handler path is required. The LSP toolkit (8 calls, 3 errors) is the top latency risk.
-
M2: sub-agent in a fresh project — natural 0/5 MCP calls vs MCP-first 9 (4 wasted on unindexed files)
2026-08-25 · mscodebase-intelligence · verdict: confirmed.
Natural agent: 0/5 MCP; MCP-first: 9 calls, ~4 idle (2 search + get_symbol_info + find_path on unindexed lab), only read_live_file (disk) productive. New files are invisible until reindex -> semantic layer on fresh code = zero. Correctness control: A found the root with exact fi…
-
M3: latency matrix — search fast 75ms HIT vs quality 3666ms MISS; get_symbol_info misses a real symbol
2026-08-25 · mscodebase-intelligence · verdict: refuted.
quality was 49x slower than fast (3666 vs 75ms) and semantically worse; get_symbol_info misses exact names carrying a trailing hint. On short queries fast search + grep + read_live_file beat quality search + symbol info on time AND accuracy.
-
E2: category pilot on the live index — fast 5/6 HIT (83%) vs quality 2/3 + leak to docs/JSON; index self-pollution discovered
2026-08-26 · mscodebase-intelligence · verdict: partial.
fast is 83% and an order of magnitude cheaper; quality saved the only fast-MISS (Q4) at ~35x cost. The leak is confirmed (Q1 quality -> JSON/CHANGELOG; Q4/Q6 quality top-1 incident datasets) but is NOT junk — semantically relevant docs; a category-priority filter is needed, not…
-
E3: category router on tasks_v3.json (30 tasks) — cascade 0.233 > fast 0.167 > quality 0.133; 4 graph classes 0.00 across all arms
2026-08-26 · mscodebase-intelligence · verdict: confirmed.
Cascade is the winner; the winner depends on klass (git_history/bug/prepare -> fast with quality 0.00 at bug; caller_callee/architecture -> quality); find_test/find_impact/modify/verify_change = 0.00 across ALL arms — search does not replace graph/AST/impact stages.
-
E4: deterministic per-class keyword router (PoC) — 0.200 < cascade 0.233, klass_acc 0.40 (NEGATIVE)
2026-08-26 · mscodebase-intelligence · verdict: refuted.
klass_acc=0.40 — keyword rules are noisy (bug tasks contain 'callers', git tasks 'why'); the union arm does not save the 4 graph classes — search cannot find what is not textually in the index (callers/callees/impact must come from the graph). The search-only ceiling on this dat…
-
E4.1: the graph stage breaks the search-only ceiling — recall 0.433 (cascade 0.267), med 177ms (Track 1)
2026-08-26 · mscodebase-intelligence · verdict: confirmed.
The graph stage added +0.166 recall on three of the four failing classes at ~15ms latency and lost nowhere (fallback to cascade by construction). Lessons: (1) has_symbol is an exact node-name — search_symbols(LIKE) + strict suffix is required; (2) graph navigation must go throug…
-
E4.2: deterministic concept resolver (no LLM) — verify_change T9/T29 HIT on real graph.db, facts 4/4 & 3/4 (Track 2)
2026-08-26 · mscodebase-intelligence · verdict: confirmed.
Both verify_change misses resolve to the correct file on the real graph.db (10748 nodes); graph rows now carry facts (graph_fact_text), not 0; regression is excluded by construction (klass-gating). A 'wordless' prompt is an anchor-resolution problem, not classification — lexical…
-
system_alerts chain (file changed -> STALE -> VOR -> one-shot alert) + idle background stale-sweep
2026-09-10 · mscodebase-intelligence · verdict: confirmed.
A one-shot alert store with atomic collect_and_clear is the right granularity: dedup by kind+payload prevents a spam-per-save alert, the atomic clear guarantees that a race between two MCP tools delivers the alert once and only once, and limit=5 caps token budget. Idle VOR is ch…
-
VOR Catch-up Rate (H1): throughput per budget and cycles-to-finish
2026-09-10 · mscodebase-intelligence · verdict: confirmed.
Real project (~247 nodes) is far from budget limits: read-path covers ~420-490 nodes/pass, background ~1600-1900 nodes/pass. N<=2000 fits one 250ms pass; N=5000 needs 2 passes. Systematic starvation (MATCHED>0/DELIVERED=0) appears only around N=5000.
-
VOR HEAD-polling catches external git changes without notify_change (H3)
2026-09-10 · mscodebase-intelligence · verdict: confirmed.
Head-invalidation per node key is an honest detector of external code drift in a git repo: the node whose file anchor was removed went REFUTED, the node whose file changed but still exists stayed VERIFIED. First exp3 run gave a false REFUTED because statuses were read from pre-r…
-
Burst-rename vs fail-closed VOR: prose-path anchoring causes false REFUTED (1-B/1-C/RT)
2026-09-11 · mscodebase-intelligence · verdict: confirmed.
Real queue is NOT huge (24 accumulated auto-refutes/month) but is DOMINATED by junk anchors from prose scanning (13/24 import:for/file/path regex hits), not by renames - renames caused only 1 FALSE_REFUTE (ADR-7232a6e2ba34: live node revoked by historical path from prose body).…
-
H3 TTL-rotation for verify-on-read: last_checked for every checked node + stale_ttl label (closes the "hangs forever" class)
2026-09-11 · mscodebase-intelligence · verdict: confirmed.
A TTL needs a trace for nodes WITHOUT verdicts, not only for VERIFIED ones: INCONCLUSIVE nodes now get last_checked and it refreshes on every idle pass (H1), so "no dates at all" is replaced by "checked every 6h, labelled stale after 30d". stale_ttl is a computational label - no…
-
H4 Agent-memory lifecycle at scale: capture latency on a dev.to graph grown 3.4x (threads 5.5x) + stale precision
2026-09-13 · mscodebase-intelligence · verdict: confirmed.
Capture latency grew with graph size (earlier own refresh on ~9k threads completed in ~2 min vs 10m38s now) even though the numeric delta between snapshots was small (threads 19,056 -> 19,101, +45). The cost is the network capture phase (134 dev.to API calls), not the local grap…
-
Bootstrap pipeline: test->function linking ? static name/import vs dynamic sys.settrace trace (89.8% vs 0%/77.9%)
2026-09-15 · mscodebase-intelligence · verdict: confirmed.
Dynamic trace (sys.settrace around pytest_runtest_call) is the only viable deterministic test->function linker: 89.8% vs 0% (name) and 77.9% (import, file-level only). Overhead +13.6% on a full run is fine for a one-off bootstrap pass. Entry points (@mcp_app.tool, 22 in src) are…
-
Bootstrap follow-up: Tarantula ranking for test->function annotation — recall refuted (22.6% rank<=3), precision confirmed; dev.to cross-check
2026-09-15 · mscodebase-intelligence · verdict: refuted.
Vanilla Tarantula cannot select the target function for most tests: recall is 22.6% rank<=3, far below the 60-70% recall target — shared utils (autouse fixtures safe_mkdir/get_data_root with 200+ callers) is the barrier, the same noise source TRUE Coverage reports (shared utilit…
-
Bootstrap A1: coverage.py sysmon driver overhead vs our sys.settrace plugin — sysmon refuted (+19.96% vs +13.6%)
2026-09-16 · mscodebase-intelligence · verdict: refuted.
sysmon is ~1.5x slower than our hand-rolled sys.settrace plugin on the full suite (+19.96% vs +13.6% in the same session, same methodology), so coverage.py is NOT adopted as the production driver. The earlier research estimate of 'sysmon ~3-7% median loss' (KNOWN_ISSUES:249) is…
-
Bootstrap Step B: Static Score Engine (AST calls/lexical/imports) vs dynamic trace ground truth
2026-09-16 · mscodebase-intelligence · verdict: refuted.
Hit >=50% confirmed (union 90.4%), but recall <=30% REFUTED: union recall is 70.0%. Static is much stronger than Exp 7 implied, because Exp 7 measured the weak name-signal (L2 reconfirms it: 17.7%). The call-level signal (L1) is a precise, narrow anchor: precision 68.0% with avg…
-
E7: lazy stat-sweep (mtime+size) vs sha256-sweep for always-fresh index (FreshnessChecker hot-reload)
2026-09-18 · mscodebase-intelligence · verdict: confirmed.
stat-sweep is 12x faster than sha256 over the same corpus and ~2% of a full reindex. (lancedb 0.34 detail: table.to_pandas(columns=[...]) crashes, table.to_lance().to_pandas(columns=[...]) is the correct API.) This made the synchronous pre-search freshness check viable: stat-fir…
-
E10: full-text chunk embedding + e5-prefix + reranker pool 50 vs pure-vector plateau
2026-09-19 · mscodebase-intelligence · verdict: refuted.
No confirmed shift within N=10 noise: quality hit@1 30%->20% (worse/noise), hit@5 30%->40% (noise). Toggles requiring a full prod reindex (~13 min) deliver zero -> pure-vector plateau reconfirmed (cf. Exp-29 search-only ceiling ~0.23). Next move: AST/Graph-hybrid re-ranking (gra…
-
E11: AST/Graph-hybrid re-ranking - symbol-lookup lift above the pure-vector plateau
2026-09-19 · mscodebase-intelligence · verdict: confirmed.
Graph lift saved 2/10 target cases (project_indexer_registry.py, indexing_tools.py) that quality-vs-baseline lost; both targets were present in search_symbols output. Symbol-lookup cost is negligible (+6ms). Signal is conditional (loose graph files add noise - MRR stays small, h…
-
E14: EmbeddingGemma 300M vs multilingual-e5-small - Hit@1 0.062->0.688 at ~4x CPU cost
2026-09-22 · mscodebase-intelligence · verdict: confirmed.
Real, reproducible quality gap on a code corpus: gemma Q8 gives +62.6pp Hit@1 and +56.4pp MRR over prod e5 at 2x RAM (176 vs 91MB) and 4.2x slower. QAT-Q4 (ggml-org) is NOT better than plain Q4_0 (unsloth): 0.653 vs 0.695 MRR - marketing claim not confirmed. Batch size does not…
-
E17: TESTS-signal in graph-stage (A/B) — covering tests added over function defs, hit@1/MRR unchanged (7/7)
2026-09-22 · mscodebase-intelligence · verdict: confirmed.
Covering tests are appended as a separate result sort (graph_score 0.4 vs def 1.0, sentinel chunk_index -20M+line) AFTER completed function defs, so def-first invariant holds (MRR 1.0 in both arms). The flag is off by default (MSCODEBASE_TESTS_SIGNAL env, like late_enrichment),…
-
E12: real-path embed throughput — chunk length sets the ~3.5k tok/s ceiling, not batch, ubatch or parallelism
2026-09-20 · mscodebase-intelligence · verdict: confirmed.
The CPU embed ceiling (e5-small Q8, 10 threads, Ryzen 5600H) is ~3.4-3.7k tok/s — model physics, independent of batch size, token budget, parallelism or truncation. The old "156 ch/s" was synthetic (~1560 tok/s on 10-token texts). 335k chunks x 203 tok = 68M tokens -> ~5.7h of p…
-
E13: text RAG (doc-chunks) vs code baseline — doc-chunks miss top-5 for 14 of 16 queries
2026-09-20 · mscodebase-intelligence · verdict: refuted.
Text RAG is far below code RAG: doc-chunks stay outside top-5 for 14 of 16 queries. The embedder packs code chunks (signatures, names) denser, doc-chunks are diffuse; the index is code-biased and queries without intent_hint='docs' route down the code path.
-
E15: bge-small-en / MiniLM / nomic in equal prod conditions — bge-small-en is quality without the 3x slowdown
2026-09-23 · mscodebase-intelligence · verdict: confirmed.
bge-small-en-v1.5 Q8 is the only "quality without losing speed" candidate: MRR 0.545 (+136% vs e5) at 5024 tok/s (1.19x FASTER than e5) and dim=384 — same scheme, no reindex, RAM halved to 46MB. MiniLM is the fastest (2.6x) but MRR 0.487. nomic does not qualify: its 8192 ctx nee…
-
E16: bootstrap-trace portability — dynamic trace transfers to foreign Python repos (97.3% / 100%)
2026-09-22 · mscodebase-intelligence · verdict: confirmed.
The dynamic trace transfers to foreign Python repos unmodified: 97.3% on gemma_agent with 2.9k tests, 100% on commit-. Overhead scales with test activity (+17.4% on gemma_agent vs +13.6% on our larger corpus), not with corpus size. Non-Python: the pytest pipeline collects 0 test…
-
F5 4-arm unit-of-return (judged reader): whole document beats top-k chunks 8x on code, direction-only on prose
2026-09-26 · mscodebase-intelligence · verdict: partial.
The unit of return acts on the reader, not on gold-file retrieval: objective retrieval was tied (A=B=5/16 top-1) yet the reader answers 8x more code questions correctly from a whole document than from chunks — confirmed for the code unit effect. For prose the effect is direction…
-
NodeRAG deterministic duel: chunked TF-IDF 8/10 vs PropertyGraph BFS 7/10 — graph does not win
2026-09-27 · mscodebase-intelligence · verdict: refuted.
No evidence graph traversal wins on this corpus: -1 hit at -43.6% tokens. The graph arm is entry-point fragile — a query whose symbols are not AST functions/classes returns nothing — while TF-IDF degrades gracefully. The token saving is real but buys lower recall here.
-
Reranker scale fix: llama.cpp raw logits vs [0,1] threshold — sigmoid in the llama_cpp branch
2026-09-27 · mscodebase-intelligence · verdict: partial.
Scale-contract violation confirmed and fixed with a 21-line single-file change, but 'threshold = root cause of P2/P3' is refuted: the cross-encoder itself rates the true files with negative logits, so no positive threshold keeps them. A threshold sweep 0.05/0.02 on the same eval…