Skip to main content

Mikhail (ManSio) — AI / Backend Engineer

Mikhail (ManSio) — MCP-Native Engineering Portfolio

Builds MCP-native tooling and AI infrastructure. Author of a production MCP server for codebase intelligence (LanceDB/BM25 hybrid search) in Zed.

Role: AI / Backend Engineer. Location: Remote-friendly.

Projects

  • MSCodeBase Intelligence

    Intelligent codebase search & indexing for Zed. Async MCP server featuring LanceDB/BM25 hybrid search, multi-bucket RAG, and autonomous self-healing workflows. High-performance, memory-safe, production-ready codebase intelligence.

    Language: Python. Stack: Python, MCP, LanceDB, BM25, RAG, Zed.

    • Hybrid search (vector + BM25) with fused ranking
    • Multi-bucket RAG over code entities
    • Autonomous self-healing: watchdog + reindex recovery
    • Memory-safe indexing pipeline (no leaks under load)

    View MSCodeBase Intelligence on GitHub

  • Gemma Agent

    Telegram assistant for a small trusted circle. Telegram assistant with memory, routing, and tools when needed — built for a small trusted circle, not as a Google Gemma product.

    Language: Python. Stack: Python, Telegram, LLM, Agents.

    • Persistent memory across sessions
    • Intent routing before tool dispatch
    • Tool use gated by permission scope

    View Gemma Agent on GitHub

  • MSPortfolio (this site)

    The MCP-native portfolio you are reading right now. A portfolio that is simultaneously a static dashboard, an MCP server about its owner's experience, and an interactive proof-of-work engine.

    Language: TypeScript. Stack: React, TypeScript, Vite, Tailwind, MCP, Fastify.

    • MCP server endpoint: any AI agent can query the CV
    • Browser agent-loop demo showing tool calls live
    • Live metrics with freshness + static fallback
    • Architecture simulator with break-it scenarios

    View MSPortfolio (this site) on GitHub

Engineering principles

  • Fail-closed by default

    When the system can't verify a condition, it must refuse rather than guess. Authorization, cache misses and unknown tool intents all default to 'no'.

    Example: In mscodebase-intelligence, a missing index returns a structured error to the agent instead of an empty result — the agent must not confuse 'nothing found' with 'index empty'.

  • Async-first I/O, offload CPU

    I/O-bound work lives on the async loop; CPU-bound work is offloaded with explicit progress reporting. No blocking calls on the request path.

    Example: The MCP tool server is fully async; reindex runs as a background task with progress notifications so the agent can poll instead of timing out.

  • Single write path for derived state

    Any state that is derived from other state (indexes, caches, fused rankings) is written through one code path. No ad-hoc mutations.

    Example: Hybrid search keeps a vector index and a BM25 index consistent because both are updated by the same ingestion pipeline.

  • Measure, don't assume

    Performance and reliability claims come from benchmarks with a command line, a thread count and raw output — not from architecture diagrams.

    Example: Every concurrency change in mscodebase-intelligence ships with a stress test that verifies both 'no exceptions' and 'correct input → correct output'.

  • Autonomous self-healing

    Where recovery is deterministic, the system does it itself: watchdogs, reindex triggers, stale-state cleanup. Humans only get involved for decisions, not for chores.

    Example: MSCodeBase Intelligence detects a corrupted/stale index and schedules a reindex automatically while serving from the last-good snapshot.

  • Agent-agnostic surfaces

    Tooling speaks protocols (MCP), not personalities. One server, consumed by Claude Code, Cursor, Copilot or any future agent.

    Example: This portfolio's own MCP server speaks standard Streamable HTTP — the same endpoint works from Claude Code, Cursor and Copilot with zero per-IDE setup.

Engineering timeline

Experiments

Every engineering experiment with its hypothesis, command and verdict (53 total), documented in the open. The interactive lab is at the lab page.

  • Memory Contamination Live-Arm: live LLM false-acceptance across 14 models

    2026-08-14 · mscodebase-intelligence · verdict: confirmed.

    The fail-closed look was anchor bias, not model paranoia: with token-string evidence qwen3.6/3.7 accepted 2-5/25 true claims (recall 0.08-0.20); with a real file fragment (V4 arm) recall jumps to 0.88 at FA 0.02-0.04. CORRECTED: the earlier 'every remaining false-accept is prese…

  • Memory Contamination proxy-control: retraction -88%, verify-on-read adoption to 0.0

    2026-08-11 · mscodebase-intelligence · verdict: confirmed.

    The proxy always decides (unknown=0) and false_accept=0 by construction - the headline numbers are a property of the heuristic, not of LLM behavior; the live arm (exp-1) measures the real gap.

  • Manifest anchoring (pkg: anchors, ADR-0005): closed-world manifest kills 7 false REFUTED

    2026-08-14 · mscodebase-intelligence · verdict: confirmed.

    A closed-world manifest is a stronger evidence source than bare tokens; the fix was verified by 1-V-REP (0 false REFUTED of TRUE facts).

  • Context aggregator vs multi-tool: 1 call instead of 4-5, -78% calls at recall parity

    2026-08-08 · mscodebase-intelligence · verdict: confirmed.

    The B-scheme (intent filter) is the optimum: same recall as multi-tool at 1 request and 275 tokens; the earlier recall gap was noise, not signal.

  • Multi-RAG ablation: FTS5-only beats the full pipeline by recall; H1 refuted

    2026-08-11 · mscodebase-intelligence · verdict: refuted.

    Recall comes from keyword tiers (FTS5/BM25), precision is bought by the reranker, and the vector tier is the weakest on symbol tasks. Production bug found: hybrid_search_async cache-hit silently skips the dense tier - cache isolation per arm is mandatory in ablations (first run…

  • Shadow-canary fail-open: 5/5 attacks passed before the fix, 0 after

    2026-08-11 · mscodebase-intelligence · verdict: confirmed.

    Relative metrics are attackable: absence of signal (empty canary) must fail closed, and a quality floor must be absolute, not relative.

  • Concurrency vs semantic correctness: '0 errors' != correct data under race

    2026-08-11 · mscodebase-intelligence · verdict: confirmed.

    VC and VOR are complementary layers: consistency (VC) catches lost writes, semantics (VOR) catches lies. A verdict that flips when one scenario variable changes (A6) is an illustration, not a law.

  • Mutation testing for the reranker grader: mutation score 8% -> 100%

    2026-08-14 · mscodebase-intelligence · verdict: confirmed.

    Validate values, not just types - NaN/Infinity otherwise silently earn the maximum score.

  • Root-cause prediction audit: top-1 accuracy 0.13 on 31 incidents

    2026-07-22 · mscodebase-intelligence · verdict: partial.

    Root-cause prediction is far from solved: the audit dataset and gold standard are the baseline, and the engine's current top-1 accuracy is 13%.

  • PID-lock self-healing: orphan lock wait 30s -> 120ms

    2026-08-08 · mscodebase-intelligence · verdict: confirmed.

    Never block boot on a lock whose holder may be a zombie: classify the holder and self-heal instead of failing closed for 30 seconds.

  • Late enrichment: imports metadata covers 0.0% of search chunks

    2026-08-08 · mscodebase-intelligence · verdict: partial.

    Index-time enrichment produced nothing for search chunks: imports coverage is 0.0, so the late path is kept behind a flag until the hypothesis is re-probed.

  • Benchmark 2.0 scaffold: 12 repository-reasoning tasks at levels L3-L5

    2026-08-08 · mscodebase-intelligence · verdict: partial.

    A reasoning-quality benchmark needs L3-L5 tasks with evidence and checks; short keyword queries are a separate failure mode (CoREB) that must be measured independently.

  • Server unavailable during reindex: sync update_all blocks the event loop (P0 root cause)

    2026-08-13 · mscodebase-intelligence · verdict: confirmed.

    552s indexing + ~220s update_all = 771s unavailability, in line with the log. Indexing-in-executor (H1) refuted; fast-fail search while is_reindexing exists; the blocker was the sync update_all (same class as BS-11: run_full_diagnostic had already been moved to asyncio.to_thread…

  • Vacuous-test scan: 1133/1143 tests are provable (hypothesis refuted in the good way)

    2026-08-11 · mscodebase-intelligence · verdict: refuted.

    The suite is almost fully provable - 1133/1143 (99.7%) contain a failing construct; the 3 'vacuous' are smoke tests that can still fail via exception propagation from helpers (discrimination weaker, not zero). MSCodeBase is NOT in the fintech state (7/40 proven) - which is preci…

  • ln.strip() bug class replayed: broken assert-extractor gives 3/8 false passes

    2026-08-11 · mscodebase-intelligence · verdict: confirmed.

    Bug class replayed: valid Python that never executes, exit 0, 'verified:true' for a wrong answer. The form count differs (3/8 vs fintech's 5/8) but the essence is identical - signature/exit-code verify the PROCESS, not the SEMANTICS ('ask for the output, not the exit code').

  • Population blind spot: '0 eligible' is indistinguishable from '0 collected'

    2026-08-11 · mscodebase-intelligence · verdict: confirmed.

    (a) and (b) share the same failure signal (passed=0, same warning class) - and (a)'s message claims 'empty/garbage chunks' when raw results were 0, a false explanation. '0 rows with 0 eligible' (healthy idle) is indistinguishable from '0 rows with N eligible' (broken collector).…

  • verify_clean_state.sh falsifiability: the drift-gate is structurally unable to fail

    2026-08-11 · mscodebase-intelligence · verdict: partial.

    (B) confirmed: the gate prints PASSED for a suite with zero asserts - semantic blindness. (A) refuted in reverse: the drift-gate is STRUCTURALLY unable to fire - its grep -iE '^pkg==' requires the pin at line start, but pins live in a TOML array (' "lancedb==0.34.0",') so PIN…

  • Evidence Ladder E1-E3: evidence form (anchor -> file fragment -> graph) vs verification quality

    2026-08-15 · mscodebase-intelligence · verdict: partial.

    Evidence form is a per-model knob, not a global answer: the file fragment is the strongest recall driver for all models; the graph closes the trap failure mode only for evidence-honest models (and on corrected labels the trap gap itself was a label issue). Fail-open models (glm…

  • Evidence Ladder E3b+E4: file+graph hybrid is NOT additive; git-provenance distinguishes existed-then vs exists-now

    2026-08-15 · mscodebase-intelligence · verdict: refuted.

    Hybridity is NOT additive: with a file fragment present, token presence dominates and the graph adds nothing (acc 0.900 < file 0.940) - VOR design must pick ONE evidence format: fragment for recall OR graph for trap-precision. Git-provenance is a cheap, powerful temporal signal…

  • ONNX embedder down: off-by-one project paths (not the model or ports)

    2026-08-03 · mscodebase-intelligence · verdict: confirmed.

    Root cause is paths, not model/ports: (1) onnx_client looked for '...src/src/core/embedder/onnx_server.py' (duplicated src), (2) onnx_server looked for the model in '.../src/.codebase_models/...' (not the root). The off-by-one was copied between files (onnx_client <- onnx_server…

  • Evidence Ladder E5: extended trap category (N=20, subject-validated labels) - present-trap is mass, graph really helps deepseek (75%->40%)

    2026-08-16 · mscodebase-intelligence · verdict: confirmed.

    Present-trap is NOT a mislabel issue: on honest subject-validated labels file_content FA is 2-15/20 (10-75%) - the earlier 'residual trap hole' was distorted by N=1. Graph evidence REALLY closes present-trap for deepseek (75%->40% FA); for glm it does not (14/20); qwen already h…

  • Pinned-rerun on the corrected dataset (fp e6ce7b90): canonical series numbers, routing band removed

    2026-08-16 · mscodebase-intelligence · verdict: confirmed.

    Routing band removed but per-model conclusions unchanged: file_content is the best recall arm for qwen (0.88), graph for glm (0.84, FA trap 1), hybrid for deepseek (0.92); temporal present-trap is universal (NOW 12/12, 9/12, 12/12), past solved by formulation (40/40). glm stays…

  • LSP live probe: call hierarchy + semantic tokens confirmed on pyright; type hierarchy / moniker NOT advertised

    2026-08-19 · mscodebase-intelligence · verdict: partial.

    Spec-present != server-implemented. On the pyright family the graph really gets call hierarchy (compiler-accurate cross-file CALLS edges) and semantic tokens (exact spans/kinds); type hierarchy (3.17) and moniker (3.16) are NOT advertised by pyright -> dormant until a server sup…

  • M1: real tool telemetry — 5/62 MCP tools ever used; the working-set boundary is single-digit

    2026-08-25 · mscodebase-intelligence · verdict: partial.

    Only 5 of 62 tools ever recorded a metric; dead-tool metrics are NOT measurable with current telemetry — a per-call counter in the error_handler path is required. The LSP toolkit (8 calls, 3 errors) is the top latency risk.

  • M2: sub-agent in a fresh project — natural 0/5 MCP calls vs MCP-first 9 (4 wasted on unindexed files)

    2026-08-25 · mscodebase-intelligence · verdict: confirmed.

    Natural agent: 0/5 MCP; MCP-first: 9 calls, ~4 idle (2 search + get_symbol_info + find_path on unindexed lab), only read_live_file (disk) productive. New files are invisible until reindex -> semantic layer on fresh code = zero. Correctness control: A found the root with exact fi…

  • M3: latency matrix — search fast 75ms HIT vs quality 3666ms MISS; get_symbol_info misses a real symbol

    2026-08-25 · mscodebase-intelligence · verdict: refuted.

    quality was 49x slower than fast (3666 vs 75ms) and semantically worse; get_symbol_info misses exact names carrying a trailing hint. On short queries fast search + grep + read_live_file beat quality search + symbol info on time AND accuracy.

  • E2: category pilot on the live index — fast 5/6 HIT (83%) vs quality 2/3 + leak to docs/JSON; index self-pollution discovered

    2026-08-26 · mscodebase-intelligence · verdict: partial.

    fast is 83% and an order of magnitude cheaper; quality saved the only fast-MISS (Q4) at ~35x cost. The leak is confirmed (Q1 quality -> JSON/CHANGELOG; Q4/Q6 quality top-1 incident datasets) but is NOT junk — semantically relevant docs; a category-priority filter is needed, not…

  • E3: category router on tasks_v3.json (30 tasks) — cascade 0.233 > fast 0.167 > quality 0.133; 4 graph classes 0.00 across all arms

    2026-08-26 · mscodebase-intelligence · verdict: confirmed.

    Cascade is the winner; the winner depends on klass (git_history/bug/prepare -> fast with quality 0.00 at bug; caller_callee/architecture -> quality); find_test/find_impact/modify/verify_change = 0.00 across ALL arms — search does not replace graph/AST/impact stages.

  • E4: deterministic per-class keyword router (PoC) — 0.200 < cascade 0.233, klass_acc 0.40 (NEGATIVE)

    2026-08-26 · mscodebase-intelligence · verdict: refuted.

    klass_acc=0.40 — keyword rules are noisy (bug tasks contain 'callers', git tasks 'why'); the union arm does not save the 4 graph classes — search cannot find what is not textually in the index (callers/callees/impact must come from the graph). The search-only ceiling on this dat…

  • E4.1: the graph stage breaks the search-only ceiling — recall 0.433 (cascade 0.267), med 177ms (Track 1)

    2026-08-26 · mscodebase-intelligence · verdict: confirmed.

    The graph stage added +0.166 recall on three of the four failing classes at ~15ms latency and lost nowhere (fallback to cascade by construction). Lessons: (1) has_symbol is an exact node-name — search_symbols(LIKE) + strict suffix is required; (2) graph navigation must go throug…

  • E4.2: deterministic concept resolver (no LLM) — verify_change T9/T29 HIT on real graph.db, facts 4/4 & 3/4 (Track 2)

    2026-08-26 · mscodebase-intelligence · verdict: confirmed.

    Both verify_change misses resolve to the correct file on the real graph.db (10748 nodes); graph rows now carry facts (graph_fact_text), not 0; regression is excluded by construction (klass-gating). A 'wordless' prompt is an anchor-resolution problem, not classification — lexical…

  • system_alerts chain (file changed -> STALE -> VOR -> one-shot alert) + idle background stale-sweep

    2026-09-10 · mscodebase-intelligence · verdict: confirmed.

    A one-shot alert store with atomic collect_and_clear is the right granularity: dedup by kind+payload prevents a spam-per-save alert, the atomic clear guarantees that a race between two MCP tools delivers the alert once and only once, and limit=5 caps token budget. Idle VOR is ch…

  • VOR Catch-up Rate (H1): throughput per budget and cycles-to-finish

    2026-09-10 · mscodebase-intelligence · verdict: confirmed.

    Real project (~247 nodes) is far from budget limits: read-path covers ~420-490 nodes/pass, background ~1600-1900 nodes/pass. N<=2000 fits one 250ms pass; N=5000 needs 2 passes. Systematic starvation (MATCHED>0/DELIVERED=0) appears only around N=5000.

  • VOR HEAD-polling catches external git changes without notify_change (H3)

    2026-09-10 · mscodebase-intelligence · verdict: confirmed.

    Head-invalidation per node key is an honest detector of external code drift in a git repo: the node whose file anchor was removed went REFUTED, the node whose file changed but still exists stayed VERIFIED. First exp3 run gave a false REFUTED because statuses were read from pre-r…

  • Burst-rename vs fail-closed VOR: prose-path anchoring causes false REFUTED (1-B/1-C/RT)

    2026-09-11 · mscodebase-intelligence · verdict: confirmed.

    Real queue is NOT huge (24 accumulated auto-refutes/month) but is DOMINATED by junk anchors from prose scanning (13/24 import:for/file/path regex hits), not by renames - renames caused only 1 FALSE_REFUTE (ADR-7232a6e2ba34: live node revoked by historical path from prose body).…

  • H3 TTL-rotation for verify-on-read: last_checked for every checked node + stale_ttl label (closes the "hangs forever" class)

    2026-09-11 · mscodebase-intelligence · verdict: confirmed.

    A TTL needs a trace for nodes WITHOUT verdicts, not only for VERIFIED ones: INCONCLUSIVE nodes now get last_checked and it refreshes on every idle pass (H1), so "no dates at all" is replaced by "checked every 6h, labelled stale after 30d". stale_ttl is a computational label - no…

  • H4 Agent-memory lifecycle at scale: capture latency on a dev.to graph grown 3.4x (threads 5.5x) + stale precision

    2026-09-13 · mscodebase-intelligence · verdict: confirmed.

    Capture latency grew with graph size (earlier own refresh on ~9k threads completed in ~2 min vs 10m38s now) even though the numeric delta between snapshots was small (threads 19,056 -> 19,101, +45). The cost is the network capture phase (134 dev.to API calls), not the local grap…

  • Bootstrap pipeline: test->function linking ? static name/import vs dynamic sys.settrace trace (89.8% vs 0%/77.9%)

    2026-09-15 · mscodebase-intelligence · verdict: confirmed.

    Dynamic trace (sys.settrace around pytest_runtest_call) is the only viable deterministic test->function linker: 89.8% vs 0% (name) and 77.9% (import, file-level only). Overhead +13.6% on a full run is fine for a one-off bootstrap pass. Entry points (@mcp_app.tool, 22 in src) are…

  • Bootstrap follow-up: Tarantula ranking for test->function annotation — recall refuted (22.6% rank<=3), precision confirmed; dev.to cross-check

    2026-09-15 · mscodebase-intelligence · verdict: refuted.

    Vanilla Tarantula cannot select the target function for most tests: recall is 22.6% rank<=3, far below the 60-70% recall target — shared utils (autouse fixtures safe_mkdir/get_data_root with 200+ callers) is the barrier, the same noise source TRUE Coverage reports (shared utilit…

  • Bootstrap A1: coverage.py sysmon driver overhead vs our sys.settrace plugin — sysmon refuted (+19.96% vs +13.6%)

    2026-09-16 · mscodebase-intelligence · verdict: refuted.

    sysmon is ~1.5x slower than our hand-rolled sys.settrace plugin on the full suite (+19.96% vs +13.6% in the same session, same methodology), so coverage.py is NOT adopted as the production driver. The earlier research estimate of 'sysmon ~3-7% median loss' (KNOWN_ISSUES:249) is…

  • Bootstrap Step B: Static Score Engine (AST calls/lexical/imports) vs dynamic trace ground truth

    2026-09-16 · mscodebase-intelligence · verdict: refuted.

    Hit >=50% confirmed (union 90.4%), but recall <=30% REFUTED: union recall is 70.0%. Static is much stronger than Exp 7 implied, because Exp 7 measured the weak name-signal (L2 reconfirms it: 17.7%). The call-level signal (L1) is a precise, narrow anchor: precision 68.0% with avg…

  • E7: lazy stat-sweep (mtime+size) vs sha256-sweep for always-fresh index (FreshnessChecker hot-reload)

    2026-09-18 · mscodebase-intelligence · verdict: confirmed.

    stat-sweep is 12x faster than sha256 over the same corpus and ~2% of a full reindex. (lancedb 0.34 detail: table.to_pandas(columns=[...]) crashes, table.to_lance().to_pandas(columns=[...]) is the correct API.) This made the synchronous pre-search freshness check viable: stat-fir…

  • E10: full-text chunk embedding + e5-prefix + reranker pool 50 vs pure-vector plateau

    2026-09-19 · mscodebase-intelligence · verdict: refuted.

    No confirmed shift within N=10 noise: quality hit@1 30%->20% (worse/noise), hit@5 30%->40% (noise). Toggles requiring a full prod reindex (~13 min) deliver zero -> pure-vector plateau reconfirmed (cf. Exp-29 search-only ceiling ~0.23). Next move: AST/Graph-hybrid re-ranking (gra…

  • E11: AST/Graph-hybrid re-ranking - symbol-lookup lift above the pure-vector plateau

    2026-09-19 · mscodebase-intelligence · verdict: confirmed.

    Graph lift saved 2/10 target cases (project_indexer_registry.py, indexing_tools.py) that quality-vs-baseline lost; both targets were present in search_symbols output. Symbol-lookup cost is negligible (+6ms). Signal is conditional (loose graph files add noise - MRR stays small, h…

  • E14: EmbeddingGemma 300M vs multilingual-e5-small - Hit@1 0.062->0.688 at ~4x CPU cost

    2026-09-22 · mscodebase-intelligence · verdict: confirmed.

    Real, reproducible quality gap on a code corpus: gemma Q8 gives +62.6pp Hit@1 and +56.4pp MRR over prod e5 at 2x RAM (176 vs 91MB) and 4.2x slower. QAT-Q4 (ggml-org) is NOT better than plain Q4_0 (unsloth): 0.653 vs 0.695 MRR - marketing claim not confirmed. Batch size does not…

  • E17: TESTS-signal in graph-stage (A/B) — covering tests added over function defs, hit@1/MRR unchanged (7/7)

    2026-09-22 · mscodebase-intelligence · verdict: confirmed.

    Covering tests are appended as a separate result sort (graph_score 0.4 vs def 1.0, sentinel chunk_index -20M+line) AFTER completed function defs, so def-first invariant holds (MRR 1.0 in both arms). The flag is off by default (MSCODEBASE_TESTS_SIGNAL env, like late_enrichment),…

  • E12: real-path embed throughput — chunk length sets the ~3.5k tok/s ceiling, not batch, ubatch or parallelism

    2026-09-20 · mscodebase-intelligence · verdict: confirmed.

    The CPU embed ceiling (e5-small Q8, 10 threads, Ryzen 5600H) is ~3.4-3.7k tok/s — model physics, independent of batch size, token budget, parallelism or truncation. The old "156 ch/s" was synthetic (~1560 tok/s on 10-token texts). 335k chunks x 203 tok = 68M tokens -> ~5.7h of p…

  • E13: text RAG (doc-chunks) vs code baseline — doc-chunks miss top-5 for 14 of 16 queries

    2026-09-20 · mscodebase-intelligence · verdict: refuted.

    Text RAG is far below code RAG: doc-chunks stay outside top-5 for 14 of 16 queries. The embedder packs code chunks (signatures, names) denser, doc-chunks are diffuse; the index is code-biased and queries without intent_hint='docs' route down the code path.

  • E15: bge-small-en / MiniLM / nomic in equal prod conditions — bge-small-en is quality without the 3x slowdown

    2026-09-23 · mscodebase-intelligence · verdict: confirmed.

    bge-small-en-v1.5 Q8 is the only "quality without losing speed" candidate: MRR 0.545 (+136% vs e5) at 5024 tok/s (1.19x FASTER than e5) and dim=384 — same scheme, no reindex, RAM halved to 46MB. MiniLM is the fastest (2.6x) but MRR 0.487. nomic does not qualify: its 8192 ctx nee…

  • E16: bootstrap-trace portability — dynamic trace transfers to foreign Python repos (97.3% / 100%)

    2026-09-22 · mscodebase-intelligence · verdict: confirmed.

    The dynamic trace transfers to foreign Python repos unmodified: 97.3% on gemma_agent with 2.9k tests, 100% on commit-. Overhead scales with test activity (+17.4% on gemma_agent vs +13.6% on our larger corpus), not with corpus size. Non-Python: the pytest pipeline collects 0 test…

  • F5 4-arm unit-of-return (judged reader): whole document beats top-k chunks 8x on code, direction-only on prose

    2026-09-26 · mscodebase-intelligence · verdict: partial.

    The unit of return acts on the reader, not on gold-file retrieval: objective retrieval was tied (A=B=5/16 top-1) yet the reader answers 8x more code questions correctly from a whole document than from chunks — confirmed for the code unit effect. For prose the effect is direction…

  • NodeRAG deterministic duel: chunked TF-IDF 8/10 vs PropertyGraph BFS 7/10 — graph does not win

    2026-09-27 · mscodebase-intelligence · verdict: refuted.

    No evidence graph traversal wins on this corpus: -1 hit at -43.6% tokens. The graph arm is entry-point fragile — a query whose symbols are not AST functions/classes returns nothing — while TF-IDF degrades gracefully. The token saving is real but buys lower recall here.

  • Reranker scale fix: llama.cpp raw logits vs [0,1] threshold — sigmoid in the llama_cpp branch

    2026-09-27 · mscodebase-intelligence · verdict: partial.

    Scale-contract violation confirmed and fixed with a 21-line single-file change, but 'threshold = root cause of P2/P3' is refuted: the cross-encoder itself rates the true files with negative logits, so no positive threshold keeps them. A threshold sweep 0.05/0.02 on the same eval…

Engineering diary

Incidents and hard bugs: symptom, root cause, fix and guard (32 entries).

  • Custom Python LSP cannot register in Zed (WONTFIX)

    2026-07-05 · mscodebase-intelligence · status: fixed.

    Zed settings.json can only override KNOWN LSPs; a pygls-based server is never started (verified on Zed 1.9.0 and 1.14.2).

  • ONNX migration: 7 critical bugs fixed while moving off LM Studio

    2026-07-08 · mscodebase-intelligence · status: fixed.

    LM Studio is an external process needing manual startup; the status API hardcoded 'lm_studio' as the provider, so it lied when port 1234 was closed.

  • Lock-zombie: MCP boot blocked up to 30s by a stale DB PID-lock

    2026-08-08 · mscodebase-intelligence · status: fixed.

    Boot failed closed on the DB PID-lock for up to 30s, exceeding the Zed timeout; the server was killed, the zombie stayed, and _is_pid_alive could not tell a healthy MCP from an orphan.

  • Probe script shadowed the production symbol build_call_graph

    2026-08-08 · mscodebase-intelligence · status: fixed.

    run_experiment_pagerank.py defined the same name as the production build_call_graph; the probe silently shadowed it in the graph module.

  • Context aggregator D1-D3 defects found by the harness

    2026-08-08 · mscodebase-intelligence · status: partial.

    Intent filters git_history/verify_change silently returned empty in the B-scheme; harness gaps (D1-D3) hid the failures.

  • Late enrichment: imports metadata = 0.0 on search chunks

    2026-08-08 · mscodebase-intelligence · status: partial.

    Chunks from the search pipeline (~186 tokens) carry no import metadata; enrichment at query time had nothing to add.

  • Memory contamination: the memory layer accepts false claims

    2026-08-11 · mscodebase-intelligence · status: fixed.

    Project memory was add-only with no retraction and no validation against the code; false claims propagated into context.

  • Shadow canary fail-open: 5/5 attacks passed

    2026-08-11 · mscodebase-intelligence · status: fixed.

    The canary trusted empty responses and a failed baseline (fail-open); the quality metric was relative, so a collapse-to-constant passed.

  • hybrid_search_async cache-hit skips the dense tier

    2026-08-11 · mscodebase-intelligence · status: partial.

    An embedding cache hit returned early, silently skipping the vector search tier; repeat queries lost vector recall (the first multi-RAG run was fully distorted).

  • Concurrency: 0 errors did not mean correct data

    2026-08-11 · mscodebase-intelligence · status: fixed.

    Replacing a thread-safety primitive created a new race surface (shared results dict, shared correlation id, result swap between inputs) that exception-free runs do not expose.

  • Memory v2: SUPERSEDED filter + false-retraction metric not committed

    2026-08-12 · mscodebase-intelligence · status: partial.

    ADR-0004 cascade behavior (REFUTED propagation) was designed and validated but left uncommitted at session end.

  • Server unreachable during/after indexing

    2026-08-13 · mscodebase-intelligence · status: fixed.

    A sync update_all in the main loop blocked the server while indexing.

  • extract_anchors produced garbage anchors -> false VOR retractions

    2026-08-13 · mscodebase-intelligence · status: fixed.

    Anchor extraction over-matched identifiers, so verify-on-read retracted valid memory.

  • Reranker offline all day: PID-reuse in _is_pid_alive

    2026-08-13 · mscodebase-intelligence · status: fixed.

    A completed process object was classified as alive via PID reuse, so the reranker process was treated as running when it was dead.

  • Duplicate servers with two Zed windows

    2026-08-13 · mscodebase-intelligence · status: fixed.

    The lock was taken before Popen, not before port readiness; the second window started a duplicate server.

  • validate_scores grader: 11 holes (NaN/Infinity pass)

    2026-08-14 · mscodebase-intelligence · status: fixed.

    Type validation without value validation: NaN/Infinity passed isinstance, clamped, and earned the maximum score.

  • stale_detector MCP tool: 11 false drifts

    2026-08-14 · mscodebase-intelligence · status: fixed.

    The stale detector reported drifts that were actually renames/reorders; severity_overrides has a Windows quirk.

  • fastmcp dist-name vs import-path: 7 false REFUTED

    2026-08-14 · mscodebase-intelligence · status: fixed.

    VOR anchors (file/import/env) could not tell a dist name from an import path; the fastmcp class was retracted as absent.

  • CoT arm: glm-4.7 upstream EMPTY_CONTENT

    2026-08-15 · mscodebase-intelligence · status: partial.

    In reasoning mode the upstream returns finish=stop with an empty body (reasoning_tokens ~6) for 16-26% of responses; the share is unstable between runs.

  • VOR showed token strings instead of file fragments -> recall 0.08 (anchor bias)

    2026-08-15 · mscodebase-intelligence · status: fixed.

    The VOR layer presented bare pattern tokens ('typesense') as evidence; a model honestly answers false/unknown because a token string proves nothing (diagnosed by the V4 arm file_content_first).

  • Idle VOR ticker plus system alerts close the file-change loop

    2026-09-09 · mscodebase-intelligence · status: fixed.

    VOR ran from exactly one call site (intel_get_project_memory); mark_stale never fired and no alert path reached the agent, so file changes never re-verified memory.

  • VOR read-path rescanned prose history and refuted live nodes (PR #34)

    2026-09-11 · mscodebase-intelligence · status: fixed.

    Read-path VOR re-scanned prose bodies for anchor-like paths even when explicit anchors existed, so stale historical paths revived at rename-sweep and falsely refuted a live node.

  • VOR timestamps every check; stale_ttl label flags long-unverified nodes

    2026-09-11 · mscodebase-intelligence · status: fixed.

    VOR dated only VERIFIED transitions, so anchorless INCONCLUSIVE nodes left no trace and hung unverified forever (70 ACTIVE nodes without any date on 2026-09-11).

  • Static test-to-code linking falls short; dynamic trace links 89.8 percent (Exp 7/38)

    2026-09-15 · mscodebase-intelligence · status: fixed.

    Name matching linked 0 percent of tests to functions and imports reached file level only (77.9 percent), so static signals cannot ground test-to-code edges.

  • sysmon tracing overhead refuted; TESTS edges land in the graph (Exp 8, A2)

    2026-09-16 · mscodebase-intelligence · status: fixed.

    The claimed 3-7 percent sysmon overhead measured plus 19.96 percent against plus 13.6 percent for settrace, so coverage-based tracing lost as the bootstrap driver.

  • AST entity detector plus automatic source-root resolution (Bootstrap Step 1)

    2026-09-17 · mscodebase-intelligence · status: fixed.

    Regex Table( matches were 90 percent open_table call noise, and the detector hardcoded src/ while real layouts use core/, libraries/, or the project name.

  • bootstrap_pipeline command wires trace-to-edges end to end (Step 3)

    2026-09-18 · mscodebase-intelligence · status: fixed.

    The trace plugin lived in experiments/ with no single entry point for an agent (MCP) or a script (CLI).

  • Silent LanceDB migration plus destructive rebuild fixed; E10 search tweak refuted

    2026-09-19 · mscodebase-intelligence · status: fixed.

    Migration helpers were imported as module functions but defined as class methods, so the ImportError path skipped migration; a schema-mismatch message was then read as a missing table, triggering a full drop and re-embed.

  • Graph-hybrid rerank probe lifts hit at 5 as signal, not proof (E11)

    2026-09-19 · mscodebase-intelligence · status: partial.

    Embedding and BM25 tiers drop exact code identifiers that the symbol index knows precisely, so deterministic targets never surface in natural-language queries.

  • Reranker scale contract fixed with sigmoid; P3 target still below threshold

    2026-09-27 · mscodebase-intelligence · status: fixed.

    The local rerank endpoint returns raw logits (about -11 to +11) while the 0.3 cutoff assumes a 0-1 range, so the filter cut 70-97 percent of candidates.

  • CI guards failed on a healthy deploy: publish race, unsatisfiable threshold, transport flake

    2026-09-28 · msp-portfolio · status: fixed.

    Four independent defects. (1) The metrics freshness check asserted wall-clock Age <= 120 min right after the deploy job, but GitHub Pages publishes 24-65 s later, so it scored the PREVIOUS run's snapshot (514 min) as stale even though the new one was on the way. (2) The bound wa…

  • Lab projections re-synced with logs; v1 paraphrase baseline moved by a real corpus entry

    2026-09-28 · msp-portfolio · status: fixed.

    lab JSON files had drifted from the markdown logs: test-suite counters were stale (101 vs the real 112 after the lab, mcp-tools and worker suites had grown) and two logs edited this session had no JSON projection. The paraphrase stage-0 test also encoded a frozen corpus: it dema…

Verifiable claims

Claims about the owner that can be grounded against the portfolio data (29 checks).

  • production MCP server for codebase intelligence — supported
  • LanceDB and BM25 hybrid search — supported
  • Telegram assistant with memory and intent routing — supported
  • React TypeScript Vite portfolio — supported
  • Hybrid search with fused ranking — supported
  • must refuse rather than guess when it cannot verify — supported
  • performance claims come from benchmarks with a command line — supported
  • derived state written through a single write path — supported
  • watchdogs and reindex triggers recover deterministically — supported
  • one server consumed by Claude Code and Cursor — supported
  • chose LanceDB and BM25 over SQLite FTS5 — supported
  • joined GitHub as ManSio — supported
  • claimed a fork as my own work — supported
  • hardcoded dark theme colors broke the light theme — supported
  • mutation testing for the reranker grader — supported
  • FTS5 only beats the full pipeline by recall — supported
  • live LLM false-acceptance across 14 models — supported
  • zero errors did not mean correct data — supported
  • reranker was offline due to PID reuse — supported
  • OpenRouter routes one prompt across multiple upstreams — supported
  • two Zed windows can start duplicate servers — supported
  • late enrichment found no imports on search chunks — supported
  • worked at Google — refused
  • led a team of engineers at Meta — refused
  • built a mobile app for iOS — refused
  • ten years at Netflix as staff engineer — refused
  • contributed to the Linux kernel — refused
  • won a Turing Award — refused
  • CTO of a Series B startup — refused

Test suites

  • Intent matching (RU/EN) — 7 tests (Agent intent routing: projects, stack filters, principles, articles, timeline, fallback)
  • MCP tools + architecture models — 20 tests (Tool surface, get_profile nextSteps (D8), filters, analyze_stack honesty, simulation (5 points, validation, llm_saturation, events), commit snapshot, antipatterns, lab tools, annotations, all architectures × all scenarios)
  • Worker /mcp integration — 28 tests (tools/list (18 tools + annotations), health without rate limit, /mcp/stats (KV), tools/call counter, 429 MCP/CHAT, security headers, CORS allow/deny, 404, adversarial (malformed JSON, unknown tool/method), 8-way concurrency with filter correctness, /resume.txt + /llms.txt + 405, anonymous quota (D4: headers + 429), /openapi.json (D5))
  • Lab data integrity + SSR render + benchmarks snapshot — 13 tests (Experiment completeness/verdicts + project tags, negative results refs, diary completeness/statuses + project tags, KI ids/temperatures + project tags, test-suite sum = total, evidence ledger ids/verdicts/summary, LabPage SSR (9 sections + filter), EN-only projections (incl. evidence claims), benchmarks.json sanity (DoD gates: llm recall ≥80%, 0 false-accepts, p95 < 3s))
  • Evidence Score v1 (D3 deterministic arm) + v2 stage-0 paraphrase set — 13 tests (verify_claim: 13th tool registration + required claim input, supported true-claim (profile evidence), project-traced claim, negative controls (Google/Meta refused), too-short/empty claim notes; computeEvidence: toolCalls/grounded/failed counts, ungrounded flag, null result = failed; evidence ledger: every canonical claim verdict matches verify_claim; v2 stage 0: 8 true paraphrases refused by v1 (recall-gap baseline, KI-017), 3 paraphrased negative controls stay refused, documented substring false-acceptance (search+engine))
  • v2 LLM arm decision logic (mocked provider) — 9 tests (verifyClaimLlmArm fail-closed guards (v2 plan §5): supported with valid cited record, refused passthrough, garbage response, supported without source, supported with source outside candidates, HTTP error, abort/timeout, too-short claim without network call, zero-overlap paraphrase reaches the LLM via padded core context (p-01 fix))
  • verify_claim + LLM arm integration (stage 2) — 5 tests (arm absent -> deterministic v1 (arm:deterministic), arm rescues miss with cited source (arm:llm), arm refusal keeps refusal, arm error fail-closed, deterministic hit never consults the arm)
  • verify_repo live GitHub verification (mocked) — 9 tests (bare name defaults to ManSio + portfolio cross-check (languageMatches), full owner/name, honest language mismatch vs curated record, 404 = not exists, 403/429 rate limit reported honestly, network failure, github.com URL normalization, empty input without network call, readme:true returns actual README text)
  • verify_article live Dev.to verification (mocked) — 4 tests (title-fragment match with platform data, honest not-found, API failure honest error, short query refused without network call)
  • verify_package live npm verification (mocked) — 4 tests (existing package + maintainer check (case-insensitive), honest 404 not-found, registry failure honest error, empty input without network call)

Known issues

Open engineering debt, tracked in the open with status and temperature.

  • KI-101 (mscodebase-intelligence) — Fix in code, temperature: watching.

    hybrid_search_async: an embedding cache-hit silently skips the dense tier - repeat queries lose vector recall (found by the multi-RAG ablation; the first run was fully distorted).

  • KI-102 (mscodebase-intelligence) — Open (by design), temperature: watching.

    Present-trap blindness is structural: claims about an existing file/import with the wrong subject or value are not caught by anchor verification (memory_first adoption 0.24 in the proxy control).

  • KI-103 (mscodebase-intelligence) — Open (research), temperature: watching.

    VOR fail-closed was anchor bias, not model paranoia: with token strings qwen3.6/3.7 accepted 2-5/25 true claims (recall 0.08-0.20); with a real file fragment (V4 arm, 2026-08-15) recall jumps to 0.88. CORRECTED (RED TEAM on ground truth, 2026-08-15): the apparent residual hole 'every remaining false-accept is present-trap (R45/R46)' was a mislabeled-data issue - R43/R45/R46/R47 are TRUE facts (th…

  • KI-104 (mscodebase-intelligence) — Open (upstream), temperature: watching.

    glm-4.7-flash in reasoning mode returns 16-26% EMPTY_CONTENT (finish=stop, empty body) - an upstream defect; its CoT numbers are qualitative only.

  • KI-105 (mscodebase-intelligence) — Open (by design), temperature: stable.

    OpenRouter routes one prompt across >=8 upstreams (Alibaba/DeepInfra/DigitalOcean/Cloudflare/Novita/Baidu/StreamLake/Bedrock): per-run verdict variance (FA +/-0.05-0.10) is expected; determinism under temp=0 is an illusion.

  • KI-106 (mscodebase-intelligence) — Open (research), temperature: watching.

    Late enrichment: imports=0.0 on search chunks (~186 tokens/chunk); the MSCODEBASE_LATE_ENRICHMENT flag stays off until the hypothesis is re-probed.

  • KI-107 (mscodebase-intelligence) — Open (by design), temperature: stable.

    The vector tier (multilingual-e5-small) is the weakest on symbol tasks (recall 0.083-0.167); recall is carried by keyword tiers, precision by the reranker.

  • KI-108 (mscodebase-intelligence) — Open (research), temperature: watching.

    Graph enrichment adds metadata (callers/callees), not text: evidence metrics by text patterns do not see it; a separate graph-contribution protocol is needed.

  • KI-109 (mscodebase-intelligence) — Open (P1), temperature: watching.

    P1: propagation_engine.py is invisible to code search and the symbol graph - the root cause is not yet established.

  • KI-110 (mscodebase-intelligence) — Open, temperature: watching.

    Path storage is scattered and there is no GC: 2481 junk folders accumulated; a cleanup task is open.

  • KI-111 (mscodebase-intelligence) — Open, temperature: stable.

    severity_overrides has a Windows quirk (stale_detector) - behavior differs between platforms.

  • KI-112 (mscodebase-intelligence) — Open (mitigated), temperature: stable.

    Two Zed windows can start duplicate servers: the lock is taken before Popen, not before port readiness (mitigated 2026-08-13).

  • KI-113 (mscodebase-intelligence) — Fixed, temperature: stable.

    Memory verification was lazy-by-hand: VOR ran only from intel_get_project_memory (1 call site); 42/136 ACTIVE nodes had no verified_at since 2026-08-11; INCONCLUSIVE nodes (no anchors) never became STALE/REFUTED. FIXED (H1, 2026-09-09): idle background VOR wired into _check_index_health (cooldown 120s) via server-injected hook set_idle_vor_callback -> layer.run_background_verify (budget=250ms, lo…

  • KI-114 (mscodebase-intelligence) — Open, temperature: stable.

    PRE-EXISTING (found 2026-09-18 during Phase 1 hot-reload): tests/test_lsp_vfs_indexing.py is 8/8 broken - a bare MagicMock().embedding_dim is truthy, so db_writer.py:59 `_target_dim = self.embedder.embedding_dim or 768` gets a MagicMock and truncates vectors to zero -> Zero vector skipping -> empty table -> all asserts fail. Invisible in CI: module is pytestmark=slow, addopts `-m not slow` -> nev…

  • KI-115 (mscodebase-intelligence) — Fixed, temperature: stable.

    F5 judge verdict parser took the first regex match, so a self-corrected verdict (incorrect ... actually correct, final answer correct) was scored inverted. Fixed 2026-09-27 with an explicit-final contract (last JSON match, else last verdict word) plus persisted judge texts and a 7-case guard.

  • KI-116 (mscodebase-intelligence) — Open, temperature: watching.

    The bge-reranker-v2-m3 model scores the true target file below the filter threshold (P3 logit -0.99 maps to 0.271 under 0.3): a model ranking gap that no positive threshold can hold, not a scale bug. Needs a holdout calibration or top-n policy.

  • KI-117 (mscodebase-intelligence) — Open, temperature: watching.

    P2 target file never reaches the final pool, yet a standalone BM25 run ranks it at 0: the loss sits inside pre-rerank pool construction or RRF merging and the root cause is unestablished. Do not assert a cause before dissecting the pool path.

  • KI-118 (mscodebase-intelligence) — Open, temperature: watching.

    Reindex ETA and progress cover only the embedding phase; parse, write, graph, and finalize phases are unmeasured (precedent: ETA 18s against 552s actual). Plan: per-phase history, per-phase models, honest None where no driver exists.

  • KI-119 (mscodebase-intelligence) — Open, temperature: watching.

    IVF finalize hang: the timeout guard cannot fire because shutdown(wait=True) joins the stuck optimize call; the fix abandons the worker via a bounded helper, but the hung native thread still leaks until process restart.

  • KI-120 (mscodebase-intelligence) — Open, temperature: watching.

    Server hard-death under two windows: fixed embedding ports are shared by two MCP servers with no ref-count, so the second server starves the first (ledger shows start without end). Recording is fixed; supervision or dynamic ports remain open.

  • KI-121 (mscodebase-intelligence) — Open, temperature: watching.

    TESTS graph edges are transitive, not tests-about-a-function: a helper carries 234 TESTS edges with zero direct calls, so a specificity filter (direct call or small coverage set) must land before the signal can rank.

  • KI-019 (msp-portfolio) — Open, temperature: stable.

    The 'chore: refresh metrics snapshot [skip ci]' commit no longer lands in the repository: main is protected and rejects the bot's push (expected and non-fatal - update-metrics.ts logs a warning and the deploy continues). Live surfaces are unaffected because the snapshot is written before the build and ships inside dist/: the freshness guard compares the deployed fetchedAt, so the site and the Wor…

  • KI-020 (msp-portfolio) — Open, temperature: watching.

    Only the mscodebase-intelligence corpus is projected into the lab pages; the msp-portfolio log entries live in the docs markdown and stay off the lab. Moving the older records over is postponed until the owner decides the boundary.

For AI agents

This portfolio exposes its content as a production MCP server. Any agent can query the same data behind this page using the standard Streamable HTTP transport.