WEAVATRIX 1.12.0
FIELD NOTES FROM THE EVIDENCE GRAPH

How agent intelligence becomes inspectable infrastructure.

Weavatrix is built in public. This journal explains the decisions behind its graphs, Search Evidence Graph, retrieval engines, memory, proof systems, trust boundaries, and research—along with the limits of every claim.

485measured rename-profile catalog tokens, o200k_base
9,331 → 485fixed MCP context on the published refactor fixture
94.8%catalog reduction; protocol runs stayed 12/12
44 toolsCore operations exist as projections, not as one dump

The catalog is not the context

MCP solved the wiring problem. Treating every attached schema as the working set is how agents run out of window before they read the code.

Compare the boundariesRename profile evidenceweavatrix-graph

The Model Context Protocol won because a coding agent can attach GitHub, a browser, an issue tracker, and a repository graph through one tool surface. The failure mode arrived immediately after the win: operators treated the catalog as the context.

By August 2026 the official MCP registry listed tens of thousands of servers. A public audit of 15,329 reachable remote servers read 140,284 tool descriptions; some instructions strings exceeded 50,000 characters and would sit in the prefix of every turn. Independent agent write-ups measured 141k tokens consumed by 172 tool definitions before a prompt was typed. Industry notes in 2026 treated five to seven connected servers as a practical ceiling. Large consumers publicly cited schema overhead as a reason not to treat MCP as a default internal bus.

The client-side answer is lazy loading. Claude Code's Tool Search defers MCP schemas until a task needs them—published figures put a heavy catalog on the order of 134k tokens down to about 5k. That is a real defense. It is also a client defense. It does not make a 44-tool server small if the session still loads all 44, and it does not stop a single search from dumping a repository into the window.

Weavatrix treats the catalog as a projection of a graph, not as the graph. Core ships 44 read-only operations because the questions differ: orientation, impact, APIs, duplicates, architecture, Git, search, memory. A rename session does not need forty-four. On the published refactor fixture, --profile=rename reduced fixed MCP context from 9,331 to 485 exact o200k_base tokens (94.8%) while keeping the TypeScript, Rust, and Python protocol runs at 12/12. The profile narrows what the model can see; write authority still requires WEAVATRIX_ALLOW_SOURCE_EDITS=1 and a plan-bound token. A smaller catalog is not a permission grant.

The second tax is output. Even a two-tool server can spend the window if every call returns unbounded source. Weavatrix tools take a token_budget. change_impact and context_bundle return a ranked workset with revision, span, extractor, and confidence—not the file tree. weavatrix-graph is why that slice can stay honest: a calls edge is a typed fact, not a cosine. Embeddings and grep remain useful for exploratory recall; they are a different kind of evidence. Weavatrix will not promote a similarity score into a caller list.

Tool Search and profiles compose. One keeps unused servers out of the prefix; the other keeps unused Weavatrix operations out of the product surface; bounded results keep the repository out of the prompt. None of this is a claim that every Weavatrix session costs 485 tokens, or that MCP itself is finished. It is a claim that the working set should be an inspectable slice of a revision-bound graph—and that a catalog which cannot name its budget is not context. It is overhead.

The site is not a crawler report

A crawler tells you what a URL returned. It does not tell you which helper emitted the title, whether a canonical was even in the crawl, or whether a Google rich-result miss is the same thing as a schema.org vocabulary miss. Ranking tools then paint the silence green.

Weavatrix SEO is a Search Evidence Graph. Live HTTP, repository routes, schema, claims, and observations bind to one revision. A finding on a city service page can name the Next.js, Nuxt, or Astro producer. A target outside the crawl stays UNMEASURED. Mixed content is a subresource problem, not a navigation problem. CLI commands and fifteen MCP tools are the same native binary: weavatrix-seo mcp is the agent socket, not a second product.

SEO does not write pages, generate articles, or apply patches. seo_plan can hand a location to Weavatrix Refactor; the write still happens only through Refactor's gates. That is the point of a fourth public product: search evidence without smuggling mutation, and without pretending a crawler dump is architecture.

Weavatrix SEOnpmcrates.io

Heterogeneous inference, measured rather than assumed

Apple Silicon offers CPU cores, a Metal GPU, and an Apple Neural Engine behind Core AI. The attractive story is that a language-model runtime should simply split work across all of them. The harder question is whether a particular split is faster without changing the answer—and whether that result survives an external control.

Weavatrix Hetero treats that question as an evidence problem. On an Apple M4 with Qwen3-0.6B, the admitted exact path routes short C32 work to Metal, C512 prefill to a Core AI M512 graph, C2048 prefill to Core AI SDPA M2048, and decode back to Metal. The compiled FP16 contract matched 26/26 reference tokens before performance could count.

Across 18 exact-prompt pairs against Ollama F16, with six balanced repetitions per context, exact prefill was 5.31% faster at C512 and 6.22% faster at C2048. C32 prefill was 36.34% slower. That loss is part of the result: the router keeps the short path on Metal instead of turning one successful graph into a universal policy.

The experimental full-Q8 path made C512 complete compute 22.03% faster than Ollama F16 in all 6/6 pairs, but it passed only 2/4 strict portable quality cases and did not clear the 20% target against Ollama Q4_K_M. It remains approximate and off by default. Performance PASS is not quality PASS.

The failed ideas matter just as much. Serial per-layer Metal-to-ANE offload amplified numerical drift; a whole Core AI decode step was exact but 1.6× slower than Metal and unstable under mixed placement; unmanaged overlap regressed foreground decode TPOT by +18%. Simulations remain explicitly non-hardware evidence until a physical run earns promotion.

This is the point of Hetero: not to claim that every available processor must be used, but to build a runtime that can say which route won, under which precision contract, on which machine—and refuse the route when any part of that evidence changes.

Read the benchmarks and limitsSee the research track

The repository should remember more than the prompt

A coding agent usually arrives with a short-lived context window and leaves with most of its investigation discarded. The next agent pays for orientation again, and architectural intent is reconstructed from filenames, search results, and confidence.

Weavatrix treats repository understanding as durable evidence instead: typed nodes, relations, exact source spans, revisions, provenance, and bounded memory. A question can begin with a small projection of that graph and expand only when the evidence requires it.

The goal is not to make the model sound more certain. It is to let the agent show what it observed, what remains unproven, and which revision made the answer true.

Explore the enginesMemory source

Three trust boundaries for agent tooling

Understanding code, changing code, and sending evidence across a network are different powers. Packaging them behind one opaque switch makes an agent convenient, but makes its authority hard to inspect.

Weavatrix keeps them separate. Core is source-read-only and network-free. SEO is a read-only Search Evidence Graph over the live origins you name plus the repository. Refactor adds reviewed local mutation with preview, confirmation, journaling, and rollback. Online owns the explicitly authorized cloud network boundary without silently bundling the others.

The separation costs a little more explanation. It also means installation, context size, and runtime configuration cannot quietly grant a capability the user never intended.

Compare the boundariesRead about SEORead about Refactor

Fast, with the scope attached

A benchmark without its fixture, platform, cache state, output contract, and losing cases is marketing—not evidence. Weavatrix publishes the boundary around its performance claims so agents and people can decide whether the comparison applies.

The resident search index is dramatically faster than starting a fresh ripgrep process on one disclosed Windows fixture; on a disclosed Ubuntu fixture, ripgrep wins. Git and filesystem results are published with exact parity contracts and reproducible sources rather than promoted as universal rankings.

Speed matters because bounded retrieval protects context and latency. Honesty about scope matters because an agent must never turn one favorable measurement into a claim about every repository.

See measured proofReproduce search evidence