Skip to content

Latest commit

 

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Status Inspired by Node.js ESM License

📚 Knowledge Forge

A persistent, compounding knowledge base maintained by LLMs.
Drop sources in. Watch a wiki build itself.

Quick Start · How It Works · Commands · Architecture · Roadmap

Knowledge Forge Web UI - Home
Dark-themed web UI with sidebar, type filters, search, and wiki link navigation

Source page with wiki links
Source pages auto-extract concepts and entities with clickable wiki links

Concept page
Concept pages accumulate cross-references from multiple sources


Inspired by Andrej Karpathy's LLM Wiki pattern.

"Instead of just retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki — a structured, interlinked collection of markdown files that sits between you and the raw sources." — Andrej Karpathy

What It Does

Knowledge Forge takes raw documents and turns them into a living, interconnected wiki. Not a one-shot RAG pipeline — a compounding knowledge base that gets richer with every source you feed it.

  • 📥 Ingest markdown/text/PDF/DOCX sources → semantically extracts grounded summaries, categories, key points, conclusions, recommendations, quotes, concepts, entities, questions, and dates
  • 🧾 OCRs scanned PDFs when pdftotext finds no embedded text
  • ♻️ Skips duplicates with a SHA-256 ingestion manifest
  • 🔗 Links related pages together with wiki-style [[links]]
  • 🗓️ Builds a source-linked timeline from relevant dates
  • 📋 Indexes everything into a navigable catalog
  • 🔍 Lints the wiki: finds orphans, dangling links, missing metadata
  • 🌐 Serves a dark-themed web UI to browse and explore
  • 💬 Queries the compiled wiki in natural language with mandatory wiki + raw-source citations
  • 🔌 Exposes MCP so coding agents can list, search, read, navigate, and ingest the wiki
  • 📝 Logs every operation chronologically

Current Status

This repo is intentionally positioned as a functional concept implementation.

That means it already proves the end-to-end pattern:

  • raw sources → wiki pages
  • cross-linking between pages
  • persistent markdown artifact
  • index + log
  • browseable UI
  • health checks / linting

But it does not yet implement the full autonomous LLM maintainer vision described by Karpathy.

What is already real

  • A working ingestion pipeline
  • Persistent wiki generation on disk
  • Concept and entity page creation
  • Incremental wiki updates from new sources
  • A usable local web UI
  • A concrete repo anyone can clone, run, and extend

What is still missing

  • Contradiction handling
    • The current version does not yet detect or annotate conflicts between sources
  • Human-in-the-loop workflows
    • No review queue, approval flow, or source triage loop yet
  • Richer search / retrieval
    • No BM25/vector search yet, only file-based navigation and simple UI filtering
  • Autonomous maintenance loop
    • No background agent that continuously ingests, revises, and improves the wiki over time

So the right framing is:

Knowledge Forge is a functional prototype of the LLM Wiki pattern, with the core architecture working today and the full LLM-native maintainer loop left as the next step.

Why Not Just RAG?

RAG Knowledge Forge
Knowledge Re-derived every query Compiled once, kept current
Cross-references Missing Built-in [[wiki links]]
Contradictions Undetected Not implemented yet
Accumulation None — each query is independent Compounds with every source
Maintenance cost Low (but shallow) Near zero (LLM does the bookkeeping)

Quick Start

git clone https://github.com/ESJavadex/knowledge-forge.git
cd knowledge-forge
npm install
npm run demo        # bootstrap + 3 sample sources
npm start           # launch web UI at http://localhost:3000

Open http://localhost:3000 and browse the wiki. The sidebar lets you filter by type, search pages, and navigate through wiki links.

Commands

node src/cli.js init              # Create folder structure + special files
node src/cli.js demo              # Create 3 sample sources and ingest them
node src/cli.js ingest <path>     # Ingest one file or every supported file in a directory
node src/cli.js ingest --all      # Ingest raw/ recursively; unchanged hashes are skipped
node src/cli.js ingest <path> --force # Reprocess even when the source hash is unchanged
node src/cli.js ingest <path> --provider openclaw --model zai/glm-5.3-flash --require-llm
                                  # Use the exact Z.AI model configured in OpenClaw; fail closed
node src/cli.js query "<question>" # Focused answer from wiki knowledge and save the cited analysis
node src/cli.js lint              # Health-check: orphans, dangling links, metadata
node src/cli.js serve             # Start the web UI (port 3000)

Or via npm scripts:

npm run init
npm run demo
npm run ingest
npm run lint
npm run mcp
npm start

Model-provider configuration

Semantic extraction supports two adapters:

  • --provider openclaw uses OpenClaw's stored provider authentication through a lean, isolated infer model run call. It sends no agent history, tools, memory, or workspace context. The recommended podcast pipeline uses zai/glm-5.3-flash through this adapter.
  • --provider openrouter uses the existing OpenRouter adapter when explicitly selected or when OPENROUTER_API_KEY is present.

Rich ingestion extracts a grounded summary, categories, key points, concepts, entities, conclusions, source-attributed recommendations, notable verbatim quotes, open questions, and relevant dates. Use --require-llm for production jobs where a heuristic fallback would be unacceptable.

For OpenClaw, authenticate and allow the selected model in OpenClaw itself; no provider credential is copied into this repository. For OpenRouter, provide OPENROUTER_API_KEY through your environment or secret manager. OPENROUTER_MODEL (or LLM_MODEL) optionally selects its model.

Without an LLM provider, ingestion can still use the original frequency/bigram heuristic, so init, demo, ingest, lint, mcp, and serve keep working offline. Query mode defaults to the lean OpenClaw adapter and zai/glm-5.3-flash on hosts where OpenClaw is configured.

When OpenRouter is enabled, source excerpts are sent to the selected model provider. Check that provider's privacy terms before ingesting sensitive personal documents.

PDF ingestion uses pdftotext (Poppler); if the PDF has no embedded text it falls back to pdftoppm + Tesseract OCR. DOCX uses unzip to read word/document.xml. None of these paths modifies the file in raw/.

On Debian/Ubuntu, the optional OCR runtime can be installed with:

sudo apt install poppler-utils unzip tesseract-ocr tesseract-ocr-spa tesseract-ocr-eng

OCR settings:

  • OCR_ENABLED=false disables OCR fallback.
  • OCR_LANGUAGES defaults to spa+eng.
  • OCR_DPI defaults to 200.

How It Works

1. Ingest

Drop a .md, .txt, .pdf, or .docx file into raw/ and run ingest. The engine:

  1. Hashes the immutable source and skips it when the same content was already ingested
  2. Reads embedded text or performs OCR, preserving page/section/paragraph locators
  3. Splits long documents on paragraph boundaries and extracts every part—no middle truncation
  4. Extracts grounded summaries, categories, key points, conclusions, source-attributed recommendations, exact quotes, open questions, concepts, entities, and dates with the configured model
  5. Creates/updates source, concept, and entity pages without changing [[wiki links]]
  6. Stores generated citation evidence under wiki/.evidence/
  7. Refreshes wiki/timeline.md, the index, manifest, and append-only log

A single source can touch 20+ wiki pages.

If OpenRouter is not configured or returns malformed structured output, ingestion falls back defensively to the original heuristic extractor.

2. Query

Use node src/cli.js query "..." or the query field in the web UI. The web UI defaults to Deep mode and also offers Fast mode.

Deep mode uses citation-preserving hierarchical RAG:

  1. Retrieve up to 20 source pages with hybrid lexical + semantic search.
  2. Select several relevant evidence chunks from each source, using the semantic result snippet as an additional chunk-ranking hint.
  3. Split the evidence into byte-bounded batches and extract atomic findings from up to three batches concurrently.
  4. Validate every finding against its exact wiki page, raw source, locator, and verbatim quote.
  5. Build an evidence ledger and run a second synthesis pass that may connect and compare only those validated findings.
  6. Expand evidence IDs back into full citations and reject unsupported quantities or unknown evidence references.

The resulting answer can contain several independently cited conclusions from the same Markdown file and is organized into useful sections such as direct answer, conclusions, actions, cross-source consensus, disagreements, and caveats. Every displayed item has one or more citations. If the final synthesis fails validation, Knowledge Forge falls back to the already validated map-stage findings instead of returning an uncited answer. Results are saved under wiki/analyses/, linked to their supporting pages, indexed, and logged.

Fast mode performs one bounded generation call over the best retrieved sources. It is useful when latency matters more than exhaustive coverage.

Hybrid search

Knowledge Forge can use a local QMD sidecar to combine BM25/full-text and semantic vector retrieval. Both the web query flow and MCP use the same adapter. Only wiki/sources/**/*.md should be indexed; navigation pages and generated analyses must stay out of the search corpus.

npm install --global @tobilu/qmd
export XDG_CONFIG_HOME="$HOME/.config/knowledge-forge-qmd"
export XDG_CACHE_HOME="$HOME/.cache/knowledge-forge-qmd"
qmd collection add "$PWD/wiki" --name forge --mask 'sources/**/*.md'
qmd update
qmd embed -c forge
qmd mcp --http --host 127.0.0.1 --port 8181

The endpoint is local-only and unauthenticated, so it must not be exposed directly to the network. Knowledge Forge connects to http://127.0.0.1:8181 by default. Override it with KNOWLEDGE_FORGE_QMD_URL, or set that variable to off to force deterministic lexical fallback. KNOWLEDGE_FORGE_QMD_REQUIRED=true makes queries fail instead of falling back when the sidecar is unavailable.

3. MCP for coding agents

Run the local stdio server with npm run mcp. Configure your coding agent with:

{
  "mcpServers": {
    "knowledge-forge": {
      "command": "node",
      "args": ["/absolute/path/to/knowledge-forge/src/mcp-server.js"]
    }
  }
}

The server exposes eight focused tools:

  • wiki_list — catalog pages by type.
  • wiki_search — deterministic text search.
  • wiki_context — compact, ranked context bundle with raw provenance for agent prompts.
  • wiki_facets — counts by type, category, podcast, and extraction model.
  • wiki_status — ingestion/model/schema status.
  • wiki_read — markdown, frontmatter, links, and raw provenance.
  • wiki_links — outgoing links and backlinks.
  • wiki_ingest — optional write tool for files/directories already under raw/, using grounded zai/glm-5.3-flash extraction. It is hidden and rejected by default; enable it only for a trusted client with KNOWLEDGE_FORGE_MCP_ALLOW_INGEST=true.

It also exposes wiki://page/{slug}, wiki://catalog, and wiki://status resources plus a grounded_wiki_research prompt. Direct arbitrary reads or writes to raw/ are intentionally not exposed.

Register the read-only server with common local clients:

openclaw mcp add knowledge-forge --command /usr/bin/node \
  --arg /absolute/path/to/knowledge-forge/src/mcp-server.js \
  --include wiki_list,wiki_search,wiki_context,wiki_facets,wiki_status,wiki_read,wiki_links \
  --parallel --approval auto

codex mcp add knowledge-forge -- \
  /usr/bin/node /absolute/path/to/knowledge-forge/src/mcp-server.js

claude mcp add --scope user knowledge-forge -- \
  /usr/bin/node /absolute/path/to/knowledge-forge/src/mcp-server.js

The MCP defaults to read-only even when a client has no tool-filter support. Set KNOWLEDGE_FORGE_MCP_ALLOW_INGEST=true only for a trusted ingestion client.

4. Lint

Run a health check to find:

  • 👻 Orphan pages — no other page links to them
  • 🔗 Dangling links[[links]] to pages that don't exist yet
  • 📋 Missing frontmatter — pages without YAML metadata

Architecture

knowledge-forge/
├── raw/                    # 📥 Immutable source documents (never modified)
│   └── *.md
├── wiki/                   # 📚 LLM-generated knowledge base
│   ├── sources/            # Summary pages for each ingested source
│   ├── concepts/           # Recurring themes and topics
│   ├── entities/           # Named things, tools, products
│   ├── analyses/           # Synthesized answers (user queries filed back)
│   ├── .evidence/          # Generated exact excerpts + page/section locators
│   ├── .ingest-manifest.json # SHA-256 deduplication and extraction metadata
│   ├── timeline.md         # Generated dated-event view
│   ├── index.md            # Catalog of all pages
│   └── log.md              # Append-only chronological record
├── schema/
│   └── AGENTS.md           # Rules for the wiki maintainer agent
├── src/
│   ├── cli.js              # CLI entry point
│   ├── ingest.js           # Source ingestion + extraction engine
│   ├── extraction.js       # Semantic extraction use case + heuristic fallback
│   ├── query.js            # Grounded query use case + citation validation
│   ├── mcp-server.js       # Local stdio MCP adapter
│   ├── wiki-reader.js      # Deterministic wiki navigation use cases
│   ├── ingest-state.js     # Manifest, evidence, and timeline persistence
│   ├── adapters/           # OpenRouter and source-reading adapters
│   ├── lint.js             # Wiki health checker
│   ├── server.js           # Express web UI + API
│   └── utils.js            # Shared utilities
├── public/
│   └── index.html          # Single-page web UI
└── package.json

Three Layers

  1. Raw sources — Your curated documents. Immutable. The LLM reads from them but never writes to them.
  2. The wiki — Structured markdown pages maintained entirely by the LLM. Source summaries, concept pages, entity pages, cross-references.
  3. The schema — Configuration (AGENTS.md) that tells the LLM how to structure, maintain, and evolve the wiki.

Wiki Link Format

Pages reference each other with Obsidian-style [[Page Name]] links. The web UI resolves these into clickable navigation. Dangling links (to pages that don't exist yet) are marked with ❓.

Web UI

The built-in UI features:

  • 🌙 Dark theme
  • 📂 Sidebar with type filters (Sources, Concepts, Entities, Analyses)
  • 🔍 Full-text search across all pages
  • 📊 Stats bar showing page counts by type
  • 🔗 SPA navigation through wiki links
  • 📱 Responsive layout

Tech Stack

  • Runtime: Node.js (ESM)
  • Server: Express.js
  • Markdown: marked (rendering) + gray-matter (frontmatter parsing)
  • UI: Vanilla HTML/CSS/JS — zero build step
  • VCS: Git (your wiki is a git repo with full history)

Demo Sources Included

Source Concepts Entities
Transformer Architecture 10 10
Retrieval-Augmented Generation 10 10
Knowledge Graphs in AI 10 10

Run npm run demo to generate all of them.

Roadmap

  • LLM-powered extraction — Structured semantic extraction through OpenRouter with offline heuristic fallback
  • Full-text search API — Integrate qmd or similar for proper search as the wiki grows
  • Query mode — Ask natural language questions and get grounded answers with wiki + raw citations
  • File-and-save — File query answers back into the wiki as linked analysis pages
  • Batch + deduplication — Recursive ingestion with SHA-256 skip logic
  • OCR fallback — Scanned PDF extraction with page provenance
  • Precise citations — Wiki + raw + locator + verbatim quote validation
  • Timeline — Persist extracted dates in an automatically linked page
  • Long-document chunking — Process all sections without middle truncation
  • MCP server — Let coding agents browse and ingest the wiki over stdio
  • Obsidian compatibility — Open the wiki folder directly in Obsidian for graph view
  • Marp export — Generate slide decks from wiki content
  • Dataview queries — YAML frontmatter + Dataview plugin integration
  • Contradiction detection — Flag when new sources contradict existing wiki claims
  • Web clipper helper — Easy ingestion from browser extensions
  • Continuous maintainer mode — Background agent loop for ingest, refinement, and linting
  • Review workflows — Human approval mode for team/internal knowledge bases

Author

Javier Santos
javadex.es · GitHub

Head of AI · Electronic Engineer · Building the future, one repo at a time.


License

MIT — use it, fork it, build on top of it.


Built with ☕ by Javier Santos · The business version of this pattern · Inspired by Andrej Karpathy

About

Persistent, compounding knowledge base maintained by LLMs. Inspired by Andrej Karpathy's LLM Wiki pattern — ingest sources, auto-generate an interlinked wiki, query with citations, and let knowledge compound over time.

Resources

Stars

33 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages