414 offline knowledge archives. πΈ Wikipedia, Khan Academy, StackOverflow, medical references, survival guides, school curricula. 54 AI models ready to search, understand, and explain any of it. No internet required. One machine. Open knowledge. Preserved for when it matters most.
Because the internet is not permanent. Services shut down. Content gets censored. Paywalls appear where free access used to be. Algorithms decide what you see. One day the infrastructure works β the next day it doesn't.
I'm an engineer from a village apartment in IaΘi County, Romania. I don't control the internet, the laws, or the infrastructure between my home and the world's knowledge. But I can control what sits on my own machine.
Everything you see on this page is open-licensed, free, and legal to preserve. Wikipedia is Creative Commons. Khan Academy is free education. Project Gutenberg is public domain. StackOverflow is CC-BY-SA. These are humanity's gifts to itself β and they deserve a backup.
This is not hoarding. This is digital preparedness.
Every archive below is served through Kiwix, the offline knowledge platform. Each ZIM file is a self-contained, searchable snapshot of an entire website or knowledge base.
Beyond Kiwix, a custom-built document processing pipeline ingests, chunks, embeds, and indexes documents for semantic search. Not just keyword matching β actual understanding.
PDFs, EPUBs, office documents β crawled, parsed, and stored in Elasticsearch. Full metadata, full text, searchable in milliseconds.
Each document split into overlapping chunks with vector embeddings. Semantic search finds meaning, not just words.
Same documents, multilingual embeddings. Search in Romanian, find results in English. Or vice versa.
Documents organized into 35,040 logical folders. Browse by topic, author, or subject area β like a real library.
Local AI models running on an RTX 4060 Ti β no cloud APIs, no subscriptions, no data leaving the machine. Ask a question about any document, any archive, any topic. Get an answer from a model that runs in your apartment.
Gemma 4, Qwen 3, GLM-4, DeepSeek. Multiple sizes from 1B to 27B parameters. Fast answers or deep reasoning β pick your trade-off.
DeepSeek Coder V2 (up to 236B parameters), Qwen Coder. Write, debug, and explain code β in any language, offline.
Custom-built models (lsn-tools-*) fine-tuned for infrastructure management, function calling, and homelab operations.
LLaVA, Llama Vision. Describe images, read text from photos, analyze diagrams. Eyes for the brain.
BGE, Nomic, MxBai embedding models + BGE Reranker. The invisible workers that make semantic search actually work.
Uncensored variants, music analysis, translation, summarization. Every niche covered β because you never know what you'll need.
One Dell Precision Tower 7910 β 88 threads πΈ, 503 GB RAM πΈ, RTX 4060 Ti πΈ. One Synology NAS with 58 TB. Everything running on K3S with 154 pods across 58 namespaces.
Elasticsearch index (304 GB) + hot AI models. The fast lane β NVMe speeds for instant search results.
AI model library (464 GB) + active data. The workhorse for model loading and inference.
Kiwix archives, books, documents, backups. 36 TB used, 23 TB free. NFS-mounted, shared across the cluster.
RAG cache, vector databases, hot data. Because even NVMe isn't fast enough for some workloads.
Censorship. Infrastructure failure. Natural disaster. War. Whatever the reason β this machine keeps running.
The entire English Wikipedia β every article, every image. Khan Academy β every lesson from arithmetic to calculus. StackOverflow β every answered programming question. Medical references β drug interactions, first aid, disease info. Survival guides β from military field manuals to self-reliance skills. 70,000+ books from Project Gutenberg. School curricula for children's education. And 54 AI models to help make sense of all of it.
No cloud. No API keys. No subscription. No permission needed. Just electricity and knowledge.
Every archive on this page uses open licenses β Creative Commons, public domain, or explicitly free to distribute. This is not piracy. This is preservation. These are humanity's gifts to itself, and they deserve a local copy.
This page β like the infrastructure it documents β was built through a conversation between a human and an AI. The data estate was audited, cataloged, and presented by Copilot CLI running on the same machine that hosts all this knowledge.
The irony isn't lost on me: an AI helped document a knowledge vault designed to survive without AI. But that's the point β the knowledge is the foundation, the AI is just a lens.