SeekEngine started from a frustration with a very specific failure mode: LLMs are excellent at sounding right and structurally incapable of knowing anything about the present, because they are frozen in their training data. Ask one for today's stock price and you will get a confident, fluent, completely fabricated number. I wanted to see whether that could be fixed by treating hallucination less as an AI problem and more as a systems problem — a model is an isolated node cut off from external state, and when it lacks real data, fluency fills the gap.
The fix I built wires together two very different upstream providers:
Google Custom Search Engine (CSE) — sparse, real, timestamped snippets from the live web;
OpenRouter — structured synthesis that is fluent and frequently fabricated when context is thin.
The two run in parallel and their outputs are merged through a fusion layer that treats factual grounding as a hard constraint. The result is not a chatbot. It is a search agent that cites sources, marks its own uncertainty, and answers "I don't know" when it genuinely has nothing.
Everything was built on free-tier APIs with no budget and no infrastructure I controlled. The system failed constantly — rate limits, dropped requests, stale results, timing mismatches. Those failures turned out to be the most useful part of the project: each one was the architecture telling me what it needed to handle.
The core claim is deliberately modest: hallucination behaves like a coordination problem. Ground the model with retrieval, verify its output against sources, and accept that accuracy has a cost — latency, bandwidth, and occasionally silence.
The usual explanations for hallucination — bad training data, weak prompting, the wrong model architecture — never matched what I saw. I tried all of those angles early and none of them was the real issue. What mattered was isolation.
An LLM generates tokens from internal priors. Search requires external state. Cut a model off from the live web and you are asking an offline node to answer time-sensitive questions from a stale snapshot of the world. In that frame, hallucination is not a bug; it is a fallback policy — when there is no data, fluency fills the void.
That observation set the project's question:
Can retrieval act as a “bootstrap node” for grounding — turning hallucination into a coordination problem rather than a probability problem?
Framing it that way pulled the project away from AI UX and toward systems thinking. The phenomena I kept hitting mapped surprisingly well onto concepts from peer-to-peer networks — and I want to be clear from the start that this is an analogy I found useful, not a claim that search engines and BitTorrent swarms are the same system. The mapping is what made the problems tractable:
LLM Search Problem
Distributed Analogy
Hallucination
Unverified piece
Retrieval
Bootstrap node
Fusion
Swarm coordination
Source citation
Piece hashing
XSS + Prompt Injection
Peer poisoning
Latency
Consistency cost
Rate limits
Network congestion
Timeout
Silent peer drop
Provider mismatch
Protocol incompatibility
Truth penalty
Distributed coordination overhead
Once I stopped treating hallucination as something to be trained away and started treating it as something to be coordinated around, the problem became tractable with tools I could actually afford.
#II. Independent Research Positioning
SeekEngine was built as a zero-budget, zero-infrastructure experiment on the open web, with no privileged access to datasets, model weights, proprietary APIs, or academic compute. I want to be direct about why that constraint appears so often in this paper: it is not a hardship narrative. Every resource I did not have forced an architectural decision that institutional projects can defer — and each of those decisions surfaced a failure mode I would otherwise have missed.
No budget meant free-tier APIs, which meant living with rate limits and instability from day one.
No private infrastructure meant a hard client/server separation, which is how credential leakage got found before deployment.
No vector database meant dynamic retrieval on every query, which exposed retrieval-starvation behavior immediately.
No observability tooling meant terminal-level logging, which made latency patterns impossible to ignore.
No protected sandbox meant the open web as the environment, which forced the threat modeling early.
Stripping away resources did not make the project harder. It made the system's real behavior visible.
#III. Research Claim (Soft)
Let me be precise about what this work does and does not claim. SeekEngine is not superior to industrial RAG pipelines, and it does not "fix" hallucination. The claims are narrower, and each one is testable:
Hallucination behaves like a coordination failure.Grounding reduces to verification.Verification has a real cost.
That cost — latency, bandwidth, occasional silence — is the limiting factor, not model creativity. Most chatbot UIs hide uncertainty behind confidence veneers; SeekEngine shows its sources and its reasoning. Most inference pipelines treat latency as a defect to optimize away; here, slower grounded answers are the honest price of not making things up. The design goal was inspectability over eloquence, and the UI decisions in this paper all follow from that.
#IV. Phase 1 — Grounding the Problem
The starting observation was simple, and it is architectural rather than moral: models are exceptionally good at sounding correct and structurally incapable of knowing whether their claims match reality, because they predict tokens from internal priors rather than from the current web. Ask a stateful question — “what is AAPL trading at right now?” — and the model manufactures a plausible number. It is not lying; it is falling back, because it has no external state.
My first framing was naive: “attach a search API and be done.” It took one prototype to see that retrieval was not an enrichment layer bolted onto the side of the model. It was the thing that gave the model a connection to reality at all — the bootstrap node, in the P2P vocabulary I was already reaching for. Without it, the model is a sealed container guessing about the world. With it, the system becomes several heterogeneous components that have to merge partial, noisy information under latency pressure.
That shift had practical consequences I kept rediscovering, so I wrote them down as working rules:
Grounding is never free — it costs latency, bandwidth, or structural complexity.
The model is a node, not an oracle — its output has to be reconciled with everything else.
Grounding beats creativity — in search, inventiveness is a liability.
The UI must show uncertainty — presenting a guess as a fact is the original sin of this whole category of tool.
#V. Phase 2 — Retrieval as Bootstrap
The BitTorrent framing deserves an honest introduction: it is an analogy, and it earned its keep. In a P2P swarm, trackers and DHT nodes are the entry point — without a bootstrap node, a peer has no swarm to join and no metadata to resolve. Retrieval plays the same role for grounded inference, and the parallels kept surfacing as I built.
I treated Google CSE as the bootstrap node. It is a frustrating provider in exactly the ways the analogy predicts: sparse, usually authoritative enough, rate-limited, non-deterministic, prone to silent drops, and adversarial at its input boundary because it indexes the whole web. Its output also resembles peer metadata more than it resembles "search results":
titles → strong anchors
snippets → partial truths
URLs → provenance
timestamps → freshness
keywords → weak alignment
ranking → heuristics, not truth
Retrieval was not the answer to hallucination. It was the substrate that made answers possible at all.
At this point the architecture formalized into a bootstrap graph:
System_Terminal
Fig. — Query Processing Pipeline
In practice this graph behaved more like:
System_Terminal
Fig. — Query Processing Pipeline
To actually watch this behavior instead of guessing about it, I built a diagnostic terminal:
System_Terminal
Fig. — Live Orchestration Diagnostics
That terminal was not decoration. It streamed the system's raw dynamics — partial information arriving, latency mounting, context missing — and it changed how I debugged everything that came after. Retrieval stopped being an API call I made and became the load-bearing part of the whole design.
#VI. Phase 3 — Parallel Orchestration & Fusion
With retrieval working, I added the second upstream node: OpenRouter as an inference relay. Now the coordination problem appeared in full: retrieval produced grounded but brittle context, inference produced fluent but ungrounded synthesis, and the two had to be merged across mismatched levels of granularity and freshness.
I tried sequential execution first:
CSE → LLM
Sequential execution — CSE first, then the LLM — produced correct facts wrapped in brittle structure: citation noise, repetitive summarization, low coherence. Running the two in parallel changed the character of the system:
CSE || LLM → Fusion Layer
Under parallel execution, ordinary failures started behaving like distributed-system failures: latency became something to negotiate, timeouts became partial failures, rate limits became congestion, and a starved retrieval meant a grounding deficit while a starved model meant a synthesis deficit. The fusion step stopped being a merge function and became a protocol in its own right — with rules.
In practice, the fusion layer was forced to operate under three constraints:
Truthfulness Constraint
Fused answers must be grounded or fail silent.
Minimality Constraint
Synthesis must be brief; verbosity dilutes claims.
Inspectability Constraint
Sources must be traceable.
The decision to surface citations and grounded snippets in the UI was not an aesthetic one. In this architecture it is a protocol requirement: the user has to be able to trace an answer back to what actually grounded it, or the verification step is theater.
To make the fusion layer visible rather than hypothetical, I built:
Next.js 14 Client
Responsive Client View
Security Proxy
API Route Handler
ORCHESTRATOR
GET /customsearchGoogle-Indexing
POST /completionOpenRouter-LLM
Data Provider A
Google CSE
Inference Provider B
OpenRouter Cloud
End-to-End Environment Encapsulation
Fig. — Core Architecture Overview
And operationally evaluated latency using:
Response Latency (ms)
Google CSE300ms
Direct LLM (OpenRouter)1200ms
SeekEngine Hybrid1500ms
The "Truth Penalty": SeekEngine trades additional latency for improved factual consistency.
Fig. — Latency Comparison Benchmark
Where a BitTorrent client pays bandwidth and time for piece verification, SeekEngine pays latency for source verification. The tradeoff is structural: grounded answers are slower, and that slowness is the point rather than the bug — the UI just has to be honest about it.
#VII. Phase 4 — Verification as Protocol
Phase 4 made grounding explicit rather than implicit. Instead of hoping the model stayed close to the retrieved context, I defined verification as a protocol with four gates that every answer must pass:
Existence Gate
Does the answer reference any retrieved sources?
Consistency Gate
Do claims align with retrieved snippets?
Temporal Gate
Are claims time-sensitive and stale?
Source Gate
Are sources adversarial or low-quality?
An answer is only synthesized after all four gates pass. Anything else is guesswork wearing a citation.
To make the verification dynamics concrete, I turned an earlier demo into a truth-vs-hallucination comparator:
Hallucination Detected
"The current stock price of Apple is $245.30, showing a strong 2% growth since this morning's opening..."
(Note: LLM is using training data from 2024 to guess 2026 prices)
Fig. — Grounded vs Ungrounded Response
In the small side-by-side comparisons I ran during development, the pattern was consistent: ungrounded inference was fluent and unreliable; grounded inference was more awkward and more often correct. The tradeoff showed up in more than latency:
grounding reduces eloquence
verification adds friction
citations expose uncertainty
sometimes the honest answer is silence
The UX consequence is one people keep tripping over: the most truthful answer is not always the most polished one. And the phase crystallized what had been nagging at me since the start — hallucination is best understood as verification failure under isolation, which is a much more tractable problem than "make the model stop being wrong."
#VIII. Phase 5 — Security & Adversarial Surface
Fusing retrieval and inference created a two-front security surface, and it took me a while to see both fronts clearly:
External adversaries — everything hostile on the open web, which is most of it: SEO poisoning, spam, XSS payloads, tracker pixels, misleading snippets, prompt-injection triggers, content farms, and stale pages dressed up as authoritative.
Internal adversaries — the LLM itself, which will hallucinate confidently, fabricate citations, and overstate its own certainty whenever context runs thin.
SeekEngine has no malicious peers in the BitTorrent sense, but it has malicious inputs on both sides of the pipe. That asymmetry shaped everything in this phase:
I adopted a zero-trust stance toward both fronts: nothing from the web and nothing from the model is trusted until it has been checked.
Retrieval Threats
Retrieval responses went through a sanitizer before anything else touched them. The scrub list was built from what the web actually sent back:
XSS
embedded scripts
base64 payloads
trackers
HTML contamination
inline injection primitives
malware URL signatures
The implementation sanitizes for all of the above:
Search Result Scrubbing
<script>alert(1)</script>
Tracking_pixel.gif
Verified Text Content only
ZOD VALIDATION
schema.parse(raw_api_response)
Fig. — Retrieval Sanitization
That step looks like hygiene; it is defense. The web returns scripts, trackers, and payloads inside what should be plain text, and the fusion layer has no business ever rendering any of it.
Inference Threats
The model received the same treatment: a potentially adversarial subsystem with its own failure modes — unsanctioned creativity, miscalibration, citation forgery, temporal guesswork, sentiment where none was asked for, and source attribution that does not survive contact with the actual page. Those risks need protocol-level guardrails, not UX hints:
unsanctioned creativity
miscalibration
citation forgery
temporal guesswork
sentimental phrasing
source attribution fakery
Boundary Security
The risk I did not anticipate was credential exposure, and it emerged exactly where the P2P analogy said it would — at a trust boundary. Both providers required API keys; the inference key was the higher-privilege one. Early prototypes shipped it inside client bundles, which is how I learned, the embarrassing way, that anything in a browser bundle is public. That forced a redesign: the execution boundary moved server-side, and only server-only handlers touch provider keys.
This surfaced the first formal trust boundary:
Client —(untrusted)→ Server —(trusted)→ Provider
To make that boundary visible in the product, I built:
Environment Encapsulation
client_side.js
const API_KEY = "sk-..." // LEAK DETECTED
server_action.ts
process.env.OPENROUTER_KEY // ENCAPSULATED
Auth Integrity: 100%
Fig. — Security Control Layers
Threat Matrix
The full set of threat classes is consolidated in the matrix below — it reads more like a web-security threat model than a typical RAG paper's, because the attack surface really is the open web:
Threat Model & Mitigations
XSS Injection
mitigated
DOMPurify sanitization
API Key Leakage
mitigated
Server-side encapsulation
Prompt Injection
partial
Input filtering (basic)
Data Persistence
mitigated
Request-scope only
Upstream Compromise
unaddressed
Outside control
Model-Level Exploits
unaddressed
Future work
MITIGATED
PARTIAL
UNADDRESSED
Fig. — Adversarial Surface Matrix
#IX. Phase 6 — Observability & Diagnostics
Once the boundaries were secure, the bottleneck moved somewhere duller and harder: I could not see what the system was doing. Failures inside the fusion layer produced no errors — they were silent, partial, or timing-based, the same class of failure the P2P literature warned about. The symptoms were things like:
retrieval starvation
inference starvation
fusion race conditions
inconsistent snippet alignment
snippet truncation
stale web results
inference guesswork
non-deterministic formatting
latency variance spikes
The fix was a diagnostic terminal that streamed the orchestration process live. It did not look like research instrumentation; it was a debug console with better posture. But that is exactly what made it work:
System_Terminal
Fig. — Live Orchestration Diagnostics
Logs had hidden these failures for weeks. The terminal made them visible: latency became structure you could watch, silence became an event instead of an absence, and degraded paths showed up as paths. The single most valuable lesson of the whole project came out of this: the absence of errors was not success — it was silent fallback, the model quietly making things up because retrieval had starved. That distinction is second nature to anyone who has operated a distributed system and nearly invisible in most AI tooling.
Observability did not make the system faster. It made the system debuggable, which is the difference between a black box and a tool.
#X. Phase 7 — Partial Failures & Silent Errors
The defining behavior of distributed systems is not crashing; it is partial failure. SeekEngine exhibited the same classes of partial failure documented in P2P and cloud systems — BitTorrent swarms, DHT tables, gossip networks, weakly consistent caches:
BitTorrent swarms
DHT peer tables
gossip networks
cloud orchestration
weakly-consistent caching systems
Failure modes included:
(a) Retrieval Starvation
CSE occasionally returned empty or stale results. The LLM compensated by fabricating plausible answers. Bootstrap failure → hallucination.
(b) Inference Starvation
OpenRouter occasionally dropped or rate-limited requests. Retrieval produced raw snippets with no synthesis. Bootstrap success → no swarm coordination.
(c) Timing Desynchronization
Parallel requests resolved in inconsistent orders. Fusion layer misaligned context and generated broken synthesis.
(d) Rate-Limit Oscillation
LLM response times oscillated under multi-query load, creating weird latency cliffs.
(e) Provider Mismatch
CSE timestamps mismatched OpenRouter’s training cutoff, producing temporal inconsistency (new vs stale knowledge).
(f) Trust Misalignment
High-ranking snippets were low-quality (SEO spam), while lower-ranked snippets were authoritative (primary sources). Retrieval ≠ trust.
These surfaced in the Limitations Matrix:
Known Limitations Matrix
No Standardized BenchmarksEvaluation
Internal testing only
high impact
Third-Party DependencyReliability
Google CSE, OpenRouter availability
medium impact
Multilingual SupportCoverage
English-primary implementation
medium impact
Temporal ConsistencyAccuracy
Real-time data freshness varies
high impact
Rate LimitingScale
Free-tier constraints
low impact
Honest assessment: These limitations are documented, not hidden.
Fig. — Limitations Assessment
SeekEngine almost never crashed. It degraded — and before the observability work, the degradation was silent. That was the finding: the interesting failures were the ones that left the process running and the answers wrong.
#XI. Phase 8 — Lessons from the System
By the time the system stabilized, it had stopped being an AI demo in my head and become a coordination problem operating across three domains:
(1) The Web as Information Substrate
→ sparse, adversarial, timestamped, unstructured
(2) The LLM as Synthesis Machine
→ structured, fluent, hallucination-prone, stochastic
(3) The UI as Epistemic Interface
→ mediates uncertainty, verification, and trust
The lessons that stuck came from the boundaries between those domains, so I wrote them down as plain claims rather than aphorisms:
Lesson 1 — retrieval alone cannot answer. Snippets are not answers; they are evidence, and evidence still has to be weighed.
Lesson 2 — verification is not free. The cost of checking an answer is real, and it reshapes the architecture — which is why the control-plane work and the UI work ended up coupled.
Lesson 3 — trust is a UI problem. A system can do everything right internally and still be untrustworthy if the interface cannot show its work.
Lesson 4 — cheap systems teach the most. With no budget to hide behind, every failure was visible, and visible failures are the best teacher this project had.
#XII. System Architecture
By Phase 3 the system had enough moving parts that informal reasoning stopped working. I needed a formal architecture — not to impress anyone, but because you cannot reason about failure modes in a system you have not drawn. Architecture, here, was an instrument for understanding rather than documentation for its own sake.
The final system decomposed into three macro-layers:
Future Development Roadmap
Q1 2026planned
FEVER-Style Factuality Benchmarks
Q2 2026research
Adaptive Retrieval Depth
Q3 2026conceptual
Cryptographic Source Signing
Q4 2026planned
Confidence Calibration UI
2027research
Prompt Injection Defense Layer
Fig. — Verification Pipeline
and a thin meta-layer:
[4] Epistemic UI (trust surface)
Layer 1 — Retrieval
Providers:
Google Custom Search Engine (CSE) — bootstrap
Web → open, adversarial, timestamped, sparse
Outputs:
snippets
urls
titles
timestamps
micro-context
Layer 2 — Inference
Provider:
OpenRouter, multi-model
Outputs:
structured synthesis
paraphrased reasoning
citation scaffolding
Layer 3 — Verification
Verification exists to resolve contradictions between three different kinds of time: the web's present state, the model's priors from the past, and the user's future-directed query. The protocol sits at the intersection of data, time, and semantics, which is exactly where the hardest failures live.
Gate conditions:
Existence Gate → do snippets exist for claim?
Consistency Gate → do claims match snippets?
Temporal Gate → are snippets stale vs query?
Source Gate → is upstream adversarial?
Layer 4 — Epistemic UI
The UI is a working surface, not decoration: it shows which sources grounded an answer, where claims are tied to evidence, where uncertainty remains, and how the synthesis happened. Users do not need to understand the pipeline; they need to be able to see it working.
#Architecture Diagram
Next.js 14 Client
Responsive Client View
Security Proxy
API Route Handler
ORCHESTRATOR
GET /customsearchGoogle-Indexing
POST /completionOpenRouter-LLM
Data Provider A
Google CSE
Inference Provider B
OpenRouter Cloud
End-to-End Environment Encapsulation
Fig. — Core Architecture Overview
In the SeekEngine implementation, architecture exists in code under:
/actions
/api
/orchestrator
/sanitizer
/components
#XIII. Operational Behavior & Performance
The Truth Penalty
Most systems work optimizes for throughput, latency, and cost. SeekEngine optimized for correctness of the output, and correctness is far more expensive than speed. What I observed, informally, was:
The "Truth Penalty": SeekEngine trades additional latency for improved factual consistency.
Fig. — Latency Comparison Benchmark
Latency breakdown:
Stage
Cost
Retrieval
network-bound
Inference
compute-bound
Fusion
synchronization-bound
Verification
consistency-bound
In my informal tests, the grounded path ran roughly 1.3–2.4× slower than ungrounded inference. I am reporting that range as an observation, not a benchmark — the exact multiplier shifts with provider latency and query type. The point is not the number; it is that grounding has a real cost, and the interface has to be honest about it.
#XIV. Threat Model & Adversarial Surface
Unlike BitTorrent, SeekEngine is not attacked by malicious peers—but it is attacked by malicious content and overconfident models.
Threat classes included:
Threat Class
Source
Mitigation
XSS Injection
Web
Sanitizer
SEO Poisoning
Web
Source Weighting
Prompt Injection
User
Input Filtering
Citation Forgery
Model
Verification
Temporal Drift
Web/Model
Timestamp Check
Credential Leakage
System
Server Actions
Upstream Collapse
Provider
Timeout + Fallback
Poisoned Snippets
Web
Snippet Consistency
Rendered as:
Threat Model & Mitigations
XSS Injection
mitigated
DOMPurify sanitization
API Key Leakage
mitigated
Server-side encapsulation
Prompt Injection
partial
Input filtering (basic)
Data Persistence
mitigated
Request-scope only
Upstream Compromise
unaddressed
Outside control
Model-Level Exploits
unaddressed
Future work
MITIGATED
PARTIAL
UNADDRESSED
Fig. — Adversarial Surface Matrix
Zero-Trust Execution
Zero trust here means something concrete: I trusted none of the four parties in the pipeline by default — providers, models, users, or the web itself. That stance is unusual in RAG prototypes and closer to how hardened web services are built, which is the direction this project needed.
#XV. Limitations (Hard & Soft)
Hard Limitations
Cannot be fixed without architectural overhaul:
no formal factuality benchmarks
no multilingual grounding
temporal inconsistency (training cutoff vs now)
dependency on hostile providers
unbounded LLM miscalibration
snippet scarcity
rate-limited retrieval API
Soft Limitations
Fixable with future work:
query expansion
snippet ranking improvement
multi-provider fusion
uncertainty calibration
timestamp weighting
Rendered as:
Known Limitations Matrix
No Standardized BenchmarksEvaluation
Internal testing only
high impact
Third-Party DependencyReliability
Google CSE, OpenRouter availability
medium impact
Multilingual SupportCoverage
English-primary implementation
medium impact
Temporal ConsistencyAccuracy
Real-time data freshness varies
high impact
Rate LimitingScale
Free-tier constraints
low impact
Honest assessment: These limitations are documented, not hidden.
Fig. — Limitations Assessment
#XVI. Future Work
The directions I would pursue next, roughly in increasing difficulty:
(1) Cryptographic Source Signing
Truth can be anchored cryptographically (web domains → signatures).
At the limit, this line of work stops being about answering questions and becomes about arbitrating claims — systems that assemble and weigh evidence across many sources rather than a single chatbot producing a single answer. A research-grade version of that idea would not generate answers so much as maps of what is known, by whom, and on what evidence. That framing is speculative, but it is where the verification protocol points.
#XVII. Conclusion: Independent Systems Research Perspective
What SeekEngine taught me is that hallucination is not really a model failure — it is a coordination failure under resource constraints. Retrieval and inference are complementary; neither is sufficient on its own. Grounding needs verification, verification costs time, and that cost reshapes the whole system around it.
It also taught me something about how research gets done. None of this required funding, institutional backing, or GPU clusters. It was built in the open on free-tier APIs, and because there was nothing to hide behind, every failure was visible — which is exactly why the failure modes got understood instead of papered over.
SeekEngine's value is not its performance; it is the framing it makes testable: hallucination behaves like a distributed coordination problem, and grounding — in latency, bandwidth, and occasional silence — is genuinely expensive. Both claims are things a future, properly benchmarked version of this system could measure.
#XVIII. Bibliographic Context & Inspirations
SeekEngine sits at the intersection of several research and engineering traditions, and it is honest about borrowing from all of them rather than claiming novelty for their parts:
Information Retrieval Research
snippet extraction
relevance ranking
query expansion
temporal freshness
semantic matching
✔ Distributed Systems & P2P
partial failure behavior
bootstrap mechanisms
adversarial assumptions
non-deterministic sequencing
swarm coordination
✔ Security Engineering
zero-trust boundaries
dominance of untrusted inputs
poisoning resistance
credential encapsulation
browser threat models
✔ LLM Research
hallucination
grounding
RAG pipelines
uncertainty calibration
prompt shaping
Institutional RAG research can assume vector databases, stable compute, and proprietary evaluation harnesses. This project could assume none of those, which is not a weakness — it forced validation to come from running the system against the real, messy web rather than against a benchmark that might not survive contact with it.
#XIX. Acknowledgments & Contributions
SeekEngine was conceived, designed, and built by Gaurav Yadav — the architecture, implementation, debugging, and the conceptual framing documented here are my own work.
Thanks to:
OpenRouter — for accessible inference
Google CSE — for the retrieval substrate
Next.js — for sane server action boundaries
Tailwind + React — for making the UI fast to build
The open web — for being exactly as adversarial as this project needed
LLMs — for their confabulation tendencies, which turned out to be the most instructive bug I have ever worked with
No institutional support, funding, or proprietary infrastructure was used.
The project's source lives in a private repository (github.com/archduke1337/SeekEngine); the layout below mirrors the architecture and security boundary described throughout this paper:
#Appendix A — Prompting & RAG Protocol Notes (Spec-Level)
SeekEngine’s prompting layer enforces invariants:
no creativity
no speculation
no sentiment
no invented citations
brief claims
explicit sourcing
failure > confabulation
Example:
<< SYSTEM >>
You are a grounding-first search agent.
If no data is retrieved, say "Unknown."
Never invent facts. Cite snippets.
Minimize fluency and avoid speculation.
This interface treats LLM synthesis as a semantic reducer, not an author.
#Appendix B — Failure Trace Catalog
Observed Failure Modes
Failure
Root Cause
Hallucination
retrieval starvation
Staleness
training cutoff mismatch
Misalignment
parallel fusion race
Speculation
inference fallback
Overconfidence
no calibration
Spam
SEO poisoning
Silence
rate limit + timeout
These traces shaped future work directions.
#Appendix C — Temporal Considerations
Temporal mismatch is a major source of epistemic error:
System_Terminal
Fig. — Query Processing Pipeline
Temporal alignment remains an open research frontier.
#Appendix D — Observability as Insight
The terminal did more than show me what the system was doing; it changed what I believed the system was doing. Instrumentation here was not an add-on — it was the difference between assuming the pipeline worked and knowing where it failed.
Diagnostic terminal:
System_Terminal
Fig. — Live Orchestration Diagnostics
Watch a system's state long enough and orchestration stops being abstract. The lesson generalizes: if you cannot see a system's decisions, you cannot audit them — and in a security-adjacent tool, an unauditable decision is indistinguishable from a wrong one.
#Appendix E — Independent Research Context
SeekEngine belongs to a tradition of independent systems work — personal DHT implementations, hobby kernels, software-defined radio stacks, homebrew BitTorrent clients — that is driven by curiosity and constraint rather than grants and hardware budgets.
This lineage includes:
personal DHT implementations
hobby kernels
SDR radio stacks
Tor middleboxes
bare-metal type systems
BitTorrent clients built from scratch
I am not claiming benchmarks are worthless; they are not. But a benchmark measures how a system performs in the environment its authors prepared, and that is a different thing from how it performs against the real web. This project chose the real web, partly by necessity and partly on purpose, and the failures it surfaced are the reason the writeup exists.
#XXII. Final Statement
SeekEngine began as a patch for hallucination and became something closer to a study in grounded inference under real constraints. The working conclusion is plain: retrieval provides grounding, verification provides validity, and the UI provides the legibility that makes both trustworthy. Hallucination, in this view, is what happens when a model is left to verify itself.
The most important limits are just as plain. The comparisons in this paper are informal, the benchmarks are still to be run, and any claim stronger than "this framing is testable" would be overreach. What exists is a working system, a documented set of failure modes, and a protocol designed so that a properly measured version of this work is a matter of effort rather than invention.