
The Problem
Traditional single-retriever RAG systems often miss relevant passages or surface irrelevant ones, especially when dealing with multiple documents and live web pages containing overlapping terminology. Keyword-only search misses semantic meaning, while dense-only retrieval can miss exact terminology, identifiers, and acronyms. Enterprise workflows require systems that combine both search paradigms, verify cross-document factual contradictions, evaluate retrieval metrics with RAGAS, and provide confidence-calibrated attributed answers — all on a 100% free open-source stack requiring zero paid API keys.
Architecture & Approach
The platform implements a multi-tier Hybrid Retrieval architecture. Ingestion processes multi-format files (PDF, DOCX, TXT, MD) and live web URLs (via BeautifulSoup4) into ChromaDB using local sentence-transformers (all-MiniLM-L6-v2) and a pure-Python BM25 Okapi keyword index. User queries and windowed history are condensed before executing min-max normalized weighted fusion (Score = α · Semantic_norm + (1-α) · BM25_norm) with per-source balancing. Generation and synthesis are handled by NVIDIA NIM (nvidia/nemotron-3-super-120b-a12b). Concurrently, a two-stage conflict detection engine runs a heuristic gate followed by targeted LLM verification; a multi-factor confidence scorer assigns 0.0–1.0 ratings with color-coded badges; and an automated RAGAS engine benchmarks Faithfulness, Relevancy, Precision, and Recall. The platform operates across a 6-tab Streamlit UI featuring Plotly analytics, full-text chunk exploration, document diffing, and session report export in Markdown and printable HTML/PDF.
About the Project
A production-grade Document Intelligence and Hybrid RAG platform fusing ChromaDB dense vector search (local all-MiniLM-L6-v2) with a pure-Python BM25 Okapi keyword index via min-max weighted score fusion, achieving 100% retrieval precision@k on multi-document evaluation sets. Features dual ingestion (local multi-format documents + live web scraping), an automated RAGAS evaluation scorecard, two-stage source conflict detection, multi-factor confidence scoring (0.0–1.0 color-coded badges), interactive Plotly analytics dashboard, chunk explorer, contract diff engine with LLM executive summaries, and session report export — all across a 6-tab Streamlit UI. Powered by NVIDIA NIM (nemotron-3-super-120b) with a 100% free open-source stack validated by 34/34 passing automated tests.
Key Metrics
100%
Retrieval Precision@k
34/34 passing
Automated Tests
NVIDIA NIM Nemotron-3
LLM Engine
100% free open-source
Stack Cost