Llm Engineering
AI RAG Knowledge Assistant
Grounded answers from retrieved knowledge, with source attribution.
A retrieval-augmented generation system that loads and chunks source documents, embeds them with OpenAI, stores vectors in FAISS, retrieves relevant context, and generates grounded answers with source metadata through FastAPI and Streamlit.
Problem
Context
Large language models can generate fluent answers without being grounded in a trusted knowledge source, creating a need for retrieval, attribution, and explicit context boundaries.
Challenge
Build a modular RAG workflow that retrieves relevant document context before generation and exposes the result through usable application interfaces.
Objective
Create an end-to-end knowledge assistant that performs document ingestion, chunking, embedding, vector retrieval, grounded generation, source attribution, and application delivery through API and UI layers.
Architecture
Documents are loaded and chunked, embedded with OpenAI text-embedding-3-small, stored in a FAISS vector index, retrieved by similarity, injected into a grounded prompt, and answered with gpt-4o-mini through API and UI layers.
- 01
Source documents
- 02
Document chunks
- 03
OpenAI embeddings
- 04
FAISS vector index
- 05
Top-k retrieved context
- 06
Grounded prompt
- 07
LLM-generated answer with source metadata
Document Ingestion
Loads source documents and prepares text for downstream chunking and retrieval.
- Python
- LangChain
Chunking Pipeline
Splits source text into retrievable units while preserving source metadata.
- Python
- LangChain
Embedding Layer
Transforms document chunks into vector representations for semantic retrieval.
- OpenAI
- text-embedding-3-small
Vector Store
Stores and retrieves document embeddings through a vector-store abstraction backed by FAISS.
- FAISS
Retrieval Layer
Selects top-k relevant chunks to provide grounded context for generation.
- FAISS
- LangChain
Generation Layer
Constructs a grounded prompt and generates answers constrained to retrieved context.
- OpenAI
- gpt-4o-mini
API Layer
Exposes retrieval-augmented question answering through an HTTP interface.
- FastAPI
User Interface
Provides an interactive front end for asking questions and reviewing grounded responses.
- Streamlit
Technology
Language
- Python 3.11
Ml Ai
- OpenAI text-embedding-3-small
- OpenAI gpt-4o-mini
- LangChain
- FAISS
Framework
- FastAPI
- Streamlit
Engineering Contribution
Built the complete RAG retrieval pipeline
Implemented document loading, chunking, embeddings, vector storage, top-k retrieval, and grounded answer generation as a modular workflow.
Added explicit grounding constraints
Constructed prompts that instruct the model to answer only from retrieved context rather than relying on unsupported external knowledge.
Preserved source attribution
Returned source and chunk metadata alongside generated answers so retrieved evidence remains inspectable.
Separated vector-store implementation behind an abstraction
Implemented FAISS through a vector-store abstraction so retrieval infrastructure is not tightly coupled to application logic.
Delivered API and interactive UI access
Exposed the RAG workflow through FastAPI and Streamlit for programmatic and interactive use.
Documented current system boundaries
Kept limitations explicit, including the local corpus, FAISS-only backend, heuristic grounding signal, and absence of authentication, streaming, formal evaluation, and production deployment.
Engineering Evidence
Architecture
End-to-End RAG Architecture
The system implements document loading, text chunking, embeddings, FAISS storage, semantic retrieval, grounded prompting, and answer generation.
Code
Semantic Vector Retrieval
Document chunks are embedded with text-embedding-3-small, stored in FAISS, and retrieved as top-k context for question answering.
Code
Grounded Answer Generation
The generation prompt instructs gpt-4o-mini to answer from retrieved context rather than inventing unsupported information.
Documentation
Source and Chunk Attribution
RAG responses include source metadata and chunk-level attribution for the retrieved material used during generation.
Architecture
Vector Store Abstraction
The current FAISS implementation is accessed through a vector-store abstraction rather than being embedded directly into higher-level application logic.
Code
FastAPI and Streamlit Interfaces
The system exposes retrieval-augmented question answering through both a FastAPI service and a Streamlit interface.
Metric
Heuristic Grounding Signal
The system includes a heuristic grounding signal to provide an additional indication of answer support, but it is not a calibrated evaluation metric.
Documentation
Documented Current Limitations
The current implementation uses a small local text corpus and FAISS backend, includes no authentication or streaming, does not yet include a formal evaluation framework, and is not presented as a production deployment.
What It Proves
Retrieval-augmented generation
Demonstrates the full RAG path from source documents through retrieval to grounded LLM generation.
LLM application engineering
Demonstrates integration of embeddings, vector search, generation, source attribution, API delivery, and an interactive UI.
Evidence-aware AI design
Demonstrates explicit grounding constraints, inspectable source metadata, and documented limits on what the current system can claim.