📘 Guide: How to Create reference_guide.md for a Project
🧠 What is reference_guide.md?
reference_guide.md is a human-written, LLM-friendly summary file that explains the structure and important components of a codebase. It’s used to help:
- Developers understand the project quickly.
- Retrieval-Augmented Generation (RAG) systems index and query the codebase more effectively.
- AI systems (like OpenAI or custom bots) retrieve relevant code chunks more accurately by understanding how your system works.
🎯 Why Is It Important?
Creating a good reference_guide.md benefits both humans and machines:
- ✅ Helps new developers onboard faster.
- ✅ Improves LLM search accuracy in RAG-based tools.
- ✅ Explains which files are most important, what they do, and how they work together.
- ✅ Reduces retrieval noise by flagging low-priority files/folders.
This is especially useful in projects where:
- The codebase is large or spans multiple services/modules.
- You use LLMs to answer technical questions, generate code, or assist debugging.
- You embed code into vector databases for retrieval-based search.
🛠️ How to Create One — Step-by-Step
1. 🗂 Describe the Folder & File Structure
List all top-level folders and summarize what each contains. Then do the same for key files.
## Project Structure
- /api/ → FastAPI route handlers and controllers
- /db/ → SQL schema, migration files
- /utils/ → Helper functions, including embedding logic
- /supabase/ → Client for calling RPC and DB from Supabase
- /tests/ → Unit tests (non-critical for RAG)
- main.py → Server entry point
2. 🧩 Document the Key Components
For each important file or module, explain: - What it does - How it's used - Any dependencies
### main.py
- Starts FastAPI server.
- Handles incoming queries.
- Uses OpenAI embeddings to embed query, then calls `match_documents` RPC function on Supabase.
### /utils/embedder.py
- Wraps OpenAI Embedding API with rate limiting.
- Used by main.py to embed query strings.
### /db/schema.sql
- Defines `code_chunks` table (stores vector, content, metadata).
- Defines `match_documents()` function for vector similarity search with optional filters.
3. 🧱 Explain Database Structure (If Applicable)
If the project stores code or data in a database (like Supabase), explain:
- The schema (e.g. code_chunks)
- What each column is for
- How metadata is used
## Database: code_chunks
| Column | Type | Description |
|-------------|--------------|--------------------------------------|
| content | TEXT | Code chunk or snippet |
| metadata | JSONB | Includes file_path, language, type |
| embedding | VECTOR(1536) | OpenAI vector for this content |
| created_at | TIMESTAMP | Insert time |
4. 🔎 Describe Retrieval Logic
Explain how the RAG system queries the database or file system, including: - The embedding model used - Filtering logic - Vector similarity calculation
## Retrieval Logic
- Embedding Model: `text-embedding-3-small` from OpenAI.
- Query is embedded and passed to `match_documents` RPC.
- Optional filters: `file_path`, `language`, `repo_name`.
- Vector similarity: `1 - (embedding <=> query_embedding)` (cosine).
- Returns top-N most similar code chunks.
5. 📌 Highlight Important Metadata
Explain what’s stored in metadata and how it's used:
## Metadata Usage
Each chunk is saved with metadata like:
- file_path: Original file source (e.g. /api/routes/query.py)
- start_line, end_line: Line range in the original file
- language: Programming language
- type: Function, class, etc.
- repo_name: Repository name
These fields can be used to filter chunks before vector search.
6. ❌ Flag Irrelevant or Noisy Files
Explicitly call out folders/files that should be ignored by retrieval or indexing.
## Irrelevant Files for Retrieval
Do not include or prioritize:
- /tests/*
- /docs/*
- /node_modules/*
- README.md (unless it contains key info)
7. ✅ Add Retrieval Prompt Tips (Optional)
Add helpful example prompts that developers or AI agents can use when interacting with the RAG system.
## Sample Useful Prompts
- "Where is the OpenAI embedding logic located?"
- "How does the FastAPI server connect to Supabase?"
- "What metadata filters are used in match_documents?"
📄 Final Example File Structure
project-root/
├── api/
├── db/
├── supabase/
├── utils/
├── main.py
├── reference_guide.md ✅ <-- This file
✅ Checklist Before Finishing
Make sure your reference_guide.md includes:
- [ ] Summary of folder structure
- [ ] Description of important files/modules
- [ ] Database schema and retrieval explanation
- [ ] Metadata field descriptions
- [ ] Retrieval logic overview
- [ ] List of ignored/low-priority files
- [ ] Sample prompts (optional)
🧑💻 Who Should Create It?
- The developer most familiar with the backend and database logic.
- When a new repo is added to the RAG system, the lead engineer should ensure this file exists.