SaigonTech Vault
Loại: knowledge-doc 2026-09-29 active #wiki-import

📘 Guide: How to Create reference_guide.md for a Project

🧠 What is reference_guide.md?

reference_guide.md is a human-written, LLM-friendly summary file that explains the structure and important components of a codebase. It’s used to help:

  • Developers understand the project quickly.
  • Retrieval-Augmented Generation (RAG) systems index and query the codebase more effectively.
  • AI systems (like OpenAI or custom bots) retrieve relevant code chunks more accurately by understanding how your system works.

🎯 Why Is It Important?

Creating a good reference_guide.md benefits both humans and machines:

  • ✅ Helps new developers onboard faster.
  • ✅ Improves LLM search accuracy in RAG-based tools.
  • ✅ Explains which files are most important, what they do, and how they work together.
  • ✅ Reduces retrieval noise by flagging low-priority files/folders.

This is especially useful in projects where:

  • The codebase is large or spans multiple services/modules.
  • You use LLMs to answer technical questions, generate code, or assist debugging.
  • You embed code into vector databases for retrieval-based search.

🛠️ How to Create One — Step-by-Step

1. 🗂 Describe the Folder & File Structure

List all top-level folders and summarize what each contains. Then do the same for key files.

## Project Structure

- /api/         → FastAPI route handlers and controllers
- /db/          → SQL schema, migration files
- /utils/       → Helper functions, including embedding logic
- /supabase/    → Client for calling RPC and DB from Supabase
- /tests/       → Unit tests (non-critical for RAG)
- main.py       → Server entry point

2. 🧩 Document the Key Components

For each important file or module, explain: - What it does - How it's used - Any dependencies

### main.py
- Starts FastAPI server.
- Handles incoming queries.
- Uses OpenAI embeddings to embed query, then calls `match_documents` RPC function on Supabase.

### /utils/embedder.py
- Wraps OpenAI Embedding API with rate limiting.
- Used by main.py to embed query strings.

### /db/schema.sql
- Defines `code_chunks` table (stores vector, content, metadata).
- Defines `match_documents()` function for vector similarity search with optional filters.

3. 🧱 Explain Database Structure (If Applicable)

If the project stores code or data in a database (like Supabase), explain: - The schema (e.g. code_chunks) - What each column is for - How metadata is used

## Database: code_chunks

| Column      | Type         | Description                          |
|-------------|--------------|--------------------------------------|
| content     | TEXT         | Code chunk or snippet                |
| metadata    | JSONB        | Includes file_path, language, type   |
| embedding   | VECTOR(1536) | OpenAI vector for this content       |
| created_at  | TIMESTAMP    | Insert time                          |

4. 🔎 Describe Retrieval Logic

Explain how the RAG system queries the database or file system, including: - The embedding model used - Filtering logic - Vector similarity calculation

## Retrieval Logic

- Embedding Model: `text-embedding-3-small` from OpenAI.
- Query is embedded and passed to `match_documents` RPC.
- Optional filters: `file_path`, `language`, `repo_name`.
- Vector similarity: `1 - (embedding <=> query_embedding)` (cosine).
- Returns top-N most similar code chunks.

5. 📌 Highlight Important Metadata

Explain what’s stored in metadata and how it's used:

## Metadata Usage

Each chunk is saved with metadata like:
- file_path: Original file source (e.g. /api/routes/query.py)
- start_line, end_line: Line range in the original file
- language: Programming language
- type: Function, class, etc.
- repo_name: Repository name

These fields can be used to filter chunks before vector search.

6. ❌ Flag Irrelevant or Noisy Files

Explicitly call out folders/files that should be ignored by retrieval or indexing.

## Irrelevant Files for Retrieval

Do not include or prioritize:
- /tests/*
- /docs/*
- /node_modules/*
- README.md (unless it contains key info)

7. ✅ Add Retrieval Prompt Tips (Optional)

Add helpful example prompts that developers or AI agents can use when interacting with the RAG system.

## Sample Useful Prompts

- "Where is the OpenAI embedding logic located?"
- "How does the FastAPI server connect to Supabase?"
- "What metadata filters are used in match_documents?"

📄 Final Example File Structure

project-root/
├── api/
├── db/
├── supabase/
├── utils/
├── main.py
├── reference_guide.md ✅ <-- This file

✅ Checklist Before Finishing

Make sure your reference_guide.md includes:

  • [ ] Summary of folder structure
  • [ ] Description of important files/modules
  • [ ] Database schema and retrieval explanation
  • [ ] Metadata field descriptions
  • [ ] Retrieval logic overview
  • [ ] List of ignored/low-priority files
  • [ ] Sample prompts (optional)

🧑‍💻 Who Should Create It?

  • The developer most familiar with the backend and database logic.
  • When a new repo is added to the RAG system, the lead engineer should ensure this file exists.