π RAG & Files
File uploads, embeddings, and retrieval are already in the chat path. Files stay on only if you keep an OpenAI key.
Upload a document, retrieve the useful bits, send them to the model. GoShipped already has that pipeline.
RAG here means: store what they uploaded, retrieve the bits that match the question, and put those bits in the model prompt. It is not a generic vector-database tutorial.
You are here to decide whether files are part of the product β and to keep an OpenAI key while they are.
User uploads a PDF
β
GoShipped stores and processes it
β
Content is chunked and embedded
β
Relevant context is retrieved
β
AI answers using the documentOne end-to-end path
- User attaches
spec.pdfin chat (apps/webβPOST /api/v1/files/upload) - Bytes go to Supabase Storage bucket
chat-filesatusers/{user}/conversations/{id}/{file_id}/spec.pdf - A row is written to
uploaded_files - PDF text is extracted, split (with page numbers), embedded with OpenAI (
text-embedding-3-small), stored inuploaded_file_chunks - The next question runs
match_uploaded_file_chunks(pgvector) - Excerpts land in the chat system prompt (
build_file_context) β the model is told to cite filename and page when useful
Small text/code files skip the vector path and inject directly. Every PDF uses extraction + chunks.
Small files vs large files
| File | What happens |
|---|---|
Text/code under FILE_SMALL_THRESHOLD_CHARS (20,000) | Full text injected (budgeted). No chunks. |
| Text/code at or above that | Chunk (~1800 chars, 200 overlap) + embed + retrieve top 8 |
| Any PDF | Extract β page-aware chunks β embed β retrieve |
| Embed/extract fails | processing_status=failed. Large text can fall back to a truncated inject. |
OpenAI while files are on
What you turn on
| Switch | Effect |
|---|---|
features.files | Product-wide uploads + file context. false β 404 on the files API, no attach button. |
Plan files_enabled / file_uploads_per_month / max_file_size_mb | Who may upload, how often, how large (plans/config.py) |
SUPABASE_SERVICE_ROLE_KEY | Server upload to Storage |
Bucket chat-files | Create in the Supabase Storage dashboard β not created by SQL migrations |
Schema: supabase/migrations/003_uploaded_files.sql, 004_uploaded_file_chunks.sql, 005_uploaded_files_pdf.sql. Commands: Database & Storage.
Limits (shipped defaults)
| Limit | Default |
|---|---|
| Extensions | .txt .md .py .js .ts .tsx .jsx .json .html .css .sql .yaml .yml .pdf |
| Text/code size | 2 MB (FILE_UPLOAD_MAX_BYTES) |
| PDF size | 10 MB (FILE_UPLOAD_PDF_MAX_BYTES) |
| Context per turn | 50k chars total; 20k direct / 25k retrieved |
Plan max file size is separate (Free 10 MB, Pro 50 MB in the shipped catalog) β the env caps still apply to the upload parser.
Tune retrieval with FILE_CHUNK_MATCH_COUNT, FILE_CHUNK_SIZE, and the context budgets in apps/api/.env.
Code map
| Piece | Path |
|---|---|
| Upload route | apps/api/app/api/v1/files.py |
| Store + retrieve | apps/api/app/services/file_service.py |
| Chunking | apps/api/app/services/file_chunking.py |
| PDF text | apps/api/app/services/pdf_extraction.py |
| Embeddings | apps/api/app/services/embedding_service.py |
| Prompt block | apps/api/app/services/file_context_hints.py |
| Attach UI | apps/web/src/hooks/use-files.ts, chat-input.tsx |
You rarely replace this pipeline. You change flags, limits, and whether your product should accept files at all.