File Upload Integration
Manually upload PDFs, Word documents, and text files into your self-healing knowledge base with visual grounding.
What it does
The File Upload integration enables manual document uploads to populate RAG databases. It leverages visual grounding models to parse unstructured files, extracting text, layout grids, tables, and infographic data inside PDFs, Word documents (`.docx`), and text files (`.txt`). This ensures your agents can query manual guides, diagrams, and corporate files.
Data Ingestion & Inferences
Uploaded documents undergo immediate, secure parsing pipelines:
- Data Read: Entire document layout, pages, tables, lists, text structures, and images/charts.
- Data Written: Text contents are converted to vector embeddings and stored inside isolated tenant vector indexes.
- Sync Frequency: One-time ingestion. Documents do not auto-update; modified versions must be re-uploaded.
Scope & Limits
- Supported Formats: `.pdf`, `.docx`, and `.txt` text documents.
- Unsupported Assets: Raw image formats (`.png`, `.jpg`), audio/video files, and structured database spreadsheets (`.xlsx`, `.csv`).
- File Limits: Individual files are capped at 25MB. Visual grounding for charts and infographics requires API key compatibility with vision-supported models. Refer to the OpenAI Files API Reference for details on file processing configurations.
Setup guide
Select File Ingestion in dashboard
Go to Knowledge → Add Source and select File Upload.
Select target documents
Drag and drop your target PDFs, DOCX, or TXT files directly into the dashed upload panel, or click browse to locate them in your local directory.
Authorize parsing
Click Upload. Verabase parses the documents, extracts text blocks and infographics via OCR, and indexes the content. View indexing logs in the Source Explorer.
Where it shows up
- Source Explorer: Appears as manual files, detailing filename, extraction indexes, size, and source type tags.
- Citation popovers: Generated agent replies cite matching PDF page numbers or manual references inside the active chat details panel.
Troubleshooting & Common Issues
- Incomplete text indexing: If text inside a PDF is not indexed, check if the file is scanned. Scanned PDFs without embedded selectable fonts require OCR visual grounding parsing to extract characters.
- Upload timeouts: Large documents exceeding the 25MB limit or containing thousands of pages can fail due to timeout limits. Split heavy guides into smaller chapters before uploading.
Frequently asked questions
Related integrations
Ready to upload manual files?
Join the waitlist for early access to the self-healing knowledge platform.
Get early access to File Upload