Self-hosted document intelligence for RAG — parse, classify, split, and extract from PDF/DOCX/PPTX and images: searchable, cited, retrieval-ready chunks and schema-driven extraction
ocr self-hosted embeddings docx pptx pdf-parsing vlm reranking rag document-ingestion vector-database paddleocr hybrid-search jinaai amazon-nova-lite
-
Updated
Aug 27, 2026 - Python