Abstract
The automated interpretation of unstructured financial records, including receipts and invoices, has become increasingly critical for intelligent document understanding. A Retrieval-Augmented Generation (RAG) framework is presented to address question answering over multilingual financial documents characterized by noisy OCR output, variable layouts, and visually embedded information. In the proposed approach, textual content is encoded using multilingual sentence encoders, while visual information is processed through multimodal language models; both representations are stored in a unified semantic index. At inference time, retrieval scores from each modality are fused to guide evidence selection, which is then provided to an instruction-tuned generator. The factuality of generated answers is subsequently validated using a lightweight verifier Large Language Model(LLM) Judge that classifies answers as grounded, partially grounded, or hallucinated. The system is trained and evaluated on a real-world dataset of 2,536 financial documents in Turkish and English, achieving 86.7% grounded answer accuracy and reducing hallucination rates by more than 50% compared to text-only retrieval. The contributions are: (i) an end-to-end multimodal RAG architecture for financial question answering,(ii)a curated multilingual benchmark dataset of real financial documents, and (iii) an efficient groundedness verification method based on LLM judgment, establishing a reproducible baseline for multimodal document understanding.
| Original language | English |
|---|---|
| Pages (from-to) | 160-165 |
| Number of pages | 6 |
| Journal | International Conference on Computer Science and Engineering, UBMK |
| Issue number | 2025 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 10th International Conference on Computer Science and Engineering, UBMK 2025 - Istanbul, Turkey Duration: 17 Sept 2025 → 21 Sept 2025 |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
Keywords
- AI-Assisted Document Management
- Embedding
- Large Language Models
- Multimodal Retrieval
- Semantic Search
Fingerprint
Dive into the research topics of 'Grounded Answer Generation over Multimodal Financial Records via Semantic Indexing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver