Semantic Relevance Detection of Visual Content in Corporate Reports Using Layout-Aware Vision-Language Models
2026 34th Signal Processing and Communications Applications Conference (SIU), İstanbul, Türkiye, 7 - 10 Temmuz 2026, ss.1-4, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1109/siu71813.2026.11636611
- Basıldığı Şehir: İstanbul
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.1-4
- Sivas Cumhuriyet Üniversitesi Adresli: Evet
Özet
This paper presents a system that automatically evaluates the semantic consistency of visual content with surrounding text in corporate sustainability reports. The proposed system consists of four stages: (1) a MinerU Parser that extracts content with layout information from PDF documents, (2) a Parquet Normalizer that converts extracted data into queryable tables, (3) an Optimistic Scanner that queries a Vision-Language Model (VLM) with each image alongside surrounding page text to produce a relevance score, and (4) an Independent Auditor that verifies these scores using only the relevant page's text. The system was tested on 30 reports, achieving 83.3% precision and an F1-Score of 0.79.