Semantic Relevance Detection of Visual Content in Corporate Reports Using Layout-Aware Vision-Language Models


Altunkaya B. S., Akgül Y. S., Yeşilyurt S., Genç Y., Aydemir S. D.

2026 34th Signal Processing and Communications Applications Conference (SIU), İstanbul, Türkiye, 7 - 10 Temmuz 2026, ss.1-4, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu71813.2026.11636611
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.1-4
  • Sivas Cumhuriyet Üniversitesi Adresli: Evet

Özet

This paper presents a system that automatically evaluates the semantic consistency of visual content with surrounding text in corporate sustainability reports. The proposed system consists of four stages: (1) a MinerU Parser that extracts content with layout information from PDF documents, (2) a Parquet Normalizer that converts extracted data into queryable tables, (3) an Optimistic Scanner that queries a Vision-Language Model (VLM) with each image alongside surrounding page text to produce a relevance score, and (4) an Independent Auditor that verifies these scores using only the relevant page's text. The system was tested on 30 reports, achieving 83.3% precision and an F1-Score of 0.79.