Interpretable Visual Feature Discovery using Multiple Instance Learning: A Case Study of Colonial Korean Print
Aron van de Pol, Jelena Prokić, Angus Mol, 2026, Computational Humanities Research
van de Pol, A., Prokić, J., & Mol, A. (2026). Interpretable Visual Feature Discovery using Multiple Instance Learning: A Case Study of Colonial Korean Print. Computational Humanities Research. https://doi.org/10.1017/chr.2026.10036
This paper explores how neural networks, particularly Multiple Instance Learning (MIL), can advance computational analysis of historical visual materials. While Distant Viewing (Arnold and Tilton, 2023a) has shown how computer vision can analyze visual materials at scale, interpreting how these models identify distinctive features remains a key methodological challenge. Although deep learning approaches achieve high classification accuracy, understanding which visual elements inform their decisions is often unclear, limiting their utility for humanities research where interpretability is crucial. We address this challenge through an innovative application of Multiple Instance Learning to historical print analysis. Using colonial Korean printshops (1910–1945) as our case study, we demonstrate how MIL’s interpretable architecture can reveal distinctive visual features while maintaining classification accuracy (Cai et al., 2024). This approach allows us to examine not just what distinguishes different printshops’ outputs, but specifically which typographic elements and patterns the model uses to make these distinctions. Drawing on a large digitized corpus from four major colonial Korean printshops, we demonstrate how MIL, originally developed for medical imaging analysis, can effectively classify and help interpret historical visual materials. The model achieves 92% accuracy while providing interpretable attention maps that highlight historically meaningful features. The results demonstrate how computational methods can reveal subtle but consistent variations in common textual elements, such as the rendering of ieung (ㅇ) in frequently occurring grammatical particles (josa 조사) or the bottom of the tigut (ㄷ). This methodological contribution extends beyond Korean print history, offering a framework for interpretable computational analysis of visual materials more broadly.