References

ellibs

Электронные библиотеки

Russian Digital Libraries Journal

1562-5419

Казанский (Приволжский) федеральный университет

10.26907/1562-5419-2025-28-6-1454-1480

ellibs-628

Research Article

Статьи

Атрибуция архивных рукописных писем с использованием сиамских нейронных сетей

Archival Handwritten Letter Attribution using Siamese Neural Networks

Пронина

Наталия Михайловна

Pronina

Nataliia Mikhailovna

natalka-pronina@mail.ru

Московский государственный университет имени М. В. ЛомоносоваLomonosov Moscow State University

2025

19122025

28614541480

2025

Пронина Н.М.

Pronina N.M.

Данная работа распространяется под лицензией Creative Commons Attribution 4.0.

This work is licensed under a Creative Commons Attribution 4.0 License.

https://ellibs.elpub.ru/jour/article/view/628

Предложен метод автоматической атрибуции архивных рукописных писем на основе сиамской нейронной сети, решающий ключевую проблему цифровой гуманитаристики – установление авторства исторических документов. Актуальность исследования обусловлена массовой оцифровкой архивов XVII–XIX вв., атрибуция которых затруднена из-за неполных исходных сведений об авторах. Метод адаптирован к работе с реальным корпусом текстов и учитывает характерные для архивов проблемы: некачественные оцифровки, значительную вариативность почерка и выраженный дисбаланс классов (от 1 до 50 и более образцов на автора). Применение сиамской архитектуры позволяет получать дискриминативные векторные представления, эмбеддинги, на основе которых выполняется не только классификация документов известных авторов, но и эффективно выявляются рукописи, не принадлежащие ни одному из них. Это сужает круг кандидатов для последующей экспертной проверки. Представлен алгоритм предобработки данных и проведено сравнительное исследование двух подходов к анализу текста: на уровне фрагментов изображения (300 × 300 пикселей) и уровне отдельных строк. Разработанный инструмент предлагает архивным работникам и филологам эффективное решение для предварительной сортировки и атрибуции крупных массивов рукописных документов.

This paper presents a method for the automated attribution of archival handwritten letters based on a Siamese neural network, addressing a key challenge in digital humanities – the authentication of historical documents. The research is motivated by the mass digitization of 17th to 19th-century archives, where attribution is often hindered by incomplete or inaccurate metadata about the authors. The method is designed for real-world document collections and accounts for challenges typical of archival materials: poor-quality scans, significant handwriting variation, and substantial class imbalance (from 1 to over 50 samples per author). The use of a Siamese network architecture enables the extraction of discriminative vector representations (embeddings). Based on these embeddings, the method not only classifies documents by known authors but also effectively identifies manuscripts that do not match any known author in the archive. This significantly narrows down the pool of candidates for subsequent expert verification. The study introduces a data preprocessing algorithm and provides a comparative analysis of two approaches to text analysis: at the image fragment level (300×300 px) and at the individual text line level. The developed tool offers archivists and philologists an effective solution for the preliminary sorting and attribution of handwritten documents large collections.

сиамская нейронная сетьидентификацияверификацияатрибуциярукописный текстархивные документысверточная нейронная сетьрекуррентная нейронная сеть

siamese neural networkidentificationverificationattributionhandwritten textarchival documentsconvolutional neural networkrecurrent neural network

References1

He K., Zhang X., Ren S., Sun J. Deep residual learning for image recognition // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016. P. 770–778. https://doi.org/10.1109/CVPR.2016.90.

Kiselev V., Kropotov D., Pronina N. Handwritten documents author verification based on the siamese network // The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences. 2024. Vol. XLVIII-2/W5-2024. P. 73–78. https://doi.org/10.5194/isprs-archives-XLVIII-2-W5-2024-73-2024

Bromley J., Bentz J., Bottou L., Guyon I., Lecun Y., Moore C., Sackinger E., Shah R. Signature verification using a "siamese" time delay neural network // International Journal of Pattern Recognition and Artificial Intelligence. 1993. Vol. 7, No. 4. P. 669–688. https://doi.org/10.1142/S0218001493000339

Solomon E., Woubie A., Emiru E.S. Deep learning-based face recognition method using siamese network. 2024. https://doi.org/10.48550/arXiv.2312.14001

Yin W., Schütze H. Convolutional neural network for paraphrase identification // Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. P. 901–911. https://doi.org/10.3115/v1/N15-1091

Koch G., Zemel R., Salakhutdinov R. et al. Siamese neural networks for one-shot image recognition // ICML Deep Learning Workshop. 2015. Vol. 2, No. 1. P. 1–30.

Chopra S., Hadsell R., LeCun Y. Learning a similarity metric discriminatively, with application to face verification // 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05). 2005. Vol. 1. P. 539–546. https://doi.org/10.1109/CVPR.2005.202

Hadsell R., Chopra S., LeCun Y. Dimensionality reduction by learning an invariant mapping // 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06). 2006. Vol. 1. P. 1735–1742. https://doi.org/10.1109/CVPR.2006.100

Schroff F., Kalenichenko D., Philbin J. Facenet: A unified embedding for face recognition and clustering // 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2015. P. 815–823. https://doi.org/10.1109/CVPR.2015.7298682

Souibgui M.A., Biswas S., Jemni S.K., Kessentini Y., Forn´es A., Llado´s J., Pal U. Docentr: An end-to-end document image enhancement transformer. 2022. P. 1699–1705. https://doi.org/10.1109/ICPR56361.2022.9956101.

Wood D.E., Salzberg S.L. Kraken: ultrafast metagenomic sequence classification using exact alignments // Genome Biology. 2014. Vol. 15, No. 1. P. R46. https://doi.org/10.1186/gb-2014-15-3-r46

Shu L., Xu H., Liu B. Doc: Deep open classification of text documents. 2017. P. 2911–2916. https://doi.org/10.18653/v1/D17-1314.

Kiselev V., Pronina N. Machine attribution of handwriting in solving source studies problems (based on the correspondence of G.N. Potanin) // Imagology and Comparative Studies. 2025. No. 24.

The authors declare that there are no conflicts of interest present.