Corvinus
Corvinus

A Proposed Framework for Robust Visual Product Identification in Dynamic Environments

Zeleny, Klaudia Éva, Ruppert, Tamás ORCID: https://orcid.org/0000-0001-9441-843X and Kő, Andrea ORCID: https://orcid.org/0000-0003-0023-1143 (2026) A Proposed Framework for Robust Visual Product Identification in Dynamic Environments. Acta Polytechnica Hungarica, 23 (8). pp. 211-240. DOI https://doi.org/10.12700/APH.23.8.2026.8.12

[img]
Preview
PDF - Requires a PDF viewer such as GSview, Xpdf or Adobe Acrobat Reader
917kB

Official URL: https://doi.org/10.12700/APH.23.8.2026.8.12


Abstract

Reliable visual product identification remains a challenging task in dynamic indoor environments, where variability in viewpoint, illumination, and object placement limits the applicability of conventional scanning-based approaches. This is particularly relevant in high-mix low-volume intralogistics settings, where incomplete observations and inconsistent acquisition conditions are common. This paper proposes a modular framework for robust visual product identification based on machine-readable codes. The approach formulates identification as a system-level reasoning problem, integrating detection, geometric normalization, decoding, validation, recovery, and temporal aggregation within a unified architecture. Frame-level observations are treated as uncertain evidence, which is accumulated and constrained at the sequence level in order to support reliable identification decisions. The individual components draw on established computer vision and decoding techniques; the contribution lies in their system-level integration with explicit uncertainty handling and temporal reasoning, rather than in new detection, decoding, or fusion algorithms. The proposed framework emphasizes robustness through structured integration of visual evidence rather than isolated recognition performance, providing a foundation for perception driven identification in visually dynamic environments. A preliminary evaluation on a controlled test set, shows that detection performance is close to ceiling, while frame-level decoding alone succeeds on only about one third of ground-truth instances (35.6%); aggregating evidence across frames within the proposed sequence-level reasoning layer raises identification success to 73.5%, providing initial evidence that the proposed architecture improves reliability over frame-level decoding alone.

Item Type:Article
Uncontrolled Keywords:computer vision; intralogistics; DataMatrix; product localization; high-mix low-volume
Divisions:Institute of Data Analytics and Information Systems
Corvinus Doctoral Schools
Subjects:Computer science
Funders:Ministry of Innovation and Technology of Hungary
Projects:2019-1.1.1-PIACI-KFI-2019-00312
DOI:https://doi.org/10.12700/APH.23.8.2026.8.12
ID Code:13304
Deposited By: MTMT SWORD
Deposited On:23 Sep 2026 09:31
Last Modified:23 Sep 2026 09:31

Repository Staff Only: item control page