- Main
How Structured Knowledge Constrains Generative Diffusion Models for Domain Visual Description
Abstract
Visual description is a fundamental yet challenging task that reflects how perceptual information is transformed into structured linguistic descriptions. Recent diffusion-based models have shown strong potential for parallel decoding and global semantic modeling. However, existing approaches largely treat generation as an unconstrained process, lacking mechanisms to incorporate structured knowledge, which often results in conceptually implausible or hallucinated descriptions, especially in domain-specific contexts. To investigate how structured knowledge constrains generative diffusion processes, we propose KenDiC, a Knowledge-enhanced Diffusion-based Captioner for visual description. KenDiC integrates a domain-adaptive visual encoder (DAVE) trained via contrastive learning to align perceptual and linguistic representations, and a domain term vocabulary (DTV) that constrains decoding to guide concept selection during generation. To support systematic analysis, we construct II2T-Bench, a domain-centric benchmark with expert annotations. Experimental results show that structured knowledge constraints significantly reduce hallucinations and improve semantic fidelity, suggesting a computational account of how knowledge guides generative visual description.