Narrative-Driven Composite Facial Sketch Generation Using Dual Semantic and Attribute Conditioning
Keywords:
composite facial sketch, narrative-to-image generation, conditional GAN, CLIP, facial attribute preservation, forensic imaging, attribute-level evaluation.Abstract
Forensic identification frequently depends on eyewitness accounts when no photograph of a suspect exists. Converting natural-language facial descriptions into composite sketches is traditionally manual, artist-dependent, and inconsistent. Although recent generative models support text-to-image synthesis and photo-to-sketch translation, narrative-driven composite sketch generation remains insufficiently addressed, particularly in terms of explicit facial-attribute preservation. This study presents an attribute-guided conditional generative adversarial network that converts eyewitness narratives into composite facial sketches through dual semantic and attribute conditioning. A frozen CLIP ViT-B/32 encoder provides a semantic representation of the narrative, while a rule-based parser extracts a 40-dimensional facial-attribute vector. Both representations jointly condition the generator, while a matching-aware discriminator evaluates image–condition compatibility. Experiments on Multi-Modal-CelebA-HQ demonstrate an attribute macro-F1 of 0.382, weighted-F1 of 0.520, CLIP narrative consistency of 0.248.Fine-grained analysis shows stronger reproduction of coarse attributes such as gender, age, and beard characteristics, whereas fine and rare attributes exhibit lower recall. The findings demonstrate the feasibility of narrative-to-composite-sketch generation with attribute-level evaluation and identify fine-attribute recall as a key challenge.





