cris.boxmetadata.label.title
Building semantic understanding beyond deep learning from sound and vision
cris.boxmetadata.label.dateissued
01 browse.startsWith.months.january 2016
cris.boxmetadata.label.accesslevel
metadata only access
cris.boxmetadata.label.resourcetype
conference paper
cris.boxmetadata.label.authors
De Souza F.
Sarkar S.
CAMARA CHAVEZ, GUILLERMO
Universidade Federal de Ouro Preto
cris.boxmetadata.label.publisher
Institute of Electrical and Electronics Engineers Inc.
cris.boxmetadata.label.abstract
Deep learning-based models have recently been widely successful at outperforming traditional approaches in several computer vision applications such as image classification, object recognition and action recognition. However, those models are not naturally designed to learn structural information that can be important to tasks such as human pose estimation and structured semantic interpretation of video events. In this paper, we demonstrate how to build structured semantic understanding of audio-video events by reasoning on multiple-label decisions of deep visual models and auditory models using Grenander's structures for imposing semantic consistency. The proposed structured model does not require joint training of the structural semantic dependencies and deep models. Instead they are independent components linked by Grenander's structures. Furthermore, we exploited Grenander's structures as a means to facilitate and enrich the model with fusion of multimodal sensory data; in particular, auditory features with visual features. Overall, we observed improvements in the quality of semantic interpretations using deep models and auditory features in combination with Grenander's structures, reflecting as numerical improvements of up to 11.5% and 12.3% in precision and recall, respectively.
cris.boxmetadata.label.citationstartpage
2097
cris.boxmetadata.label.citationendpage
2102
cris.boxmetadata.label.volume
0
cris.boxmetadata.label.language
English
cris.boxmetadata.label.ocdeknowledgeArea
Educación general (incluye capacitación, pedadogía) Ciencias de la computación Lingüística
cris.boxmetadata.label.doi
cris.boxmetadata.label.scopusidentifier
2-s2.0-85019084617
cris.boxmetadata.label.isbn
9781509048472
cris.boxmetadata.label.containerissn
10514651
cris.boxmetadata.label.containerisbn
978-150904847-2
cris.boxmetadata.label.conference
Institute of Electrical and Electronics Engineers Inc. - 23rd International Conference on Pattern Recognition, ICPR 2016
cris.boxmetadata.label.sponsor
This research was supported in part by NSF grants 1217676. The authors would like to thank the Brazilian National Research Council - CNPq (Grant # 234272/2014-7)
peru-layout.shadow-copies Directorio de Producción Científica Scopus