Document Type

Article

Publication Date

4-4-2022

Publication Source

IEEE Transactions on Neural Networks and Learning Systems

Abstract

In this article, we adopt the maximizing mutual information (MI) approach to tackle the problem of unsupervised learning of binary hash codes for efficient cross-modal retrieval. We proposed a novel method, dubbed cross-modal info-max hashing (CMIMH). First, to learn informative representations that can preserve both intramodal and intermodal similarities, we leverage the recent advances in estimating variational lower bound of MI to maximizing the MI between the binary representations and input features and between binary representations of different modalities. By jointly maximizing these MIs under the assumption that the binary representations are modeled by multivariate Bernoulli distributions, we can learn binary representations, which can preserve both intramodal and intermodal similarities, effectively in a mini-batch manner with gradient descent. Furthermore, we find out that trying to minimize the modality gap by learning similar binary representations for the same instance from different modalities could result in less informative representations. Hence, balancing between reducing the modality gap and losing modality-private information is important for the cross-modal retrieval tasks. Quantitative evaluations on standard benchmark datasets demonstrate that the proposed method consistently outperforms other state-of-the-art cross-modal retrieval methods.

Inclusive pages

2023-01-14

ISBN/ISSN

2162-237X

Comments

The article available for download is the authors' accepted manuscript, provided in compliance with the publisher's policy on self-archiving. Permission documentation is on file.

Publisher

IEEE

Peer Reviewed

true

Keywords

Task analysis, Semantics, Binary codes, Correlation, Representation learning, Training, Matrix decomposition, Cross-modal retrieval, multi-modal, mutual information (MI), representation learning, unsupervised hashing, variational information maximization

Link to published version

Share

COinS