Computing Similarity Queries for Correlated Gaussian Sources

01/22/2020
by   Hanwei Wu, et al.
0

Among many current data processing systems, the objectives are often not the reproduction of data, but to compute some answers based on the data resulting from queries. The similarity identification task is to identify the items in a database that are similar to a given query item for a given metric. The problem of compression for similarity identification has been studied in arXiv:1307.6609 [cs.IT]. Unlike classical compression problems, the focus is not on reconstructing the original data. Instead, the compression rate is determined by the desired reliability of the answers. Specifically, the information measure identification rate characterizes the minimum rate that can be achieved among all schemes which guarantee reliable answers with respect to a given similarity threshold. In this paper, we propose a component-based model for computing correlated similarity queries. The correlated signals are first decorrelated by the KLT transform. Then, the decorrelated signal is processed by a distinct D-admissible system for each component. We show that the component-based model equipped with KLT can perfectly represent the multivariate Gaussian similarity queries when optimal rate-similarity allocation applies. Hence, we can derive the identification rate of the multivariate Gaussian signals based on the component-based model. We then extend the result to general Gaussian sources with memory. We also study the models equipped with practical compone systems. We use TC- schemes that use type covering signatures and triangle-inequality decision rules as our component systems. We propose an iterative method to numerically approximate the minimum achievable rate of the TC- scheme. We show that our component-based model equipped with TC- schemes can achieve better performance than the TC- scheme unaided on handling the multivariate Gaussian sources.

READ FULL TEXT
research
04/29/2019

Learning Image Information for eCommerce Queries

Computing similarity between a query and a document is fundamental in an...
research
05/16/2022

Characterization of the Gray-Wyner Rate Region for Multivariate Gaussian Sources: Optimality of Gaussian Auxiliary RV

Examined in this paper, is the Gray and Wyner achievable lossy rate regi...
research
07/18/2018

Robust Distributed Compression of Symmetrically Correlated Gaussian Sources

Consider a lossy compression system with ℓ distributed encoders and a ce...
research
09/16/2021

SEACOW: Synopsis Embedded Array Compression using Wavelet Transform

Recently, multidimensional data is produced in various domains; because ...
research
11/23/2018

Selected Methods for non-Gaussian Data Analysis

The basic goal of computer engineering is the analysis of data. Such dat...
research
06/25/2019

Coding for Crowdsourced Classification with XOR Queries

This paper models the crowdsourced labeling/classification problem as a ...
research
07/18/2022

Don't Be a Tattle-Tale: Preventing Leakages through Data Dependencies on Access Control Protected Data

We study the problem of answering queries when (part of) the data may be...

Please sign up or login with your details

Forgot password? Click here to reset