Probing the Geometry of Data with Diffusion Fréchet Functions

Many complex ecosystems, such as those formed by multiple microbial taxa, involve intricate interactions amongst various sub-communities. The most basic relationships are frequently modeled as co-occurrence networks in which the nodes represent the various players in the community and the weighted edges encode levels of interaction. In this setting, the composition of a community may be viewed as a probability distribution on the nodes of the network. This paper develops methods for modeling the organization of such data, as well as their Euclidean counterparts, across spatial scales. Using the notion of diffusion distance, we introduce diffusion Frechet functions and diffusion Frechet vectors associated with probability distributions on Euclidean space and the vertex set of a weighted network, respectively. We prove that these functional statistics are stable with respect to the Wasserstein distance between probability measures, thus yielding robust descriptors of their shapes. We apply the methodology to investigate bacterial communities in the human gut, seeking to characterize divergence from intestinal homeostasis in patients with Clostridium difficile infection (CDI) and the effects of fecal microbiota transplantation, a treatment used in CDI patients that has proven to be significantly more effective than traditional treatment with antibiotics. The proposed method proves useful in deriving a biomarker that might help elucidate the mechanisms that drive these processes.


page 1

page 2

page 3

page 4


Permutation invariant networks to learn Wasserstein metrics

Understanding the space of probability measures on a metric space equipp...

Design and analysis of computer experiments with both numeral and distribution inputs

Nowadays stochastic computer simulations with both numeral and distribut...

The Shape of Data and Probability Measures

We introduce the notion of multiscale covariance tensor fields (CTF) ass...

Local Network Community Detection with Continuous Optimization of Conductance and Weighted Kernel K-Means

Local network community detection is the task of finding a single commun...

Local Distribution Obfuscation via Probability Coupling

We introduce a general model for the local obfuscation of probability di...

PhD dissertation to infer multiple networks from microbial data

The interactions among the constituent members of a microbial community ...

Exchange-Based Diffusion in Hb-Graphs: Highlighting Complex Relationships

Most networks tend to show complex and multiple relationships between en...

Please sign up or login with your details

Forgot password? Click here to reset