Research

Understanding the limits of data science and designing efficient algorithms.

My research lies at the intersection of information theory, machine learning, and computational biology. I am interested in what can be learned or recovered from limited observations, and how to use those observations efficiently.

Efficient and adaptive algorithms for data science

How can we make learning algorithms faster by looking at less data? My work explores randomized sampling and adaptive computation as tools for reducing the cost of core data science and machine learning tasks. The goal is to spend computation where it is most informative, while retaining the accuracy and reliability that make an algorithm useful. Applications include clustering, similarity search, graph learning, and eigenvector estimation.

Three stages of adaptive sampling in the US airport network. Node size and color indicate estimated eigenvector centrality, with major hubs becoming prominent as sampling progresses.
Identifying hubs in the US airport network with the Adaptive Power Method. Each iteration samples new edges using the current centrality estimates, shown by node size and color, to focus observations on the most informative parts of the network.

Selected papers

Information theory for DNA storage, genomics, and unordered data

DNA storage and genome sequencing share an unusual feature: information is observed through a set of out-of-order fragments. What can be recovered from these observations, and how much information can such systems reliably preserve? I am interested in mathematical models of fragmentation, sampling, noise, and loss of order that allow us to study fundamental limits and guide the design of coding and reconstruction methods.

Encoding digital data into DNA molecules, storing them in a pool, then sampling and decoding the molecules.
Writing and reading data in DNA: digital information is encoded into molecules, stored, and recovered through sequencing and decoding.

Selected papers

Computational genomics and biological data analysis

How can we turn large collections of biological measurements into useful information? I am interested in developing computational methods for comparing, aligning, assembling, and grouping DNA sequences, as well as analyzing data from different biological data modalities. A recurring theme is connecting the structure of the data to efficient algorithms: using statistical and information-theoretic insights to decide which computations are needed and which can be avoided.

Selected papers

Statistical inference and learning with structure

Observations often carry structure: neighboring variables may be correlated, samples may come from distinct clusters, or a sequence may be visible only through incomplete traces. I study how these dependencies affect what can be inferred and how many observations are needed. This direction connects fundamental recovery limits with algorithm design, with problems ranging from group testing to clustering and sequence reconstruction.

Selected papers

Information theory and wireless communications

My earlier research studied the fundamental limits of communication in wireless networks. How should information be relayed across multiple hops, how can interference be managed, and what rates are achievable when several users communicate simultaneously? This work examines network capacity, degrees of freedom, compression, and energy efficiency, using mathematical models to understand both the possibilities and the limitations of communication systems.

Selected papers