Direct estimation of genotype fitness from time series

Avatar
Poster
Voice is AI-generated
Connected to paperThis paper is a preprint and has not been certified by peer review

Direct estimation of genotype fitness from time series

Authors

Mohanty, V.; Shakhnovich, E.

Abstract

Heterogeneous adapting populations, whether in laboratory evolution experiments or global-scale pandemics, experience complex evolutionary dynamics due to the interplay of selection, mutation, and stochasticity. Inference of individual genotypes' fitnesses therefore becomes difficult, especially when many lineages are competing and data are noisy. Existing fitness inference methods tend to rely on assumptions on the fitness landscape's maximum order of epistasis, or they require complicated iterative optimization algorithms to converge on fitness estimates. Here, we show that fitness landscapes can be computed from time series data, without any restrictions on epistatic order or iterative optimization, using a simple, closed-form mathematical expression that is easily implemented with standard matrix operations used commonly in linear algebra. We demonstrate successful fitness inference from noisy in silico evolutionary dynamics from four different noisy microscopic processes, including Wright-Fisher, Moran, ProSeD (serial dilution), and barcoded passage simulations. Then, we illustrate the broad applicability of the equation to five experimental time series datasets, including barcoded yeast evolution experiments, murine norovirus-1 serial passage experiments, and SARS-CoV-2 global genomic prevalence data. Our formula successfully infers fitnesses for even for rare genotypes several orders of magnitude less prevalent than top lineages, works with both laboratory evolution and epidemiological data, and can be implemented in most modern scientific programming languages.

Follow Us on

0 comments

Add comment