Sequenziergerät und Proben im Genomiklabor

Projekt

Artificial life as a benchmark for molecular evolution and genomics

We propose to implement an original computer simulation test principle for evolutionary genomic inference methods. These methods, although used daily in areas as diverse as health, agriculture, biodiversity protection or justice, make historical inferences difficult to test experimentally. The weakness of the present…

We propose to implement an original computer simulation test principle for evolutionary genomic inference methods. These methods, although used daily in areas as diverse as health, agriculture, biodiversity protection or justice, make historical inferences difficult to test experimentally. The weakness of the present evaluation systems is to insert in simulations the same simplifying hypotheses as in the inference methods, because they are developed by the same designers, for validation purposes. For example, genes are defined a priori as evolutionary units, which makes their annotation and classification trivial. Extinct species are usually not simulated if no measurement is made from them, even if they interfere through hybridization or horizontal gene transfer. We propose a new test bench approach rather than a validation approach, where the development teams and the test team will be separate. We are bringing together two teams, one in phylogeny, the other in artificial life, to build "blind" simulations of inference methods. Preliminary tests have proved that this universally recognized scientific principle (blind tests) but never used in evolutionary studies, could reveal unexpected pitfalls of the methods, and correct them. We will make the results of the simulations available to the community as a validation standard. This proof of principle is encouraging for the effectiveness of our approach. We will generalize this principle to evolutionary studies by adapting a program resulting from artificial life, Aevol. Importantly Aevol was not designed to generate a test bench, which makes it paradoxically interesting to be used as such. We will organize the collaboration between the two teams in a mode of mutually addressed challenges, respecting a communication on biological processes, and avoiding communication on computer models, which should as far as possible remain separate. We will organize an international competition to promote this approach and test a large number of methods. The project will therefore also have the effect of enhancing cooperation between teams and best practices of method comparisons. In particular, we will apply the benchmark to modern phylogeny methods developed in the team, integrating several evolutionary scales and their interactions.