Mining for Putative Regulatory Elements in the Yeast Genome Using Gene Expression Data

Authors

Jaak Vilo and Alvis Brazma

European Bioinformatics Institute; Inge Jonassen

University of Bergen; Alan Robinson

European Bioinformatics Institute; Esko Ukkonen

University of Helsinki

Proceedings:

Proceedings of the Twentieth International Conference on Machine Learning, 2000

Volume

Issue:

Proceedings of the Twentieth International Conference on Machine Learning, 2000

Track:

Contents

Downloads:

Download PDF

Abstract:

We have developed a set of methods and tools for automatic discovery of putative regulatory signals in genome sequences. The analysis pipeline consists of gene expression data clustering, sequence pattern discovery from upstream sequences of genes, a control experiment for pattern significance threshold limit detection, selection of interesting patterns, grouping of these patterns, representing the pattern groups in a concise form and evaluating the discovered putative signals against existing databases of regulatory signals. The pattern discovery is computationally the most expensive and crucial step. Our tool performs a rapid exhaustive search for a priori unknown statistically significant sequence patterns of unrestricted length. The statistical significance is determined for a set of sequences in each cluster with respect to a set of back-ground sequences allowing the detection of subtle regulatory signals specific for each cluster. The potentially large number of significant patterns is reduced to a small number of groups by clustering them by mutual similarity. Automatically derived consensus patterns of these groups represent the results in a comprehensive way for a human investigator. We have performed a systematic analysis for the yeast Sac-charomyces cerevisiae. We created a large number of inde-pendent clusterings of expression data simultaneously assessing the goodness of each cluster. For each of the over 52000 clusters acquired in this way we discovered significant pat-terns in the upstream sequences of respective genes. We selected nearly 1500 significant patterns by formal criteria and matched them against the experimentally mapped transcription factor binding sites in the SCPD database. We clustered the 1500 patterns to 62 groups for which we derived automat-ically alignments and consensus patterns. Of these 62 groups 48 had patterns that have matching sites in SCPD database.

ISMB

Proceedings of the Twentieth International Conference on Machine Learning, 2000

Cookie	Duration	Description
cookielawinfo-checkbox-analytics	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
cookielawinfo-checkbox-functional	11 months	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
cookielawinfo-checkbox-necessary	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
cookielawinfo-checkbox-others	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
cookielawinfo-checkbox-performance	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
viewed_cookie_policy	11 months	The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.