Serveur d'exploration MERS

Attention, ce site est en cours de développement !
Attention, site généré par des moyens informatiques à partir de corpus bruts.
Les informations ne sont donc pas validées.

Integrating genomic data to predict transcription factor binding.

Identifieur interne : 002292 ( PubMed/Corpus ); précédent : 002291; suivant : 002293

Integrating genomic data to predict transcription factor binding.

Auteurs : Dustin T. Holloway ; Mark Kon ; Charles Delisi

Source :

RBID : pubmed:16362910

English descriptors

Abstract

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

PubMed: 16362910

Links to Exploration step

pubmed:16362910

Le document en format XML

<record>
<TEI>
<teiHeader>
<fileDesc>
<titleStmt>
<title xml:lang="en">Integrating genomic data to predict transcription factor binding.</title>
<author>
<name sortKey="Holloway, Dustin T" sort="Holloway, Dustin T" uniqKey="Holloway D" first="Dustin T" last="Holloway">Dustin T. Holloway</name>
<affiliation>
<nlm:affiliation>Molecular Biology Cell Biology and Biochemistry, Boston University, Boston, MA 02215, USA. dth128@bu.edu</nlm:affiliation>
</affiliation>
</author>
<author>
<name sortKey="Kon, Mark" sort="Kon, Mark" uniqKey="Kon M" first="Mark" last="Kon">Mark Kon</name>
</author>
<author>
<name sortKey="Delisi, Charles" sort="Delisi, Charles" uniqKey="Delisi C" first="Charles" last="Delisi">Charles Delisi</name>
</author>
</titleStmt>
<publicationStmt>
<idno type="wicri:source">PubMed</idno>
<date when="2005">2005</date>
<idno type="RBID">pubmed:16362910</idno>
<idno type="pmid">16362910</idno>
<idno type="wicri:Area/PubMed/Corpus">002292</idno>
<idno type="wicri:explorRef" wicri:stream="PubMed" wicri:step="Corpus" wicri:corpus="PubMed">002292</idno>
</publicationStmt>
<sourceDesc>
<biblStruct>
<analytic>
<title xml:lang="en">Integrating genomic data to predict transcription factor binding.</title>
<author>
<name sortKey="Holloway, Dustin T" sort="Holloway, Dustin T" uniqKey="Holloway D" first="Dustin T" last="Holloway">Dustin T. Holloway</name>
<affiliation>
<nlm:affiliation>Molecular Biology Cell Biology and Biochemistry, Boston University, Boston, MA 02215, USA. dth128@bu.edu</nlm:affiliation>
</affiliation>
</author>
<author>
<name sortKey="Kon, Mark" sort="Kon, Mark" uniqKey="Kon M" first="Mark" last="Kon">Mark Kon</name>
</author>
<author>
<name sortKey="Delisi, Charles" sort="Delisi, Charles" uniqKey="Delisi C" first="Charles" last="Delisi">Charles Delisi</name>
</author>
</analytic>
<series>
<title level="j">Genome informatics. International Conference on Genome Informatics</title>
<idno type="ISSN">0919-9454</idno>
<imprint>
<date when="2005" type="published">2005</date>
</imprint>
</series>
</biblStruct>
</sourceDesc>
</fileDesc>
<profileDesc>
<textClass>
<keywords scheme="KwdEn" xml:lang="en">
<term>Algorithms</term>
<term>Base Sequence</term>
<term>Bayes Theorem</term>
<term>Binding Sites</term>
<term>Chromatin Immunoprecipitation</term>
<term>Cluster Analysis</term>
<term>Computational Biology</term>
<term>Evolution, Molecular</term>
<term>Gene Expression Profiling</term>
<term>Gene Expression Regulation, Fungal</term>
<term>Genes, Fungal</term>
<term>Genome, Fungal</term>
<term>Phylogeny</term>
<term>Promoter Regions, Genetic</term>
<term>Protein Binding</term>
<term>Saccharomyces cerevisiae (genetics)</term>
<term>Transcription Factors (genetics)</term>
<term>Transcription Factors (metabolism)</term>
</keywords>
<keywords scheme="MESH" type="chemical" qualifier="genetics" xml:lang="en">
<term>Transcription Factors</term>
</keywords>
<keywords scheme="MESH" qualifier="genetics" xml:lang="en">
<term>Saccharomyces cerevisiae</term>
</keywords>
<keywords scheme="MESH" type="chemical" qualifier="metabolism" xml:lang="en">
<term>Transcription Factors</term>
</keywords>
<keywords scheme="MESH" xml:lang="en">
<term>Algorithms</term>
<term>Base Sequence</term>
<term>Bayes Theorem</term>
<term>Binding Sites</term>
<term>Chromatin Immunoprecipitation</term>
<term>Cluster Analysis</term>
<term>Computational Biology</term>
<term>Evolution, Molecular</term>
<term>Gene Expression Profiling</term>
<term>Gene Expression Regulation, Fungal</term>
<term>Genes, Fungal</term>
<term>Genome, Fungal</term>
<term>Phylogeny</term>
<term>Promoter Regions, Genetic</term>
<term>Protein Binding</term>
</keywords>
</textClass>
</profileDesc>
</teiHeader>
<front>
<div type="abstract" xml:lang="en">Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.</div>
</front>
</TEI>
<pubmed>
<MedlineCitation Status="MEDLINE" Owner="NLM">
<PMID Version="1">16362910</PMID>
<DateCompleted>
<Year>2006</Year>
<Month>01</Month>
<Day>18</Day>
</DateCompleted>
<DateRevised>
<Year>2008</Year>
<Month>11</Month>
<Day>21</Day>
</DateRevised>
<Article PubModel="Print">
<Journal>
<ISSN IssnType="Print">0919-9454</ISSN>
<JournalIssue CitedMedium="Print">
<Volume>16</Volume>
<Issue>1</Issue>
<PubDate>
<Year>2005</Year>
</PubDate>
</JournalIssue>
<Title>Genome informatics. International Conference on Genome Informatics</Title>
<ISOAbbreviation>Genome Inform</ISOAbbreviation>
</Journal>
<ArticleTitle>Integrating genomic data to predict transcription factor binding.</ArticleTitle>
<Pagination>
<MedlinePgn>83-94</MedlinePgn>
</Pagination>
<Abstract>
<AbstractText>Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.</AbstractText>
</Abstract>
<AuthorList CompleteYN="Y">
<Author ValidYN="Y">
<LastName>Holloway</LastName>
<ForeName>Dustin T</ForeName>
<Initials>DT</Initials>
<AffiliationInfo>
<Affiliation>Molecular Biology Cell Biology and Biochemistry, Boston University, Boston, MA 02215, USA. dth128@bu.edu</Affiliation>
</AffiliationInfo>
</Author>
<Author ValidYN="Y">
<LastName>Kon</LastName>
<ForeName>Mark</ForeName>
<Initials>M</Initials>
</Author>
<Author ValidYN="Y">
<LastName>DeLisi</LastName>
<ForeName>Charles</ForeName>
<Initials>C</Initials>
</Author>
</AuthorList>
<Language>eng</Language>
<PublicationTypeList>
<PublicationType UI="D003160">Comparative Study</PublicationType>
<PublicationType UI="D016428">Journal Article</PublicationType>
</PublicationTypeList>
</Article>
<MedlineJournalInfo>
<Country>Japan</Country>
<MedlineTA>Genome Inform</MedlineTA>
<NlmUniqueID>101280573</NlmUniqueID>
<ISSNLinking>0919-9454</ISSNLinking>
</MedlineJournalInfo>
<ChemicalList>
<Chemical>
<RegistryNumber>0</RegistryNumber>
<NameOfSubstance UI="D014157">Transcription Factors</NameOfSubstance>
</Chemical>
</ChemicalList>
<CitationSubset>IM</CitationSubset>
<MeshHeadingList>
<MeshHeading>
<DescriptorName UI="D000465" MajorTopicYN="N">Algorithms</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D001483" MajorTopicYN="N">Base Sequence</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D001499" MajorTopicYN="N">Bayes Theorem</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D001665" MajorTopicYN="N">Binding Sites</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D047369" MajorTopicYN="N">Chromatin Immunoprecipitation</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D016000" MajorTopicYN="N">Cluster Analysis</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D019295" MajorTopicYN="N">Computational Biology</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D019143" MajorTopicYN="N">Evolution, Molecular</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D020869" MajorTopicYN="N">Gene Expression Profiling</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D015966" MajorTopicYN="N">Gene Expression Regulation, Fungal</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D005800" MajorTopicYN="N">Genes, Fungal</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D016681" MajorTopicYN="Y">Genome, Fungal</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D010802" MajorTopicYN="N">Phylogeny</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D011401" MajorTopicYN="N">Promoter Regions, Genetic</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D011485" MajorTopicYN="N">Protein Binding</DescriptorName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D012441" MajorTopicYN="N">Saccharomyces cerevisiae</DescriptorName>
<QualifierName UI="Q000235" MajorTopicYN="Y">genetics</QualifierName>
</MeshHeading>
<MeshHeading>
<DescriptorName UI="D014157" MajorTopicYN="N">Transcription Factors</DescriptorName>
<QualifierName UI="Q000235" MajorTopicYN="N">genetics</QualifierName>
<QualifierName UI="Q000378" MajorTopicYN="Y">metabolism</QualifierName>
</MeshHeading>
</MeshHeadingList>
</MedlineCitation>
<PubmedData>
<History>
<PubMedPubDate PubStatus="pubmed">
<Year>2005</Year>
<Month>12</Month>
<Day>20</Day>
<Hour>9</Hour>
<Minute>0</Minute>
</PubMedPubDate>
<PubMedPubDate PubStatus="medline">
<Year>2006</Year>
<Month>1</Month>
<Day>19</Day>
<Hour>9</Hour>
<Minute>0</Minute>
</PubMedPubDate>
<PubMedPubDate PubStatus="entrez">
<Year>2005</Year>
<Month>12</Month>
<Day>20</Day>
<Hour>9</Hour>
<Minute>0</Minute>
</PubMedPubDate>
</History>
<PublicationStatus>ppublish</PublicationStatus>
<ArticleIdList>
<ArticleId IdType="pubmed">16362910</ArticleId>
<ArticleId IdType="pii">161083</ArticleId>
</ArticleIdList>
</PubmedData>
</pubmed>
</record>

Pour manipuler ce document sous Unix (Dilib)

EXPLOR_STEP=$WICRI_ROOT/Sante/explor/MersV1/Data/PubMed/Corpus
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 002292 | SxmlIndent | more

Ou

HfdSelect -h $EXPLOR_AREA/Data/PubMed/Corpus/biblio.hfd -nk 002292 | SxmlIndent | more

Pour mettre un lien sur cette page dans le réseau Wicri

{{Explor lien
   |wiki=    Sante
   |area=    MersV1
   |flux=    PubMed
   |étape=   Corpus
   |type=    RBID
   |clé=     pubmed:16362910
   |texte=   Integrating genomic data to predict transcription factor binding.
}}

Pour générer des pages wiki

HfdIndexSelect -h $EXPLOR_AREA/Data/PubMed/Corpus/RBID.i   -Sk "pubmed:16362910" \
       | HfdSelect -Kh $EXPLOR_AREA/Data/PubMed/Corpus/biblio.hfd   \
       | NlmPubMed2Wicri -a MersV1 

Wicri

This area was generated with Dilib version V0.6.33.
Data generation: Mon Apr 20 23:26:43 2020. Site generation: Sat Mar 27 09:06:09 2021