ABSTRACT
The continued integration of the computational and biological sciences has revolutionized genomic and proteomic studies. However, efficient collaboration between these fields requires the creation of shared standards. A common problem arises when biological input does not properly fit the expectations of the algorithm, which can result in misinterpretation of the output. This potential confounding of input/output is a drawback especially when regarding motif finding software. Here we propose a method for improving output by selecting input based upon evolutionary distance, domain architecture, and known function. This method improved detection of both known and unknown motifs in two separate case studies. By standardizing input considerations, both biologists and bioinformaticians can better interpret and design the evolving sophistication of bioinformatic software.
- Hedges SB, J Blair, M Venturi, J Shoe. A molecular timescale of eukaryote evolution and the rise of complex multicellular life. BMC Evolutionary Biology. 2004; 4:2.Google Scholar
- Kumar S, Filipski A, Swarna V, Walker A, Hedges SB. Placing confidence limits on the molecular age of the human-chimpanzee divergence. Proceedings of the Natural Academy of Sciences. 27 Dec 2005; 102(52):18842--18847.Google ScholarCross Ref
- Quest D, K Dempsey, M Shafiullah, D Bastola, and H Ali. MTAP: A Motif Tool Assessment Pipeline for Automated Assessment of De Novo Regulatory Motif Discovery Tool. BMC Bioinformatics. 2008 Aug 12; 9 Suppl 9:S6.Google Scholar
- Tompa M, N Li, T Bailey, G Church, B DeMoor, E Eskin, A Favorov, M Frith, Y Fu, W Kent, V Makeev, A Mironov, W Noble, G Pavesi, G Pesole, M Regnier, N Simonis, S Sinha, G Thijs, J. van Helden, M Vandenbogaert, Z Weng, C Workman, C Ye, and Z Zhu. Assessing Computational Tools for the Discovery of Transcription Factor Binding Sites. Z Nature Biotechnology. 1 Jan 2005; 23(1):137--144.Google Scholar
- Zheng, J., et al., Prestin is the motor protein of cochlear outer hair cells. Nature, 2000. 405(6783): p. 149--55.Google Scholar
- Dorwart, M. R., et al., The solute carrier 26 family of proteins in epithelial ion transport. Physiology (Bethesda), 2008. 23: p. 104--14.Google ScholarCross Ref
- Yarov-Yarovoy V, Baker D, Catterall WA. Voltage sensor conformations in the open and closed states in ROSETTA structural models of K(+) channels. Proc Natl Acad Sci USA, 2006 May 9; 103(19):7292--7. Epub 2006 Apr 28.Google ScholarCross Ref
- Haitin Y, Yisharel I, Malka E, Shamgar L, Schottelndreier H, Peretz A, Paas Y, Attali B. S1 constraints in the voltage sensor domain of Kv7.1 K+ channels. PLoS One. 2008 Apr 9; 3(4):e1935.Google Scholar
- Heginbotham K, Lu Z, Abramson T, MacKinnon R. Mutations in the K+ channel signature sequence. Biophys J. 1994 Apr; 66(4):1061--7.Google Scholar
- Crooks GE, Hon G, Chandonia JM, Brenner SE WebLogo: A sequence logo generator, Genome Res, 14:1188--1190, (2004)Google ScholarCross Ref
- An intelligent data-centric approach toward identification of conserved motifs in protein sequences
Recommendations
Finding motifs for insufficient number of sequences with strong binding to transcription facto
RECOMB '04: Proceedings of the eighth annual international conference on Research in computational molecular biologyFinding motifs is an important problem in computational biology. Our paper makes two major contributions to this problem. Firstly, we better characterize the types of problem instances that cannot be solved by most existing methods of finding motifs. ...
Identification of Context-Dependent Motifs by Contrasting ChIP Binding Data
Motivation: DNA binding proteins play crucial roles in the regulation of gene expression. Transcription factors (TFs) activate or repress genes directly while other proteins influence chromatin structure for transcription. Binding sites of a TF ...
Comments