Non-overlapping Common Substrings Allowing Mutations

Chan, H. L.; Lam, T. W.; Sung, W. K.; Wong, P. W. H.; Yiu, S. M.

doi:10.1007/s11786-007-0030-6

Non-overlapping Common Substrings Allowing Mutations

Published: 01 April 2008

Volume 1, pages 543–555, (2008)
Cite this article

Mathematics in Computer Science Aims and scope Submit manuscript

H. L. Chan¹,
T. W. Lam²,
W. K. Sung³,
P. W. H. Wong⁴ &
…
S. M. Yiu²

70 Accesses
1 Citation
Explore all metrics

Abstract.

This paper studies several combinatorial problems arising from finding the conserved genes of two genomes (i.e., the entire DNA of two species). The input is a collection of n maximal common substrings of the two genomes. The problem is to find, based on different criteria, a subset of such common substrings with maximum total length. The most basic criterion requires that the common substrings selected have the same ordering in the two genomes and they do not overlap among themselves in either genome. To capture mutations (transpositions and reversals) between the genomes, we do not insist the substrings selected to have the same ordering. Conceptually, we allow one ordering to go through some mutations to become the other ordering. If arbitrary mutations are allowed, the problem of finding a maximum-length, non-overlapping subset of substrings is found to be NP-hard. However, arbitrary mutations probably overmodel the problem and are likely to find more noise than conserved genes. We consider two criteria that attempt to model sparse and non-overlapping mutations. We show that both can be solved in polynomial time using dynamic programming.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Author information

Authors and Affiliations

Department of Computer Science, University of Pittsburgh, Room 6406, Sennott Square Building, 210 South Bouquet Street, Pittsburgh, PA, 15260, USA
H. L. Chan
Department of Computer Science, The University of Hong Kong, Chow Yei Ching Building, Pokfulam Road, Hong Kong, China
T. W. Lam & S. M. Yiu
Department of Computer Science, School of Computing, National University of Singapore, Computing 1, Law Link, Singapore, 117590, Republic of Singapore
W. K. Sung
Department of Computer Science, The University of Liverpool, Ashton Building, Ashton Street, Liverpool, L69 3BX, United Kingdom
P. W. H. Wong

Authors

H. L. Chan
View author publications
You can also search for this author in PubMed Google Scholar
T. W. Lam
View author publications
You can also search for this author in PubMed Google Scholar
W. K. Sung
View author publications
You can also search for this author in PubMed Google Scholar
P. W. H. Wong
View author publications
You can also search for this author in PubMed Google Scholar
S. M. Yiu
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding authors

Correspondence to H. L. Chan or S. M. Yiu.

Rights and permissions

Reprints and permissions

About this article

Cite this article

Chan, H.L., Lam, T.W., Sung, W.K. et al. Non-overlapping Common Substrings Allowing Mutations. Math.comput.sci. 1, 543–555 (2008). https://doi.org/10.1007/s11786-007-0030-6

Download citation

Received: 08 June 2007
Revised: 03 September 2007
Accepted: 15 October 2007
Published: 01 April 2008
Issue Date: June 2008
DOI: https://doi.org/10.1007/s11786-007-0030-6

Mathematics Subject Classification (2000).

Keywords.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Non-overlapping Common Substrings Allowing Mutations

Abstract.

Access this article

Similar content being viewed by others

The Complexity of Finding Common Partitions of Genomes with Predefined Block Sizes

Duplication-Loss Genome Alignment: Complexity and Algorithm

Aligning and Labeling Genomes under the Duplication-Loss Model

Author information

Authors and Affiliations

Corresponding authors

Rights and permissions

About this article

Cite this article

Mathematics Subject Classification (2000).

Keywords.

Navigation

Non-overlapping Common Substrings Allowing Mutations

Abstract.

Access this article

Similar content being viewed by others

The Complexity of Finding Common Partitions of Genomes with Predefined Block Sizes

Duplication-Loss Genome Alignment: Complexity and Algorithm

Aligning and Labeling Genomes under the Duplication-Loss Model

Author information

Authors and Affiliations

Corresponding authors

Rights and permissions

About this article

Cite this article

Share this article

Mathematics Subject Classification (2000).

Keywords.

Search

Navigation