Best Paper Award
Efficient Correlated Sequence Queries for Large-scale DNA Databases[IEICE TRANS. INF. & SYST.,Vol. J108–D No.5 2025]
Correlated sequence queries, which extract subsequences that frequently co-occur with a given DNA sequence from large-scale DNA databases, constitute a fundamental technique in genome sequence assembly and biomarker discovery. Conventional approaches have primarily relied on string similarity searches based on edit distance or sequence alignment. However, since these methods evaluate only syntactic similarity, they lack robustness when sequence lengths differ significantly and are inherently limited in their ability to directly capture correlations grounded in biological function.
This paper addresses these challenges by newly formalizing the correlation query problem and proposing an efficient method for retrieving the top-k subsequences with the highest correlation values to a given query from large-scale DNA databases. A particularly noteworthy contribution lies in the reduction of the problem to a graph search formulation. By representing DNA sequences as graph structures, the method leverages pruning techniques developed in frequent subgraph mining to eliminate, at an early stage of the search, candidates that are theoretically guaranteed not to be included in the top-k results. This leads to a substantial reduction in computational cost. Furthermore, the paper demonstrates through experiments on real datasets that the proposed pruning strategy preserves exactness without compromising result accuracy.
The primary contribution of this work is the realization of an exact and efficient top-k search for computationally challenging correlation queries, achieved without resorting to approximation or heuristic methods. This is accomplished through the rigorous development and application of theoretical upper bounds on correlation values, ensuring that no false negatives are introduced. In light of the rapid advancement of data-driven genome analysis, the mathematical rigor and robustness of the proposed method, which guarantees the completeness of analytical results, are of particular significance. For these reasons, this paper represents an outstanding contribution to the field and is deemed worthy of the Best Paper Award.