Data clustering is an important task in the field of data mining. In many real applications, clustering algorithms must consider the order of data, resulting in the problem of clustering sequential data. For instance, analyzing the moving pattern of an object and detecting community structure in a complex network are related to sequential data clustering. The constraint of the continuous region prevents previous clustering algorithms from being directly applied to the problem. A dynamic programming algorithm was proposed to address the issue, which returns the optimal sequential data clustering. However, it is not scalable and hence the practicality is limited. This paper revisits the solution and enhances it by introducing a greedy stopping condition. This condition halts the algorithm’s search process when it is likely that the optimal solution has been found. Experimental results on multiple datasets show that the algorithm is much faster than its original solution while the optimality gap is negligible.
Data Availability
The Abdominal and Direct Fetal ECG Database analyzed during the current study is available in the https://physionet.org/content/adfecgdb/1.0.0/, the Brno University of Technology ECG Signal Database is available in the https://physionet.org/content/but-pdb/1.0.0/, the MIT-BIH arrhythmia database is available in the https://physionet.org/content/mitdb/1.0.0/, and the QT database is available in https://physionet.org/content/qtdb/1.0.0/ repository.
