{"entity": "publication", "iuid": "d680cb93ceb34a73bb9413d7e2a9b0e6", "timestamp": "2026-09-28T23:04:47.229Z", "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/publication/d680cb93ceb34a73bb9413d7e2a9b0e6.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/publication/d680cb93ceb34a73bb9413d7e2a9b0e6"}}, "title": "microTaboo: a general and practical solution to the k-disjoint problem.", "authors": [{"family": "Al-Jaff", "given": "Mohammed", "initials": "M"}, {"family": "Sandstr\u00f6m", "given": "Eric", "initials": "E"}, {"family": "Grabherr", "given": "Manfred", "initials": "M", "orcid": "0000-0001-8792-6508", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/b7d96bfeead545498f188dae001abcef.json"}}], "type": "journal article", "published": "2017-05-02", "journal": {"title": "BMC Bioinformatics", "issn": "1471-2105", "volume": "18", "issue": "1", "pages": "228", "issn-l": "1471-2105"}, "abstract": "A common challenge in bioinformatics is to identify short sub-sequences that are unique in a set of genomes or reference sequences, which can efficiently be achieved by k-mer (k consecutive nucleotides) counting. However, there are several areas that would benefit from a more stringent definition of \"unique\", requiring that these sub-sequences of length W differ by more than k mismatches (i.e. a Hamming distance greater than k) from any other sub-sequence, which we term the k-disjoint problem. Examples include finding sequences unique to a pathogen for probe-based infection diagnostics; reducing off-target hits for re-sequencing or genome editing; detecting sequence (e.g. phage or viral) insertions; and multiple substitution mutations. Since both sensitivity and specificity are critical, an exhaustive, yet efficient solution is desirable.\n\nWe present microTaboo, a method that allows for efficient and extensive sequence mining of unique (k-disjoint) sequences of up to 100 nucleotides in length. On a number of simulated and real data sets ranging from microbe- to mammalian-size genomes, we show that microTaboo is able to efficiently find all sub-sequences of a specified length W that do not occur within a threshold of k mismatches in any other sub-sequence. We exemplify that microTaboo has many practical applications, including point substitution detection, sequence insertion detection, padlock probe target search, and candidate CRISPR target mining.\n\nmicroTaboo implements a solution to the k-disjoint problem in an alignment- and assembly free manner. microTaboo is available for Windows, Mac OS X, and Linux, running Java 7 and higher, under the GNU GPLv3 license, at: https://MohammedAlJaff.github.io/microTaboo.", "doi": "10.1186/s12859-017-1644-6", "pmid": "28464826", "labels": [], "xrefs": [{"db": "pmc", "key": "PMC5414201"}, {"db": "pii", "key": "10.1186/s12859-017-1644-6"}], "notes": [], "created": "2026-09-23T14:56:36.413Z", "modified": "2026-09-23T14:56:36.429Z"}