Desbordante project announces version 2.5.0, a significant release approximately one year after the previous stable build. The delay allowed contributors to accumulate a substantial set of additions across dependency mining, graph algorithms, and Python integration.

Event sequences and episode mining

Desbordante 2.5.0 introduces a new data type: event sequences. This format opens the tool to problem formulations where the order and timing of events matters. Three mining variants ship with this release: frequent episode mining, maximal episode mining, and top-k episode mining. Each variant carries its own algorithm, and the distribution includes working examples to help users get started.

Conditional Inclusion Dependencies with Cinderella and PLI-CIND

Among the pattern-based features, Conditional Inclusion Dependencies (CIND) returns as a refined take on standard Inclusion Dependencies. Two search algorithms—Cinderella and PLI-CIND—accompany a new validator. CIND extends IND by revealing the specific conditions under which an inclusion dependency holds true when restricted to a subset of the data, giving analysts more granular control over dependency discovery.

Approximate and Sequential Order Dependencies

The release also broadens the section on order-related dependencies. Approximate Order Dependencies (set-based) now offers both mining and validation: if data is sorted by one attribute set, the resulting order will be nearly identical when sorted by another. A separate Sequential Dependency validator operates on the same principle but adds a numeric constraint—after sorting by one attribute set, values in the second set must differ by no more than a defined threshold. Mining for this second pattern remains under active development and is slated for the next release.

Graph mining: gSpan and Graph Differential Dependencies

Graph-pattern capabilities see two additions. Frequent subgraph mining now uses the gSpan algorithm, a well-established method for discovering recurring substructures within graph data. Additionally, support for Graph Differential Dependencies (GDD) includes a validator that checks differential constraints across graph subsets. The maintainers note that GDD validation performance will be addressed in a future release, with speed improvements planned.

Domain PACs for region estimation

Desbordante adds its first PAC-pattern type with Domain PAC. Though only the validator is available in this release, it serves as a powerful tool for estimating the proportion of data points falling within a specified region and for exploring that region’s neighborhood. Domain PAC is the first of three PAC patterns planned for the project.

Differential Dependency validator with user-defined metrics

The Differential Dependency validator receives a notable enhancement: users can now supply a user-defined metric (UDM) directly from Python. The bindings and an accompanying example have been updated to demonstrate how to incorporate a custom metric, expanding the validator’s flexibility for domain-specific workloads.

FastADC: parallelized Approximate Denial Constraints

FastADC, the algorithm for discovering Approximate Denial Constraints, has been significantly sped up. Parallelization is among the changes contributing to performance gains; in testing, speedups of up to three times were observed, though the exact factor depends on the dataset characteristics.

Conditional Functional Dependencies: unified mining and verification

Conditional Functional Dependencies (CFD) mining and verification have been unified under a single representation. The mining example has been revised to use the verifier, streamlining the workflow. Future releases aim to broaden the selection of CFD discovery algorithms beyond what is available today.

Classical association rules validator

Support for association rules expands with the addition of a validator for classical association rules, completing a set of rule-types that had previously been partially covered.

Approximate Functional Dependencies: expanded metric support

Work on Approximate Functional Dependencies (AFDs) continues. When the project began, approximate dependencies were limited to the g1 metric. Recent publications have demonstrated that other metrics carry meaningful information, and this release expands Desbordante’s FD mining capabilities to support additional metrics. While the feature set is still maturing—for example, TANE output now reports error values and the example has been updated—support for the remaining metrics is planned for forthcoming versions.

Python bindings: equality, hashing, and compatibility

Python bindings see targeted improvements. Implementations of __eq__ and __hash__ now appear in the bindings for IND, MD, NAR, and SFD patterns. Additionally, compatibility has been added between the validator and the object returned by the miner for IND and UCC patterns. No examples accompany these bindings yet.

Per-column statistical summaries

The release includes several simple per-column statistical summaries. Users can obtain counts of entries in a column that satisfy various conditions, including filters for string values, character values, and boolean values.

Internal cleanup and logging improvements

Under the hood, internal dependencies have been cleaned up. This refactoring enables logging messages to be forwarded into Python idiomatically, so the library’s logging API—named "desbordante"—can be used directly from Python code.