Continuous Post-Mining of Association Rules in a Data Stream Management System

Hetal Thakkar; Barzan Mozafari; Carlo Zaniolo

doi:10.4018/978-1-60566-404-0.ch007

Special Offers
- IGI Global’s New Emerging Topic e-Book Collections
  Acquire highly focused and affordable Cutting-Edge Peer-Reviewed Research Content through a selection of 17 topic-focused e-Book Collections discounted up to 90%, compared to list prices. Collection topics include Artificial Intelligence, Data Science, Language Learning, Marketing and Customer Relations, Sustainability, and many more. Hosted on the InfoSci^® platform, these collections feature no DRM, no additional cost for multi-user licensing, no embargo of content, full-text PDF & HTML format, and more.
  Learn More
- Open Access Book (Free Access) - Encyclopedia of Information Science and Technology, Sixth Edition (ISBN: 9781668473665)
  The Encyclopedia of Information Science and Technology, Sixth Edition) continues the legacy set forth by the first five editions by providing comprehensive coverage and up-to-date definitions of the most important issues, concepts, and trends pertaining to technological advancements and information management within a variety of settings and industries. The entire book is being published under open access.
  Read Now
- Open Access Book (Free Access) - Food Sustainability, Environmental Awareness, and Adaptation and Mitigation Strategies for Developing Countries (ISBN: 9781668456293)
  Food Sustainability, Environmental Awareness, and Adaptation and Mitigation Strategies for Developing Countries provides information on the recent technology, mitigation, and environmental protection that must be applied for food sustainability in developing countries. This book is being published under Platinum Open Access through funding from Diponegoro University, Indonesia.
  Read Now
- Open Access Book (Free Access) - New Models of Higher Education: Unbundled, Rebundled, Customized, and DIY (ISBN: 9781668438091)
  The Walmart Corporation and the Lumina Foundation have provided funding to make New Models of Higher Education: Unbundled, Rebundled, Customized, and DIY fully open access, completely removing any paywall between scholars in education and the latest research on new models for the future of higher education.
  Read Now
- Open Access Book (Free Access) - Handbook of Research on the Global View of Open Access and Scholarly Communications (ISBN: 9781799898054)
  Through a collaboration between IGI Global and the University of North Texas, the Handbook of Research on the Global View of Open Access and Scholarly Communications has been published as fully open access, completely removing any paywall between researchers of any field, and the latest research on the equitable and inclusive nature of Open Access and all of its complications.
  Read Now
Books
- - Books by Subject
  - Business, Administration, & Management
  - Scientific, Technical, & Medical (STM)
  - Education
  - Books by Field
Journals
- - Journals
  - OnDemand Journal Articles
  - Journals by Subject
  - Business, Administration, & Management
  - Scientific, Technical, & Medical (STM)
  - Education
  - Journals by Field
e-Collections
Open Access
- View All Open Access Opportunities
  Search across all of IGI Global’s available open access publishing opportunities to unleash your research potential.
  Find an Open Access Journal for Your Next Manuscript
  Search across all of IGI Global’s available open access publishing opportunities to unleash your research potential.
  Submit an Open Access Book Proposal
  Learn more about open access book publishing and how it can propel your research forward in the field.
  Convert Your Work to Open Access
  Already published? You can convert your work to open access to increase its impact through IGI Global’s Restrospective Open Access Program.
  Utilize Open Access Collection Database
  Open up your research potential by utilizing our open access content or integrating the open access collection into your library
  Consider Open Access Agreements
  For Libraries: consider no-cost or investment-level open access agreements with IGI Global to support your faculty's research endeavors.
  Search Funding Resources
  Looking for additional funding resources to support your open accesss endeavors? View industry resources compiled by our open access team.
  Review Open Access Policies & Ethical Guidelines
  Considering IGI Global to publish your work under open access? Review IGI Global’s open access policies and ethical guidelines
Publish with Us
Resources
- - Instructors
  - Course Adoption
  - Teaching Cases
  - K-12 Online Learning Collection
  - Authors and Editors
  - eEditorial Discovery^® System
  - Peer Review Process
  - Ethics and Malpractice
  - COPE Membership
  - Fair Use Policy
  - Open Access Publishing
  - FAQ
Catalogs
About Us
Newsroom

Continuous Post-Mining of Association Rules in a Data Stream Management System

Hetal Thakkar, Barzan Mozafari, Carlo Zaniolo

Source Title: Post-Mining of Association Rules: Techniques for Effective Knowledge Extraction

DOI: 10.4018/978-1-60566-404-0.ch007

OnDemand:

(Individual Chapters)

Available

$37.50

Current Special Offers

No Current Special Offers

Abstract

The real-time (or just-on-time) requirement associated with online association rule mining implies the need to expedite the analysis and validation of the many candidate rules, which are typically created from the discovered frequent patterns. Moreover, the mining process, from data cleaning to post-mining, can no longer be structured as a sequence of steps performed by the analyst, but must be streamlined into a workflow supported by an efficient system providing quality of service guarantees that are expected from modern Data Stream Management Systems (DSMSs). This chapter describes the architecture and techniques used to achieve this advanced functionality in the Stream Mill Miner (SMM) prototype, an SQL-based DSMS designed to support continuous mining queries.

Chapter Preview

Top

Introduction

Driven by the need to support a variety of applications, such as click stream analysis, intrusion detection, and web-purchase recommendation systems, much of recent research work has focused on the difficult problem of mining data streams for association rules. The paramount concern in previous works was how to devise frequent itemset algorithms that are fast and light enough for mining massive data streams continuously with real-time or quasi real-time response (Jiang, 2006). The problem of post-mining the association rules, so derived from the data streams, has so far received much less attention, although it is rich in practical importance and research challenges. Indeed, the challenge of validating the large number of generated rules is even harder in the time-constrained environment of on-line data mining, than it is in the traditional off-line environment. On the other hand, data stream mining is by nature a continuous and incremental process, which makes it possible to apply application-specific knowledge and meta-knowledge acquired in the past, to accelerate the search for new rules. Therefore, previous post-mining results can be used to prune and expedite both (i) the current search for new frequent patterns and (ii) the post processing of the candidate rules thus derived. These considerations have motivated the introduction of efficient and tightly-coupled primitives for mining and post-mining association rules in Stream Mill Miner (SMM), a DSMS designed for mining applications (Thakkar, 2008). SMM is the first of its kind and thus must address a full gamut of interrelated challenges pertaining to (i) functionality, (ii) performance, and (iii) usability. Toward that goal, SMM supports

•
A rich library of mining methods and operators that are fast and light enough to be used for online mining of massive and often bursty data streams,
•
The management of the complete DM process as a workflow, which (i) begins with the preprocessing of data (e.g., cleaning and normalization), (ii) continues with the core mining task (e.g., frequent pattern extraction), and (iii) completes with post-mining tasks for rule extraction, validation, and historical preservation.
•
Usability based on high-level, user-friendly interfaces, but also customizability and extensibility to meet the specific demands of different classes of users.

Performing these tasks efficiently on data streams has proven difficult for all mining methods, but particularly so, for association rule mining. Indeed, in many applications, such as click stream analysis and intrusion detection, time is of the essence, and new rules must be promptly deployed, while current ones must be revised in a timely manner to adapt to concept shifts and drifts. However, many of the generated rules are either trivial or nonsensical, and thus validation by the analyst is required, before they can be applied in the field. While this human validation step cannot be completely skipped, it can be greatly expedited by the approach taken in SMM where the bulk of the candidate rules are filtered out by the system, so that only a few highly prioritized rules are sent to analyst for validation. Figure 1 shows the architecture used in SMM to support the rule mining and post-mining process.

Figure 1.

Post-mining flow

As shown in Figure 1, the first step consists in mining the data streams for frequent itemsets using an algorithm called SWIM (Mozafari, 2008) that incrementally mines the data stream partitioned into slides. As shown in Figure 1, SWIM can be directed by the analyst to (i) accept/reject interesting and uninteresting items, and (ii) monitor particular patterns of interest while avoiding the generation of others. Once SWIM generates these frequent patterns, a Rule Extractor module derives interesting association rules based on these patterns. Furthermore, SMM supports the following customization options to expedite the process:

Complete Chapter List

Search this Book:

Reset

MLA

APA

Chicago

Export Reference

Continuous Post-Mining of Association Rules in a Data Stream Management System

Abstract

Introduction

Complete Chapter List