Interactive Quality-Oriented Data Warehouse Development

Maurizio Pighin; Lucio Ieronutti

doi:10.4018/978-1-60566-232-9.ch004

Special Offers
- IGI Global’s New Emerging Topic e-Book Collections
  Acquire highly focused and affordable Cutting-Edge Peer-Reviewed Research Content through a selection of 17 topic-focused e-Book Collections discounted up to 90%, compared to list prices. Collection topics include Artificial Intelligence, Data Science, Language Learning, Marketing and Customer Relations, Sustainability, and many more. Hosted on the InfoSci^® platform, these collections feature no DRM, no additional cost for multi-user licensing, no embargo of content, full-text PDF & HTML format, and more.
  Learn More
- Open Access Book (Free Access) - Encyclopedia of Information Science and Technology, Sixth Edition (ISBN: 9781668473665)
  The Encyclopedia of Information Science and Technology, Sixth Edition) continues the legacy set forth by the first five editions by providing comprehensive coverage and up-to-date definitions of the most important issues, concepts, and trends pertaining to technological advancements and information management within a variety of settings and industries. The entire book is being published under open access.
  Read Now
- Open Access Book (Free Access) - Food Sustainability, Environmental Awareness, and Adaptation and Mitigation Strategies for Developing Countries (ISBN: 9781668456293)
  Food Sustainability, Environmental Awareness, and Adaptation and Mitigation Strategies for Developing Countries provides information on the recent technology, mitigation, and environmental protection that must be applied for food sustainability in developing countries. This book is being published under Platinum Open Access through funding from Diponegoro University, Indonesia.
  Read Now
- Open Access Book (Free Access) - New Models of Higher Education: Unbundled, Rebundled, Customized, and DIY (ISBN: 9781668438091)
  The Walmart Corporation and the Lumina Foundation have provided funding to make New Models of Higher Education: Unbundled, Rebundled, Customized, and DIY fully open access, completely removing any paywall between scholars in education and the latest research on new models for the future of higher education.
  Read Now
- Open Access Book (Free Access) - Handbook of Research on the Global View of Open Access and Scholarly Communications (ISBN: 9781799898054)
  Through a collaboration between IGI Global and the University of North Texas, the Handbook of Research on the Global View of Open Access and Scholarly Communications has been published as fully open access, completely removing any paywall between researchers of any field, and the latest research on the equitable and inclusive nature of Open Access and all of its complications.
  Read Now
Books
- - Books by Subject
  - Business, Administration, & Management
  - Scientific, Technical, & Medical (STM)
  - Education
  - Books by Field
Journals
- - Journals
  - OnDemand Journal Articles
  - Journals by Subject
  - Business, Administration, & Management
  - Scientific, Technical, & Medical (STM)
  - Education
  - Journals by Field
e-Collections
Open Access
- View All Open Access Opportunities
  Search across all of IGI Global’s available open access publishing opportunities to unleash your research potential.
  Find an Open Access Journal for Your Next Manuscript
  Search across all of IGI Global’s available open access publishing opportunities to unleash your research potential.
  Submit an Open Access Book Proposal
  Learn more about open access book publishing and how it can propel your research forward in the field.
  Convert Your Work to Open Access
  Already published? You can convert your work to open access to increase its impact through IGI Global’s Restrospective Open Access Program.
  Utilize Open Access Collection Database
  Open up your research potential by utilizing our open access content or integrating the open access collection into your library
  Consider Open Access Agreements
  For Libraries: consider no-cost or investment-level open access agreements with IGI Global to support your faculty's research endeavors.
  Search Funding Resources
  Looking for additional funding resources to support your open accesss endeavors? View industry resources compiled by our open access team.
  Review Open Access Policies & Ethical Guidelines
  Considering IGI Global to publish your work under open access? Review IGI Global’s open access policies and ethical guidelines
Publish with Us
Resources
- - Instructors
  - Course Adoption
  - Teaching Cases
  - K-12 Online Learning Collection
  - Authors and Editors
  - eEditorial Discovery^® System
  - Peer Review Process
  - Ethics and Malpractice
  - COPE Membership
  - Fair Use Policy
  - Open Access Publishing
  - FAQ
Catalogs
About Us
Newsroom

Interactive Quality-Oriented Data Warehouse Development

Maurizio Pighin, Lucio Ieronutti

Source Title: Progressive Methods in Data Warehousing and Business Intelligence: Concepts and Competitive Analytics

DOI: 10.4018/978-1-60566-232-9.ch004

OnDemand:

(Individual Chapters)

Available

$37.50

Current Special Offers

No Current Special Offers

Abstract

Data Warehouses are increasingly used by commercial organizations to extract, from a huge amount of transactional data, concise information useful for supporting decision processes. However, the task of designing a data warehouse and evaluating its effectiveness is not trivial, especially in the case of large databases and in presence of redundant information. The meaning and the quality of selected attributes heavily influence the data warehouse’s effectiveness and the quality of derived decisions. Our research is focused on interactive methodologies and techniques targeted at supporting the data warehouse design and evaluation by taking into account the quality of initial data. In this chapter we propose an approach for supporting the data warehouses development and refinement, providing practical examples and demonstrating the effectiveness of our solution. Our approach is mainly based on two phases: the first one is targeted at interactively guiding the attributes selection by providing quantitative information measuring different statistical and syntactical aspects of data, while the second phase, based on a set of 3D visualizations, gives the opportunity of run-time refining taken design choices according to data examination and analysis. For experimenting proposed solutions on real data, we have developed a tool, called ELDA (EvaLuation DAta warehouse quality), that has been used for supporting the data warehouse design and evaluation.

Chapter Preview

Top

Introduction

Data Warehouses are widely used by commercial organizations to extract from an huge amount of transactional data concise information useful for supporting decision processes. For example, organization managers greatly benefit from the availability of tools and techniques targeted at deriving information on sale trends and discovering unusual accounting movements. With respect to the entire amount of data stored into the initial database (or databases, hereinafter DBs), such analysis is centered on a limited subset of attributes (i.e., datawarehouse measures and dimensions). As a result, the datawarehouse (hereinafter DW) effectiveness and the quality of related decision is strongly influenced by the semantics of selected attributes and the quality of initial data. For example, information on customers and suppliers as well as products ordered and sold are very meaningful from data analysis point of view due to their semantics. However, the availability of information measuring and representing different aspect of data can make easier the task of selecting DW attributes, especially in presence of multiple choices (i.e., redundant information) and in the case of DBs characterized by an high number of attributes, tables and relations. Quantitative measurements allows DW engineers to better focus their attention towards the attributes characterized by the most desirable features, while qualitative data representations enables one to interactively and intuitively examine the considered data subset, allowing one to reduce the time required for the DW design and evaluation.

Our research is focused on interactive methodologies and techniques aimed at supporting the DW design and evaluation by taking into account the quality of initial data. In this chapter we propose an approach supporting the DW development and refinement, providing practical examples demonstrating the effectiveness of our solution. Proposed methodology can be effectively used (i) during the DW construction phase for driving and interactively refining the attributes selection, and (ii) at the end of the design process, to evaluate the quality of taken DW design choices.

While most solutions that have been proposed in the literature for assessing data quality are related with semantics, our goal is to propose an interactive approach focused on statistical aspects of data. The approach is mainly composed by two phases: an analytical phase based on a set of metrics measuring different data features (quantitative information), and an exploration phase based on an innovative graphical representation of DW ipercubes that allows one to navigate intuitively through the information space to better examine the quality and data distribution (qualitative information). The interaction is one of the most important feature of our approach: the designer can incrementally defines the DW measures and dimensions and both quality measurements and data representations change according to such modifications. This solution allows one to evaluate rapidly and intuitively the effects of alternative design choices. For example, the designer can immediately discover that the inclusion of an attribute negatively influences the global DW quality. If the quantitative evaluation does not convince the designer, he can explore the DW ipercubes to better understand relations among data, data distributions and behaviors.

In a real world scenario, DW engineers greatly benefit from the possibility of obtaining concise and easy-to-understand information describing the data actually stored into the DB, since they typically have a partial knowledge and vision of a specific operational DB (e.g., how an organization really uses the commercial system). Indeed, different commercial organizations can use the same information system, but each DB instantiation stores data that can be different from the point of view of distribution, correctness and reliability (e.g., an organization never fills a particular field of the form). As a result, the same DW design choices can produce different informative effects depending on the data actually stored into the DB. Then, although the attributes selection is primarily based on data semantics, the availability of both quantitative and qualitative information on data could greatly support the DW design phase. For example, in the presence of alternative choices (valid from semantic point of view), the designer can select the attribute characterized by the most desirable syntactical and statistical features. On the other hand, the designer can decide to change his design choice if he discovers that the selected attribute is characterized by undesirable features (for instance, an high percentage of null values).

Complete Chapter List

Search this Book:

Reset

MLA

APA

Chicago

Export Reference

Interactive Quality-Oriented Data Warehouse Development

Abstract

Introduction

Complete Chapter List