Tampilkan postingan dengan label Data Warehousing. Tampilkan semua postingan
Tampilkan postingan dengan label Data Warehousing. Tampilkan semua postingan

Rabu, 15 Juni 2011

Oracle Warehouse Builder 11g R2: Getting Started 2011 (with code)


Oracle Warehouse Builder 11g R2: Getting Started 2011 (with code) Summary:
Publisher: Pa,ckt Pub,lishi,ng 2011 | 424 Pages | ISBN: 1849683441 | PDF | 10 MB

In today’s economy, businesses and IT professionals cannot afford to lag behind the latest technologies. Data warehousing is a critical area to the success of many enterprises, and Oracle Warehouse Builder is a powerful tool for building data warehouses. It comes free with the latest version of the Oracle database. Written in an accessible, informative, and focused manner, this book will teach you to use Oracle Warehouse Builder to build your data warehouse. Covering warehouse design, the import of source data, the ETL cycle and more, this book will have you up and running in next to no time. This book will walk you through the complete process of planning, building, and deploying a data warehouse using Oracle Warehouse Builder. By the book’s end, you will have built your own data warehouse from scratch. Starting with the installation of the Oracle Database and Warehouse Builder software, this book then covers the analysis of source data, designing a data warehouse, and extracting, transforming, and loading data from the source system into the data warehouse. You’ll follow the whole process with detailed screenshots of key steps along the way that have all been updated for the new Fusion Client Platform interface in 11gR2, alongside numerous tips and hints not covered by the official documentation. You’ll finish up with a brand new chapter on code templates where you’ll implement a complete mapping using JDBC connectivity and code template mappings.
  • Configure your Oracle database to communicate with a non-Oracle database using Oracle Heterogeneous Services
  • Build, view, and edit a cube and its dimensions, using Warehouse Builder Wizards and the Data Object Editor
  • Discover the underlying star schema relational structure that Warehouse Builder will build to implement your cube and dimensions
  • Recognize what a staging area is, and build a staging table and mapping to load it from SQL Server database tables
  • Master various operators for use in mappings, for the transformation and flow of data
  • Build and execute mappings to load dimensions and a cube using the Warehouse Builder Mapping Editor and Control Center Manager
  • Detailed explanation of the new interface in 11gR2
  • Implement JDBC connectivity to a remote database and access source data using a code template mapping
or
or
or
or

Jumat, 25 Desember 2009

Oracle DBA Guide to Data Warehousing and Star Schemas


Oracle DBA Guide to Data Warehousing and Star Schemas Summary:
Prentice Hall | ISBN: 0130325848 | 2003 | CHM | 240 pages | 1.53 MB


This book addresses all aspects of constructing star schemas within Oracle data warehouses, from modeling and design through high-speed loads and lightning fast queries. The book delivers meaningful examples complemented by empirical samples and benchmarks,
such that readers will learn more than just the mechanics. This book transforms readers into subject matter experts for dimensional modeling, star schemas and data warehousing in general for the Oracle database environment. This book is based on research conducted for the multi-terabyte data warehouse for the 7-Eleven Corporation. Star schema: a data warehouse design that enhances the performance of multidimensional queries on traditional relational databases.


or

Jumat, 30 Oktober 2009

Professional SQL Server 2000 Data Warehousing with Analysis Services



Chris Graves,Mark Scott,Mike Benkovich,Paul Turley,Robert Skoglund,Robin Dewson,Sakhr Youness,Denny Lee,Sam Ferguson,Tony Bain,Terrence Joubert, "Professional SQL Server 2000 Data Warehousing with Analysis Services"
Publisher: Peer Information Inc.; 1st edition | 2001 | 701 Pages | ISBN: 1861005407 | PDF | 14 MB

Data warehouses have evolved to cope with the huge volumes of data flowing through the workplace by separating the data used for reporting and decision making from the operational systems. The purpose of the data warehouse is simply to store the raw data, and in combination with Microsoft SQL Server 2000 Analysis Services, this data can be transformed into accessible information that reflects the real factors affecting the enterprise.

In this book, we introduce the key concepts of data warehousing, OLAP, and data mining. In addition to coverage of Data Transformation Services (DTS) and MDX, this book also demonstrates how to develop an Analysis Services client application, and how to secure and optimize your data warehouse. There is also an in-depth discussion of the exciting new topic of Web Housing.

By reading this book, you will learn how best to employ data warehousing and OLAP in your business, and how to leverage it to provide your organization with improved revenue and profitability.

This book covers:

Understanding Analysis Services Architecture
Designing Data Warehouses and Data Marts
Using Data Transformation Services (DTS) in Data Warehousing
Techniques for Data Mining and Analysis
Building OLAP cubes with Analysis Manager, and programmatically through DSO
Securing, Administrating and Optimizing a Data Warehouse and OLAP system
Using Multidimensional Expressions (MDX) to query OLAP cubes
Building OLAP client applications with Visual Basic and ASP
English Query, PivotTable Service

Download:
Link_1
Or
Link_2

Kamis, 29 Oktober 2009

Data Warehousing Design and Advanced Engineering Applications: Methods for Complex Construction



Ladjel Bellatreche, "Data Warehousing Design and Advanced Engineering Applications: Methods for Complex Construction" 
Information Science Reference | 2009 | ISBN: 1605667560 | 362 pages | PDF | 10 MB

Data warehousing and online analysis technologies have shown their effectiveness in managing and analyzing a large amount of disparate data, attracting much attention from numerous research communities. Data Warehousing Design and Advanced Engineering Applications: Methods for Complex Construction covers the complete process of analyzing data to extract, transform, load, and manage the essential components of a data warehousing system. A defining collection of field discoveries, this advanced title provides significant industry solutions for those involved in this distinct research community.

Download:
Link_1
Or
Link_2

Sabtu, 24 Oktober 2009

Strategic Data Warehousing: Achieving Alignment with Business By Neera Bhansali



Strategic Data Warehousing: Achieving Alignment with Business By Neera Bhansali
Publisher: AUERBACH 2009-07-29 | 224 Pages | ISBN: 1420083945 | PDF | 2.2 MB

Strategic Data Warehousing: Achieving Alignment with Business provides an integrated approach to achieving successful and sustainable alignment of data warehouses and business goals. It details the roles and responsibilities of the data warehouse and business managers in achieving strategic alignment, technical integration, and improved flexibility. Complete with case studies depicting real-world scenarios, the text examines the organizational, user, data, and technological factors proven to promote successful data warehousing, provides actionable solutions for achieving strategic alignment, and includes a model that readers can apply in aligning their own data warehouse needs and business goals.

Download:
Link_1
Or
Link_2

Senin, 12 Oktober 2009

Data Warehousing: Architecture and Implementation



Data Warehousing: Architecture and Implementation 
Prentice Hall PTR | January 9, 1999 | ISBN-10: 0130809020 | 360 pages | PDF | 2.7 mb

This book is intended for Information Technology (IT) professionals who have been hearing about or have been tasked to evaluate, learn or implement data warehousing technologies. Far from being just a passing fad, data warehousing technology has grown much in scale and reputation in the past few years, as evidenced by the increasing number of products, vendors, organizations, and yes, even books, devoted to the subject. Enterprises that have successfully implemented data warehouses find it strategic and often wonder how they ever managed to survive without it in the past. As early as 1995, a Gartner Group survey of Fortune 500 IT managers found that 90 percent of all organizations had planned to implement data warehouses by 1998. Virtually all Top-100 US banks will actively use a data warehouse-based profitability application by 1998. Nearly 30 percent of companies that actively pursue this technology have created a permanent or semipermanent unit to plan, create, maintain, promote, and support the data warehouse. If you are an IT professional who has been tasked with planning, managing, designing, implementing, supporting, or maintaining your organization's data warehouse, then this book is intended for you. The first section introduces the Enterprise Architecture and Data Warehouse concepts, the basis of the reasons for writing this book. The second section of this book focuses on three of the key People in any data warehousing initiative: the Project Sponsor, the CIO, and the Project Manager. This section is devoted to addressing the primary concerns of these individuals.

Download:
Link_1
Or
Link_2

Selasa, 29 September 2009

Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development


Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development Summary:
Publisher: Information Science Reference | 426 pages | 2009-07-31 | ISBN: 160566748X | English | PDF | 16.25 MB


Recently, researchers have focused on challenging problems facing the development of data warehousing, knowledge discovery, and data mining applications.


Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development: Innovative Methods and Applications provides a comprehensive analysis of current issues and trends in retrieval expansion. Containing research from leading international experts, this book presents future challenges and opportunities in the field valuable to academicians, researchers, and practitioners.


Table of Contents:


Section I: DWH Architectures & Fundamentals:


Three chapters in Section I, Data Warehouse Architectures & Fundamentals present the current trends of research on Data Warehouse architecture, storage and implementations which towards on improving performance and response time.


Chapter I: The LBF R-tree: Scalable Indexing and Storage for Data Warehousing Systems


In Chapter I, the authors propose a LBF R-tree framework for effective indexing mechanisms in multi-dimensional database environment. The proposed framework addresses not only improves performance on common user-defined range queries, but also gracefully degrades to a linear scan of the data on pathologically large queries. Experimental results demonstrating both efficient disk access on the LBF R-tree, as well as impressive compression ratios for data and indexes.


Chapter II: Dynamic Workload for Schema Evolution in Data Warehouses: a Performance Issue


Chapter II addresses the issues related to the workload’s evolution and maintenance in data warehouse systems in response to new requirements modeling resulting from users’ personalized analysis needs. The proposed workload management system assists the administrator to maintain and adapt dynamically the workload according to changes arising on the data warehouse schema by improving two types of workload updates: (1) maintaining existing queries consistent with respect to the new data warehouse schema and (2) creating new queries based on the new dimension hierarchy levels.


Chapter III: Preview: Optimizing View Materialization Cost in Spatial Data Warehouses


Chapter III presents an optimization approach for materialized view implementation in Spatial Data Warehouse. Due to the fact that spatial data are larger in size and spatial operations are more complex than the traditional relational operations, both the view materialization cost and the on-the-fly computation cost are often extremely high. The authors propose a new notion, called preview, for which both the materialization and on-the-fly costs are significantly smaller than those of the traditional views, so that the total cost is optimized.


Section II: Multidimensional Data and OLAP


Section II consists of three chapters discussing related issues and challenges in multidimensional database and Online Analytical Processing (OLAP) environment.


Chapter IV: Decisional Annotations: Integrating and Preserving Decision-Makers’ Expertise in Multidimensional Systems


Chapter IV deals with an annotation-based decisional system. The decisional system is based on multidimensional databases, which are composed of facts and dimensions. The expertise of decision-makers is modeled, shared and stored through annotations. Every piece of multidimensional data can be associated with zero or more annotations which allow decision-makers to carry on active analysis and to collaborate with other decision-makers on a common analysis.


Chapter V: Federated Data Warehouses


Chapter V discuss Federated Data Warehouse Systems which consist of a collection of Data Marts provided by different enterprises or public organizations and widen the knowledge base for business analysts, thus enabling better founded strategic decisions. The authors argue that the integration of heterogeneous Data Marts at the logical schema level is preferable to the migration of data into a physically new system if the involved organizations remain autonomous. They present a federated architecture that provides a global multi-dimensional schema to which the Data Mart schemas are semantically mapped, repairing all heterogeneities.


Chapter VI: Built-In Indicators to Support Business Intelligence in OLAP Databases


Chapter VI describes algorithms to support business intelligence in OLAP databases applying Data Mining methods to the multidimensional environment. Those methods help end-users’ analysis in two ways. First, they identify the most interesting dimensions to expand in order to explore the data. Then, they automatically detect interesting cells among a user selected ones. The solution is based on a tight coupling between OLAP tools and statistical methods, based upon the built-in indicators computed instantaneously during the end-users’ exploration of the data cube.


Section III: DWH and OLAP Applications


The next three chapters in Section III, DWH & OLAP Applications, present some typical applications using Data Warehouse and OLAP technology as well as the challenges and issues facing in the real practice.


Chapter VII: Conceptual Data Warehouse Design Methodology for Business Process Intelligence


Chapter VII presents a conceptual framework for adopting the data warehousing technology for business process analysis, with Surgical Workflows Analysis as a challenging real-world application. Deficiencies of the conventional OLAP approach are overcome by proposing an extended multidimensional data model, which enables adequate capturing of flow-oriented data. The model supports a number of advanced properties, such as non-quantitative and heterogeneous facts, many-to-many relationships between facts and dimensions, full and partial dimension sharing, dynamic specification of new measures, and interchangeability of fact and dimension roles.


Chapter VIII: Data Warehouse Facilitating Evidence-Based Medicine


Deployment of a federated data warehouse approach for the integration of the wide range of different medical data sources and for distribution of evidence-based clinical knowledge, to support clinical decision makers, primarily clinicians at the point of care is the main topic of chapter VIII: Data Warehouse Facilitating Evidence-Based Medicine. A real-world scenario is used to illustrate the possible application field in the area of emergency and intensive care in which the evidence-based medicine merges data originating in a pharmacy database, a social insurance company database and diverse clinical DWHs with the minimized administration effort.


Chapter IX: Deploying Data Warehouses in Grids with Efficiency and Availability


Chapter IX, Deploying Data Warehouses in Grids with Efficiency and Availability, discusses the deployment of data warehouses over Grids. The authors present the Grid-NPDW architecture, which aims at providing high throughput and data availability in grid-based warehouses. High efficiency in situations with site failure is also achieved with the use of on-demand query scheduling and data partitioning and replication. The Chapter also describes the main components of the Grid-NPDW Scheduler and presents some experimental results of proposed strategies.


Section IV: Data Mining Techniques


Section IV, Data Mining Techniques, consists of three chapters discussing a variety of traditional data mining techniques such as clustering, ranking, classification but towards the efficiency and performance improvement.


Chapter X: MOSAIC: Agglomerative Clustering with Gabriel Graphs


Chapter X, MOSAIC: Agglomerative Clustering with Gabriel Graphs, introduces MOSAIC, a post-processing technique that has been designed to overcome these disadvantages. MOSAIC is an agglomerative technique that greedily merges neighboring clusters maximizing an externally given fitness function; Gabriel graphs are used to determine which clusters are neighboring, and non-convex shapes are approximated as the unions of small convex clusters. Experimental results are presented that show that using MOSAIC leads to clusters of higher quality compared to running a representative clustering algorithm stand-alone.


Chapter XI: Ranking Gradients in Multi-Dimensional Spaces


Chapter XI investigates how to mine and rank the most interesting changes in a multi-dimensional space applying a promising TOP-K gradient strategy. Interesting changes in customer behavior are usually discovered by gradient queries which are particular cases of multi-dimensional data analysis on large data warehouses. The main problem, however, arises from the fact that more interesting changes should be those ones having more dimensions in the gradient query (the curse-of-dimensionality dilemma). Besides, the number of interesting changes should be of a large amount (the preference selection criteria).


Chapter XII: Simultaneous Feature Selection and Tuple Selection for Efficient Classification


In Chapter XII, a method is proposed to combine feature selection and tuple selection to improve classification accuracy. Although feature selection and tuple selection have been studied earlier in various research areas such as machine learning, data mining, and so on, they have rarely been studied together. Feature selection and tuple selection help the classifier to focus better. The method is based on the principle that a representative subset has similar histogram as the full set. The proposed method uses this principle both to choose a subset of features and also to choose a subset of tuples. The empirical tests show that the proposed method performs better than several existing feature selection methods.


Section V: Advanced Mining Applications


The last four chapters in Section V, Advanced Mining Applications introduces innovative algorithms and applications in some emerging application fields in Data Mining and Knowledge Discovery, especially continuous data stream mining, which could not be solved by traditional mining technology.


Chapter XIII: Learning Cost-Sensitive Decision Trees to Support Medical Diagnosis


In Chapter XIII, Learning cost-sensitive decision trees to support medical, the authors discuss about diagnosis a cost-sensitive learning method. The chapter aims to enhance the understand of cost-sensitive learning problems in medicine and presents a strategy for learning and testing cost-sensitive decision trees, while considering several types of costs associated with problems in medicine. It begins with a contextualization and a discussion of the main types of costs. Then, reviews related work and presents a discussion about the evaluation of classifiers as well as explains a cost-sensitive decision tree strategy and presents some experimental results.


Chapter XIV: An Approximate Approach for Maintaining Recent Occurrences of Itemsets in a Sliding Window over Data Streams


Chapter XIV discuss about catching the recent trend of data when mining frequent itemsets over data streams. A data representation method, named frequency changing point (FCP), is introduced for monitoring the recent occurrence of itemsets over a data stream to prevent from storing the whole transaction data within a sliding window. The effect of old transactions on the mining result of recently frequent itemsets is diminished by performing adjusting rules on the monitoring data structure. Accordingly, the recently frequent itemsets or representative patterns are discovered from the maintained structure approximately. Experimental studies demonstrate that the proposed algorithms achieve high true positive rates and guarantees no false dismissal to the results yielded.


Chapter XV: Protocol Identification of Encrypted Network Streams


Chapter XV proposes a simple machine learning approach for protocol identification in network streams that have been encrypted, such that the only information available for identifying the underlying protocol of a connection was the size, timing and direction of packets. With very little information available from the network stream, it is possible to pinpoint potentially inappropriate activities for a workplace, institution or research center, such as using BitTorrent, GMail or MSN, and not confuse them with other common protocols such as HTTP and SSL.


Chapter XVI: Exploring Calendar-Based Pattern Mining in Data Streams


Chapter XVI introduces a calendar-based pattern mining aims at identifying patterns on specific calendar partitions in continuous data streams. The authors present how a data warehouse approach can be applied to leverage calendar-based pattern mining in data streams and how the framework of the DWFIST approach can cope with tight time constraints imposed by data streams, keep storage requirements at a manageable level and, at the same time, support calendar-based frequent itemset mining. The minimum granularity of analysis, parameters of the data warehouse (e.g. mining minimum support) and parameters of the database (e.g. extent size) provide ways to tune the load performance.


Data Quality and High-Dimensional Data Analysis




Data Quality and High-Dimensional Data Analysis Summary:
World Scientific Publishing Company | 2009 | ISBN: 9814273481 | 106 pages | PDF | 1,8 MB


Poor data quality is known to compromise the credibility and efficiency of commercial and public endeavours. Also, the importance of managing data quality has increased manifold as the diversity of sources, formats and volume of data grows. This volume targets the data quality in the light of collaborative information systems where data creation and ownership is increasingly difficult to establish.


Download 
or
or 

Senin, 28 September 2009

SAP Business Information Warehouse Reporting: Building Better BI with SAP BI 7.0



Peter Jones, "SAP Business Information Warehouse Reporting: Building Better BI with SAP BI 7.0" 
MgH | 2008 | ISBN: 0071496165 | 894 pages | PDF | 38,4 MB

Your Hands-On Guide to SAP Business Information Warehouse

Give your company the competitive edge by delivering up-to-date, pertinent business reports to users inside and outside your enterprise. SAP Business Information Warehouse Reporting shows you how to construct Enterprise Data Warehouses, create workbooks and queries, analyze and format results, and supply meaningful reports. Learn how to use the BEx and Web Analyzers, Web Application Designer, Visual Composer, and Information Broadcaster. You will also find out how to forecast future business trends, build enterprise portals and websites, and tune performance.
Group data into InfoCubes and DataStore Objects and generate reports using queries and workbooks
Work with the BEx Analyzer, Web Analyzer, and Query Designer
Build queries and reports using the Business Administration Workbench
Add attachments and drill-through using Document Integration and RRI
Format and distribute results using Report Designer and Information Broadcaster
Extend functionality with Enterprise Portal, Data Modeling, and Visual Composer
Deploy charts, maps, diagrams, and unit of measure conversions
Predict trends and possible outcomes using SBC and Integrated Planning
Generate HTML pages using Enterprise Reporting and Web Application Designer
Create BI-based corporate Web and intranet sites using SAP Enterprise Portal

Download:
Link_1
Or
Link_2

SAP® NetWeaver Portal Technology: The Complete Reference



Rabi Jay, "SAP® NetWeaver Portal Technology: The Complete Reference"
MgH | 2008 | ISBN: 007154853X | 735 pages | PDF | 25,2 MB

Your Hands-on Guide to SAP NetWeaver Portal Technology

Master SAP NetWeaver Portal with the most comprehensive, step-by-step reference available on the entire portal implementation life cycle. Written by SAP architect Rabi Jay, this book provides everything you need to plan, design, install, configure, and administer SAP NetWeaver Portal, including SAP NetWeaver Application Server Java.

SAP NetWeaver Portal Technology: The Complete Reference is filled with detailed descriptions, numerous illustrations, and hundreds of expert tips. Design and deploy portals with high availability, scalability, and performance. Implement single sign-on to backend systems and integrate SAP and non-SAP applications. Configure reliable J2EE engine and portal security, and devise a flawless portal backup and restore strategy. Improve performance using portal workload, GC, thread dump, and HTTP analysis.
Plan futuristically using PAM, release planning, and maintenance strategy
Design global portals using federated portal networks and external-facing portals
Implement self-registration and delegated user and content administration
Enable authorization using security zones, UME actions, and ACL permissions
Manage users centrally using LDAP, UME, and Identity Management
Implement user-, type-, and attribute-based authentication
Brand your portal using portal desktop rules, themes, and framework pages
Configure portal transports, and deploy patches and business packages using JSPM
Monitor your portal using CCMS and GRMG Availability Monitoring
Manage your portal centrally using NWA and maintain systems using SLD

Download:
Or

Selasa, 22 September 2009

Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development


Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development Summary:
Publisher: Information Science Reference | 426 pages | 2009-07-31 | ISBN: 160566748X | English | PDF | 16.25 MB


Recently, researchers have focused on challenging problems facing the development of data warehousing, knowledge discovery, and data mining applications.


Complex Data Warehousing and Knowledge Discovery for Advanced Retrieval Development: Innovative Methods and Applications provides a comprehensive analysis of current issues and trends in retrieval expansion. Containing research from leading international experts, this book presents future challenges and opportunities in the field valuable to academicians, researchers, and practitioners.


Table of Contents:


Section I: DWH Architectures & Fundamentals:


Three chapters in Section I, Data Warehouse Architectures & Fundamentals present the current trends of research on Data Warehouse architecture, storage and implementations which towards on improving performance and response time.


Chapter I: The LBF R-tree: Scalable Indexing and Storage for Data Warehousing Systems


In Chapter I, the authors propose a LBF R-tree framework for effective indexing mechanisms in multi-dimensional database environment. The proposed framework addresses not only improves performance on common user-defined range queries, but also gracefully degrades to a linear scan of the data on pathologically large queries. Experimental results demonstrating both efficient disk access on the LBF R-tree, as well as impressive compression ratios for data and indexes.


Chapter II: Dynamic Workload for Schema Evolution in Data Warehouses: a Performance Issue


Chapter II addresses the issues related to the workload’s evolution and maintenance in data warehouse systems in response to new requirements modeling resulting from users’ personalized analysis needs. The proposed workload management system assists the administrator to maintain and adapt dynamically the workload according to changes arising on the data warehouse schema by improving two types of workload updates: (1) maintaining existing queries consistent with respect to the new data warehouse schema and (2) creating new queries based on the new dimension hierarchy levels.


Chapter III: Preview: Optimizing View Materialization Cost in Spatial Data Warehouses


Chapter III presents an optimization approach for materialized view implementation in Spatial Data Warehouse. Due to the fact that spatial data are larger in size and spatial operations are more complex than the traditional relational operations, both the view materialization cost and the on-the-fly computation cost are often extremely high. The authors propose a new notion, called preview, for which both the materialization and on-the-fly costs are significantly smaller than those of the traditional views, so that the total cost is optimized.


Section II: Multidimensional Data and OLAP


Section II consists of three chapters discussing related issues and challenges in multidimensional database and Online Analytical Processing (OLAP) environment.


Chapter IV: Decisional Annotations: Integrating and Preserving Decision-Makers’ Expertise in Multidimensional Systems


Chapter IV deals with an annotation-based decisional system. The decisional system is based on multidimensional databases, which are composed of facts and dimensions. The expertise of decision-makers is modeled, shared and stored through annotations. Every piece of multidimensional data can be associated with zero or more annotations which allow decision-makers to carry on active analysis and to collaborate with other decision-makers on a common analysis.


Chapter V: Federated Data Warehouses


Chapter V discuss Federated Data Warehouse Systems which consist of a collection of Data Marts provided by different enterprises or public organizations and widen the knowledge base for business analysts, thus enabling better founded strategic decisions. The authors argue that the integration of heterogeneous Data Marts at the logical schema level is preferable to the migration of data into a physically new system if the involved organizations remain autonomous. They present a federated architecture that provides a global multi-dimensional schema to which the Data Mart schemas are semantically mapped, repairing all heterogeneities.


Chapter VI: Built-In Indicators to Support Business Intelligence in OLAP Databases


Chapter VI describes algorithms to support business intelligence in OLAP databases applying Data Mining methods to the multidimensional environment. Those methods help end-users’ analysis in two ways. First, they identify the most interesting dimensions to expand in order to explore the data. Then, they automatically detect interesting cells among a user selected ones. The solution is based on a tight coupling between OLAP tools and statistical methods, based upon the built-in indicators computed instantaneously during the end-users’ exploration of the data cube.


Section III: DWH and OLAP Applications


The next three chapters in Section III, DWH & OLAP Applications, present some typical applications using Data Warehouse and OLAP technology as well as the challenges and issues facing in the real practice.


Chapter VII: Conceptual Data Warehouse Design Methodology for Business Process Intelligence


Chapter VII presents a conceptual framework for adopting the data warehousing technology for business process analysis, with Surgical Workflows Analysis as a challenging real-world application. Deficiencies of the conventional OLAP approach are overcome by proposing an extended multidimensional data model, which enables adequate capturing of flow-oriented data. The model supports a number of advanced properties, such as non-quantitative and heterogeneous facts, many-to-many relationships between facts and dimensions, full and partial dimension sharing, dynamic specification of new measures, and interchangeability of fact and dimension roles.


Chapter VIII: Data Warehouse Facilitating Evidence-Based Medicine


Deployment of a federated data warehouse approach for the integration of the wide range of different medical data sources and for distribution of evidence-based clinical knowledge, to support clinical decision makers, primarily clinicians at the point of care is the main topic of chapter VIII: Data Warehouse Facilitating Evidence-Based Medicine. A real-world scenario is used to illustrate the possible application field in the area of emergency and intensive care in which the evidence-based medicine merges data originating in a pharmacy database, a social insurance company database and diverse clinical DWHs with the minimized administration effort.


Chapter IX: Deploying Data Warehouses in Grids with Efficiency and Availability


Chapter IX, Deploying Data Warehouses in Grids with Efficiency and Availability, discusses the deployment of data warehouses over Grids. The authors present the Grid-NPDW architecture, which aims at providing high throughput and data availability in grid-based warehouses. High efficiency in situations with site failure is also achieved with the use of on-demand query scheduling and data partitioning and replication. The Chapter also describes the main components of the Grid-NPDW Scheduler and presents some experimental results of proposed strategies.


Section IV: Data Mining Techniques


Section IV, Data Mining Techniques, consists of three chapters discussing a variety of traditional data mining techniques such as clustering, ranking, classification but towards the efficiency and performance improvement.


Chapter X: MOSAIC: Agglomerative Clustering with Gabriel Graphs


Chapter X, MOSAIC: Agglomerative Clustering with Gabriel Graphs, introduces MOSAIC, a post-processing technique that has been designed to overcome these disadvantages. MOSAIC is an agglomerative technique that greedily merges neighboring clusters maximizing an externally given fitness function; Gabriel graphs are used to determine which clusters are neighboring, and non-convex shapes are approximated as the unions of small convex clusters. Experimental results are presented that show that using MOSAIC leads to clusters of higher quality compared to running a representative clustering algorithm stand-alone.


Chapter XI: Ranking Gradients in Multi-Dimensional Spaces


Chapter XI investigates how to mine and rank the most interesting changes in a multi-dimensional space applying a promising TOP-K gradient strategy. Interesting changes in customer behavior are usually discovered by gradient queries which are particular cases of multi-dimensional data analysis on large data warehouses. The main problem, however, arises from the fact that more interesting changes should be those ones having more dimensions in the gradient query (the curse-of-dimensionality dilemma). Besides, the number of interesting changes should be of a large amount (the preference selection criteria).


Chapter XII: Simultaneous Feature Selection and Tuple Selection for Efficient Classification


In Chapter XII, a method is proposed to combine feature selection and tuple selection to improve classification accuracy. Although feature selection and tuple selection have been studied earlier in various research areas such as machine learning, data mining, and so on, they have rarely been studied together. Feature selection and tuple selection help the classifier to focus better. The method is based on the principle that a representative subset has similar histogram as the full set. The proposed method uses this principle both to choose a subset of features and also to choose a subset of tuples. The empirical tests show that the proposed method performs better than several existing feature selection methods.


Section V: Advanced Mining Applications


The last four chapters in Section V, Advanced Mining Applications introduces innovative algorithms and applications in some emerging application fields in Data Mining and Knowledge Discovery, especially continuous data stream mining, which could not be solved by traditional mining technology.


Chapter XIII: Learning Cost-Sensitive Decision Trees to Support Medical Diagnosis


In Chapter XIII, Learning cost-sensitive decision trees to support medical, the authors discuss about diagnosis a cost-sensitive learning method. The chapter aims to enhance the understand of cost-sensitive learning problems in medicine and presents a strategy for learning and testing cost-sensitive decision trees, while considering several types of costs associated with problems in medicine. It begins with a contextualization and a discussion of the main types of costs. Then, reviews related work and presents a discussion about the evaluation of classifiers as well as explains a cost-sensitive decision tree strategy and presents some experimental results.


Chapter XIV: An Approximate Approach for Maintaining Recent Occurrences of Itemsets in a Sliding Window over Data Streams


Chapter XIV discuss about catching the recent trend of data when mining frequent itemsets over data streams. A data representation method, named frequency changing point (FCP), is introduced for monitoring the recent occurrence of itemsets over a data stream to prevent from storing the whole transaction data within a sliding window. The effect of old transactions on the mining result of recently frequent itemsets is diminished by performing adjusting rules on the monitoring data structure. Accordingly, the recently frequent itemsets or representative patterns are discovered from the maintained structure approximately. Experimental studies demonstrate that the proposed algorithms achieve high true positive rates and guarantees no false dismissal to the results yielded.


Chapter XV: Protocol Identification of Encrypted Network Streams


Chapter XV proposes a simple machine learning approach for protocol identification in network streams that have been encrypted, such that the only information available for identifying the underlying protocol of a connection was the size, timing and direction of packets. With very little information available from the network stream, it is possible to pinpoint potentially inappropriate activities for a workplace, institution or research center, such as using BitTorrent, GMail or MSN, and not confuse them with other common protocols such as HTTP and SSL.


Chapter XVI: Exploring Calendar-Based Pattern Mining in Data Streams


Chapter XVI introduces a calendar-based pattern mining aims at identifying patterns on specific calendar partitions in continuous data streams. The authors present how a data warehouse approach can be applied to leverage calendar-based pattern mining in data streams and how the framework of the DWFIST approach can cope with tight time constraints imposed by data streams, keep storage requirements at a manageable level and, at the same time, support calendar-based frequent itemset mining. The minimum granularity of analysis, parameters of the data warehouse (e.g. mining minimum support) and parameters of the database (e.g. extent size) provide ways to tune the load performance.


Rabu, 16 September 2009

Mastering Data Warehouse Aggregates: Solutions for Star Schema Performance


Christopher Adamson, Mastering Data Warehouse Aggregates: Solutions for Star Schema Performance
Wiley | ISBN: 0471777099 | Year 2006 | 345 Pages | PDF | 6.08 MB

The first book to offer in-depth coverage of star schema aggregate tables. Dubbed by Ralph Kimball as the most effective technique for maximizing star schema performance, dimensional aggregates are a powerful and efficient tool that can accelerate data warehouse queries more dramatically than any other technology. After you ensure that a database is properly designed, configured, and tuned, any measures you take to address data warehouse performance should begin with aggregates. Yet, many businesses ignore aggregates, instead turning to specialized, proprietary hardware and software products to solve performance problems. This book fills the knowledge gap that has led businesses on this expensive and risky path.

Download:
Or

Minggu, 13 September 2009

Oracle 10g Data Warehousing


Lilian Hobbs, Oracle 10g Data Warehousing
Digital Press | ISBN: 1555583229 | 2004 | PDF | 872 pages | 18.11 MB

Oracle 10g Data Warehousing is a guide to using the Data Warehouse features in the latest version of Oracle Oracle Database 10g. Written by people on the Oracle development team that designed and implemented the code and by people with industry experience implementing warehouses using Oracle technology, this thoroughly updated and extended edition provides an insiders view of how the Oracle Database 10g software is best used for your application.

Download: