E-Book, Englisch, 864 Seiten
Nisbet / Elder / Miner Handbook of Statistical Analysis and Data Mining Applications
1. Auflage 2009
ISBN: 978-0-08-091203-5
Verlag: Elsevier Science & Techn.
Format: EPUB
Kopierschutz: 6 - ePub Watermark
E-Book, Englisch, 864 Seiten
ISBN: 978-0-08-091203-5
Verlag: Elsevier Science & Techn.
Format: EPUB
Kopierschutz: 6 - ePub Watermark
The Handbook of Statistical Analysis and Data Mining Applications is a comprehensive professional reference book that guides business analysts, scientists, engineers and researchers (both academic and industrial) through all stages of data analysis, model building and implementation. The Handbook helps one discern the technical and business problem, understand the strengths and weaknesses of modern data mining algorithms, and employ the right statistical methods for practical application. Use this book to address massive and complex datasets with novel statistical approaches and be able to objectively evaluate analyses and solutions. It has clear, intuitive explanations of the principles and tools for solving problems using modern analytic techniques, and discusses their application to real problems, in ways accessible and beneficial to practitioners across industries - from science and engineering, to medicine, academia and commerce. This handbook brings together, in a single resource, all the information a beginner will need to understand the tools and issues in data mining to build successful data mining solutions.Written 'By Practitioners for Practitioners' Non-technical explanations build understanding without jargon and equations Tutorials in numerous fields of study provide step-by-step instruction on how to use supplied tools to build models Practical advice from successful real-world implementationsIncludes extensive case studies, examples, MS PowerPoint slides and datasetsCD-DVD with valuable fully-working 90-day software included: 'Complete Data Miner - QC-Miner - Text Miner' bound with book
Dr. Robert Nisbet was trained initially in Ecology and Ecosystems Analysis. He has over 30 years of experience in complex systems analysis and modeling, most recently as a Researcher (University of California, Santa Barbara). In business, he pioneered the design and development of configurable data mining applications for retail sales forecasting, and Churn, Propensity-to-buy, and Customer Acquisition in Telecommunications, Insurance, Banking, and Credit industries. In addition to data mining, he has expertise in data warehousing technology for Extract, Transform, and Load (ETL) operations, Business Intelligence reporting, and data quality analyses. He is lead author of the 'Handbook of Statistical Analysis & Data Mining Applications” (Academic Press, 2009), and a co-author of 'Practical Text Mining' (Academic Press, 2012), and co-author of 'Practical Predictive Analytics and Decisioning Systems for Medicine (Academic Press, 2015). Currently, he serves as an Instructor in the University of California, Irvine Predictive Analytics Certificate Program, teaching online and on-campus courses in Effective Data preparation, and Applications of Predictive Analytics. Additionally Bob is in the last stages of writing another book on 'Data Preparation for Predictive Analytic Modeling.
Autoren/Hrsg.
Weitere Infos & Material
1;Front Cover;1
2;Handbook of Statistical Analysis and Data Mining Applications;4
3;Copyright Page;5
4;Table of Contents;6
5;Foreword 1;16
6;Foreword 2;18
7;Preface;20
8;Introduction;24
9;List of Tutorials by Guest Authors;30
10;Part 1: History of Phases of Data Analysis, Basic Theory, and the Data Mining Process;36
10.1;Chapter 1: The Background for Data Mining Practice;38
10.1.1;Preamble;38
10.1.2;A Short History of Statistics and Data Mining;39
10.1.3;Modern Statistics: A Duality?;40
10.1.3.1;Assumptions of the Parametric Model;41
10.1.4;Two Views of Reality;43
10.1.4.1;Aristotle;43
10.1.4.2;Plato;44
10.1.5;The Rise of Modern Statistical Analysis: The Second Generation;45
10.1.5.1;Data, Data Everywhere;46
10.1.6;Machine Learning Methods: The Third Generation;46
10.1.7;Statistical Learning Theory: The Fourth Generation;47
10.1.8;Postscript;48
10.1.9;References;49
10.2;Chapter 2: Theoretical Considerations for Data Mining;50
10.2.1;Preamble;50
10.2.2;The Scientific Method;51
10.2.3;What Is Data Mining?;52
10.2.4;A Theoretical Framework for the Data Mining Process;53
10.2.4.1;Microeconomic Approach;54
10.2.4.2;Inductive Database Approach;54
10.2.5;Strengths of the Data Mining Process;54
10.2.6;Customer-Centric Versus Account-Centric: A New Way to Look at Your Data;55
10.2.6.1;The Physical Data Mart;55
10.2.6.2;The Virtual Data Mart;56
10.2.6.3;Householded Databases;56
10.2.7;The Data Paradigm Shift;57
10.2.8;Creation of the CAR;57
10.2.9;Major Activities of Data Mining;58
10.2.10;Major Challenges of Data Mining;60
10.2.11;Examples of Data Mining Applications;61
10.2.12;Major Issues in Data Mining;61
10.2.13;General Requirements for Success in a Data Mining Project;63
10.2.14;Example of a Data Mining Project: Classify a Bat's Species by Its Sound;63
10.2.15;The Importance of Domain Knowledge;65
10.2.16;Postscript;65
10.2.16.1;Why Did Data Mining Arise?;65
10.2.16.2;Some Caveats with Data Mining Solutions;66
10.2.17;References;67
10.3;Chapter 3: The Data Mining Process;68
10.3.1;Preamble;68
10.3.2;The Science of Data Mining;68
10.3.3;The Approach to Understanding and Problem Solving;69
10.3.3.1;CRISP-DM;70
10.3.4;Business Understanding (Mostly Art);71
10.3.4.1;Define the Business Objectives of the Data Mining Model;71
10.3.4.2;Assess the Business Environment for Data Mining;72
10.3.4.3;Formulate the Data Mining Goals and Objectives;72
10.3.5;Data Understanding (Mostly Science);74
10.3.5.1;Data Acquisition;74
10.3.5.2;Data Integration;74
10.3.5.3;Data Description;75
10.3.5.4;Data Quality Assessment;75
10.3.6;Data Preparation (A Mixture of Art and Science);75
10.3.7;Modeling (A Mixture of Art and Science);76
10.3.7.1;Steps in the Modeling Phase of CRISP-DM;76
10.3.8;Deployment (Mostly Art);80
10.3.9;Closing the Information Loop (Art);81
10.3.10;The Art of Data Mining;81
10.3.10.1;Artistic Steps in Data Mining;82
10.3.11;Postscript;82
10.3.12;References;83
10.4;Chapter 4: Data Understanding and Preparation;84
10.4.1;Preamble;84
10.4.2;Activities of Data Understanding and Preparation;85
10.4.2.1;Definitions;85
10.4.3;Issues That Should be Resolved;86
10.4.3.1;Basic Issues That Must Be Resolved in Data Understanding;86
10.4.3.2;Basic Issues That Must Be Resolved in Data Preparation;86
10.4.4;Data Understanding;86
10.4.4.1;Data Acquisition;86
10.4.4.2;Data Extraction;88
10.4.4.3;Data Description;89
10.4.4.4;Data Assessment;91
10.4.4.5;Data Profiling;91
10.4.4.6;Data Cleansing;91
10.4.4.7;Data Transformation;92
10.4.4.8;Data Imputation;94
10.4.4.9;Data Weighting and Balancing;97
10.4.4.10;Data Filtering and Smoothing;99
10.4.4.11;Data Abstraction;101
10.4.4.12;Data Reduction;104
10.4.4.13;Data Sampling;104
10.4.4.14;Data Discretization;108
10.4.4.15;Data Derivation;108
10.4.5;Postscript;110
10.4.6;References;110
10.5;Chapter 5: Feature Selection;112
10.5.1;Preamble;112
10.5.2;Variables as Features;113
10.5.3;Types of Feature Selections;113
10.5.4;Feature Ranking Methods;113
10.5.4.1;Gini Index;113
10.5.4.2;Bi-variate Methods;115
10.5.4.3;Multivariate Methods;115
10.5.4.4;Complex Methods;117
10.5.5;Subset Selection Methods;117
10.5.5.1;The Other Two Ways of Using Feature Selection in STATISTICA: Interactive Workspace;128
10.5.5.2;STATISTICA DMRecipe Method;128
10.5.6;Postscript;131
10.5.7;References;132
10.6;Chapter 6: Accessory Tools for Doing Data Mining;134
10.6.1;Preamble;134
10.6.2;Data Access Tools;135
10.6.2.1;Structured Query Language (SQL) Tools;135
10.6.2.2;Extract, Transform, and Load (ETL) Capabilities;135
10.6.3;Data Exploration Tools;136
10.6.3.1;Basic Descriptive Statistics;136
10.6.3.2;Combining Groups (Classes) for Predictive Data Mining;140
10.6.3.3;Slicing/Dicing and Drilling Down into Data Sets/Results Spreadsheets;141
10.6.4;Modeling Management Tools;142
10.6.4.1;Data Miner Workspace Templates;142
10.6.5;Modeling Analysis Tools;142
10.6.5.1;Feature Selection;142
10.6.5.2;Importance Plots of Variables;143
10.6.6;In-Place Data Processing (IDP);148
10.6.6.1;Example: The IDP Facility of STATISTICA Data Miner;149
10.6.6.2;How to Use the SQL;149
10.6.7;Rapid Deployment of Predictive Models;149
10.6.8;Model Monitors;151
10.6.9;Postscript;152
10.6.10;Bibliography;152
11;Part 2: The Algorithms in Data Mining and Text Mining, the Organization of the Three most common Data Mining Tools, and Selected Speci...;154
11.1;Chapter 7: Basic Algorithms for Data Mining: A Brief Overview;156
11.1.1;Preamble;156
11.1.1.1;STATISTICA Data Miner Recipe (DMRecipe);158
11.1.1.2;KXEN;159
11.1.2;Basic Data Mining Algorithms;161
11.1.2.1;Association Rules;161
11.1.2.2;Neural Networks;163
11.1.2.3;Radial Basis Function (RBF) Networks;171
11.1.2.4;Automated Neural Nets;173
11.1.3;Generalized Additive Models (GAMs);173
11.1.3.1;Outputs of GAMs;174
11.1.3.2;Interpreting Results of GAMs;174
11.1.4;Classification and Regression Trees (CART);174
11.1.4.1;Recursive Partitioning;179
11.1.4.2;Pruning Trees;179
11.1.4.3;General Comments about CART for Statisticians;179
11.1.4.4;Advantages of CART over Other Decision Trees;180
11.1.4.5;Uses of CART;181
11.1.5;General Chaid Models;181
11.1.5.1;Advantages of CHAID;182
11.1.5.2;Disadvantages of CHAID;182
11.1.6;Generalized EM and k-Means Cluster Analysis-An Overview;182
11.1.6.1;k-Means Clustering;182
11.1.6.2;EM Cluster Analysis;183
11.1.6.3;Processing Steps of the EM Algorithm;184
11.1.6.4;V-fold Cross-Validation as Applied to Clustering;184
11.1.7;Postscript;185
11.1.8;References;185
11.1.9;Bibliography;185
11.2;Chapter 8: Advanced Algorithms for Data Mining;186
11.2.1;Preamble;186
11.2.2;Advanced Data Mining Algorithms;189
11.2.2.1;Interactive Trees;189
11.2.2.2;Multivariate Adaptive Regression Splines (MARSplines);193
11.2.2.3;Statistical Learning Theory: Support Vector Machines;197
11.2.2.4;Sequence, Association, and Link Analyses;199
11.2.2.5;Independent Components Analysis (ICA);203
11.2.2.6;Kohonen Networks;204
11.2.2.7;Characteristics of a Kohonen Network;204
11.2.2.8;Quality Control Data Mining and Root Cause Analysis;204
11.2.3;Image and Object Data Mining: Visualization and 3D-Medical and Other Scanning Imaging;205
11.2.4;Postscript;206
11.2.5;References;206
11.3;Chapter 9: Text Mining and Natural Language Processing;208
11.3.1;Preamble;208
11.3.2;The Development of Text Mining;209
11.3.3;A Practical Example: NTSB;210
11.3.3.1;Goals of Text Mining of NTSB Accident Reports;219
11.3.3.2;Drilling into Words of Interest;223
11.3.3.3;Means with Error Plots;224
11.3.3.4;Feature Selection Tool;225
11.3.3.5;A Conclusion: Losing Control of the Aircraft in Bad Weather Is Often Fatal;226
11.3.3.6;Summary;229
11.3.4;Text Mining Concepts Used in Conducting Text Mining Studies;229
11.3.5;Postscript;229
11.3.6;References;230
11.4;Chapter 10: The Three Most Common Data Mining Software Tools;232
11.4.1;Preamble;232
11.4.2;SPSS Clementine Overview;232
11.4.2.1;Overall Organization of Clementine Components;233
11.4.2.2;Organization of the Clementine Interface;234
11.4.2.3;Clementine Interface Overview;234
11.4.2.4;Setting the Default Directory;236
11.4.2.5;SuperNodes;236
11.4.2.6;Execution of Streams;237
11.4.3;SAS-Enterprise Miner (SAS-EM) Overview;238
11.4.3.1;Overall Organization of SAS-EM Version 5.3 Components;238
11.4.3.2;Layout of the SAS-Enterprise Miner Window;239
11.4.3.3;Various SAS-EM Menus, Dialogs, and Windows Useful During the Data Mining Process;240
11.4.3.4;Software Requirements to Run SAS-EM 5.3 Software;241
11.4.4;STATISTICA Data Miner, QC-Miner, and Text Miner Overview;249
11.4.4.1;Overall Organization and Use of STATISTICA Data Miner;249
11.4.4.2;Three Formats for Doing Data Mining in STATISTICA;265
11.4.5;Postscript;269
11.4.6;References;269
11.5;Chapter 11: Classification;270
11.5.1;Preamble;270
11.5.2;What Is Classification?;270
11.5.3;Initial Operations in Classification;271
11.5.4;Major Issues with Classification;271
11.5.4.1;What Is the Nature of the Data Set to Be Classified?;271
11.5.4.2;How Accurate Does the Classification Have to Be?;271
11.5.4.3;How Understandable Do the Classes Have to Be?;272
11.5.5;Assumptions of Classification Procedures;272
11.5.5.1;Numerical Variables Operate Best;272
11.5.5.2;No Missing Values;272
11.5.5.3;Variables Are Linear and Independent in Their Effects on the Target Variable;272
11.5.6;Methods for Classification;273
11.5.6.1;Nearest-Neighbor Classifiers;274
11.5.6.2;Analyzing Imbalanced Data Sets with Machine Learning Programs;275
11.5.6.3;CHAID;281
11.5.6.4;Random Forests and Boosted Trees;283
11.5.6.5;Logistic Regression;285
11.5.6.6;Neural Networks;286
11.5.6.7;Naive Bayesian Classifiers;288
11.5.7;What Is the Best Algorithm for Classification?;291
11.5.8;Postscript;292
11.5.9;References;293
11.6;Chapter 12: Numerical Prediction;294
11.6.1;Preamble;294
11.6.2;Linear Response Analysis and the Assumptions of the Parametric Model;295
11.6.3;Parametric Statistical Analysis;296
11.6.4;Assumptions of the Parametric Model;297
11.6.4.1;The Assumption of Independency;297
11.6.4.2;The Assumption of Normality;297
11.6.4.3;Normality and the Central Limit Theorem;298
11.6.4.4;The Assumption of Linearity;299
11.6.5;Linear Regression;299
11.6.5.1;Methods for Handling Variable Interactions in Linear Regression;300
11.6.5.2;Collinearity among Variables in a Linear Regression;300
11.6.5.3;The Concept of the Response Surface;301
11.6.6;Generalized Linear Models (GLMs);305
11.6.7;Methods for Analyzing Nonlinear Relationships;306
11.6.8;Nonlinear Regression and Estimation;306
11.6.8.1;Logit and Probit Regression;307
11.6.8.2;Poisson Regression;307
11.6.8.3;Exponential Distributions;307
11.6.8.4;Piecewise Linear Regression;308
11.6.9;Data Mining and Machine Learning Algorithms Used in Numerical Prediction;309
11.6.9.1;Numerical Prediction with C&RT;309
11.6.9.2;Model Results Available in C&RT;311
11.6.10;Advantages of Classification and Regression Trees (C&RT) Methods;312
11.6.10.1;General Issues Related to C&RT;314
11.6.11;Application to Mixed Models;315
11.6.12;Neural Nets for Prediction;315
11.6.12.1;Manual or Automated Operation?;315
11.6.12.2;Structuring the Network for Manual Operation;315
11.6.12.3;Modern Neural Nets Are "Gray Boxes";316
11.6.12.4;Example of Automated Neural Net Results;316
11.6.13;Support Vector Machines (SVMs) and Other Kernel Learning Algorithms;317
11.6.14;Postscript;319
11.6.15;References;319
11.7;Chapter 13: Model Evaluation and Enhancement;320
11.7.1;Preamble;320
11.7.2;Introduction;321
11.7.3;Model Evaluation;321
11.7.3.1;Splitting Data;322
11.7.3.2;Avoiding Overfit Through Complexity Regularization;323
11.7.3.3;Error Metric: Estimation;326
11.7.3.4;Error Metric: Classification;326
11.7.3.5;Error Metric: Ranking;328
11.7.3.6;Cross-Validation to Estimate Error Rate and Its Confidence;330
11.7.3.7;Bootstrap;331
11.7.3.8;Target Shuffling to Estimate Baseline Performance;332
11.7.4;Re-Cap of the Most Popular Algorithms;335
11.7.4.1;Linear Methods (Consensus Method, Stepwise Is Variable-Selecting);335
11.7.4.2;Decision Trees (Consensus Method, Variable-Selecting);335
11.7.4.3;Neural Networks (Consensus Method);336
11.7.4.4;Nearest Neighbors (Contributory Method);336
11.7.4.5;Clustering (Consensus or Contributory Method);337
11.7.5;Enhancement Action Checklist;337
11.7.6;Ensembles of Models: The Single Greatest Enhancement Technique;339
11.7.6.1;Bagging;340
11.7.6.2;Boosting;340
11.7.6.3;Ensembles in General;341
11.7.7;How to Thrive as a Data Miner;342
11.7.7.1;Big Picture of the Project;342
11.7.7.2;Project Methodology and Deliverables;343
11.7.7.3;Professional Development;344
11.7.7.4;Three Goals;345
11.7.8;Postscript;346
11.7.9;References;346
11.8;Chapter 14: Medical Informatics;348
11.8.1;Preamble;348
11.8.2;What Is Medical Informatics?;348
11.8.3;How Data Mining and Text Mining Relate to Medical Informatics ;349
11.8.3.1;XplorMed;351
11.8.3.2;ABView: HivResist;352
11.8.4;3D Medical Informatics;352
11.8.4.1;What Is 3D Informatics?;352
11.8.4.2;Future and Challenges of 3D Medical Informatics;353
11.8.4.3;Journals and Associations in the Field of Medical Informatics;353
11.8.5;Postscript;353
11.8.6;References;354
11.8.7;Bibliography;354
11.9;Chapter 15: Bioinformatics;356
11.9.1;Preamble;356
11.9.2;What Is Bioinformatics?;358
11.9.3;Data Analysis Methods in Bioinformatics;361
11.9.3.1;ClustalW2: Sequence Alignment;361
11.9.3.2;Searching Databases for RNA Molecules;362
11.9.4;Web Services in Bioinformatics;362
11.9.5;How Do We Apply Data Mining Methods to Bioinformatics?;364
11.9.6;Postscript;367
11.9.6.1;Tutorial Associated with This Chapter on Bioinformatics;367
11.9.6.2;Books, Associations, and Journals on Bioinformatics, and Other Resources, Including Online;367
11.9.7;References;368
11.9.8;Bibliography;369
11.10;Chapter 16: Customer Response Modeling;370
11.10.1;Preamble;370
11.10.2;Early CRM Issues in Business;371
11.10.3;Knowing How Customers Behaved Before They Acted;371
11.10.3.1;Transforming Corporations into Business Ecosystems: The Path to Customer Fulfillment;372
11.10.4;CRM in Business Ecosystems;373
11.10.4.1;Differences Between Static Measures and Evolutionary Measures;373
11.10.4.2;How Can Human Nature as Viewed Through Plato Help Us in Modeling Customer Response?;374
11.10.4.3;How Can We Reorganize Our Data to Reflect Motives and Attitudes?;374
11.10.4.4;What Is a Temporal Abstraction?;375
11.10.5;Conclusions;379
11.10.6;Postscript;380
11.10.7;References;380
11.11;Chapter 17: Fraud Detection;382
11.11.1;Preamble;382
11.11.2;Issues with Fraud Detection;383
11.11.2.1;Fraud Is Rare;383
11.11.2.2;Fraud Is Evolving;383
11.11.2.3;Large Data Sets Are Needed;383
11.11.2.4;The Fact of Fraud Is Not Always Known During Modeling;383
11.11.2.5;When the Fraud Happened Is Very Important to Its Detection;384
11.11.2.6;Fraud Is Very Complex;384
11.11.2.7;Fraud Detection May Require the Formulation of Rules Based on General Principles,"Red Flags," Alerts, and Profiles;384
11.11.2.8;Fraud Detection Requires Both Internal and External Business Data;384
11.11.2.9;Very Few Data Sets and Modeling Details Are Available;385
11.11.3;How Do You Detect Fraud?;385
11.11.4;Supervised Classification of Fraud;386
11.11.5;How Do You Model Fraud?;387
11.11.6;How Are Fraud Detection Systems Built?;388
11.11.7;Intrusion Detection Modeling;390
11.11.8;Comparison of Models with and Without Time-Based Features;390
11.11.9;Building Profiles;395
11.11.10;Deployment of Fraud Profiles;395
11.11.11;Postscript and Prolegomenon;396
11.11.12;References;396
12;Part 3: Tutorials-Step-by-Step Case Studies as a Starting Point to learn how to do Data Mining Analyses;398
12.1;Guest Authors of the Tutorials;400
12.1.1;Tutorial A: How to Use Data Miner Recipe STATISTICA Data Miner Only;402
12.1.1.1;What Is STATISTICA Data Miner Recipe (DMR)?;408
12.1.1.2;Core Analytic Ingredients;408
12.1.2;Tutorial B: Data Mining for Aviation Safety Using Data Mining Recipe
"Automatized Data Mining" from
STATISTICA;412
12.1.2.1;Airline Safety;413
12.1.2.2;SDR Database;414
12.1.2.3;Preparing the Data for Our Tutorial;417
12.1.2.4;Data Mining Approach;418
12.1.2.5;Data Mining Algorithm Error Rate;421
12.1.2.6;Conclusion;422
12.1.2.7;References;424
12.1.3;Tutorial C: Predicting Movie Box-Office Receipts Using SPSS Clementine Data
Mining Software;426
12.1.3.1;Introduction;426
12.1.3.2;Data and Variable Definitions;427
12.1.3.3;Getting to Know the Workspace of the Clementine Data Mining Toolkit;428
12.1.3.4;Results;431
12.1.3.5;Publishing and Reuse of Models and Other Outputs;439
12.1.3.6;References;450
12.1.4;Tutorial D: Detecting Unsatisfied Customers: A Case Study Using SAS Enterprise Miner
Version 5.3 for the Analysis;452
12.1.4.1;Introduction;453
12.1.4.1.1;The Data;453
12.1.4.1.2;The Objectives of the Study;453
12.1.4.1.3;SAS-EM 5.3 Interface;454
12.1.4.2;A Primer of SAS-EM Predictive Modeling;455
12.1.4.2.1;Homework 1;465
12.1.4.2.2;Discussions;466
12.1.4.2.3;Homework 2;466
12.1.4.2.4;Homework 3;466
12.1.4.3;Scoring Process and the Total Profit;467
12.1.4.3.1;Homework 4;473
12.1.4.3.2;Discussions;474
12.1.4.4;Oversampling and Rare Event Detection;474
12.1.4.4.1;Discussion;481
12.1.4.5;Decision Matrix and the Profit Charts;481
12.1.4.5.1;Discussions;488
12.1.4.6;Micro-Target the Profitable Customers;488
12.1.4.7;Appendix;490
12.1.4.8;Reference;493
12.1.5;Tutorial E: Credit Scoring Using STATISTICA Data
Miner;494
12.1.5.1;Introduction: What Is Credit Scoring?;494
12.1.5.2;Credit Scoring: Business Objectives;495
12.1.5.3;Case Study: Consumer Credit Scoring;496
12.1.5.3.1;Description;496
12.1.5.3.2;Data Preparation;497
12.1.5.3.3;Feature Selection;497
12.1.5.3.4;STATISTICA Data Miner: "Workhorses" or Predictive Modeling;498
12.1.5.3.5;Overview: STATISTICA Data Miner Workspace;499
12.1.5.4;Analysis and Results;500
12.1.5.4.1;Decision Tree: CHAID;500
12.1.5.4.2;Classification Matrix: CHAID Model;502
12.1.5.5;Comparative Assessment of the Models (Evaluation);502
12.1.5.5.1;Classification Matrix: Boosting Trees with Deployment Model (Best Model);504
12.1.5.6;Deploying the Model for Prediction;504
12.1.5.7;Conclusion;505
12.1.6;Tutorial F: Churn Analysis With SPSS-Clementine;506
12.1.6.1;Objectives;506
12.1.6.2;Steps;507
12.1.7;Tutorial G: Text Mining: Automobile Brand Review Using STATISTICA Data
Miner and Text Miner;516
12.1.7.1;Introduction;516
12.1.7.2;Text Mining;517
12.1.7.2.1;Input Documents;517
12.1.7.2.2;Selecting Input Documents;517
12.1.7.2.3;Stop Lists, Synonyms, and Phrases;517
12.1.7.2.4;Stemming and Support for Different Languages;518
12.1.7.2.5;Indexing of Input Documents: Scalability of STATISTICA Text Mining and Document Retrieval;518
12.1.7.2.6;Results, Summaries, and Transformations;518
12.1.7.3;Car Review Example;519
12.1.7.3.1;Saving Results into Input Spreadsheet;533
12.1.7.4;Interactive Trees (C&RT, CHAID);538
12.1.7.5;Other Applications of Text Mining;547
12.1.7.6;Conclusion;547
12.1.8;Tutorial H: Predictive Process Control: QC-Data Mining Using STATISTICA Data Miner
and QC-Miner;548
12.1.8.1;Predictive Process Control Using STATISTICA and STATISTICA QC-Miner;548
12.1.8.2;Case Study: Predictive Process Control;549
12.1.8.2.1;Understanding Manufacturing Processes;549
12.1.8.2.2;Data File: ProcessControl.sta;550
12.1.8.2.3;Variable Information;550
12.1.8.2.4;Problem Definition;550
12.1.8.2.5;Design Approaches;550
12.1.8.3;Data Analyses with STATISTICA;552
12.1.8.3.1;Split Input Data into the Training and Testing Sample;552
12.1.8.3.2;Stratified Random Sampling;552
12.1.8.3.3;Feature Selection and Root Cause Analyses;552
12.1.8.3.4;Different Models Used for Prediction;553
12.1.8.3.5;Compute Overlaid Lift Charts from All Models: Static Analyses;555
12.1.8.3.6;Classification Trees: CHAID;556
12.1.8.3.7;Compute Overlaid Lift/Gain Charts from All Models: Dynamic Analyses;558
12.1.8.3.8;Cross-Tabulation Matrix;559
12.1.8.3.9;Comparative Evaluation of Models: Dynamic Analyses;561
12.1.8.3.10;Gains Analyses by Deciles: Dynamic Analyses;561
12.1.8.3.11;Transformation of Change;562
12.1.8.3.12;Feature Selection and Root Cause Analyses;563
12.1.8.3.13;Interactive Trees: C&RT;563
12.1.8.4;Conclusion;564
12.1.9;Tutorials I, J, AND K: Three Short Tutorials Showing the Use of Data Mining and Particularly C&RT to Predict and Display Possible Structural Relationships among Data;566
12.1.9.1;Tutorial I: Business Administration in a Medical Industry: Determining Possible
Predictors for Days with
Hospice Service for Patients
with Dementia;568
12.1.9.2;Tutorial J: Clinical Psychology: Making Decisions about Best Therapy for a Client: Using Data Mining to Explore the
Structure of a Depression
Instrument;602
12.1.9.3;Tutorial K: Education-Leadership Training for Business and Education Using C&RT to Predict and
Display Possible Structural
Relationships;622
12.1.9.3.1;References;656
12.1.10;Tutorial L: Dentistry: Facial Pain Study Based on 84 Predictor Variables
(Both Categorical and
Continuous);658
12.1.11;Tutorial M: Profit Analysis of the German Credit Data Using SAS-EM Version 5.3;686
12.1.11.1;Introduction;686
12.1.11.2;Modeling Strategy;688
12.1.11.3;SAS-EM 5.3 Interface;689
12.1.11.4;A Primer of SAS-EM Predictive Modeling;689
12.1.11.4.1;Discussions;704
12.1.11.5;Advanced Techniques of Predictive Modeling;704
12.1.11.5.1;Conclusion;711
12.1.11.6;Micro-Target the Profitable Customers;711
12.1.11.7;Appendix;713
12.1.11.8;References;715
12.1.12;Tutorial N: Predicting Self-Reported Health Status Using Artificial Neural Networks;716
12.1.12.1;Background;716
12.1.12.2;Data;717
12.1.12.2.1;Preprocessing and Filtering;718
12.1.12.2.2;Part 1: Using a Wrapper Approach in Weka to Determine the Most Appropriate Variables for Your Neural Network Model;719
12.1.12.2.3;Part 2: Taking the Results from the Wrapper Approach in Weka into STATISTICA Data Miner to do Neural Network Analyses;726
12.1.12.3;References;738
13;Part 4: Measuring true complexity, the "Right Model for the Right Use," Top Mistakes, and the Future of Analytics;740
13.1;Chapter 18: Model Complexity (and How Ensembles Help);742
13.1.1;Preamble;742
13.1.2;Model Ensembles;743
13.1.3;Complexity;745
13.1.4;Generalized Degrees of Freedom;748
13.1.5;Examples: Decision Tree Surface with Noise;749
13.1.6;Summary and Discussion;754
13.1.7;Postscript;755
13.1.8;References;755
13.2;Chapter 19: The Right Model for the Right Purpose: When Less Is Good Enough;758
13.2.1;Preamble;758
13.2.2;More Is Not Necessarily Better: Lessons from Nature and Engineering;759
13.2.3;Embrace Change Rather Than Flee from It;760
13.2.4;Decision Making Breeds True in the Business Organism;760
13.2.4.1;Muscles in the Business Organism;761
13.2.4.2;What Is a Complex System?;761
13.2.5;The 80:20 Rule in Action;763
13.2.6;Agile Modeling: An Example of How to Craft Sufficient Solutions;763
13.2.7;Postscript;765
13.2.8;References;766
13.3;Chapter 20: Top 10 Data Mining Mistakes;768
13.3.1;Preamble;768
13.3.2;Introduction;769
13.3.3;0. Lack Data;769
13.3.4;1. Focus on Training;770
13.3.5;2. Rely on One Technique;771
13.3.6;3. Ask the Wrong Question;773
13.3.7;4. Listen (Only) to the Data;774
13.3.8;5. Accept Leaks from the Future;777
13.3.9;6. Discount Pesky Cases;778
13.3.10;7. Extrapolate;779
13.3.11;8. Answer Every Inquiry;782
13.3.12;9. Sample Casually;785
13.3.13;10. Believe the Best Model;787
13.3.14;How Shall We Then Succeed?;788
13.3.15;Postscript;788
13.3.16;References;788
13.4;Chapter 21: Prospects for the Future of Data Mining and Text Mining as Part of Our Everyday Lives;790
13.4.1;Preamble;790
13.4.2;RFID;791
13.4.3;Social Networking and Data Mining;792
13.4.3.1;Example 1;793
13.4.3.2;Example 2;794
13.4.3.3;Example 3;795
13.4.3.4;Example 4;796
13.4.4;Image and Object Data Mining;796
13.4.4.1;Visual Data Preparation for Data Mining: Taking Photos, Moving Pictures, and Objects into Spreadsheets Representing the Photos...;800
13.4.5;Cloud Computing;804
13.4.5.1;The Next Generation of Data Mining;807
13.4.5.2;From the Desktop to the Clouds;813
13.4.6;Postscript;813
13.4.7;References;813
13.5;Chapter 22: Summary: Our Design;816
13.5.1;Preamble;816
13.5.2;Beware of Overtrained Models;817
13.5.3;A Diversity of Models and Techniques Is Best;818
13.5.4;The Process Is More Important Than the Tool;818
13.5.5;Text Mining of Unstructured Data Is Becoming Very Important;819
13.5.6;Practice Thinking about Your Organization as Organism Rather Than as Machine;819
13.5.7;Good Solutions Evolve Rather Than Just Appear after Initial Efforts;820
13.5.8;What You Don't Do Is Just as Important as What You Do;820
13.5.9;Very Intuitive Graphical Interfaces Are Replacing Procedural Programming;821
13.5.10;Data Mining Is No Longer a Boutique Operation; It Is Firmly Established in the Mainstream of Our Society;821
13.5.11;"Smart" Systems Are the Direction in which Data Mining Technology is Going;822
13.5.12;Postscript;822
13.5.13;References;823
14;Glossary;824
15;Index;836
16;DVD Install Instructions
;858
16.1;Installing STATISTICA;858




