Define Noisy data while doing data pre-processing. Delete the noise with Binning smoothing techniques for the following details using partition in Bins (Equal-frequency) : 4, 2, 6, 10, 8, 16, 12, 24, 22, 14, 26 stored price details (in dollars).
Appeared in 8 out of 9 exams — high recurrence pattern.
(d) What is a star-schema dimensional modeling ? What are its characteristics ? Draw a star schema for four dimensions–Time, Item, Branch, Location with 2 measures namely ‘Units-sold’ and ‘Amount-sold’ and ‘sales’ is the Fact Table. Also, explain the star-schema diagram. Assumptions can be made wherever necessary.
Appeared in 7 out of 9 exams — high recurrence pattern.
(b) What is frequent pattern mining ? Briefly explain the following classifications of frequent pattern mining along with an example for each : (i) Based on the levels of abstraction involved in the rule-set. (ii) Based on the types of values handled in the rule. (iii) Based on the kinds of rules to be mined.
Appeared in 6 out of 9 exams — high recurrence pattern.
b) Define classification. With the help of an example for each, explain the following classification models : (i) Descriptive modeling (ii) Predictive modeling
Appeared in 5 out of 9 exams — high recurrence pattern.
Discuss ETL and its need. Explain in detail, all the steps involved in ETL with the help of a suitable diagram.
Appeared in 4 out of 9 exams — high recurrence pattern.
(a) With reference to data warehousing, explain the following terms : (i) Metadata and Data warehousing (ii) Data Granularity (iii) Operational data store (iv) Data Mart
Appeared in 4 out of 9 exams — high recurrence pattern.
(a) What is text mining ? With reference to text mining, explain the following techniques : (i) Information Extraction (ii) Text Summarization (iii) Text Categorization (iv) Text Clustering Also, mention any four applications of text mining.
Appeared in 4 out of 9 exams — high recurrence pattern.
d) Discuss vector space modeling for representing text documents. With reference to this modeling, explain TF-IDF and Inverse Document Frequency (IDF).
Appeared in 3 out of 9 exams — high recurrence pattern.
Write short notes on any four of the following : (a) OLTP (b) Data Marts (c) Web structure mining (d) Cloud Data Warehousing (e) Data Granularity
Appeared in 3 out of 9 exams — high recurrence pattern.
d) Discuss the following categories of Data Mining Issues : (i) Mining Methodology and User Iteration Issues (ii) Performance-based Issues (iii) Diverse Data Types Issues
Appeared in 3 out of 9 exams — high recurrence pattern.