Font size
WorksheetsDWBI
Total questions: 81
Worksheet time: 41mins
Non-volatile in the context of a Data Warehouse means...
Data only persists during the active session
Data does not change after being loaded into the warehouse
Data cannot be analyzed again
Data is frequently updated to maintain accuracy
An analyst selects only sales data for the region "West Java" and the product "Beverages". This is an example of...
Pivot
Drill-down
Dice
Roll-up
What is the primary definition of a Data Warehouse?
A data storage system used for analysis and decision-making
A system for daily transactions in a company
A place to store project document files
Hardware for storing customer data
Transaction data from various branches is uploaded to the data warehouse every night. This process is part of the stage…
Integration
Transformation
Extraction
Load
When a customer changes address, the system keeps the old version and the new version in two separate rows. There are “Start Date” and “End Date” columns to track the active period. This is an example of…
SCD Type 0
SCD Type 1
SCD Type 3
SCD Type 2
In a school’s analytics system, the fact table “Student Grades” is connected to the dimensions “Student”, “Subject”, and “Teacher”. The “Subject” dimension has subdimensions “Subject Group” and “Field of Study” in separate tables. This reflects the schema…
Star Schema
Snowflake Schema
Flat Schema
Fact Constellation
If a dimension table has sub-dimensions or hierarchical relationships, the schema used is...
Flat Schema
Star Schema
Cube Schema
Snowflake Schema
A company stores the current version information in the main row and keeps all historical data in a separate history table. This is the method...
SCD Type 3
SCD Type 2
SCD Type 4
SCD Type 0
A manufacturing company builds a data warehouse schema that displays hierarchical relationships in the location dimension, where City is linked to Province and Country as sub-dimensions in separate table structures. Such a schema indicates characteristics of...
Star Schema
Hybrid Schema
Snowflake Schema
Conformed Dimension
A university manages two fact tables: "Student Enrollment" and "Tuition Payment". Both tables are linked to the same dimension tables, such as Student, Study Program, and Semester. Which schema model is most suitable for this situation?
Normalized Schema
Fact Constellation Schema
Independent Star Schema
Star Schema
During the ETL process, the same transaction data appears twice due to input errors. The required transformation technique is...
Data Deduplication
Data Mining
Sorting
Aggregation
In an OLAP report, a user chooses to view only data from the year 2024. This operation is called...
Slice
Pivot
Drill-down
Roll-up
An analyst wants to filter data only for January 2025 across all regions and all products. The operation performed is...
Slice
Dice
Roll-up
Drill-down
In a Star Schema, dimension tables are typically...
Denormalized for query efficiency
Do not contain attributes
Unrelated to the fact table
Highly structured and normalized
A financial system only records the latest account status without keeping previous versions. This is classified as…
SCD Type 0
SCD Type 2
SCD Type 3
SCD Type 1
A retail company builds several data marts separately for the sales, finance, and logistics divisions. Each is developed by different teams without coordination and they are not connected to each other. What architecture are they using?
Independent Data Marts
Data Mart Bus Architecture
Federated Data Warehouse
Hub-and-Spoke Architecture
Which of the following schemas usually requires more JOIN operations due to complex dimension structures?
Snowflake Schema
Star Schema
Dimensional Flat Schema
Fact Constellation
A fact table "Daily Sales" at a retail company is directly connected to the dimensions "Product," "Time," and "Store." All dimensions consist of a single table without subdimensions. This approach describes…
Star Schema
Federated Schema
Fact Constellation
Snowflake Schema
The main characteristics of a Data Warehouse are, except
Time-variant
Integrated
Volatile
Subject-oriented
What happens when data from various operational systems enters the Data Warehouse and becomes integrated
The data is immediately displayed on a dashboard
The data is stored in its original format
The data is converted to a standard format and consolidated
The data is stored only if it comes from the main system
Non-volatile in the context of Data Warehouse means...
Data only persists during an active session
Data does not change after being loaded into the warehouse
Data cannot be analyzed again
Data is frequently updated to maintain accuracy
An analyst selects only sales data for the region "West Java" and the product "Beverages". This is an example of...
Pivot
Drill-down
Dice
Roll-up
What is the primary definition of a Data Warehouse?
A data storage system used for analysis and decision-making
A system for daily transactions in a company
A place to store project document files
Hardware for storing customer data
Transaction data from various branches is uploaded to the data warehouse every night. This process is part of the stage...
Integration
Transformation
Extraction
Load
When a customer changes address, the system stores the old and new versions in two separate rows. There are “Start Date” and “End Date” columns to track the active period. This is an example of...
SCD Type 0
SCD Type 1
SCD Type 3
SCD Type 2
In a school analytics system, the fact table “Student Grades” is connected to the dimensions “Student,” “Subject,” and “Teacher.” The “Subject” dimension has subdimensions “Subject Group” and “Field of Study” in separate tables. This reflects the schema...
Star Schema
Snowflake Schema
Flat Schema
Fact Constellation
If a dimension table has sub-dimensions or hierarchical relationships, the appropriate schema is
Flat Schema
Star Schema
Cube Schema
Snowflake Schema
A company stores the current version of information in the main row and keeps all historical data in a separate history table. This method is
SCD Type 3
SCD Type 2
SCD Type 4
SCD Type 0
A manufacturing company designs a data warehouse schema showing hierarchical relationships in the location dimension, where City connects to Province and Country as subdimensions in separate tables. This schema is characteristic of
Star Schema
Hybrid Schema
Snowflake Schema
Conformed Dimension
A university manages two fact tables: Student Enrollment and Tuition Payment. Both fact tables link to the same dimension tables, such as Student, Program, and Semester. The most suitable schema model for this situation is
Normalized Schema
Fact Constellation Schema
Independent Star Schema
Star Schema
During ETL, the same transaction data appears twice due to an input error. The required transformation technique is...
Data Deduplication
Data Mining
Sorting
Aggregation
In an OLAP report, a user chooses to view only data from the year 2024. This operation is called...
Slice
Pivot
Drill-down
Roll-up
An analyst wants to filter data only for January 2025 across all regions and products. The operation performed is...
Slice
Dice
Roll-up
Drill-down
In a Star Schema, dimension tables are typically...
Denormalized for query efficiency
Do not contain attributes
Not related to the fact table
Highly structured and normalized
A finance system only records the latest account status without storing previous versions. This is classified as...
SCD Type 0
SCD Type 2
SCD Type 3
SCD Type 1
A retail company builds several data marts separately for sales, finance, and logistics. Each is developed by different teams without coordination and they are not interconnected. Which architecture are they using?
Independent Data Marts
Data Mart Bus Architecture
Federated Data Warehouse
Hub-and-Spoke Architecture
Which of the following schemas typically requires more JOIN operations due to complex dimension structures?
Snowflake Schema
Star Schema
Dimensional Flat Schema
Fact Constellation
A "Daily Sales" fact table in a retail company connects directly to the dimensions "Product", "Time", and "Store". Each dimension consists of a single table without subdimensions. This approach describes...
Star Schema
Federated Schema
Fact Constellation
Snowflake Schema
Which of the following is NOT a primary characteristic of a Data Warehouse?
Time-variant
Integrated
Volatile
Subject-oriented
When data from various operational systems enters the Data Warehouse and becomes integrated, what happens to the data?
It is immediately displayed on a dashboard
It is stored in its original format
It is converted to a standard format and consolidated
It is stored only if it comes from the main system
Regression is a widely used statistical analysis technique. One of the benefits of regression is
Proving a hypothesis
Showing a depiction of data spread
Prediction/forecasting (predicting the value of the dependent variable when all independent variables’ values are known)
Depicting the spread or distribution of data
A diagram that maps hierarchical data using nested shapes (typically rectangles), displaying the hierarchy as a set of layered rectangles, is called
Histogram
Bullet
Tree map
Heatmap
Main purpose of data visualization is
To make it easier for stakeholders to make decisions based on the displayed data
To help study data or information
To make data or information look more attractive
To communicate information clearly and efficiently to users through information graphics
The type of chart suitable for visualizing words that are frequently discussed in a social media conversation topic is
Gantt chart
Bar chart
Word cloud
Pie chart
In a medical dataset of laboratory examination results, numeric values must be recorded to an appropriate number of decimal places for accurate interpretation of test outcomes. The data readiness characteristic in which data values are defined at a detailed level for the intended use is called
Data granularity
Data consistency
Data relevancy
Data accuracy
Consider variables such as age, number of children, total household income (in Rupiah), travel speed (in Km/hour), and temperature (in degrees Celsius). The taxonomy of these data is called
Nominal
Ratio
Numeric
Categorical
The mathematical method used to estimate or describe the degree of variation in a variable is called. Select one.
Variance
Measures of Dispersion
Standard Deviation
Quartiles and interquartile range
Measurement variables commonly found in physics and engineering—such as mass, length, time, plane angle, energy, and electric charge—are examples of the data taxonomy called. Select one.
Interval data
Ordinal data
Numeric data
Ratio data
Select one. Measurement of dispersion is a mathematical method used to estimate or describe the level of variation in a particular variable of interest and represents the numerical spread of a given dataset. The following is not a measure of dispersion:
Variance
Skewness
Standard Deviation
Quartiles and interquartile range
Select one. The representation of data that is changed into groups without considering the order among the groups, for example "marital status," in data taxonomy is called:
Numerical
Ordinal
Text
Interval
Reducing the range of data in each numeric variable to a standard range using normalization or scaling is the process. Select one.
Reduce dimension
Reduce noise
Normalize data
Balance data
Hypothesis testing is an example of implementation of. Select one.
OLAP
OLTP
Descriptive Statistics
Inferential Statistics
Scaling values to the range 0–1 in data transformation is known as. Select one.
Discretize Data
Reduce Attribute
Normalize Data
Balance Skewed Data
Select one: Business reports can be produced in the following formats, except
Metric management reports
Dashboard-type reports
Balanced scorecard reports
Database management
Select one: A Metric Management Report that presents an integrated view of organizational success covering financial, customer, business process, and learning-and-growth perspectives is called
Six Sigma
Balanced Scorecard
Total Quality Management
Key performance indicators
Select one: Sales data in a supermarket are presented as an example above. Data like that are called
Structured data
Semi-structured data
Unstructured data
All of the above
Select one: Celsius temperature is an example of …
Ordinal Data
Nominal Data
Interval Data
Ratio Data
Select one: Which is an example of ordinal data?
Temperature measurement data
Population count data
Binomial data (yes/no)
Range of students’ course grades
Select one: Raw Data Source → (1) Data Transformation → (2) Data Cleaning → (3) Data Reduction → (4) Data Consolidation → Well Formed Data. The correct order of Data Preprocessing steps is …
3-2-4-1
4-2-1-3
4-1-3-2
2-3-4-1
Select one: PT XYZ wants to monitor the relationship between total product sales and the costs incurred for product advertising. Based on this case, the most appropriate type of chart to use is
Bar Chart
Scatterplot
Treemap
Boxplot
Select one: In linear regression, the dispersion of errors from predicted values must be consistent; this assumption is known as
Linearity
Constant Variance
Multicollinearity
Independence
Select one: A person’s marital status is an example of
Ordinal Data
Nominal Data
Interval Data
Ratio Data
Statistical methods in business reports can be categorized into. Select one.
Descriptive or inferential
Descriptive or narrative
Narrative or inferential
Inferential or differential
Stunting diagnosis in children can be performed through physical examinations, such as measuring body weight, height, and head circumference. The taxonomy of data obtained in that diagnosis is called … Select one.
Nominal
Ordinal
Interval
Ratio
In linear regression assumptions, the assumption stating that the errors of the response variable are not correlated with each other is the assumption of. Select one.
Linearity
Independence (of error)
Normality (of error)
Constant variance (of error)
The head of the community health center wants to compare the number of toddlers with stunting across 5 Posyandu to identify which has the fewest to the most cases. The most appropriate chart for this visualization is
Line Chart
Bar Chart
Bubble Chart
Pie Chart
The term used to analyze, characterize, and summarize structured data stored in an organization’s database using cubes is
Statistics
Descriptive
Inferential
OLAP
Which of the following is not a key to successful reporting?
Clarity
Brevity
Open
Correctness
If imbalanced data are found, which data reduction process should be performed?
Construct Attribute
Reduce Attribute
Normalize Data
Oversampling Data
One characteristic of data readiness for an analytic study is that every element of the data is available in the dataset; this is called
Data source reliability
Data source accuracy
Data accessibility
Data richness
Select one. Simple, unorganized and unprocessed facts are characteristics of
Data
Information
Knowledge
Policy
Select one. The processes included in data preprocessing are
Data collection – data cleaning – data transformation – data reduction
Data consolidation – data cleaning – data transformation – data reduction
Data collection – data cleaning – data normalization – data reduction
Data consolidation – data cleaning – data normalization – data reduction
Select one. Reducing the range of existing data to a standard range (0–1) in data preprocessing is called
Data normalization
Data cleaning
Data reduction
Data collecting
Select one: Yes/No and True/False belong to which type of data?
Nominal data
Ordinal data
Textual data
Unstructured data
Select one: A district head claims that the poverty rate in his region is very low. To verify this, a survey of household income and expenditure is conducted, which theoretically can yield the poverty rate. Considering time and cost, a sample of 10,000 households is chosen from a total population of 100,000 households. The suitable statistical modeling for this case is
Descriptive statistics
Central tendency statistics
Inferential statistics
Distribution statistics
Select one: The appropriate diagram to illustrate comparison of continuous values across two categories using color, so users can quickly see where category intersections are strongest and weakest based on numeric measurements, is —
Histogram
Bullet
Tree map
Heatmap
Select one: Which activities are part of the data cleaning process?
Impute values, reduce noise, eliminate duplicates
Normalize data, discretize data, create attributes
Reduce dimension, reduce volume, balance data
Normalize data, reduce noise, eliminate duplicates
Select one. Data based on context refers to data suitable within a specific domain, and subject refers to data arranged according to the relevant subject. This data readiness characteristic is called
Data granularity
Data consistency
Data relevancy
Data accuracy
Select one. Which of the following are statistical measures that describe the dispersion of data?
Standard deviation and interquartile range
Mode and variance
Range and mean
Quartiles and median
A measurement scale used to determine the ranking of a particular group. In this ranking, only the order of objects is considered from the largest to the smallest or from the highest to the lowest. For example, students’ knowledge about Covid-19 (1 = poor, 2 = adequate, 3 = good). The data taxonomy for such an example is called
Nominal
Ordinal
Interval
Ratio
In the Kimball approach, the commonly used data structure is
XML
Star Schema and Snowflake Schema
3rd Normal Form (3NF)
Hierarchical database
