WorksheetsTEK615_Descriptive_Analytics_Mock_Exam
Total questions: 14
Worksheet time: 14mins
Q1-Q2. Dr. Quick and Dr. Quack are both in the business of selling diets, and they both claim that their programs have indeed helped people reduce their weights successfully. They collected a random sample of their customers after following their programs within one month and conducted a statistical mean testing using the following null hypothesis
H0: The average weight change of the population of interest (their dieters) = zero
Q1: Looking at the statistical test result for Dr. Quack’s program (N=500), which of the following statements is correct?
Given the sample data, the probability that the null hypothesis is true is 0.0367.
Since the p-value is smaller than 0.05, it means that the null hypothesis has been proven to be false, that is, the average weight change of the population of interest is not zero.
If one assumes that the null hypothesis is true, the probability of observing an average weight change of -0.91 or less in the population of interest is 0.0184.
Given the sample data, the probability that the null hypothesis is true is 0.0184.
Q1-Q2. Dr. Quick and Dr. Quack are both in the business of selling diets, and they both claim that their programs have indeed help people reduce their weights successfully. They collected a random sample of their customers after following their programs within one month and conducted a statistical mean testing using the following null hypothesis
H0: The average weight change of the population of interest (their dieters) = zero
Q2: Looking at the statistical test result for Dr. Quick’s program (N=20), which of the following statements is correct?
If we were to repeat the experiment over and over, then 95% of the time the true average weight change of their dieters falls between -7.96 and 2.505.
The true average weight change of their dieters is zero because the 95% confidence interval of the sample mean includes zero.
We can be 95% confident on the (repeated) sampling method that the true average weight change of their dieters falls between -7.96 and 2.505.
There is a 95% probability that the true average weight change of their dieters lies between these two values (-7.96 and 2.505).
Q3. A business analyst is hired to analyze customer profitability of a bank. The bank has approximately 5 million customers. With the assistance of IT department, he managed to get 31634 random samples out of those 5 million customers. He wants to know whether those customers who use online banking are more profitable than those who do not (i.e. offline). He uses linear regression method with profit as the Y variable (name: 9Profit, type: continuous, value: real numbers) and whether the customer uses online banking (1) or not (0) as the X variable (name: 9Online, type: continuous, value: 0 or 1). The results are shown as follow.
The difference between the profit mean of the online customers and that of the offline ones is inconclusive.
There is a meaningful difference between the profit mean of the online customers and the offline ones as indicated by the p-value of the slope (0.2099).
Since one cannot reject the null hypothesis, this means that it is now proven by the data that the profit mean of the online customers is the same as that of the offline customers.
Based on the sample data, the estimated profit mean for the offline customers is $116.67.
Q4. In order to better explain the variation in profit, he looks at the demographic data of the customers. He hypothesized that younger customers may be more likely both to be online and to be less profitable. Therefore, the age of the customer may be of help in explaining the profit variation. Here is the result when the variable age is added together with whether the customer uses online banking or not (Y=9Profit, X1=9Online, X2=9Age). Note that the age variable’s values are coded numerically (i.e. 1 = less than 15 years; 2 = 15-24 years; 3 = 25-34 years; 4 = 35-44 years; 5 = 45-54 years; 6 = 55-64 years; 7 = 65 years and older).
What can he say about the role of age in explaining the variation in profit?
Two customers, one is in age group 1 and uses online banking, the other is in age group 2 and does not use online banking; the one who uses online banking is, on average, around $27.19 more profitable than the one who does not use online banking.
The result is inconclusive. This means that there is a lack of evidence that the age of the customer can explain the profit variation between the online customers and the offline ones.
Two customers, one is in age group 1 and uses online banking, the other is in age group 2 and uses online banking; the one in age group 2 is, on average, around $25.86 more profitable than the one in age group 1.
None of the other alternatives
Q5-Q6. Kang Corporation is considering 3 options for managing its data processing operation: continuing with its own staff, outsource it to an external vendor, or using a combination of own staff and outsourcing. The cost of the operation depends on future demand. The probability of a high demand is 0,2, the probability of a medium demand is 0,5, and the probability of a low demand is 0,3. The advantage of using its own staff is that it costs the same to manage the data process for a high demand or a medium demand: this cost is $650 000. If the demand is low, the company would only save $50 000 compared to a high demand. If the company decides to outsource the data processing operation the cost is $900 000 for a high demand, $600 000 for a medium demand, and $300 000 for a low demand. A combination of outsourcing and its own staff costs $500 000 for a low demand, $650 000 for a medium demand and $800 000 for a high demand.
Q5. Which of the following options would you recommend for managing the data processing?
Own Staff
Outsource
Combination
Q5-Q6. Kang Corporation is considering 3 options for managing its data processing operation: continuing with its own staff, outsource it to an external vendor, or using a combination of own staff and outsourcing. The cost of the operation depends on future demand. The probability of a high demand is 0,2, the probability of a medium demand is 0,5, and the probability of a low demand is 0,3. The advantage of using its own staff is that it costs the same to manage the data process for a high demand or a medium demand: this cost is $650 000. If the demand is low, the company would only save $50 000 compared to a high demand. If the company decides to outsource the data processing operation the cost is $900 000 for a high demand, $600 000 for a medium demand, and $300 000 for a low demand. A combination of outsourcing and its own staff costs $500 000 for a low demand, $650 000 for a medium demand and $800 000 for a high demand.
Q6. What is the expected cost of your recommendation?
$635 000
$520 000
$570 000
$610 000
Q7. What is the difference between ratio and ordinal variables?
For ordinal variables categories can be rank ordered, while for ratio variables they cannot.
Distance between categories is equal for ratio variables, while for ordinal variables distance between categories can be different.
Distance between categories is equal for ordinal variables, while for ratio variables distance between categories can be different.
Ratio variables have two categories, while ratio variable have more than two categories.
Q8. You are preparing a dashboard to support your diagnosis of shipments punctuality. You want to highlight that the size of the transport operator’s fleet and the number of warehouses in the country seem to affect the KPI that measures the punctuality of shipments. Choose one type of chart to include in your dashboard:
A stacked 100% area chart
A scatter plot bubble size
A variable width chart
A bar histogram
Q9. What is FALSE about unsupervised learning algorithms?
The data does not have a dependent variable.
The goal of these algorithms is to capture patterns based on a sample, similar to a process of “reverse engineering.”
They aim at classifying objects.
The dependent variable should be correlated with the independent variables.
Q10. What is NOT a desirable characteristic of a cluster assignment?
Each cluster contains only members of a single class.
Members of a given class are assigned to the same cluster.
A homogeneous cluster contains only correlated variables.
Clusters have high cohesion and there is high separation among them
Q11. Which of the following examples of supply chain management problems could potentially be addressed using a k-means algorithm?
Optimizing inventory levels in a retail store chain based on historical sales data.
Identifying patterns in customer purchasing behavior to segment markets.
Predicting the demand for a new product based on demographic data.
Selecting the most cost-effective transportation routes for delivering goods to customers.
Q12-Q13. You are a supply chain analytics consultant for a small company that commercializes construction supplies. You have information about INVOICES from CUSTOMERS that demanded items from different LINES of PRODUCTS that the company bought from VENDORS. The following graph represents the relational database.
Q12. To obtain a dataframe containing information about the line of products for the product reference P_CODE=54778-2T, sold to the customer identified with CUS_CODE = 10014, while excluding any unnecessary data, which merging function would be the most appropriate?
Inner join.
Left join.
Right join.
Full outer join.
Q12-Q13. You are a supply chain analytics consultant for a small company that commercializes construction supplies. You have information about INVOICES from CUSTOMERS that demanded items from different LINES of PRODUCTS that the company bought from VENDORS. The following graph represents the relational database.
Q13. Based on the relational database graph, indicate which of the following statements is incorrect.
A vendor may supply many products.
One product is supplied by only one single vendor.
An invoice contains only one line of products.
A line of products can’t be in many invoices.
Q14. When conducting an RFM analysis, which of the following queries generates a report sorting the most recent customers, without duplicates, who have purchased your company's products?
SELECT P_DESCRIPT, P_PRICE, V_NAME, V_CONTACT, V_AREACODE,V_PHONE
FROM PRODUCT, VENDOR
WHERE PRODUCT.V_CODE = VENDOR.V_CODE
AND P_INDATE > '2024-05-01'
SELECT CUSTOMER.CUS_CODE, INVOICE.INV_DATE, CUSTOMER.CUS_LNAME
FROM INVOICE, CUSTOMER
WHERE INVOICE.CUS_CODE = CUSTOMER.CUS_CODE
ORDER BY INV_DATE DESC
SELECT DISTINCT *
From CUSTOMER
ORDER BY INV_DATE DESC"""
SELECT DISTINCT CUSTOMER.CUS_CODE, CUSTOMER.CUS_LNAME, CUSTOMER.CUS_FNAME
From INVOICE, CUSTOMER
WHERE INVOICE.CUS_CODE = CUSTOMER.CUS_CODE
ORDER BY INV_DATE DESC
