WorksheetsUAS Kelompok 2
Total questions: 10
Worksheet time: 10mins
What is different between clustering and classification?
Classification where each training data instance is a global category. In clustering the data is unlabeled and the process is unsupervised
Classification where each training data instance belongs to a particular class. In clustering the data is unlabeled and the process is supervised
Classification where each training data instance belongs to a particular class. In clustering the data is unlabeled and the process is unsupervised
Classification where each training data instance belongs to a global class. In clustering the data is unlabeled and the process is supervised.
They are the same but different framework
If you look around, you can find many other applications of clustering, but generally, clustering can be used for one of the following purposes except:
Exploratory data analysis
Summary generation or reducing the scale
Outlier detection, especially to be used for fraud detection, or noise removal
Finding duplicates in datasets
Classifying customers
What is simple linear regression?
Simple linear regression is when one independent variable is used to estimate a dependent variable
Simple linear regression is when one dependent variable is used to estimate an independent variable
Simple linear regression is when one independent variable is used to estimate an independent variable
Simple linear regression is when one dependent variable is used to estimate an dependent variable
Simple linear regression is when one dependent variable is used to predict a value
Which of the following is correct about multiple linear regression?
Unlike the case with simple linear regression, multiple linear regression is a method of predicting a continuous variable
It uses multiple variables, called independent variables, or predictors, that best predict the value of the target variable, which is also called the dependent variable
In multiple linear regression, the target value, x, is a linear combination of independent variables, x
It can’t be used when we would like to identify the strength of the effect that the independent variables have on a dependent variable
It can be used to predict the impact of changes
What is Training accuracy?
The percentage of correct predictions that the model makes when using the test dataset
How accurate is the data
Result of trained supervised data that has been processed as data set
Result of trained unsupervised data that has been processed as data set
Result of correct model that has been processed and proved to be a non-over-fit data model
Why a high training accuracy isn’t necessarily a good thing?
Accuracy can be manipulated by the data engineer
Its calculation can be modified at any time
Accuracy result is not that important to be counted
Having a high training accuracy may result in an ‘over-fit’ of the data
Accuracy is not important as true positive value
What does over-fit mean?
The model is overly trained to the dataset, which may capture noise and produce a non-generalized model
The model is overly trained to the unsupervised dataset, which may capture noise and produce a non-generalized model
The model is overly trained to the supervised dataset, which may capture noise and produce a non-generalized model
The model has low accuracy as the result of less data checked in the process
The model has low accuracy as the result of too much data checked in the process
Why data privacy is important? What do company face if they don’t protect customer’s data?
They face potential financial and legal repercussions
They want a longer relationship
They will be given a new contract
They wanted a better money
They will be run out money
Enterprises must learn how to use data responsibly and transparently with:
Stricter regulations and build relationship with customer
New leadership and employee
New board of directions and coordinator
Modifying vision and mission
Generating more report and decision making tools
Companies that don’t properly protect and account for data risk abusing it unwittingly, which can damage a company’s:
Physical data storage such as servers
Data and technical problems
Business process and systems
Reputation and erode its relationships with its customers
Income line
