NEW
Font size
WorksheetsData Science
Total questions: 30
Worksheet time: 27mins
Q1. What is meant by a correlation?
A. The relationship between two or more variables
B. When there is an upward trend in a graph
C. When there is a set of data that doesn’t lie in the normal or expected range
D. When data is placed in a graph
Q6. The following graph shows the annual average temperatures recorded by a weather station over a period of 20 years. Identify the outlier in the data.
A. Point A
B. Point B
C. Point C
D. Point D
Q14. What was the difference between the cars sold on Monday and Tuesday than the cars sold on Friday and Saturday? (1 mark)
4
3
2
1
Sequential Modelling is done on
ANN
RNN
KNN
CNN
Data has been collected on visitors' viewing habits at a bank's website. Which technique is used to identify pages commonly viewed during the same visit to the website?
Clustering
Association Rules
Classification
Regression
You have been assigned to run a logistic regressionmodel for each of 100 countries, and all the data is currently stored in a PostgreSQL database. Which tool/library would you use to produce these models with the least effort?
MADlib
Mahout
RStudio
HBase
Which of the following of a random variable is a measure of spread?
Empirical mean
Standard deviation
Variance
All of the above
Which of the following testing is concerned with making decisions using data?
Probability
Casual
Hypothesis
Non of the above
Which of the following model is usually gold standard for data analysis?
Inferential
Casual
Descriptive
All of the above
Weighted Average is used in
Forecasting
Classification
Regression
All of the above
Which of the following is characteristic of Processed Data?
All steps should be noted
Hard to use for data analysis
Data is not ready for analysis
None of the above
In Supervised Learning :
The algorithm finds the right way to deal with a problem in successive iterations
Algorithm derives rules from unqualified data
The algorithm learns from qualified data
Steps in Data Science
Data Modeling ->Data Acquisition -> Clean Data ->Data Analysis ->Deployment and optimization
Data Acquisition -> Clean Data ->Data Analysis -> Data Modeling ->Deployment and optimization
Clean Data ->Data Analysis -> Data Modeling ->Deployment and optimization -> Data Acquisition
Data Modeling ->Data Acquisition -> Clean Data ->Data Analysis ->Deployment and optimization
Data from Facebook, Instagram, Twitter, Blogs are considered __________ data.
Unstructured/Internal
Unstructured/External
___________ is the last step of Data Science Project Lifecycle (DSPLC).
Data acquisition
Data preparation
Deployment
Optimization
A sample in which each individual or object in the entire population has an equal chance of being selected.
Non Probability Random Sample
Probability Random Sample
Biased Sample
Population
The value appearing at the center of a sorted version of a list, or the mean of the two central values.
Median
Mean
Range
Modus
Which of the following are "Measures of Central Tendency"?
Range, Standard Deviation, Variance
Mean,Range, Mode
Mean, Standard Deviation, Range
Mode, Mean, Median
Will filters work when we do data blending?
True
False
Data Science is
The science of creating data.
It is a branch of Social Studies.
Multidisciplinary study of data collections for analysis, prediction, learning and prevention.
It is a specialized field of study under Artificlal Intelligence
Which of the following are Continuous Quantitative Data?
Observations & Errors
Observations & Time
Time & Height
Time & Errors
Which of the following is true?
Means are correct!
Means are lies!
Means with confidence intervals are lies!
Means with confidence intervals are incorrect!
A hypothesis associated with a contradiction to a theory one would like to prove.
Null hypothesis
Alternative hypothesis
One hypothesis
Alpha
Can be described as inconsistencies and uncertainty in data
Volume
Value
Veracity
Variety
Velocity
Binary Logistic Regression following distribution?
Normal
Poisson
Bernouli
Binomial
Which of the following characteristic of big data is relatively more concerned to data science ?
Velocity
Variety
Volume
None of the Mentioned
Which of the following step is performed by data scientist after acquiring the data ?
Data Integration
Data Replication
Data Cleansing
All of the Mentioned
Which of the following of a random variable is a measure of spread?
Empirical mean
Standard deviation
Variance
All of the above
Weighted Average is used in
Forecasting
Classification
Regression
All of the above
Which of the following diagram is used to view correlation?
Corrgram
Triangle
Boxplot
Histogram
