wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Web Mining 1

Total questions: 82

Worksheet time: 41mins

Name
Class
Date
1.
According to the lecture, the Web is an important channel for which activity?
a)
Operating system scheduling
b)
CPU instruction decoding
c)
Analog signal filtering
d)
Memory garbage collection
e)
Transactions such as buying products online
2.
The Web is based on which architecture?
a)
Peer-to-peer only architecture
b)
Client-server architecture
c)
Single-machine architecture
d)
Offline batch-only architecture
e)
Sensor-network-only architecture
3.
Who invented the Web according to the lecture?
a)
Alan Turing
b)
Bill Gates
c)
Vint Cerf
d)
Tim Berners-Lee
e)
Claude Shannon
4.
Which early network is shown in the Web history timeline as a predecessor of the Internet?
a)
ARPANET
b)
Ethernet
c)
Bluetooth
d)
LTE
e)
Wi-Fi
5.
Mosaic is described as the first browser with what key feature?
a)
Built-in spreadsheet editing
b)
Full virtual reality support
c)
A graphical interface with mouse interaction
d)
Hardware-accelerated 3D gaming
e)
Encrypted email by default
6.
According to the lecture, which protocol suite allows networks to connect?
a)
HTML
b)
JPEG
c)
USB
d)
SQL
e)
TCP/IP
7.
What is one main role of the W3C mentioned in the lecture?
a)
Manufacturing web browsers
b)
Building standards for the Web
c)
Selling domain names
d)
Providing internet access
e)
Hosting all websites
8.
Which item appears as a subsection under the data mining part of the lecture?
a)
Compiler optimization
b)
Cache coherence
c)
Computer vision models
d)
Data types
e)
Digital signal modulation
9.
Data mining is also known as what in the lecture?
a)
Knowledge Discovery in Databases
b)
Web page rendering
c)
Network routing
d)
Image compression
e)
Process scheduling
10.
Why is preprocessing needed before applying data mining techniques?
a)
Raw data is always perfect and needs no changes
b)
Preprocessing replaces the need for evaluation
c)
Raw data is typically unsuitable and needs cleansing
d)
Preprocessing is only for creating web pages
e)
Preprocessing is only for encrypting data
11.
Which query language is listed as part of DBMS history in the lecture?
a)
HTML
b)
CSS
c)
HTTP
d)
SMTP
e)
SQL
12.
Which is an example of complex data management mentioned for advanced DBMS?
a)
Only plain text emails
b)
Spatial data
c)
Only binary executables
d)
Only printed paper forms
e)
Only hand-written notes
13.
Which capability is listed under advanced data analytics in the lecture?
a)
CPU micro-architecture design
b)
Circuit board soldering
c)
File system journaling
d)
Data warehouse and OLAP
e)
Kernel interrupt handling
14.
In the phases of data mining shown, which phase is the step of applying techniques to discover patterns?
a)
Mining
b)
Cleansing
c)
Integration
d)
Selection
e)
Representation
15.
In a relational database table, what identifies an object uniquely?
a)
A random file name
b)
A screen resolution
c)
A unique key
d)
A browser version
e)
A video codec
16.
A data warehouse is modeled by a multidimensional structure called what?
a)
A linked list
b)
A stack
c)
A binary heap
d)
A register file
e)
A data cube
17.
In OLAP, what does drill-down enable a user to do?
a)
Encrypt data for storage
b)
View data at a more detailed level
c)
Delete duplicate records
d)
Convert images to text
e)
Replace missing values with zeros
18.
Transactional data mining in the lecture focuses mainly on detecting what?
a)
Compiler warnings
b)
CPU pipeline hazards
c)
Image edges
d)
Frequent item sets
e)
Audio frequencies
19.
Which is an example of temporal data mentioned in the lecture?
a)
Stock data
b)
Static HTML template
c)
A single PDF page
d)
A fixed dictionary
e)
A one-time screenshot
20.
The lecture states there are two types of data mining tasks: descriptive and what?
a)
Procedural
b)
Administrative
c)
Predictive
d)
Decorative
e)
Mechanical
21.
Data discrimination is described as doing what?
a)
Compressing images into smaller files
b)
Encrypting network traffic
c)
Rendering web pages in a browser
d)
Scheduling processes on a CPU
e)
Comparing general features of one class with other classes
22.
In association analysis, confidence represents what?
a)
The total number of database tables
b)
The probability of Y given X
c)
The size of a data warehouse
d)
The number of web browsers installed
e)
The speed of a network link
23.
An association rule is considered invalid in the lecture if it is below which thresholds?
a)
Maximum speed threshold and minimum latency threshold
b)
Minimum voltage threshold and maximum current threshold
c)
Minimum font size threshold and maximum page width threshold
d)
Minimum support threshold and minimum confidence threshold
e)
Maximum file size threshold and minimum color depth threshold
24.
Classification builds a model from which type of training data?
a)
Items with class labels
b)
Only unlabeled items
c)
Only encrypted items
d)
Only image pixels without context
e)
Only random noise
25.
Clustering aims to do what according to the lecture?
a)
Maximize inter-cluster similarity and minimize intra-cluster similarity
b)
Sort web pages alphabetically
c)
Maximize intra-cluster similarity and minimize inter-cluster similarity
d)
Encrypt all database tables
e)
Convert all numbers to text
26.
In some applications like fraud detection, anomalies are considered what?
a)
Always irrelevant noise
b)
Guaranteed correct patterns
c)
Required training labels
d)
A type of web browser
e)
Important signals rather than just noise
27.
Which objective metric for association rules is explicitly mentioned in the lecture?
a)
Frame rate
b)
Support
c)
Battery level
d)
Screen brightness
e)
Disk rotation speed
28.
Which field is shown as related to data mining in the techniques overview?
a)
Thermodynamics
b)
Organic chemistry
c)
Classical sculpture
d)
Machine learning
e)
Marine biology
29.
A statistical hypothesis test can be used to do what with data mining output?
a)
Verify that a pattern is statistically significant
b)
Render a web page faster
c)
Increase network bandwidth
d)
Convert PDF files to images
e)
Install a database server
30.
Semi-supervised learning is described as using what kind of data?
a)
Only labeled data
b)
Only unlabeled data
c)
Both labeled and unlabeled data
d)
Only synthetic data
e)
Only encrypted data
31.
A data point that does not fit common behavior is often called what in the lecture context?
a)
Index
b)
Dimension
c)
Hyperlink
d)
Schema
e)
Outlier
32.
The lecture states that a data warehouse supports which kind of operations?
a)
CPU instruction decoding
b)
OLAP operations
c)
Wireless signal modulation
d)
Keyboard input processing
e)
Video frame rendering
33.
In the lecture, information retrieval typically assumes a query consists of what?
a)
Full database schemas with constraints
b)
Machine code instructions
c)
Encrypted binary blobs only
d)
Keywords without complex structure
e)
Hand-drawn diagrams
34.
In business intelligence, OLAP tools are based on what according to the lecture?
a)
A data warehouse
b)
A single local text file
c)
Only a web browser cache
d)
Only a spreadsheet macro
e)
Only a network switch
35.
The lecture notes that in web search systems, which part is processed online?
a)
All model training
b)
All crawler design decisions
c)
User queries
d)
All hardware manufacturing
e)
All web page authoring
36.
Which issue can negatively affect data mining and lead to wrong patterns, as stated in the lecture?
a)
Too many keyboard shortcuts
b)
Too many monitor pixels
c)
Too many browser tabs
d)
Too many USB ports
e)
Noise, uncertainty, and missing data
37.
The lecture says domain knowledge should be integrated into a data mining system for what purpose?
a)
To increase screen resolution
b)
To evaluate patterns and guide the mining process
c)
To reduce file compression ratio
d)
To change CPU clock speed
e)
To block internet access
38.
What does incremental mining enable according to the lecture?
a)
Replacing OLAP with OLTP
b)
Eliminating the need for preprocessing
c)
Guaranteeing zero noise in data
d)
Updating with new data without restarting the mining process
e)
Making all web data structured
39.
One effect-related recommendation in the lecture is to evaluate data sensitivity and ensure what?
a)
Data policy
b)
Video codec
c)
Browser theme
d)
Keyboard layout
e)
CPU cache
40.
Web data is described as heterogeneous. What does this mean in the lecture?
a)
All pages use an identical schema
b)
All web content is guaranteed correct
c)
Similar information can appear in different formats across pages
d)
All web pages are stored in one database
e)
All web pages are static and never change
41.
The lecture states that web information can be misleading mainly because the Web has what property?
a)
A single global editor approves all content
b)
Only experts can post online
c)
All pages are peer-reviewed journals
d)
All content is automatically verified
e)
No built-in content proof and anyone can publish anything
42.
Which activity is listed as a focus of web data mining in the lecture?
a)
Design CPU instruction sets
b)
Monitor web evolution
c)
Build mechanical engines
d)
Develop analog radios
e)
Manufacture storage disks
43.
A browser is described as sending a request, getting a response, compiling HTML, then displaying content. Which option best matches this flow?
a)
Compile HTML, request to server, display content, response from server
b)
Display content, compile HTML, request to server, response from server
c)
Request to server, response from server, compile HTML, display content
d)
Response from server, display content, request to server, compile HTML
e)
Request to server, compile HTML, response from server, display content
44.
Hypertext is described as allowing authors to link documents and embed multimedia. Which feature is a direct consequence of this idea?
a)
A document can include links to other documents and embed images, audio, or video
b)
A document must be stored only in a relational database table
c)
A document can be accessed only by a single user at a time
d)
A document can be read only on one operating system
e)
A document can be searched only by file name
45.
The lecture contrasts old information retrieval with the Internet era. Which change is emphasized?
a)
Information retrieval requires visiting a physical library more often
b)
Information retrieval depends primarily on television broadcasts
c)
Information retrieval becomes impossible without a printed index
d)
Information retrieval becomes possible with a few clicks from home or office
e)
Information retrieval becomes limited to a single company intranet
46.
The initial components of the Web include server, browser, HTTP, HTML, and URL. Which component primarily provides a way to address and locate resources?
a)
HTML
b)
URL
c)
HTTP
d)
Browser
e)
Server
47.
A web developer wants different web applications to interoperate using shared specifications. Which organization mentioned in the lecture is responsible for web standards?
a)
ARPA
b)
CERN
c)
Microsoft MSN
d)
Illinois University
e)
W3C
48.
The lecture states that discovered patterns should be correct, useful, and understandable. Which option is NOT one of these stated criteria?
a)
Correct
b)
Useful
c)
Visually colorful
d)
Understandable
e)
Potentially usable
49.
The data mining process is described as repeating until satisfied. If the mined output is not useful, what is the most consistent next step?
a)
Iterate the process by adjusting selection, preprocessing, or evaluation
b)
Skip evaluation and immediately deploy results
c)
Change the browser used for accessing data
d)
Replace all data with random samples without checking relevance
e)
Disable preprocessing to speed up mining
50.
When a dataset is too large or contains irrelevant attributes, which preprocessing step is suggested in the lecture?
a)
HTML compilation
b)
URL rewriting
c)
Transaction rollback
d)
Sampling or feature selection
e)
Hyperlink removal
51.
Traditional mining techniques focused on structured data. With the development of the Web, which data types become more important?
a)
Only binary machine code
b)
Semi-structured and unstructured data
c)
Only handwritten documents
d)
Only numerical matrices
e)
Only encrypted archives
52.
An organization needs summarized historical information for decision support, collected from multiple sources. Which storage concept best fits the lecture description?
a)
OLTP transaction database only
b)
Browser cache
c)
Single XML file
d)
Search engine index only
e)
Data warehouse
53.
A data cube is described as having dimensions and cells that store aggregates like counts or sums. Why does this help decision support tasks?
a)
It guarantees every web page is correct
b)
It eliminates the need for data integration
c)
It enables pre-computation and quick access to summarized views
d)
It converts unstructured text into fully labeled data
e)
It forces all queries to be processed without optimization
54.
In OLAP, drill-down and roll-up allow viewing different summarization levels. Which option is an example of roll-up consistent with the lecture?
a)
Observing country-level data from province-level data
b)
Observing monthly data from quarterly data
c)
Detecting anomalies using density-based methods
d)
Predicting labels from unlabeled items
e)
Crawling web pages to build an index
55.
A retailer asks: 'Which products are often bought together in one purchase?' According to the lecture, which mining focus best answers this?
a)
Regression modeling of continuous functions
b)
Taxonomy construction by clustering
c)
Web page classification by hyperlink count
d)
Frequent item set mining in transaction data
e)
SQL query optimization in DBMS
56.
Customers often buy a computer, then a camera, then a memory card in that order. Which pattern type from the lecture matches this example?
a)
Frequent item set
b)
Frequent sequential pattern
c)
Classification rule
d)
Data cube dimension
e)
Outlier cluster
57.
An association rule has high confidence but very low support. What is the best interpretation based on the lecture definitions?
a)
The rule applies to most transactions, but is rarely correct
b)
The rule is guaranteed to be correct in all future data
c)
The rule is a regression model for continuous outputs
d)
The rule does not require any minimum thresholds
e)
The rule is usually true when it applies, but it applies to a small fraction of transactions
58.
You have labeled training examples and need to predict class labels for new items. Which main data mining task fits best?
a)
Clustering
b)
Association rule mining
c)
Classification
d)
Data cleansing
e)
Roll-up
59.
A model is needed to predict a numeric value such as a continuous outcome. Which task from the lecture is most appropriate?
a)
Regression
b)
Clustering
c)
Classification
d)
Hypertext linking
e)
Transaction recovery
60.
A fraud detection system looks for transactions far from normal clusters. Which anomaly detection approach from the lecture matches this idea?
a)
Density-based methods
b)
Only hypothesis testing
c)
Only OLAP roll-up
d)
Distance-based methods
e)
Only association rule confidence
61.
A user believes a discovered pattern is interesting because it was not expected and provides strategic insight. Which type of evaluation metric does this align with?
a)
Objective metrics based on random variables
b)
Subjective metrics based on user belief
c)
Indexing metrics based on SQL optimization
d)
Hyperlink metrics based on URL length
e)
Graphical metrics based on screen pixels
62.
Information retrieval assumes data is mostly non-structured and queries are keyword-based. Which scenario fits this assumption best?
a)
Updating rows in a relational table using a primary key
b)
Executing a transaction with concurrency control
c)
Aggregating a data cube cell by sum
d)
Training a neural network using labeled images only
e)
Searching a large collection of text documents using keywords
63.
You have a small labeled dataset and a large unlabeled dataset. Which learning setup in the lecture is designed for this situation?
a)
Supervised learning only
b)
Unsupervised learning only
c)
Semi-supervised learning
d)
OLTP processing
e)
Crawling and indexing
64.
A learning approach is described where users participate by labeling limited data to optimize model quality under a labeling budget. What is this called?
a)
Active learning
b)
Incremental mining
c)
Attribute oriented inference
d)
Roll-up
e)
Ranking query processing
65.
The lecture notes that search models and query classifiers are built offline while queries are processed online. Which item is built offline?
a)
All user queries
b)
All web page downloads
c)
All advertisements displayed
d)
Search models and query classifiers
e)
All hyperlinks created by authors
66.
The lecture notes that web pages are linked and densely linked pages tend to have higher quality or impact. What web data characteristic supports link-based importance estimation?
a)
Only the volume of images on pages
b)
Hyperlink connectivity between web pages
c)
Uniform content formatting across all sites
d)
Guaranteed content correctness
e)
Complete absence of duplicates on the Web
67.
A company analyzes a very large network and wants groups of similar nodes, but also needs an importance ordering of nodes within and across groups. Which challenge-oriented approach from the lecture best matches this need?
a)
Use only roll-up and drill-down in OLAP
b)
Use only SQL indexing without mining
c)
Use only hypertext hyperlinks as final results without analysis
d)
Use only data cleansing and stop before mining
e)
Combine clustering and ranking to form high-quality clusters and rankings
68.
You want to discover patterns that depend jointly on multiple attributes (for example, age, income, and location) rather than analyzing each attribute separately. Which mining direction mentioned in the lecture best fits?
a)
Restrict mining to one dimension to avoid complexity
b)
Mining in multi-dimensional space by combining attribute dimensions
c)
Replace mining with web page rendering
d)
Use only client-server communication logs without features
e)
Use only transaction recovery mechanisms
69.
A system mines knowledge in an environment where items are semantically linked, and knowledge about one item should influence mining on related items. Which lecture idea is being used?
a)
Assuming all web content is structurally identical
b)
Treating all anomalies as irrelevant noise in every application
c)
Avoiding integration because heterogeneity is not an issue
d)
Mining in a semantically linked environment using related-item knowledge
e)
Using only hierarchical directory browsing for retrieval
70.
A dataset has noise, missing values, and uncertainty that could lead to wrong patterns. Which combination of remedies is most aligned with the lecture?
a)
Data cleansing and preprocessing, plus anomaly detection and uncertainty handling
b)
Skip preprocessing and rely on final visualization only
c)
Use only drill-down and roll-up operations
d)
Use only browser-side caching to remove noise
e)
Use only URL rewriting to remove missing values
71.
An association rule meets minimum confidence but fails minimum support. Based on the lecture, what is the correct decision about this rule under threshold-based mining?
a)
It is accepted because confidence is the only required threshold
b)
It is accepted and treated as statistically significant by default
c)
It is rejected because it does not meet the minimum support threshold
d)
It is converted into a regression model automatically
e)
It becomes valid only if the URL is shortened
72.
A discovered pattern has high support and confidence, but domain experts say it is obvious and provides no strategic value. Which evaluation perspective from the lecture best explains why it may be considered low value?
a)
Objective evaluation based on support only
b)
Query optimization based on indexing
c)
Client-server evaluation based on HTTP status codes
d)
Graphical evaluation based on browser rendering speed
e)
Subjective evaluation based on user belief and usefulness
73.
A web mining pipeline must handle pages that mix main content with ads, redirection, and policy text, and also handle the fact that the same information appears in different formats across sites. Which pair of web-data characteristics drives these needs?
a)
High-quality proof and uniform formatting
b)
Misleading content and heterogeneity
c)
Complete stability and perfect correctness
d)
Single data type and zero hyperlinks
e)
Only structured tables and no multimedia
74.
A fraud detection task must run continuously on streaming transactions and respond quickly. Which lecture-aligned system design choice best supports this requirement?
a)
Use only a static data warehouse snapshot with no updates
b)
Rely only on manual inspection without automation
c)
Use only web page crawling frequency as the signal
d)
Integrate data mining functions into DBMS and use incremental or stream-oriented mining
e)
Avoid preprocessing because it increases latency
75.
A mining algorithm must process a very large volume of data with short and predictable execution time. Which scalability strategy from the lecture is most directly applicable?
a)
Parallel and distributed computing with data partitioning and result combining
b)
Single-threaded execution on one machine with no partitioning
c)
Increasing query complexity to slow down processing
d)
Avoiding integration across data sources
e)
Replacing mining with static reports only
76.
Web content changes continuously, and an application must track changes without rerunning the entire mining process from scratch. Which concept from the lecture best addresses this?
a)
OLTP transaction rollback to undo changes
b)
Hypertext linking to embed multimedia
c)
Incremental mining to update results with new data
d)
Attribute oriented inference without any updates
e)
SQL joins to create larger tables
77.
A dataset contains anomalies located in a region where a global statistical distribution model performs poorly, but local neighborhood density reveals unusual behavior. Which anomaly detection approach from the lecture is most suitable?
a)
Distance-based methods only
b)
Treat anomalies as noise and remove them without analysis
c)
Use only OLAP roll-up
d)
Use only association rule support
e)
Density-based methods
78.
A learning team has a strict budget on how many items can be labeled, and wants to choose which items to label to maximize model quality. Which learning approach from the lecture is most appropriate?
a)
Unsupervised learning
b)
Active learning
c)
Data cube pre-computation
d)
Hyperlink-based page ranking only
e)
Transaction recovery
79.
A document collection is mostly unstructured text, and you want to identify the main themes that recur across documents without predefined labels. Which method mentioned in the lecture fits best?
a)
OLTP concurrency control
b)
HTTP request scheduling
c)
ER modeling in DBMS
d)
Topic models
e)
URL shortening
80.
A BI team wants fast interactive exploration across many dimensions, including drill-down and roll-up, without recomputing aggregates for every query. Which data-warehouse concept from the lecture best enables this?
a)
A data cube with pre-computation of summarized data
b)
A single flat table with no aggregation
c)
A browser cache of HTML pages
d)
A set of unrelated text files
e)
A web crawler schedule only
81.
A commercial website wants to automate suggestions and customer care to improve effectiveness. Which lecture view of the Web motivates applying data mining here?
a)
The Web is only a static library with fixed content
b)
The Web is proof-checked so mining is unnecessary
c)
The Web is a business and commercial channel that benefits from automation
d)
The Web contains only structured tables so integration is trivial
e)
The Web has no user interactions, only documents
82.
A team plans to deploy data mining inside an existing service to improve quality for users who may not understand mining techniques, while also reducing societal risks. Which combined consideration matches the lecture?
a)
Avoid integration and ignore data policy because it slows development
b)
Use only W3C standards and skip evaluation
c)
Use only OLAP drill-down and ignore user interaction
d)
Use only web crawling and skip preprocessing
e)
Integrate mining into systems and ensure data sensitivity and data policy are addressed