Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Lecture 1 and 2

Total questions: 100

Worksheet time: 50mins

Name
Class
Date
1.

What is the output of the following code? t = (1, 2, 3) + (4, 5) print(t)

a)

(1, 2, 3, 4, 5)

b)

(5, 4, 3, 2, 1)

c)

[1, 2, 3, 4, 5]

d)

(1, 2, 3), (4, 5)

2.

What method is used to remove a specific value from a list?

a)

remove()

b)

pop()

c)

delete()

d)

del

3.

The _____ function returns a new sorted list from any sequence.

a)

sorted

b)

sort

c)

order

d)

arrange

4.

The update method is used to merge one dictionary into another.

a)

True

b)

False

5.

Which method is used to check if a dictionary contains a specific key?

a)

in

b)

contains()

c)

has_key()

d)

key_check()

6.

The pandas library provides high-level data structures for working with structured and tabular data.

a)

True

b)

False

7.

Which of the following is a mutable object in Python?

a)

String

b)

Tuple

c)

List

d)

None

8.

What will be the output of this code? x = 10; y = 3; print(x // y)

a)

3.33

b)

3

c)

3.0

d)

4

9.

Output of this code: nums = [1,2,3] print(sum(nums))

a)

6

b)

5

c)

3

d)

Error

10.

What does the following code print? a = [1,2,3] print(a[::-1])

a)

A) [3,2,1]

b)

B) [1,2,3]

c)

C) [1,3,2]

d)

D) Error

11.

Output of this code: d = {"a":1, "b":2} print(d.get("c", 0))

a)

0

b)

None

c)

Error

d)

2

12.

What does the following code return? lst = [1,2,3,4] print(lst.index(3))

a)

A) 2

b)

B) 3

c)

C) 1

d)

D) Error

13.

The with statement is used to handle ______ safely and ensure they are closed properly.

a)

files

b)

loops

c)

variables

d)

exceptions

14.

Which of the following is **not** a valid input for creating a DataFrame?

a)

A single integer

b)

A dictionary of lists

c)

A NumPy array

d)

A list of dictionaries

15.

What is the output of the following code? import numpy as np a = np.array([1, 2, 3]) print(a[1])

a)

1

b)

2

c)

3

d)

Error

16.

What does this code print? import numpy as np a = np.array([[1, 2], [3, 4]]) print(a.shape)

a)

(2, 2)

b)

(2,)

c)

(4,)

d)

Error

17.

Output of the following code? import numpy as np a = np.array([1,2,3]) print(a.dtype)

a)

int64

b)

float64

c)

object

d)

Error

18.

What is the result? import numpy as np a = np.zeros((2,3))

a)

0 0 0 0 0 0

b)

0 0 0 0 0 0

c)

0 0 0 0 0 0

d)

0 0 0 0 0

19.

Output of this code? import numpy as np a = np.array([1,2,3,4]) print(np.where(a>2))

a)

(array([2,3]),)

b)

(array([0,1]),)

c)

[3 4]

d)

Error

20.

What does this print? import numpy as np a = np.array([1,2,3,4]) print(np.unique([1,2,2,3,3,4]))

a)

A) [1 2 3 4]

b)

B) [1 2 2 3 3 4]

c)

C) [2 3 4]

d)

D) Error

21.

What is a key difference between a pandas `Series` and a `DataFrame`?

a)

A) `Series` is one-dimensional, while a `DataFrame` is two-dimensional.

b)

B) `Series` can contain multiple data types, while a `DataFrame` cannot.

c)

C) `DataFrame` does not have indexes, while a `Series` always does.

d)

D) Both `Series` and `DataFrame` are always empty by default.

22.

Which method allows you to select data from a DataFrame by row and column labels?

a)

`loc`

b)

`iloc`

c)

`index`

d)

`slice`

23.

What will the following code output? ```python import pandas as pd data = {"A": [1, 2], "B": [3, 4]} df = pd.DataFrame(data) print(df.iloc[1]) ```

a)

A) A 2, B 4

b)

B) A 1, B 3

c)

C) [1, 3]

d)

D) Error: No row with index 1

24.

What will the following code output? ``` python import pandas as pd s = pd.Series([1, 2, 3], index=["a", "b", "c"]) print(s["b"]) ```

a)

A) 2

b)

B) 1

c)

C) "b"

d)

D) Error: Invalid index access

25.

What does this code return? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df[df['A']>1])

a)

A) A B 1 2 5 2 3 6

b)

B) A B 0 1 4

c)

C) [2,3]

d)

D) Error

26.

Output of this code? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df.iloc[1,1])

a)

5

b)

2

c)

4

d)

Error

27.

What does the following code produce? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df.loc[0,'B'])

a)

4

b)

1

c)

0

d)

Error

28.

Which library is typically used to read and parse Excel files in pandas?

a)

openpyxl

b)

pickle

c)

lxml

d)

json

29.

Which pandas method writes a DataFrame to pickle format?

a)

to_pickle()

b)

write_pickle()

c)

to_binary()

d)

pickle_dump()

30.

What will the following code do? import pandas as pd df = pd.read_excel('data.xlsx', sheet_name='Sheet1') print(df.head())

a)

Reads the Excel file 'data.xlsx' from 'Sheet1' and prints the first 5 rows.

b)

Writes the DataFrame to a new Excel file.

c)

Displays the entire DataFrame in the console.

d)

Imports the pandas library.

31.

What does this code output?

a)

Converts numbers like '1,000' to 1000

b)

Treats ',' as delimiter

c)

Reads all values as strings

d)

Error

32.

Output of this code? import pandas as pd df = pd.read_csv('data.csv', skip_blank_lines=True)

a)

Ignores blank lines in the CSV

b)

Reads blank lines as NaN

c)

Error

d)

Deletes CSV

33.

What is the output of this code? import pandas as pd df = pd.read_csv('data.csv', low_memory=False)

a)

A) Prevents dtype guessing and ensures proper memory usage

b)

B) Reads CSV in chunks

c)

C) Converts everything to string

d)

D) Error

34.

What does this code do? import pandas as pd df = pd.read_sql('SELECT * FROM table1', conn, index_col='ID')

a)

Reads SQL table and sets 'ID' as index

b)

Writes SQL table

c)

Converts SQL table to CSV

d)

Creates a new SQL table

35.

What does this code do?

a)

Reads CSV without using the first row as header

b)

Reads CSV using the first row as header

c)

Reads only the first row

d)

Writes CSV without header

36.

How do you remove rows with all NaN values?

a)

Drops rows where all values are NaN

b)

Drops any row with NaN

c)

Drops column with NaN

d)

Fills NaN

37.

How can you strip special characters from a string column?

a)

Removes all non-alphanumeric characters

b)

Converts to lowercase

c)

Converts to uppercase

d)

Replaces spaces with underscores

38.

Output of this code? import pandas as pd df = pd.DataFrame({'A':[1,2,3,4,5]}) Q1 = df['A'].quantile(0.25) Q3 = df['A'].quantile(0.75) IQR = Q3 - Q1 df_filtered = df[(df['A'] >= Q1 - 1.5*IQR) & (df['A'] <= Q3 + 1.5*IQR)] print(df_filtered)

a)

Filters out outliers based on IQR

b)

Keeps only outliers

c)

Drops all rows

d)

Error

39.

How do you replace all infinite values with NaN?

a)

Replaces +inf/-inf with NaN

b)

Drops infinite values

c)

Converts to zero

d)

Raises error

40.

What does this code do?

a)

Converts all strings in the column to lowercase

b)

Converts to uppercase

c)

Strips spaces

d)

Deletes column

41.

How can you remove columns with more than 50% missing values?

a)

Drops columns with more than 50% NaN

b)

Drops rows with >50% NaN

c)

Replaces NaN with 0

d)

Keeps only rows with >50% non-NaN

42.

What does this code produce? import pandas as pd df1 = pd.DataFrame({'key':[1,2,3],'A':[10,20,30]}) df2 = pd.DataFrame({'key':[3,4,5],'B':[300,400,500]}) pd.concat([df1, df2], axis=0, ignore_index=True)

a)

Stacks df1 and df2 vertically and resets the index

b)

Stacks horizontally

c)

Performs a merge

d)

Produces an error

43.

What is the result of this code? import pandas as pd df = pd.DataFrame({'X':[1,2],'Y':[3,4]}) pd.melt(df, id_vars=['X'])

a)

Keeps 'X' fixed and unpivots 'Y' into long format

b)

Keeps 'Y' fixed and unpivots 'X'

c)

Drops column 'Y'

d)

Produces an error

44.

What does this code do?

a)

Outer join: keeps all keys from both DataFrames, fills missing with NaN

b)

Left join

c)

Inner join

d)

Right join

45.

Output of this code? import pandas as pd df = pd.DataFrame({'id':[1,1,2,2],'variable':['X','Y','X','Y'],'value':[10,20,30,40]}) df.pivot(index='id', columns='variable', values='value')

a)

Reshapes from long to wide format with 'id' as index

b)

Reshapes from wide to long

c)

Drops column 'value'

d)

Produces an error

46.

Left join: keeps all rows from df1, adds matching rows from df2

a)

Inner join

b)

Right join

c)

Outer join

47.

What is the difference between merge and concat?

a)

merge joins based on columns or keys, concat stacks DataFrames along axis

b)

merge concatenates, concat joins

c)

Both do the same thing

d)

merge deletes duplicates, concat does not

48.

How can you reorder levels in a MultiIndex?

a)

Switches the positions of the levels

b)

Sorts the index

c)

Drops a level

d)

Creates a new column

49.

What is the key difference between reorder_levels() and swaplevel() in hierarchical indexing?

a)

reorder_levels() allows arbitrary reordering of index levels, while swaplevel() only swaps two levels

b)

swaplevel() can reorder multiple levels, while reorder_levels() swaps levels

c)

reorder_levels() can be used on non-hierarchical indexes, while swaplevel() cannot

d)

swaplevel() requires specifying all levels, while reorder_levels() defaults to the first two

50.

In a merge() operation, what does the how='outer' parameter do?

a)

Performs a union of the keys from both DataFrames, including all rows from both

b)

Includes only rows with keys present in both DataFrames

c)

Includes rows from the left DataFrame only

d)

Includes rows from the right DataFrame only

51.

Output of this code? import pandas as pd arrays = [['A','A','B','B'], [1,2,1,2]] index = pd.MultiIndex.from_arrays(arrays, names=('letter','num')) df = pd.DataFrame({'val':[10,20,30,40]}, index=index) df.loc['A']

a)

Selects all rows where first level of MultiIndex is 'A'

b)

Selects rows where second level is 'A'

c)

Returns columns named 'A'

d)

Produces an error

52.

Output of this code? import pandas as pd df = pd.DataFrame({'id':[1,1,2,2],'variable':['X','Y','X','Y'],'value':[10,20,30,40]}) df.pivot(index='id', columns='variable', values='value')

a)

Reshapes from long to wide format with 'id' as index

b)

Reshapes from wide to long

c)

Drops column 'value'

d)

Produces an error

53.

How do you rotate y-axis tick labels?

a)

Rotates y-axis labels by 90 degrees

b)

Rotates x-axis

c)

Rotates plot

d)

Produces an error

54.

How do you change bar width in a bar plot?

a)

Sets bar width to 0.3

b)

Sets spacing

c)

Changes color

d)

Produces an error

55.

To plot a histogram with normalized frequencies, which parameter should be set in the plotting function?

a)

Set the 'density' parameter to True

b)

Set the 'bins' parameter to False

c)

Set the 'color' parameter to 'blue'

d)

Set the 'alpha' parameter to 0.5

56.

df['A'].plot.hist(density=True) Which of the following does this code do?

a)

Plots histogram normalized to form a probability density

b)

Plots raw counts

c)

Plots line plot

d)

Produces an error

57.

Which parameters control the spacing between subplots in matplotlib?

a)

wspace and hspace

b)

width_space and height_space

c)

width_space and height_space

d)

padding_x and padding_y

58.

What is the default behavior for connecting points in a matplotlib line plot?

a)

Linear interpolation

b)

Cubic interpolation

c)

Step-wise connection

d)

No connection

59.

Which function is used to set the x-axis label in matplotlib?

a)

set_xlabel()

b)

set_xlim()

c)

set_xticks()

d)

set_title()

60.

Which function is used to create a legend in a plot?

a)

plt.legend()

b)

plt.show_legend()

c)

plt.add_legend()

d)

plt.legend_box()

61.

How can you customize the row and column variables in a seaborn FacetGrid?

a)

Pass them to the row and col parameters of sns.FacetGrid()

b)

Set them using grid.set_rows() and grid.set_cols()

c)

Define them in the facet_vars parameter of FacetGrid()

d)

Use the rows and columns parameters in FacetGrid()

62.

What happens to missing values in a group key during a GroupBy operation?

a)

They are excluded from the result

b)

They are replaced with zeros

c)

They are included as a separate group

d)

An error is raised

63.

When grouping by multiple keys, the type of the first element in the group tuple is:

a)

A tuple containing all grouping keys

b)

A list of grouped values

c)

A dictionary of key-value pairs

d)

A single grouping key

64.

What does the GroupBy object return when iterated over?

a)

Tuples containing the group key and group data

b)

Only the group key

c)

Only the group data

d)

Lists of grouped rows

65.

What does grouping by a dictionary in pandas achieve?

a)

Maps specific values to group names

b)

Creates hierarchical groups

c)

Applies multiple aggregation functions

d)

Splits groups by index levels

66.

How do you apply different aggregations to different columns? df.groupby('Category').agg({'Value':'sum','Score':'max'})

a)

Aggregates Value by sum and Score by max per group

b)

Aggregates all by sum

c)

Aggregates all by max

d)

Produces an error

67.

How do you pivot with multiple index and columns?

a)

Creates hierarchical index and columns, aggregating sum

b)

Creates flat index

c)

Drops Region

d)

Produces an error

68.

How do you count occurrences in crosstab? pd.crosstab(df['Category'], df['SubCategory'])

a)

Counts frequency of each Category/SubCategory combination

b)

Sums values

c)

Computes mean

d)

Produces an error

69.

What does pd.groupby return?

a)

A tuple of key values

b)

A single key value

c)

A pandas DataFrame

d)

A list of grouped rows

70.

df.groupby('Category')['SubCategory'].nunique() What does this code do?

a)

Returns number of unique SubCategory values per Category

b)

Returns total count

c)

Returns mean

d)

Produces an error

71.

Output of this code? df.groupby('Category')['Value'].agg(['sum','count','mean'])

a)

Returns sum, count, and mean for each category

b)

Returns sum only

c)

Returns mean only

d)

Produces an error

72.

What does this code produce? import pandas as pd df = pd.DataFrame({'Category':['A','B','A','B'],'Value':[10,20,30,40]}) df.groupby('Category').sum()

a)

Sums 'Value' for each category

b)

Counts rows per category

c)

Averages 'Value' per category

d)

Produces an error

73.

How do you apply multiple aggregations and rename columns? df.groupby('Category')['Value'].agg(Total='sum', Average='mean')

a)

Returns grouped aggregation with renamed columns

b)

Produces error in older pandas versions

c)

Aggregates sum only

d)

Aggregates mean only

74.

What will pd.to_datetime(["2018-02-29"]) return?

a)

ValueError due to an invalid date

b)

NaT for the invalid date

c)

A datetime object representing February 28, 2018

d)

An empty DataFrame

75.

When assembling datetime objects using pd.to_datetime(df), what happens if a column is missing (e.g., "hour")?

a)

The missing column is filled with default values.

b)

An error is raised and the operation fails.

c)

The missing column is ignored and not included in the datetime object.

d)

The missing column is replaced with NaN values.

76.

What happens if a column is missing when performing an operation on a DataFrame?

a)

The missing column is filled with default values (e.g., 0 for hours)

b)

An error is raised due to the missing column

c)

The operation skips rows with missing columns

d)

The DataFrame is converted without the missing field

77.

Which parameter in pd.to_datetime() allows you to set the timezone of the resulting datetime objects?

a)

utc

b)

tz

c)

timezone

d)

localize

78.

What happens if the freq="infer" parameter fails to determine a consistent frequency?

a)

A ValueError is raised

b)

The index is created without a frequency

c)

The freq parameter defaults to daily

d)

A warning is issued

79.

Which format string correctly parses "12-11-2010 00:00"?

a)

%d-%m-%Y %H:%M

b)

%Y-%m-%d %H:%M

c)

%d/%m/%Y %H:%M

d)

%m-%d-%Y %H:%M

80.

What does the freq="M" parameter specify in pd.date_range()?

a)

Monthly frequency

b)

Minute-based frequency

c)

Monday-based weekly frequency

d)

Milliseconds frequency

81.

Which method converts a pandas period series back to timestamps?

a)

to_timestamp()

b)

to_period()

c)

to_datetime()

d)

to_dates()

82.

How do you forward-fill missing timestamps after resampling?

a)

Fills NaNs using previous available value

b)

Fills with zero

c)

Drops missing

d)

Produces error

83.

Computing rolling correlation between two series involves:

a)

Calculating the correlation coefficient over a moving window

b)

Summing the values of both series

c)

Multiplying the values of both series

d)

Finding the maximum value in each window

84.

df['value'].rolling(5).corr(df['other']) What does this code do?

a)

Returns rolling correlation over 5-row window

b)

Returns covariance

c)

Returns sum

d)

Produces error

85.

df['value'].shift(freq=pd.DateOffset(days=3)) What does this code do?

a)

Shifts values 3 days along the datetime index

b)

Shifts by 3 rows

c)

Drops rows

d)

Produces error

86.

pd.date_range('2020-01-01','2020-01-10', freq='B') What does this code do?

a)

Returns dates skipping weekends

b)

Returns all days

c)

Returns only weekends

d)

Produces error

87.

df.index + pd.offsets.MonthEnd() What does this code do?

a)

Shifts each date to the end of month

b)

Shifts to month start

c)

Produces error

d)

Drops index

88.

Shifting values using a time offset is done by:

a)

Applying a lag or lead function to the data.

b)

Multiplying the values by a constant factor.

c)

Sorting the data in ascending order.

d)

Filtering out null values from the dataset.

89.

Backward-filling missing timestamps after upsampling is done by:

a)

Using the bfill method to fill missing values with the next valid observation

b)

Using the ffill method to fill missing values with the previous valid observation

c)

Dropping all missing timestamps

d)

Interpolating missing timestamps with linear values

90.

df.resample('H').bfill() What does this operation do?

a)

Fills NaNs using next valid value

b)

Fills with zero

c)

Drops rows

d)

Produces an error

91.

How do you shift values by 2 periods with shift? df['value'].shift(2) What does this operation do?

a)

Moves values down 2 rows

b)

Shifts index

c)

Drops rows

d)

Produces an error

92.

s = pd.Series([1,2,3], index=pd.date_range('2023-01-01', periods=3)) s.rolling(2, min_periods=1).sum() What will be the output?

a)

2023-01-01 1.0 2023-01-02 3.0 2023-01-03 5.0 dtype: float64

b)

2023-01-01 NaN 2023-01-02 3 2023-01-03 5 dtype: float64

c)

2023-01-01 1 2023-01-02 2 2023-01-03 3 dtype: int64

d)

Produces an error

93.

df = pd.DataFrame({'value':[1,2,3]}, index=pd.period_range('2023-01-01','2023-03', freq='M')) df.index.to_timestamp() What will be the output?

a)

2023-01-01 1 2023-02-01 2 2023-03-01 3 Freq: MS, Name: value, dtype: int64

b)

2023-01-31 1 2023-02-28 2 2023-03-31 3 dtype: int64

c)

Produces NaNs

d)

Produces an error

94.

s = pd.Series([10,20,30], index=pd.date_range('2023-01-01', periods=3)) s.ewm(span=2).mean() What will be the output?

a)

2023-01-01 10.000000 2023-01-02 16.666667 2023-01-03 26.666667 dtype: float64

b)

2023-01-01 10 2023-01-02 20 2023-01-03 30 dtype: int64

c)

2023-01-01 10 2023-01-02 15 2023-01-03 25 dtype: int64

d)

Produces an error

95.

Which of the following is a valid input for constructing a pd.DatetimeIndex?

a)

A list of ISO-8601 date strings

b)

A dictionary of date strings

c)

An integer index

d)

A set of strings

96.

How do you convert a column to datetime in pandas?

a)

Converts column to string

b)

Converts to numeric

c)

Produces an error

97.

How do you set a datetime column as index?

a)

Sets 'date' column as index

b)

Drops the column

c)

Converts index to numeric

d)

Produces an error

98.

Output of this code? df = pd.DataFrame({'A':[1,2,3],'B':[4,5,6]}) df.applymap(lambda x: x**2)

a)

A B 0 1 16 1 4 25 2 9 36

b)

A B 0 1 4 1 2 5 2 3 6

c)

Produces error

99.

Output of this code? df = pd.DataFrame({'A':[1,2,3]}) df.transform(lambda x: x*3)

a)

A 0 3 1 6 2 9

b)

A 0 1 1 2 2 3

c)

Produces error

d)

Returns a Series

100.

What does transform return?

a)

Same shape as input DataFrame or Series

b)

Always a single value per column

c)

Always a single scalar

d)

Produces error