WorksheetsLecture 1 and 2
Total questions: 100
Worksheet time: 50mins
What is the output of the following code? t = (1, 2, 3) + (4, 5) print(t)
(1, 2, 3, 4, 5)
(5, 4, 3, 2, 1)
[1, 2, 3, 4, 5]
(1, 2, 3), (4, 5)
What method is used to remove a specific value from a list?
remove()
pop()
delete()
del
The _____ function returns a new sorted list from any sequence.
sorted
sort
order
arrange
The update method is used to merge one dictionary into another.
True
False
Which method is used to check if a dictionary contains a specific key?
in
contains()
has_key()
key_check()
The pandas library provides high-level data structures for working with structured and tabular data.
True
False
Which of the following is a mutable object in Python?
String
Tuple
List
None
What will be the output of this code? x = 10; y = 3; print(x // y)
3.33
3
3.0
4
Output of this code: nums = [1,2,3] print(sum(nums))
6
5
3
Error
What does the following code print? a = [1,2,3] print(a[::-1])
A) [3,2,1]
B) [1,2,3]
C) [1,3,2]
D) Error
Output of this code: d = {"a":1, "b":2} print(d.get("c", 0))
0
None
Error
2
What does the following code return? lst = [1,2,3,4] print(lst.index(3))
A) 2
B) 3
C) 1
D) Error
The with statement is used to handle ______ safely and ensure they are closed properly.
files
loops
variables
exceptions
Which of the following is **not** a valid input for creating a DataFrame?
A single integer
A dictionary of lists
A NumPy array
A list of dictionaries
What is the output of the following code? import numpy as np a = np.array([1, 2, 3]) print(a[1])
1
2
3
Error
What does this code print? import numpy as np a = np.array([[1, 2], [3, 4]]) print(a.shape)
(2, 2)
(2,)
(4,)
Error
Output of the following code? import numpy as np a = np.array([1,2,3]) print(a.dtype)
int64
float64
object
Error
What is the result? import numpy as np a = np.zeros((2,3))
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0
Output of this code? import numpy as np a = np.array([1,2,3,4]) print(np.where(a>2))
(array([2,3]),)
(array([0,1]),)
[3 4]
Error
What does this print? import numpy as np a = np.array([1,2,3,4]) print(np.unique([1,2,2,3,3,4]))
A) [1 2 3 4]
B) [1 2 2 3 3 4]
C) [2 3 4]
D) Error
What is a key difference between a pandas `Series` and a `DataFrame`?
A) `Series` is one-dimensional, while a `DataFrame` is two-dimensional.
B) `Series` can contain multiple data types, while a `DataFrame` cannot.
C) `DataFrame` does not have indexes, while a `Series` always does.
D) Both `Series` and `DataFrame` are always empty by default.
Which method allows you to select data from a DataFrame by row and column labels?
`loc`
`iloc`
`index`
`slice`
What will the following code output? ```python import pandas as pd data = {"A": [1, 2], "B": [3, 4]} df = pd.DataFrame(data) print(df.iloc[1]) ```
A) A 2, B 4
B) A 1, B 3
C) [1, 3]
D) Error: No row with index 1
What will the following code output? ``` python import pandas as pd s = pd.Series([1, 2, 3], index=["a", "b", "c"]) print(s["b"]) ```
A) 2
B) 1
C) "b"
D) Error: Invalid index access
What does this code return? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df[df['A']>1])
A) A B 1 2 5 2 3 6
B) A B 0 1 4
C) [2,3]
D) Error
Output of this code? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df.iloc[1,1])
5
2
4
Error
What does the following code produce? import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[4,5,6]}) print(df.loc[0,'B'])
4
1
0
Error
Which library is typically used to read and parse Excel files in pandas?
openpyxl
pickle
lxml
json
Which pandas method writes a DataFrame to pickle format?
to_pickle()
write_pickle()
to_binary()
pickle_dump()
What will the following code do? import pandas as pd df = pd.read_excel('data.xlsx', sheet_name='Sheet1') print(df.head())
Reads the Excel file 'data.xlsx' from 'Sheet1' and prints the first 5 rows.
Writes the DataFrame to a new Excel file.
Displays the entire DataFrame in the console.
Imports the pandas library.
What does this code output?
Converts numbers like '1,000' to 1000
Treats ',' as delimiter
Reads all values as strings
Error
Output of this code? import pandas as pd df = pd.read_csv('data.csv', skip_blank_lines=True)
Ignores blank lines in the CSV
Reads blank lines as NaN
Error
Deletes CSV
What is the output of this code? import pandas as pd df = pd.read_csv('data.csv', low_memory=False)
A) Prevents dtype guessing and ensures proper memory usage
B) Reads CSV in chunks
C) Converts everything to string
D) Error
What does this code do? import pandas as pd df = pd.read_sql('SELECT * FROM table1', conn, index_col='ID')
Reads SQL table and sets 'ID' as index
Writes SQL table
Converts SQL table to CSV
Creates a new SQL table
What does this code do?
Reads CSV without using the first row as header
Reads CSV using the first row as header
Reads only the first row
Writes CSV without header
How do you remove rows with all NaN values?
Drops rows where all values are NaN
Drops any row with NaN
Drops column with NaN
Fills NaN
How can you strip special characters from a string column?
Removes all non-alphanumeric characters
Converts to lowercase
Converts to uppercase
Replaces spaces with underscores
Output of this code? import pandas as pd df = pd.DataFrame({'A':[1,2,3,4,5]}) Q1 = df['A'].quantile(0.25) Q3 = df['A'].quantile(0.75) IQR = Q3 - Q1 df_filtered = df[(df['A'] >= Q1 - 1.5*IQR) & (df['A'] <= Q3 + 1.5*IQR)] print(df_filtered)
Filters out outliers based on IQR
Keeps only outliers
Drops all rows
Error
How do you replace all infinite values with NaN?
Replaces +inf/-inf with NaN
Drops infinite values
Converts to zero
Raises error
What does this code do?
Converts all strings in the column to lowercase
Converts to uppercase
Strips spaces
Deletes column
How can you remove columns with more than 50% missing values?
Drops columns with more than 50% NaN
Drops rows with >50% NaN
Replaces NaN with 0
Keeps only rows with >50% non-NaN
What does this code produce? import pandas as pd df1 = pd.DataFrame({'key':[1,2,3],'A':[10,20,30]}) df2 = pd.DataFrame({'key':[3,4,5],'B':[300,400,500]}) pd.concat([df1, df2], axis=0, ignore_index=True)
Stacks df1 and df2 vertically and resets the index
Stacks horizontally
Performs a merge
Produces an error
What is the result of this code? import pandas as pd df = pd.DataFrame({'X':[1,2],'Y':[3,4]}) pd.melt(df, id_vars=['X'])
Keeps 'X' fixed and unpivots 'Y' into long format
Keeps 'Y' fixed and unpivots 'X'
Drops column 'Y'
Produces an error
What does this code do?
Outer join: keeps all keys from both DataFrames, fills missing with NaN
Left join
Inner join
Right join
Output of this code? import pandas as pd df = pd.DataFrame({'id':[1,1,2,2],'variable':['X','Y','X','Y'],'value':[10,20,30,40]}) df.pivot(index='id', columns='variable', values='value')
Reshapes from long to wide format with 'id' as index
Reshapes from wide to long
Drops column 'value'
Produces an error
Left join: keeps all rows from df1, adds matching rows from df2
Inner join
Right join
Outer join
What is the difference between merge and concat?
merge joins based on columns or keys, concat stacks DataFrames along axis
merge concatenates, concat joins
Both do the same thing
merge deletes duplicates, concat does not
How can you reorder levels in a MultiIndex?
Switches the positions of the levels
Sorts the index
Drops a level
Creates a new column
What is the key difference between reorder_levels() and swaplevel() in hierarchical indexing?
reorder_levels() allows arbitrary reordering of index levels, while swaplevel() only swaps two levels
swaplevel() can reorder multiple levels, while reorder_levels() swaps levels
reorder_levels() can be used on non-hierarchical indexes, while swaplevel() cannot
swaplevel() requires specifying all levels, while reorder_levels() defaults to the first two
In a merge() operation, what does the how='outer' parameter do?
Performs a union of the keys from both DataFrames, including all rows from both
Includes only rows with keys present in both DataFrames
Includes rows from the left DataFrame only
Includes rows from the right DataFrame only
Output of this code? import pandas as pd arrays = [['A','A','B','B'], [1,2,1,2]] index = pd.MultiIndex.from_arrays(arrays, names=('letter','num')) df = pd.DataFrame({'val':[10,20,30,40]}, index=index) df.loc['A']
Selects all rows where first level of MultiIndex is 'A'
Selects rows where second level is 'A'
Returns columns named 'A'
Produces an error
Output of this code? import pandas as pd df = pd.DataFrame({'id':[1,1,2,2],'variable':['X','Y','X','Y'],'value':[10,20,30,40]}) df.pivot(index='id', columns='variable', values='value')
Reshapes from long to wide format with 'id' as index
Reshapes from wide to long
Drops column 'value'
Produces an error
How do you rotate y-axis tick labels?
Rotates y-axis labels by 90 degrees
Rotates x-axis
Rotates plot
Produces an error
How do you change bar width in a bar plot?
Sets bar width to 0.3
Sets spacing
Changes color
Produces an error
To plot a histogram with normalized frequencies, which parameter should be set in the plotting function?
Set the 'density' parameter to True
Set the 'bins' parameter to False
Set the 'color' parameter to 'blue'
Set the 'alpha' parameter to 0.5
df['A'].plot.hist(density=True) Which of the following does this code do?
Plots histogram normalized to form a probability density
Plots raw counts
Plots line plot
Produces an error
Which parameters control the spacing between subplots in matplotlib?
wspace and hspace
width_space and height_space
width_space and height_space
padding_x and padding_y
What is the default behavior for connecting points in a matplotlib line plot?
Linear interpolation
Cubic interpolation
Step-wise connection
No connection
Which function is used to set the x-axis label in matplotlib?
set_xlabel()
set_xlim()
set_xticks()
set_title()
Which function is used to create a legend in a plot?
plt.legend()
plt.show_legend()
plt.add_legend()
plt.legend_box()
How can you customize the row and column variables in a seaborn FacetGrid?
Pass them to the row and col parameters of sns.FacetGrid()
Set them using grid.set_rows() and grid.set_cols()
Define them in the facet_vars parameter of FacetGrid()
Use the rows and columns parameters in FacetGrid()
What happens to missing values in a group key during a GroupBy operation?
They are excluded from the result
They are replaced with zeros
They are included as a separate group
An error is raised
When grouping by multiple keys, the type of the first element in the group tuple is:
A tuple containing all grouping keys
A list of grouped values
A dictionary of key-value pairs
A single grouping key
What does the GroupBy object return when iterated over?
Tuples containing the group key and group data
Only the group key
Only the group data
Lists of grouped rows
What does grouping by a dictionary in pandas achieve?
Maps specific values to group names
Creates hierarchical groups
Applies multiple aggregation functions
Splits groups by index levels
How do you apply different aggregations to different columns? df.groupby('Category').agg({'Value':'sum','Score':'max'})
Aggregates Value by sum and Score by max per group
Aggregates all by sum
Aggregates all by max
Produces an error
How do you pivot with multiple index and columns?
Creates hierarchical index and columns, aggregating sum
Creates flat index
Drops Region
Produces an error
How do you count occurrences in crosstab? pd.crosstab(df['Category'], df['SubCategory'])
Counts frequency of each Category/SubCategory combination
Sums values
Computes mean
Produces an error
What does pd.groupby return?
A tuple of key values
A single key value
A pandas DataFrame
A list of grouped rows
df.groupby('Category')['SubCategory'].nunique() What does this code do?
Returns number of unique SubCategory values per Category
Returns total count
Returns mean
Produces an error
Output of this code? df.groupby('Category')['Value'].agg(['sum','count','mean'])
Returns sum, count, and mean for each category
Returns sum only
Returns mean only
Produces an error
What does this code produce? import pandas as pd df = pd.DataFrame({'Category':['A','B','A','B'],'Value':[10,20,30,40]}) df.groupby('Category').sum()
Sums 'Value' for each category
Counts rows per category
Averages 'Value' per category
Produces an error
How do you apply multiple aggregations and rename columns? df.groupby('Category')['Value'].agg(Total='sum', Average='mean')
Returns grouped aggregation with renamed columns
Produces error in older pandas versions
Aggregates sum only
Aggregates mean only
What will pd.to_datetime(["2018-02-29"]) return?
ValueError due to an invalid date
NaT for the invalid date
A datetime object representing February 28, 2018
An empty DataFrame
When assembling datetime objects using pd.to_datetime(df), what happens if a column is missing (e.g., "hour")?
The missing column is filled with default values.
An error is raised and the operation fails.
The missing column is ignored and not included in the datetime object.
The missing column is replaced with NaN values.
What happens if a column is missing when performing an operation on a DataFrame?
The missing column is filled with default values (e.g., 0 for hours)
An error is raised due to the missing column
The operation skips rows with missing columns
The DataFrame is converted without the missing field
Which parameter in pd.to_datetime() allows you to set the timezone of the resulting datetime objects?
utc
tz
timezone
localize
What happens if the freq="infer" parameter fails to determine a consistent frequency?
A ValueError is raised
The index is created without a frequency
The freq parameter defaults to daily
A warning is issued
Which format string correctly parses "12-11-2010 00:00"?
%d-%m-%Y %H:%M
%Y-%m-%d %H:%M
%d/%m/%Y %H:%M
%m-%d-%Y %H:%M
What does the freq="M" parameter specify in pd.date_range()?
Monthly frequency
Minute-based frequency
Monday-based weekly frequency
Milliseconds frequency
Which method converts a pandas period series back to timestamps?
to_timestamp()
to_period()
to_datetime()
to_dates()
How do you forward-fill missing timestamps after resampling?
Fills NaNs using previous available value
Fills with zero
Drops missing
Produces error
Computing rolling correlation between two series involves:
Calculating the correlation coefficient over a moving window
Summing the values of both series
Multiplying the values of both series
Finding the maximum value in each window
df['value'].rolling(5).corr(df['other']) What does this code do?
Returns rolling correlation over 5-row window
Returns covariance
Returns sum
Produces error
df['value'].shift(freq=pd.DateOffset(days=3)) What does this code do?
Shifts values 3 days along the datetime index
Shifts by 3 rows
Drops rows
Produces error
pd.date_range('2020-01-01','2020-01-10', freq='B') What does this code do?
Returns dates skipping weekends
Returns all days
Returns only weekends
Produces error
df.index + pd.offsets.MonthEnd() What does this code do?
Shifts each date to the end of month
Shifts to month start
Produces error
Drops index
Shifting values using a time offset is done by:
Applying a lag or lead function to the data.
Multiplying the values by a constant factor.
Sorting the data in ascending order.
Filtering out null values from the dataset.
Backward-filling missing timestamps after upsampling is done by:
Using the bfill method to fill missing values with the next valid observation
Using the ffill method to fill missing values with the previous valid observation
Dropping all missing timestamps
Interpolating missing timestamps with linear values
df.resample('H').bfill() What does this operation do?
Fills NaNs using next valid value
Fills with zero
Drops rows
Produces an error
How do you shift values by 2 periods with shift? df['value'].shift(2) What does this operation do?
Moves values down 2 rows
Shifts index
Drops rows
Produces an error
s = pd.Series([1,2,3], index=pd.date_range('2023-01-01', periods=3)) s.rolling(2, min_periods=1).sum() What will be the output?
2023-01-01 1.0 2023-01-02 3.0 2023-01-03 5.0 dtype: float64
2023-01-01 NaN 2023-01-02 3 2023-01-03 5 dtype: float64
2023-01-01 1 2023-01-02 2 2023-01-03 3 dtype: int64
Produces an error
df = pd.DataFrame({'value':[1,2,3]}, index=pd.period_range('2023-01-01','2023-03', freq='M')) df.index.to_timestamp() What will be the output?
2023-01-01 1 2023-02-01 2 2023-03-01 3 Freq: MS, Name: value, dtype: int64
2023-01-31 1 2023-02-28 2 2023-03-31 3 dtype: int64
Produces NaNs
Produces an error
s = pd.Series([10,20,30], index=pd.date_range('2023-01-01', periods=3)) s.ewm(span=2).mean() What will be the output?
2023-01-01 10.000000 2023-01-02 16.666667 2023-01-03 26.666667 dtype: float64
2023-01-01 10 2023-01-02 20 2023-01-03 30 dtype: int64
2023-01-01 10 2023-01-02 15 2023-01-03 25 dtype: int64
Produces an error
Which of the following is a valid input for constructing a pd.DatetimeIndex?
A list of ISO-8601 date strings
A dictionary of date strings
An integer index
A set of strings
How do you convert a column to datetime in pandas?
Converts column to string
Converts to numeric
Produces an error
How do you set a datetime column as index?
Sets 'date' column as index
Drops the column
Converts index to numeric
Produces an error
Output of this code? df = pd.DataFrame({'A':[1,2,3],'B':[4,5,6]}) df.applymap(lambda x: x**2)
A B 0 1 16 1 4 25 2 9 36
A B 0 1 4 1 2 5 2 3 6
Produces error
Output of this code? df = pd.DataFrame({'A':[1,2,3]}) df.transform(lambda x: x*3)
A 0 3 1 6 2 9
A 0 1 1 2 2 3
Produces error
Returns a Series
What does transform return?
Same shape as input DataFrame or Series
Always a single value per column
Always a single scalar
Produces error
