Font size
WorksheetsPySpark Quiz Round
Total questions: 11
Worksheet time: 6mins
Which of the following is a transformation operation in PySpark?
count()
filter()
reduce()
collect()
Which of the following is true for RDD?
RDD is programming paradigm
RDD in Apache Spark is an immutable collection of objects
It is a database
None of the above
words_list = sc.parallelize ( ["pyspark", "quiz", "questions", "at", "quiz.com"] )
filtered_words = words_list.filter(lambda x: 'quiz' in x)
matched_words= filtered_words.collect()
print(matched_words)
[ "quiz", "quiz.com" ]
[ "quiz" ]
["quiz.com" ]
Error
Let us consider, we have a data frame "df". Then what does the expression '[.]{2,}' signify for the following transformation?
df = df.withColumn('var_addrss', sf.regexp_replace('var_addrss', '[.]{2,}', ''))
A single dot (".") followed by 2 integers
A single dot (".") followed by the integer '2'
Single dot (".") appearing twice consecutively
None of these
Let us consider, we have a data frame "df". Then what does the expression '^[0]*' signify for the following transformation?
df = df.withColumn('var_addrss', sf.regexp_replace('var_addrss', '^[0]*', ''))
The value starts with 0 OR followed by a sequence of 0s
The value starts with 0 and ends with 0
The value starts with 0 and followed by a sequence of 0s
The value starts with anything other than 0
Let's assume we have the following data frame "df".
How to display the 'age' column in descending order?
display(df.orderBy(df.age.desc()))
display(df.sort(df.age.desc()))
display(df.orderBy(df.age, sort = desc()))
None of these
What will the data type of the columns for the following PySpark data frame "df"?
df = spark.read.format("csv").option("header", "true").option("inferSchema", "false").option("delimeter", ",").load("/mnt/temp/test.csv")
Data types of columns will be int
Data types of columns will be read as per the data types defined in the file
Data types of all columns will be string
None of the above
Let's consider, we have this data frame "df".
How to find the sum of column "aggregation" w.r.t each partition?
df = df.withColumn('sum_total', sum('aggregation').over(Window.partitionBy('partition'))
display(df)
df.groupBy("partition").sum("aggregation").show()
display(df.withColumn('sum_total', sum('aggregation').over(Window.partitionBy('partition')))
None of the above
What is the function to convert column data type from Unix Time Seconds to Date and Timestamp?
None of these
Both the commands are correct
unix_timestamp()
from_unixtime()
Consider, we have this data frame "df" (as shown in the pic).
How to replace the 1st 3 values of column 'alchohol' as "Nan"?
import pandas as pd
import numpy as np
df.iloc[0:3, 0] = np.nan
df
import pandas as pd
import numpy as np
df.loc[0:3, 0] = np.nan
df
import pandas as pd
import numpy as np
df.iloc[0:3] = np.nan
df
import pandas as pd
import numpy as np
df.iloc[0:3] = np.nan
df.show()
Which of the following options will remove duplicates from the array column "address_struct_str"?
1) from pyspark.sql.functions import array_distinct
df = df.withColumn("address_dict", array_distinct("address_struct_str"))
2) from pyspark.sql.functions import udf
dist_addr = udf(lambda row: list(set(row)), ArrayType(StringType()))
df = df.withColumn("address_dict", dist_addr("address_struct_str"))
Option-1 is correct
Option-2 is correct
Both options- 1 & 2 are correct
None of these are correct
