Font size
WorksheetsEDU215_Item Analysis and Validation
Total questions: 45
Worksheet time: 18mins
It refers to the consistency of the scores obtained.
Validity
Reliability
Practicality
Authenticity
It is a measure of internal consistency, that is, how closely related a set of items are as a group.
Cronbach's Alpha
Stability Test
Internal Consistency
Pearson Product Moment Correlation
The researcher determines the validity by looking at the features of the instrument.
Construct Validity
Content Validity
Face Validity
Predictive Validity
It measures by subjecting the instrument to an analysis by a group of experts who are knowledgeable about the subject both in theory and practices.
Construct Validity
Content Validity
Face Validity
Predictive Validity
This refers whether the test corresponds to its theoretical construct.
Construct Validity
Content Validity
Face Validity
Divergent Validity
It is determined by administering both the new test and the standardized test to a group of respondents, then finding the correlation between the two sets of the scores.
Concurrent
Predictive
Consistency
Convergent
It refers to how well the test predicts some future behavior of the examinees.
Consistency
Concurrent
Predictive
Convergent
The ability of the instrument to measure what it intends to measure.
Validity
Reliability
Item Difficulty
Discrimination Index
The same test is given to a group of respondents twice.
Internal Consistency
Test-retest or Stability Test
Equivalence Test
Inter-Rater
Items sought must be correlated with each other and the test should be internally consistent.
Internal Consistency
Test-retest or Stability Test
Parallel Test
Spli-Half
Which of the following are ways of assessing reliability?
Test-retest reliability
Content Reliability
Face reliability
Concurrent reliability
Which of the following are ways of assessing validity?
Test-retest validity
Inter-observer validity
Face validity
Concurrent validity
When correlating results to assess reliability, what correlation coefficient must be obtained for the data to be considered reliable?
0.8
0.6
0.9
0.7
Which of the following statements are true?
Inter-observer reliability must involve the use of one observer.
Inter-observer reliability must involve participants taking part in the research more than once.
Test-retest reliability must involve participants taking part in the research more than once.
Test-retest reliability must involve the use of more than one observer.
When using a statistical test to assess test-retest or inter-observer reliability, which of the following criteria MUST that test meet?
It must be a test of difference
It must be a test of correlation
It must be appropriate for nominal data
It must be appropriate for a repeated measures design
Which is a way in establishing test reliability?
The test is examined if free from errors and properly administered.
Scores in a test with different versions are correlated to test if they are parallel.
The components or factors of the test contain items that are strongly uncorrelated.
Two or more measures are correlated to show the same characteristics of the examinee.
What is being established if items in the test are consistently answered by the students?
Internal Consistency
Inter-rater Reliability
Test-retest
Split-half
Which type of validity was established if the components or factors of a test are hypothesized to have a negative correlation?
Construct Validity
Predictive Validity
Content Validity
Divergent Validity
How do we determine if an item is easy or difficult?
An item is easy if majority of students are not able to provide the correct answer. The item is difficult if majority of the students are able to answer correctly.
An item is difficult if majority of the students are not able to provide the correct answer. The item is easy if majority of the students are able to answer correctly.
An item can be determined difficult if the examinees who are high in the test can answer more the items correctly than the examinees who got lows scores. If not, the item is easy.
An item can be determined easy if the examinees who are high in the test can answer more the items correctly than the examinees who got low scores. If not, the item is difficult.
Which is used when the scores of the two variables measured by a test taken at two different times by the same participants are correlated?
Pearson r correlation
Linear Regression
Significance of the Correlation
Cronbach Alpha
A school psychologist administers the same intelligence test to a student on two different occasions, three months apart. The scores are highly consistent. Which type of reliability was established?
Test-retest reliability
Split-half reliability
Inter-rater reliability
Internal Consistency
Two different clinicians independently score a patient's behavioral observation checklist for anxiety symptoms. Their scores are very similar. What is being established?
Construct Validity
Test-retest Reliability
Parallel-Form Reliability
Inter-rater Reliability
A professor creates a final exam for a course on Philippine history. The exam questions cover the entirety of the course material, including topics from every lecture and reading assignment. Which type of validity is being established?
Predictive Validity
Construct Validity
Face Validity
Content Validity
An educational psychologist develops two different versions of a math skills test, Test A and Test B, that are designed to be equivalent in content and difficulty. A group of students takes both tests, and their scores are highly correlated. What kind of reliability is being demonstrated?
Internal Consistency
Parallel-Form Reliability
Test-retest Reliability
Split-half Reliability
A company uses a pre-employment test to screen job applicants. After six months, they compare the test scores of new hires to their job performance ratings. They find a strong positive correlation. What type of validity is this a measure of?
Predictive Validity
Construct Validity
Face Validity
Concurrent Validity
A researcher is developing a new measure of 'extroversion'. She finds that the scores on her test are highly correlated with scores on a well-established and validated measure of extroversion. What type of validity is this demonstrating?
Predictive Validity
Content Validity
Divergent Validity
Convergent Validity
A test has 100 items. To quickly estimate its reliability, a researcher divides the test into two equal halves, one with odd-numbered items and the other with even-numbered items. She then correlates the scores from the two halves. Which method is being used?
Test-retest reliability
Parallel-forms reliability
Split-half reliability
Inter-rater reliability
A new self-esteem questionnaire is created. The developers want to ensure that it is not just measuring depression. They administer the new questionnaire and a well-established depression scale to the same group of people. They hope to find a very low or negative correlation between the two. What kind of validity are they trying to establish?
Predictive validity
Divergent validity
Convergent validity
Content validity
A test for 'math anxiety' is given to a group of students. The scores are then compared to the students' grades in their current math class. The scores and grades are found to be highly correlated. Which type of validity is being established?
Content validity
Concurrent validity
Predictive validity
Face validity
A test question is answered correctly by 95% of the students. What can be said about this item's difficulty index?
The item is too difficult.
The item is invalid.
The item is too easy.
The item has a moderate difficulty.
A test item has a difficulty index of P=0.50 and a discrimination index of D=0.45. What is the best conclusion about this item?
The item is too easy and effectively discriminates.
The item is too easy and does not discriminate.
The item is too difficult and does not discriminate.
The item is of ideal difficulty and effectively discriminates.
A researcher is developing a new anxiety scale for adolescents. He wants to demonstrate that the scale is related to other established measures of anxiety (like clinical diagnoses) and unrelated to measures of different constructs, like social desirability. Which type of validity is he trying to establish?
Content validity
Face validity
Predictive validity
Construct validity
A teacher wants to make sure her new math quiz has items that 'look' like they are measuring math skills, even to her students. Which type of validity is she most concerned with?
Face validity
Content validity
Construct validity
Predictive validity
An item on a test is answered correctly by most of the high-scoring students and incorrectly by most of the low-scoring students. What does this indicate about the item?
The item has a poor difficulty index.
The item has a good discrimination index.
The item has a negative discrimination index.
The item has a low discrimination index.
A test developer calculates a Cronbach's alpha coefficient for a new scale. The alpha coefficient is very high (>0.90). Which type of reliability is the developer measuring?
Inter-rater reliability
Parallel-forms consistency
Internal consistency
Test-retest reliability
A university is developing a new admissions test. They want to know if the test scores correlate with students' first-year college GPAs. Which type of validity are they most interested in?
Concurrent validity
Predictive validity
Face validity
Content validity
A test item's difficulty index is found to be 0.05. What is the most appropriate action to take with this item?
Keep the item as is, it's a good discriminator.
Move the item to the beginning of the test.
Make the item more difficult.
Revise or remove the item.
A test has a negative discrimination index. What does this mean?
The item is answered correctly by more low-scoring students than high-scoring students.
The item is of ideal difficulty.
The item is answered correctly by more high-scoring students than low-scoring students.
The item is too easy for the test takers.
A test is designed to measure 'creativity,' a concept that is not directly observable. What type of validity is most crucial to establish for this test?
Construct validity
Face validity
Content validity
Predictive validity
A teacher wants to know if a new math test she created is a good measure of student achievement. She compares the scores on her test to the students' recent scores on a standardized math test. She finds a high correlation. Which type of validity is she demonstrating?
Face validity
Construct validity
Predictive validity
Concurrent validity
A test item on a history test is correctly answered by 85% of the students. The upper-group students (top 27%) answered it correctly at a rate of 90%, while the lower-group students (bottom 27%) answered it correctly at a rate of 80%. What is the discrimination index for this item?
D=0.80
D=0.85
D=0.90
D=0.10
A researcher creates two different forms of a vocabulary test and administers them to the same group of students on the same day. The scores are not correlated. Which type of reliability is low?
Test-retest reliability
Inter-rater reliability
Internal consistency
Parallel-forms reliability
A test item has a negative discrimination index. What should be done with this item?
Move the item to the beginning of the test.
Revise or remove the item.
Keep the item as is, it's a good discriminator.
Make the item more difficult.
A teacher wants to create a parallel test to her current history quiz. What is the most important characteristic of the new test?
The new test must have the same form, but different content.
The new test must have the same questions as the original test.
The new test must have the same number of questions.
The new test must be administered at the same time as the original test.
A test item has a difficulty index of P=0.95 and a discrimination index of D=0.05. What is the most appropriate conclusion?
The item is too easy but has a good discrimination.
The item is too easy and has a low discrimination.
The item is of ideal difficulty and has a good discrimination.
The item is too difficult and has a low discrimination.
