Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Educational Evaluation

Total questions: 65

Worksheet time: 34mins

Name
Class
Date
1.

Which type of assessment is most suitable for evaluating a student's level of knowledge? (a)  

Choose from the below words
process checklist
product review
questionnaire
paper-and-pencil test
2.

Assessments used to monitor student progress during instruction are called

a)

placement assessments

b)

criterion assessments

c)

formative assessments

d)

summative assessments

3.

Summative assessments are concerned with

a)

certification of mastery

b)

diagnosis of errors

c)

entry learning skills

d)

progress during learning

4.

Formative testing is used primarily for

a)

grading students

b)

monitoring student progress

c)

forming student groups

d)

placing students in groups

5.

The extent to which achievement tests contribute to improved learning and instruction is determined largely by the principles underlying their

a)

development and evaluation

b)

development and use

c)

research and development

d)

use and evaluation

6.

A summative test evaluates terminal achievement of students.

a)

TRUE

b)

FALSE

7.

A formative test is designed to cover a wide range of content in a unit.

a)

TRUE

b)

FALSE

8.

An achievement test used to measure entry performance is called a diagnostic test.

a)

TRUE

b)

FALSE

9.

Formative testing is used primarily for grading students.

a)

TRUE

b)

FALSE

10.

Which of the following is a TRUE statement?

a)

Paper-and-pencil tests are low in realism of task but HIGH in assessment time needed.

b)

Performance assessments are low in complexity of task BUT high in assessment time needed.

c)

Paper-and-pencil tests are high in realism of task AND high in judgment in scoring.

d)

Performance assessments are high in complexity of task AND high in assessment time needed.

11.

Criterion-referenced assessments provide an indication of a student's

a)

relative ranking with other students

b)

ability to perform a task at a specified standard

c)

attitude toward the training program

d)

level of knowledge in the subject matter compared with other students

12.

Which assessment method is highest in realism?

a)

extended performance

b)

supply response

c)

selected response

d)

restricted performance

13.

If you have the fourth highest score in the class, which interpretation of your score has been made?

a)

criterion-referenced

b)

norm-referenced

c)

domain-referenced

d)

performance-referenced

14.

Rating scales or scoring rubrics are used when judging performance on a task.

a)

TRUE

b)

FALSE

15.

Stating that a student completed 8 out of 10 steps in a task correctly would be considered a norm-referenced interpretation.

a)

TRUE

b)

FALSE

16.

Both a norm-referenced and criterion-referenced interpretation may be used with the same assessment.

a)

TRUE

b)

FALSE

17.

Selected-response items are highest in task complexity.

a)

TRUE

b)

FALSE

18.

Scores from standardized tests typically are assessed using a criterion-referenced approach.

a)

TRUE

b)

FALSE

19.

The two most important characteristics of a well-constructed achievement test are validity and

a)

reliability

b)

usability

c)

difficulty

d)

maintainability

20.

The revised Taxonomy of Educational Objectives is particularly useful when

a)

administering a paper-and-pencil test

b)

determining the reliability of an assessment

c)

developing scoring rubrics

d)

preparing instructional objectives

21.

The selection and use of assessment methods should be planned

a)

during the planning of instructional activities

b)

after the instructional activities have been developed

c)

after the instruction has been delivered to the students

d)

prior to the beginning of instructional activities

22.

Test items that are unrelated to the intended learning outcomes is a concern related to

a)

vailidity

b)

difficulty

c)

maintainability

d)

reliability

23.

To improve the validity of an assessment, test items should be related to intended learning outcomes.

a)

TRUE

b)

FALSE

24.

An assessment must be valid to be reliable.

a)

TRUE

b)

FALSE

25.

Paper-and-pencil tests are the best method to determine how well a student can perform a task.

a)

TRUE

b)

FALSE

26.

The cognitive process dimension of Bloom's revised taxonomy labeled "Remember" requires students to only retrieve information from long-term memory.

a)

TRUE

b)

FALSE

27.

Solving an algebraic equation using a procedure learned in class would be considered an example of Application.

a)

TRUE

b)

FALSE

28.

The purpose of a table of specifications is to

a)

ensure that the test measures a sample of the learning outcomes

b)

determine how difficult the test items should be

c)

identify learning outcomes and content areas measured by the test

d)

ensure that the test items obtain a spread of scores

29.

Test items used to measure the lowest level of the cognitive taxonomy are

a)

analysis

b)

application

c)

understand

d)

remember

30.

Identifies, names, defines, describes, lists, matches, selects, and outlines are verbs used for which level of the cognitive taxonomy?

a)

analysis

b)

application

c)

understand

d)

remember

31.

Which of the following would indicate the lowest level of learning for a student?

a)

applies a principle

b)

gives a textbook definition of a principle

c)

explains the principle in his own words

d)

states an example of the principle

32.

The weight assigned to each topic area in a table of specifications should be determined by the

a)

complexity of the topic area

b)

instructional time devoted to the topic area

c)

emphasis given in leading published tests

d)

method used to evaluate the topic area

33.

The ideal difficulty for an item in a standardized, norm-referenced multiple-choice test that has four alternatives with 15 in the high group and 15 in the low group would range between

a)

1 and 25 percent

b)

30 and 40 percent

c)

60 and 65 percent

d)

80 and 100 percent

34.

The appropriateness and meaningfulness of the inferences we make from assessment results refers to a test's

a)

reliability

b)

validity

c)

objectivity

d)

difficulty

35.

To obtain evidence of validity based on content considerations, you would examine the

a)

expectancy table

b)

size of the correlation coefficient

c)

type of criterion used

d)

table of specifications

36.

Interpreting a student's chances of success in college based on the Scholastic Aptitude Test (SAT) requires a

a)

predictive study

b)

criterion studty

c)

construct study

d)

concurrent study

37.

External influences such as disruptions during testing that may lower the scores of the students are referred to as

a)

systematic errors

b)

standard errors

c)

random errors

d)

reliability errors

38.

Ensuring an assessment has a good representative sample of relevant test items is a characteristic of

a)

validity

b)

referencing

c)

reliablity

d)

weighting

39.

Follow all of the following for constructing multiple-choice items EXCEPT

a)

design each item to measure an important learning outcome

b)

put as little wording as possible in the stem of the item

c)

state the stem of the item in positive form whenever possible

d)

ensure that each item is independent of other items in the test

40.

The lack of plausible, but incorrect, alternatives will cause the greatest difficulty when constructing

a)

listing items

b)

multiple-choice items

c)

short-answer items

d)

true-false items

41.

Multiple-choice items designed to measure complex achievement contain

a)

obscure content to increase difficulty

b)

new or novel material

c)

alternatives at the evaluate level

d)

different plausible distractors

42.

Increasing the number of alternatives in the items of a test produces the same effect as

a)

confusing the answer

b)

giving the correct clue

c)

ensuring the objectives are tested

d)

lengthening the test

43.

It is a good strategy to include answers to earlier items that help students answer subsequent items.

a)

TRUE

b)

FALSE

44.

Lack of plausible, but incorrect, alternatives should cause the test author to consider changing the test item from a multiple-choice to a true-false item.

a)

TRUE

b)

FALSE

45.

A limitation of multiple-choice items is that they can only measure simple learning outcomes.

a)

TRUE

b)

FALSE

46.

Scores from multiple-choice items are influenced more by guessing than true-false items.

a)

TRUE

b)

FALSE

47.

To increase the difficulty of multiple-choice items, it is a good strategy to state the stem in negative form.

a)

TRUE

b)

FALSE

48.

Which of the following is the best-stated true-false item?

a)

A barometer may be helpful in predicting weather

b)

A rising barometer forecasts fair weather

c)

All barometers give precise measures of air pressure

d)

The barometer is the most useful weather instrument

49.

The matching item consists of

a)

stems and distracters

b)

premises and responses

c)

stems and responses

d)

premises and distracters

50.

Asking students to judge each statement as true or false, and then to change the false statements so they are true increases the item's level of

a)

validity

b)

reliability

c)

plausibility

d)

difficulty

51.

A series of selection-type test items based on introductory materials such as a graph is an example of

a)

a restricted-response test

b)

an interpretive exercise

c)

an extended-response essay

d)

a student portfolio

52.

A statement of opinion, by itself, cannot be marked true or false.

a)

TRUE

b)

FALSE

53.

In writing true-false items, one useful rule. is to include absolute terms like "always" or "never."

a)

TRUE

b)

FALSE

54.

Scores from true-false items are more likely to be influenced by student guessing than matching items.

a)

TRUE

b)

FALSE

55.

A limitation of the interpretive exercise is that scoring is highly subjective.

a)

TRUE

b)

FALSE

56.

A true-false item with the correct answer being "false" provides evidence that the student knows the correct answer.

a)

TRUE

b)

FALSE

57.

For which of the following general outcomes is the essay item least appropriate?

a)

remembering

b)

creating

c)

application

d)

evaluation

58.

Essay questions are more appropriate than multiple-choice items when the specific outcome calls for

a)

identification of concepts

b)

identification of data

c)

supplying the answer

d)

matching terms to definitions

59.

Short-answer items are typically limited to measuring a student's ability to

a)

evaluate the merits of a product

b)

carry out a procedure

c)

reorganize elements

d)

remember information

60.

A weakness of short-answer questions is that they

a)

are too easy to answer

b)

can potentially have several answers

c)

don't require student differentiation

d)

require time-consuming computation

61.

Having two or more persons grade each essay question is the best way to check the reliability of scoring.

a)

TRUE

b)

FALSE

62.

Only grade spelling, grammar, and punctuation of supply-type answers when they relate to the intended learning outcome.

a)

TRUE

b)

FALSE

63.

According to the "Rules for Scoring Essay Answers" in your textbook, it is best to grade essay tests questions by question, rather than student by student.

a)

TRUE

b)

FALSE

64.

According to the "Rules for Scoring Essay Answers" in your textbook, it is best to limit the amount of time a student has to answer each essay.

a)

TRUE

b)

FALSE

65.

Asking a student to defend a position or point of view would best be assessed with an extended-response essay question.

a)

TRUE

b)

FALSE