"Reliability, Validity, and Holistic Scoring: What We Know and What We Need to Know" -- Brian Huot
Brian Huot, in his article, “Reliability, Validity, and Holistic Scoring” (1990), examines the issues of reliability and validity within portfolio assessment. An early emphasis in writing assessment on reliability fueled the drive to develop a more consistent means of evaluating student writing. From this, arose holistic scoring. Holistic scoring employs "a rater's full impression of a text without trying to reduce her judgment to a set of recognizable skills" (Huot 201). With the development and improvement of holistic scoring procedures, composition researchers combated the most serious barrier to the widespread use of a direct measurement of writing. Huot points out that prior to the development of standardized means for the use of holistic scoring, which includes rater-training, and the use of scoring rubrics and anchor papers, interrater reliability for most direct writing assessment practices was dismal. For example, in 1961, Paul Diederich, John French, and Sydell Carlton reported that in their study of the evaluation of 300 student essays, by untrained raters, evaluated on a nine-point scale, 94% of the essays received at least seven different scores, equating to a correlation coefficient of .31, far below any acceptable standard (qtd. in Huot 1990). However, at this point, after 30 years of refining and redefining approaches to scoring in writing assessment, testing professionals are enjoying a high degree of reliability with the testing procedures that have been developed, but Huot argues that the emphasis on reliability in written assessment has overlooked an issue assumed to be acceptable -- validity.
Huot takes issue with researchers who assume blindly that simply because what raters are evaluating when they rate writing is writing, that the measure is valid. Huot quotes Charles Cooper's article, "Holistic Evaluation of Writing," in which Cooper states, "Since holistic evaluation can be as reliable as multiple choice testing and since it is always more valid . . . " (qtd in Huot 1990). Huot contends that Cooper's assertion is nothing more than hearsay, and a common hearsay among testing professionals at that, since Cooper provides no empirical evidence to back his assertion of validity in holistic writing assessment.
Huot finds a number of objections have been raised to the validity of holistic scoring as an assessment procedure:
1. holistic ratings correlate with appearance and length (Markham, Sloan and McGinnis)
2. the product orientation of holistic ratings is unsuitable for informed decisions about student writing (Faigley et. al.; Gere; Odell and Cooper)
3. Holistic ratings cannot be used beyond the population that generated them, so this method is useless an an overall indicator of writing quality
4. Holistic training procedures alter the process of scoring and reading and distort the rater's ability to make sound choices regarding writing ability (Charney; Gere; Huot)
Furthmore, Huot argues that because holistic scoring contains a very high face validity -- the test appears to test what it purports to test -- it is often accepted at face value; however, according to the American Psychological Association's publication, Standards for Educational and Psychological Tests, face validity is the least important of all measures of validity and should never be used as the sole basis for determining the validity of a test. The APA's publication lists four main types of validity that should be considered: predictive, concurrent, content, and construct.
Predictive validity assures that the rating will be of value in predicting the test-taker's future success with the item measured. So, for example, upon entering college, students are often given an English placement test, and based upon their score with this test, they are placed in the appropriate English course. A test with high predictive validity would correlate highly with the student's actual performance in the class he or she was placed in.
Concurrent validity is (look up the definition in the text). In writing assessment, this means that if a test-taker is given an essay test, his score on this essay test should correlate highly with his or her score on another measure of the same trait, so, for example, his or her score on an indirect measure of writing ability, like a syntax or reading test (research has shown a very high correlation between a person's reading scores and their writing ability).
Content validity suggests the assessment measure contains the necessary protocols to actually measure what it intends to measure. In the case of writing and holistic scoring, this means that the rubric developed and used by the raters actually allows for the true measurement of writing ability.
Finally, construct validity assumes the theoretical efficacy of the measurement procedure. According to Anne Anastasi, "the construct validity of a test is the extent to which the test may be said to measure a theoretical construct or trait" (151). This means that a student who scored well in a holistic scoring of his writing actually is a competent writer, and contrarily, if a student scores poorly, he or she lacks aptitude in writing.
Huot argues that there has not been adequate attention paid to the issue of validity in holistic scoring, and he goes on to argue that the questions of validity in writing assessment change based on what the purpose of the assessment is. So, for example, in placement testing, in which the test-user is trying to determine the relative ability of a writer in comparison to other writers in a particular college class, the scoring is really a blend of criterion and norm-referenced testing. Certainly, there are criteria that a writer must meet in order to advance to a beginning college writing course, and Huot argues that raters necessarily bring into the assessment situation preconceptions of what student needs to pass his or her course, so their decisions may be based more upon whether or not they feel the writing sample represents a student who is ready for their course, and in that way, the scoring is more norm-referenced because it is based on the reader's past experience and understanding of writer's who have been successful in their particular course. Because of this, the ratings are site-specific and non-generalizable. And, in this case, validity suffers, and, more specifically, predictive validity. In this situation, predictive validity is paramount and needs to examined closely to ensure substance and significance of the testing procedure. And, moreover, based upon what the assessment is intended for, other forms of validity may have to be examined to ensure the success and validity of the procedure as a whole.
Huot takes issue with researchers who assume blindly that simply because what raters are evaluating when they rate writing is writing, that the measure is valid. Huot quotes Charles Cooper's article, "Holistic Evaluation of Writing," in which Cooper states, "Since holistic evaluation can be as reliable as multiple choice testing and since it is always more valid . . . " (qtd in Huot 1990). Huot contends that Cooper's assertion is nothing more than hearsay, and a common hearsay among testing professionals at that, since Cooper provides no empirical evidence to back his assertion of validity in holistic writing assessment.
Huot finds a number of objections have been raised to the validity of holistic scoring as an assessment procedure:
1. holistic ratings correlate with appearance and length (Markham, Sloan and McGinnis)
2. the product orientation of holistic ratings is unsuitable for informed decisions about student writing (Faigley et. al.; Gere; Odell and Cooper)
3. Holistic ratings cannot be used beyond the population that generated them, so this method is useless an an overall indicator of writing quality
4. Holistic training procedures alter the process of scoring and reading and distort the rater's ability to make sound choices regarding writing ability (Charney; Gere; Huot)
Furthmore, Huot argues that because holistic scoring contains a very high face validity -- the test appears to test what it purports to test -- it is often accepted at face value; however, according to the American Psychological Association's publication, Standards for Educational and Psychological Tests, face validity is the least important of all measures of validity and should never be used as the sole basis for determining the validity of a test. The APA's publication lists four main types of validity that should be considered: predictive, concurrent, content, and construct.
Predictive validity assures that the rating will be of value in predicting the test-taker's future success with the item measured. So, for example, upon entering college, students are often given an English placement test, and based upon their score with this test, they are placed in the appropriate English course. A test with high predictive validity would correlate highly with the student's actual performance in the class he or she was placed in.
Concurrent validity is (look up the definition in the text). In writing assessment, this means that if a test-taker is given an essay test, his score on this essay test should correlate highly with his or her score on another measure of the same trait, so, for example, his or her score on an indirect measure of writing ability, like a syntax or reading test (research has shown a very high correlation between a person's reading scores and their writing ability).
Content validity suggests the assessment measure contains the necessary protocols to actually measure what it intends to measure. In the case of writing and holistic scoring, this means that the rubric developed and used by the raters actually allows for the true measurement of writing ability.
Finally, construct validity assumes the theoretical efficacy of the measurement procedure. According to Anne Anastasi, "the construct validity of a test is the extent to which the test may be said to measure a theoretical construct or trait" (151). This means that a student who scored well in a holistic scoring of his writing actually is a competent writer, and contrarily, if a student scores poorly, he or she lacks aptitude in writing.
Huot argues that there has not been adequate attention paid to the issue of validity in holistic scoring, and he goes on to argue that the questions of validity in writing assessment change based on what the purpose of the assessment is. So, for example, in placement testing, in which the test-user is trying to determine the relative ability of a writer in comparison to other writers in a particular college class, the scoring is really a blend of criterion and norm-referenced testing. Certainly, there are criteria that a writer must meet in order to advance to a beginning college writing course, and Huot argues that raters necessarily bring into the assessment situation preconceptions of what student needs to pass his or her course, so their decisions may be based more upon whether or not they feel the writing sample represents a student who is ready for their course, and in that way, the scoring is more norm-referenced because it is based on the reader's past experience and understanding of writer's who have been successful in their particular course. Because of this, the ratings are site-specific and non-generalizable. And, in this case, validity suffers, and, more specifically, predictive validity. In this situation, predictive validity is paramount and needs to examined closely to ensure substance and significance of the testing procedure. And, moreover, based upon what the assessment is intended for, other forms of validity may have to be examined to ensure the success and validity of the procedure as a whole.

0 Comments:
Post a Comment
<< Home