EdREM 6709

Thursday, March 03, 2005

Summative Assessment of Portfolios: an examination of different approaches to agreement over outcomes -- Brenda Johnston

Johnston looks at portfolio assessment from a number of different angles, examining how different theoretical assumptions impact the effectiveness of portfolio assessment. Johnston looks at the positivist perspective, and two post-structuralist positions -- interprevist and feminist. She concludes that the positivist approach falls short because of its belief in one true score, and its absolute focus on reliability, which tends to simplify complex tasks too much, although, while the post-structuralist approaches focus less on true score theory, and more on the development, through negotiation, of a community standard, there is little research to support it as the answer for portfolio assessment.

Positivist Approaches to Assessment --

Positivist researchers "assume it should be possible to reach one ideal, objective assessment of a portfolio through appropriate training of assessors, construction of clear guidelines and other such measures" (Johnston 397). Johnston quotes studies by Miller & Legg, and LeMahieu et. al., who find that inter-rater reliability in portfolio assessment is quite high as long as the assessment task is clear and straighforward; however, as the level of complexity rises with the assessment task, inter-rater reliability decreases accordingly.

The post-structural approach to assessment grounds itself in very different principles than the positivist approach. Post-structuralist doctrine challenges the notion of absolute truth, or, in this case, a true identifiable score. Post-structuralists "focus instead on the notion of competing discourses, conflicting scripts, and the socially contigent nature of knowledge" (Johnston 398).

One facet of the post-structuralist camp in assessment is the interpretivist approach. The interpretivist sees "Truth" as a "matter of consensus among informed and sophisticated constructors, not of correspondence with an objective reality." (Johnston 398). The interpretivist establishes reliability through negotiation of the community standard with others. In this case, assessment is not, and cannot be, acontextual, but it must take the social context into account when evaluating a portfolio. Johnston summarizes, "interpretivist approaches stress the importance of local context, connection and holistic integration instead of the distance, independent observation and aggregated scores from separate assessment which are utilized in psychometrically-based assessment" (399).

Another facet of post-structuralist assessment theory , the feminist approach, suggests that "readers brings their own cultural and gender constructs to the assessment process" (Johnston 400), and therefore, any formal system of assessment must take into account these differences in order to establish an assessment that is culture and gender-fair. The feminist approach examines the entire process with the understanding that men and women, and possibly races as well, embody different ways of knowing, and a fair assessment procedure must account for this.

Johnston goes on to examine the literature, mostly positivist, surrounding inter-rater reliability, and what issues have a significant impact. Johnston explains that how the grade is calculated plays a significant role in reliability; that is, defining "agreement" -- is a one-step difference in scoring still considered agreement on a 4-point scale, versus a 2-step difference on a 6-point scale. Obviously, how one defines agreement matters, and Johnston cites a study conducted by Applebee et. al. that reveals agreement levels between 25% and 58%, depending on the interpretation of agreement on the 4-point scale used.

Moreover, scores range widely, and reliability coeeficients as well, based on whether individual elements, or an aggragate, are assessed. A study conducted by Baume and York (2002) found that in a study of an Open University portfolio-based course the inter-rater agreement for individual elements of the portfolios was 85%, however, when assessing the portfolios as a whole, agreement fell to 61% (402). These sorts of discrepancies hound positivist researchers, however, contrarily, post-structuralists argue that "grading systems are constructed instruments leading to different assessment outcomes, rather than objective tools and that, therefore, we should probe who is favored, or otherwise, under different systems" (Johnston 402), and, additionally, examine the social and educational context for the assessment and the procedures used.

Wednesday, March 02, 2005

Todd Lieber -- Portfolio-Based Exit Assessment: A Progress Report

Lieber describes the development of a portfolio-based exit assessment Simpson college instituted to replace their old system of large-scale, institutional writing assessment. Prior to the development of the new program, Simpson students submitted an 8-12 page paper at the end of the 7th semester, which was certified by a faculty member as having been written for a particular class. This essay was assessed by two anonymous faculty members from any department.

The college desired a less antiquated system, and one based in current research and pedagogy.

Simpson developed a portfolio system that required students to submit a collection of four essays, drawn from at least two different disciplines, evaluated by faculty from across disciplines. Faculty were resistant at first, but in the end, the conversations and the community writing standard that developed improved faculty's teaching of writing overall, and, thus, student's learning in writing.

Lieber also argues that the entire process lays bare the process of evaluating writing, which makes for a fairer, more even-handed evaluation of student writing across the board. Lieber (1997) says, "Recognizing differences in grading and acknowledging their legitimacy while working toward a degree of commonality has been a largely positive experience" (28).

Lieber (1997) also argues that the portfolio assessment system has a high degree of face validity for students and faculty: "They feel, as we do, that the portfolio is fairer to them and that it allows them to present themselves [as writers] more accurately . . . . They feel less at the mercy of a particular reader's whims on a particular day . . . " (28).

Sunday, February 27, 2005

"Reliability, Validity, and Holistic Scoring: What We Know and What We Need to Know" -- Brian Huot

Brian Huot, in his article, “Reliability, Validity, and Holistic Scoring” (1990), examines the issues of reliability and validity within portfolio assessment. An early emphasis in writing assessment on reliability fueled the drive to develop a more consistent means of evaluating student writing. From this, arose holistic scoring. Holistic scoring employs "a rater's full impression of a text without trying to reduce her judgment to a set of recognizable skills" (Huot 201). With the development and improvement of holistic scoring procedures, composition researchers combated the most serious barrier to the widespread use of a direct measurement of writing. Huot points out that prior to the development of standardized means for the use of holistic scoring, which includes rater-training, and the use of scoring rubrics and anchor papers, interrater reliability for most direct writing assessment practices was dismal. For example, in 1961, Paul Diederich, John French, and Sydell Carlton reported that in their study of the evaluation of 300 student essays, by untrained raters, evaluated on a nine-point scale, 94% of the essays received at least seven different scores, equating to a correlation coefficient of .31, far below any acceptable standard (qtd. in Huot 1990). However, at this point, after 30 years of refining and redefining approaches to scoring in writing assessment, testing professionals are enjoying a high degree of reliability with the testing procedures that have been developed, but Huot argues that the emphasis on reliability in written assessment has overlooked an issue assumed to be acceptable -- validity.

Huot takes issue with researchers who assume blindly that simply because what raters are evaluating when they rate writing is writing, that the measure is valid. Huot quotes Charles Cooper's article, "Holistic Evaluation of Writing," in which Cooper states, "Since holistic evaluation can be as reliable as multiple choice testing and since it is always more valid . . . " (qtd in Huot 1990). Huot contends that Cooper's assertion is nothing more than hearsay, and a common hearsay among testing professionals at that, since Cooper provides no empirical evidence to back his assertion of validity in holistic writing assessment.

Huot finds a number of objections have been raised to the validity of holistic scoring as an assessment procedure:

1. holistic ratings correlate with appearance and length (Markham, Sloan and McGinnis)
2. the product orientation of holistic ratings is unsuitable for informed decisions about student writing (Faigley et. al.; Gere; Odell and Cooper)
3. Holistic ratings cannot be used beyond the population that generated them, so this method is useless an an overall indicator of writing quality
4. Holistic training procedures alter the process of scoring and reading and distort the rater's ability to make sound choices regarding writing ability (Charney; Gere; Huot)

Furthmore, Huot argues that because holistic scoring contains a very high face validity -- the test appears to test what it purports to test -- it is often accepted at face value; however, according to the American Psychological Association's publication, Standards for Educational and Psychological Tests, face validity is the least important of all measures of validity and should never be used as the sole basis for determining the validity of a test. The APA's publication lists four main types of validity that should be considered: predictive, concurrent, content, and construct.

Predictive validity assures that the rating will be of value in predicting the test-taker's future success with the item measured. So, for example, upon entering college, students are often given an English placement test, and based upon their score with this test, they are placed in the appropriate English course. A test with high predictive validity would correlate highly with the student's actual performance in the class he or she was placed in.

Concurrent validity is (look up the definition in the text). In writing assessment, this means that if a test-taker is given an essay test, his score on this essay test should correlate highly with his or her score on another measure of the same trait, so, for example, his or her score on an indirect measure of writing ability, like a syntax or reading test (research has shown a very high correlation between a person's reading scores and their writing ability).

Content validity suggests the assessment measure contains the necessary protocols to actually measure what it intends to measure. In the case of writing and holistic scoring, this means that the rubric developed and used by the raters actually allows for the true measurement of writing ability.

Finally, construct validity assumes the theoretical efficacy of the measurement procedure. According to Anne Anastasi, "the construct validity of a test is the extent to which the test may be said to measure a theoretical construct or trait" (151). This means that a student who scored well in a holistic scoring of his writing actually is a competent writer, and contrarily, if a student scores poorly, he or she lacks aptitude in writing.

Huot argues that there has not been adequate attention paid to the issue of validity in holistic scoring, and he goes on to argue that the questions of validity in writing assessment change based on what the purpose of the assessment is. So, for example, in placement testing, in which the test-user is trying to determine the relative ability of a writer in comparison to other writers in a particular college class, the scoring is really a blend of criterion and norm-referenced testing. Certainly, there are criteria that a writer must meet in order to advance to a beginning college writing course, and Huot argues that raters necessarily bring into the assessment situation preconceptions of what student needs to pass his or her course, so their decisions may be based more upon whether or not they feel the writing sample represents a student who is ready for their course, and in that way, the scoring is more norm-referenced because it is based on the reader's past experience and understanding of writer's who have been successful in their particular course. Because of this, the ratings are site-specific and non-generalizable. And, in this case, validity suffers, and, more specifically, predictive validity. In this situation, predictive validity is paramount and needs to examined closely to ensure substance and significance of the testing procedure. And, moreover, based upon what the assessment is intended for, other forms of validity may have to be examined to ensure the success and validity of the procedure as a whole.

Saturday, February 19, 2005

Kathleen Blake Yancey - "Historicizing Writing Assessment"

Yancey takes readers through the development of written assessment, dividing the chronology into three stages: 1950's - 1970, 1970-1986, 1986 - present.

In the first stage, writing assessment for placement and program review was completed through the use of multiple choice reading and editing tests, which while creating a high degree 0f reliability, were, rightly, of questionable validity. The editing test is an indirect measure of writing ability, although Paul Diederich, an ETS researcher, during the 50's, characterized the correlation between a good reading test and a good teacher's estimate of their own student's writing ability at .65. Diederich characteriZed a good objective test at .60 with such judgements, and, actually, found a single two-hour essay, graded twice by expert college board readers at .45 to .50.

In this case, the actual "essay" test fares poorly in contrast to the other supposed measure of writing ability. This is, however, probably indicative of problems with test design and scoring rather than an indictment of using WRITING to assess writing, which is, obviously, the most reliable (testing what it purports to measure) method/

With the second wave, rather than focusing on reliability, which the multiple choice tests were excellent for, the testers, administration and faculty switched their focus to establishing validity with written assessment. The result of this push was holistic scored essay exams. This push was spearheaded by Edward White in his position as the director of the California State University Freshman English Equivalency Examination program.

In 1971, the CSU chancellor adopted a plan to essentially eliminate the first-year of college by replacing it with standardized test which, if a student passed. would replace those required courses. Faculty resisted vehemently because of what such a move said about those courses they were teaching, e.g., that they were largely remedial and unnecessary for the most part, for many students. It was due to this impetus that White and others organized, and rather than defeat the measure outright, compromised (wisely) with the chancellor and agreed to take the issue of equivalency up themselves, devising their own means for establishing equivalency -- the holistically scored essay test.

In this case, reliability was established by adhereing to stringent procedures:

1. using writing prompts that directed students;
2. selecting "achor" papers and scoring guides that directed teacher-readers who rated; and
3. devising methods of calculating "acceptable agreement. (Yancey 490).

The third wave of writing assessment developed out of the second waves focus and need to establish validity: the writing portfolio. Gordon Brossell summarizes:

we know that for a valid test of writing performance, multiple writing samples written on different occassions and in various rhetorical modes are preferable to single document drawn from an isolated writing instance (qtd in Yancey 491).

This method, collecting a number of pieces written in real rhetorical situations, rated by teachers of writing as either pass / fail, rather than the holistic scale of, for example, 1-6, improved the reliability of the testing by redefining the term within this context. According to Yancey, the raters "were not trained to agree, as in the holistic scoring model, but rather released to read, to negotiate among themselves, 'hammering out an agreeable compromise.' (492).

In this way reliability is increased through the negotiation of what Peter Elbow terms "community standards" (Elbow, Portfolios, xxi). Throughout the development of writing assessment, classroom teachers had lamented the fact that what was done with assessment did not relate to what they were doing in their classes creating dissonance between testing and practice. Finally, with Elbow and Belanoff and others developing the portfolio as a model for placement and program assessment, as well as classroom assessment, the two were merged: knowledge is created about assessment, but also about, and informing practice as well.

At this point, I might introduce information from SWIC's basic writing assessment program.

Friday, February 18, 2005

Edward White - "The Opening Of the Modern Era of Writing Assessment"

White explains how the modern era of writing assessment came to fuition as a result of the California State University system's attempt to replace much of the traditional, first-year, freshman curriculum with equivalency exams, developed and administered by ETS.

Such a move is offensive because it defines that coursework as essentially remedial, and especially with writing, reduces the complex process of writing -- one that defines who we are, what we know and understand, and can help us to know and understand better -- to "a merely utilitarian view of writing as simple communication in the service of business interests" (319).

The English departments, in response to this initiative, organized in protest, and drafted a resolution that, among other things, agreed that some students came to college knowing what was taught in first-year writing already, but, it should be the English faculty assessing students and awarded college credit where necessary, and that the only way to assess writing is by having the students write, not take a multiple choice exam that asks about writing.

CSU's response was to develop a faculty-driven, large scale writing assessment program that relied on real and substantial samples of student writing, and faculty trained and accustomed to evaluating writing completing the assessment, using multiple readers rather than single readers to establish reliability and validity for the process.

College English, Volume 63.3, January 2001, (306-320).