Summative Assessment of Portfolios: an examination of different approaches to agreement over outcomes -- Brenda Johnston
Johnston looks at portfolio assessment from a number of different angles, examining how different theoretical assumptions impact the effectiveness of portfolio assessment. Johnston looks at the positivist perspective, and two post-structuralist positions -- interprevist and feminist. She concludes that the positivist approach falls short because of its belief in one true score, and its absolute focus on reliability, which tends to simplify complex tasks too much, although, while the post-structuralist approaches focus less on true score theory, and more on the development, through negotiation, of a community standard, there is little research to support it as the answer for portfolio assessment.
Positivist Approaches to Assessment --
Positivist researchers "assume it should be possible to reach one ideal, objective assessment of a portfolio through appropriate training of assessors, construction of clear guidelines and other such measures" (Johnston 397). Johnston quotes studies by Miller & Legg, and LeMahieu et. al., who find that inter-rater reliability in portfolio assessment is quite high as long as the assessment task is clear and straighforward; however, as the level of complexity rises with the assessment task, inter-rater reliability decreases accordingly.
The post-structural approach to assessment grounds itself in very different principles than the positivist approach. Post-structuralist doctrine challenges the notion of absolute truth, or, in this case, a true identifiable score. Post-structuralists "focus instead on the notion of competing discourses, conflicting scripts, and the socially contigent nature of knowledge" (Johnston 398).
One facet of the post-structuralist camp in assessment is the interpretivist approach. The interpretivist sees "Truth" as a "matter of consensus among informed and sophisticated constructors, not of correspondence with an objective reality." (Johnston 398). The interpretivist establishes reliability through negotiation of the community standard with others. In this case, assessment is not, and cannot be, acontextual, but it must take the social context into account when evaluating a portfolio. Johnston summarizes, "interpretivist approaches stress the importance of local context, connection and holistic integration instead of the distance, independent observation and aggregated scores from separate assessment which are utilized in psychometrically-based assessment" (399).
Another facet of post-structuralist assessment theory , the feminist approach, suggests that "readers brings their own cultural and gender constructs to the assessment process" (Johnston 400), and therefore, any formal system of assessment must take into account these differences in order to establish an assessment that is culture and gender-fair. The feminist approach examines the entire process with the understanding that men and women, and possibly races as well, embody different ways of knowing, and a fair assessment procedure must account for this.
Johnston goes on to examine the literature, mostly positivist, surrounding inter-rater reliability, and what issues have a significant impact. Johnston explains that how the grade is calculated plays a significant role in reliability; that is, defining "agreement" -- is a one-step difference in scoring still considered agreement on a 4-point scale, versus a 2-step difference on a 6-point scale. Obviously, how one defines agreement matters, and Johnston cites a study conducted by Applebee et. al. that reveals agreement levels between 25% and 58%, depending on the interpretation of agreement on the 4-point scale used.
Moreover, scores range widely, and reliability coeeficients as well, based on whether individual elements, or an aggragate, are assessed. A study conducted by Baume and York (2002) found that in a study of an Open University portfolio-based course the inter-rater agreement for individual elements of the portfolios was 85%, however, when assessing the portfolios as a whole, agreement fell to 61% (402). These sorts of discrepancies hound positivist researchers, however, contrarily, post-structuralists argue that "grading systems are constructed instruments leading to different assessment outcomes, rather than objective tools and that, therefore, we should probe who is favored, or otherwise, under different systems" (Johnston 402), and, additionally, examine the social and educational context for the assessment and the procedures used.
Positivist Approaches to Assessment --
Positivist researchers "assume it should be possible to reach one ideal, objective assessment of a portfolio through appropriate training of assessors, construction of clear guidelines and other such measures" (Johnston 397). Johnston quotes studies by Miller & Legg, and LeMahieu et. al., who find that inter-rater reliability in portfolio assessment is quite high as long as the assessment task is clear and straighforward; however, as the level of complexity rises with the assessment task, inter-rater reliability decreases accordingly.
The post-structural approach to assessment grounds itself in very different principles than the positivist approach. Post-structuralist doctrine challenges the notion of absolute truth, or, in this case, a true identifiable score. Post-structuralists "focus instead on the notion of competing discourses, conflicting scripts, and the socially contigent nature of knowledge" (Johnston 398).
One facet of the post-structuralist camp in assessment is the interpretivist approach. The interpretivist sees "Truth" as a "matter of consensus among informed and sophisticated constructors, not of correspondence with an objective reality." (Johnston 398). The interpretivist establishes reliability through negotiation of the community standard with others. In this case, assessment is not, and cannot be, acontextual, but it must take the social context into account when evaluating a portfolio. Johnston summarizes, "interpretivist approaches stress the importance of local context, connection and holistic integration instead of the distance, independent observation and aggregated scores from separate assessment which are utilized in psychometrically-based assessment" (399).
Another facet of post-structuralist assessment theory , the feminist approach, suggests that "readers brings their own cultural and gender constructs to the assessment process" (Johnston 400), and therefore, any formal system of assessment must take into account these differences in order to establish an assessment that is culture and gender-fair. The feminist approach examines the entire process with the understanding that men and women, and possibly races as well, embody different ways of knowing, and a fair assessment procedure must account for this.
Johnston goes on to examine the literature, mostly positivist, surrounding inter-rater reliability, and what issues have a significant impact. Johnston explains that how the grade is calculated plays a significant role in reliability; that is, defining "agreement" -- is a one-step difference in scoring still considered agreement on a 4-point scale, versus a 2-step difference on a 6-point scale. Obviously, how one defines agreement matters, and Johnston cites a study conducted by Applebee et. al. that reveals agreement levels between 25% and 58%, depending on the interpretation of agreement on the 4-point scale used.
Moreover, scores range widely, and reliability coeeficients as well, based on whether individual elements, or an aggragate, are assessed. A study conducted by Baume and York (2002) found that in a study of an Open University portfolio-based course the inter-rater agreement for individual elements of the portfolios was 85%, however, when assessing the portfolios as a whole, agreement fell to 61% (402). These sorts of discrepancies hound positivist researchers, however, contrarily, post-structuralists argue that "grading systems are constructed instruments leading to different assessment outcomes, rather than objective tools and that, therefore, we should probe who is favored, or otherwise, under different systems" (Johnston 402), and, additionally, examine the social and educational context for the assessment and the procedures used.