Monday, 15 May 2017

Three types of task planning on Fluency, Accuracy and Complexity in L2 oral production

"The Differential Effects of Three Types of Task Planning on the Fluency, Complexity, and Accuracy in L2 Oral Production"- Paper by Rod Ellis, published in Applied Linguistics 30/4, 2009. Available at this LINK to download.

This paper by Rod Ellis is an excellent summary of research on the effects of planning time in L2 oral production till 2009. Ellis puts the research till date in perspective and outlines questions for future research. A summary with comments is given below. I am adding my comments in Italics wherever appropriate. Relevant references are copied from the original paper  and given at the end of this post for your quick reference. 

Research on the effects of planning time on L2 oral production informs the methodology of task-based teaching for improvement.

Three kinds of planning are distinguished by Ellis (2005). They are Rehearsal planning, Strategic planning and Within-task planning. 
Rehearsal and Strategic are both Pre-task planning which is done before the learner actually does the task.
Rehearsal planning is when the learner rehearses the entire task before performing the task a second time. 
Strategic planning lets the learner plan what language and content to use during the task, but doesn't let the learner rehearse the task.
Within task planning is when planning takes place during the performance of the task. There are two kinds- Pressurised and Unpressurised. Pressurised planning has a specific time limit. Therefore, learners must strive to finish the task within this constraint. This is otherwise known as pressurised online planning. Unpressurised planning provides unlimited time to complete the task. This leads to careful online planning. 

Three aspects of language production are looked at to see the the effect of planning time on them. They are, accuracy, fluency and complexity. They constitute a learner's language proficiency. These three terms are defined by Skehan (1998) and Skehan and Foster (1999). Fluency is the capacity to use language in real time, complexity is the capacity to use more advanced language and Accuracy is the ability to avoid error in performance. In practice, however, researchers develop and use operational definitions to serve the goals of their research. This is problematic when it comes to comparison of different studies. 

The framework developed by Skehan (1998) speaks of two systems possessed by language speakers. They are rule-based and memory-based systems. Rule-based system consists of the underlying rules and patterns of language. Memory-based system consists of chunks of language or formulaic sequences which are readily accessible for use. The difference between these two is in terms of processing load. The former requires large amounts of attentional resources to process, while the latter requires much less resources. Language performance involves the use of both in a variable manner depending on the requirements and demands of language processing at the time of production. This has three implications. Firstly, this implies that proficiency is a flexible phenomenon which adapt according to the conditions available during performance. Secondly, if one learner lacks in one of the systems, he/she can compensate that shortcoming using the other. Thirdly, when the situation demands greater dependence on one system, some aspects of performance might be affected. For example, when rule-based system is heavily used during a performance for the sake of structural accuracy, fluency of speech might get affected. Individual differences can also affect performance in this manner. 

1. Studies in Rehearsal Panning Time

Rehearsal planning give the learner an opportunity to perform the entire task prior to actual performance that is counted. For the same reason, rehearsal planning time is not practical in testing practice. Ellis cites three studies: Bygate (1996), Gass et.al. (1999) and Bygate (2001). Two major questions are asked in these studies.
i. Does task repetition have any effect on the performance of the same task?
All three studies produced evidence for beneficial effect of task repetition on the performance of the same task.
ii. Does task repetition have any effect on the performance of a new task?
All studies reported that that there is no transference of effects to new tasks, even when the new tasks are of the same type as the rehearsal task.
Which aspect of performance is affected by rehearsal?
Fluency and complexity are the most influenced aspects as reported by all three studies. Bygate used 10 week gaps to see if the effect is due to immediate recall, which is not. Accuracy was found to have no/little effect on performance.
Acquisition-Performance connection
Since the studies showed that there is no transference of the effect to new tasks, it implies that performance in L2 does not lead to acquisition. Or, to lead to acquisition, learners need to get feedback on the initial performance to enable 'noticing' and successive acquisition. An interesting study reported is Sheppard (2006)- an unpublished PhD thesis from the University of Auckland. This study reports the feedback intervention mentioned above.

Thus we could conclude that task repetition affects performance. But when used along with other tools like feedback, it could even lead to acquisition, and positive influences on all three aspects of performance. 

2. Studies in Strategic planning time

Ellis cites 19 studies here. He uses four parametres to talk about the results of these studies. They are: learners, settings, tasks and planning.

Learner Variables:
  1. Second and Foreign language learners 
  2. Proficiency level of learners (most studies looked at intermediate level learners). Wigglesworth (1997), Kawauchi (2005) and Tavakoli and Skehan (2005) manipulated proficiency as a variable. 
  3. Learner's orientation to planning 
  4. Individual learner difference factors
Settings Variables:
  1. Classroom
  2. Laboratory
  3. Testing
These variables are connected to length of planning time. Compared to testing context, the other two settings provided longer planning time to learners. 

Task Variables:
  1. Interactive Variables (Monologic vs. Dialogic tasks)
  2. Task Complexity (Simple vs. Complex tasks)
All the testing studies included in Ellis' list used monologic tasks in order to control the multitude of influences that come in with an interlocutor. All initial testing studies used monologues. But there are dialogic testing studies available. 

Task complexity or task difficulty is a difficult topic due to many variables involved in it. The factors affecting complexity, and used in the studies listed here are: degree of familiarity with the task context, the degree of structure in the information to be communicated, the number of distinct referents to be encoded and temporal reference (Here-and-Now vs. There-and-Then).

Planning:

  1. Length of planning time
  2. Guided vs. Unguided planning
  3. Form-focus vs. Meaning-focus
Most teaching research used ten minute planning time. Testing research uses lesser duration. Wigglesworth (1997) used 1 minute, Elder and Iwashita (2005) used 3 minutes, and Tavakoli and Skehan (2005) used 5 minutes. Mehnert (1998) studied the effect of varied planning times and found that longer planning time had more effect on performance. 

Guided planning asks learners to do particular things while planning, while unguided condition doesn't. Kawauchi (2005) studied three kinds of guided planning: writing what was planned, rehearsal of what is to say, and read a model of what to say. 

How does the above variables affect Fluency, Complexity and Accuracy?

Fluency

Operationalisation of Fluency: a. Measure of temporal aspects of fluency (number of words/syllables per minute), b. measure of repair phenomena (false starts, repetitions, reformulations). 

General finding in teaching studies is that strategic planning has a positive effect on fluency in both temporal and repair fluency aspects. But in testing, this is different. Elder and Iwashita (2005) found no effect, and Wigglesworth (2001) found negative effect. 

Second and foreign language learners experienced positive effects of planning on fluency.

Influence of proficiency level on the effect of planning on fluency was found to be varied. Wigglesworth (1997) found greater effects on the fluency of high proficiency test-takers. Kawauchi (2005) reports positive effect on low and high proficiency learners, but not on advanced learners (probably because they did not need planning time to speak well). Tavakoli and Skehan found that high proficient learners performed better than low proficient learners. Thus with the available data, no conclusion can be made on the effect of planning time for learners of different proficiency levels.

Learner's attitude towards the opportunity to plan found varied results in different studies. 

Learner's Memory played a role in fluency in some studies. Guara-Taveres (2008) found that learner's working memory and measures of fluency correlated well when planning time was available. 

Effect of Settings have clear conclusions. In laboratory and classroom settings, planning time has positive effects on performance. But in testing context, it seems to have lesser or no effect. 

With the available studies, one cannot make a conclusion regarding fluency about the participatory structure of tasks. Majority of studies involving dialogic tasks showed positive effect. Only two studies in monologic tasks did not show an effect. 

Task complexity was well studied in Foster and Skehan's studies. They used personal information, narrative and decision making tasks. The first was deemed easy because no external information was needed. The third was deemed more difficult because of the unfamiliarity of information to be communicated, and its lack of availability of structure. Mehnert (1998) found more effect on fluency in more structured tasks. Tavakoli and Skehan (2005) found more effect on fluency on structured tasks. In short, planning interacts with task complexity/difficulty, and the effect on fluency is more in case of less complex tasks. 

Length of Planning seems to have a clear effect on fluency. Mehnert (1998) found that longer the planning time, greater the effect. 

Type of Planning has an effect. Guided planning had more effect on fluency than unguided planning. But more studies are needed to find when guided planning works better than unguided planning- whether it depends on the nature of the task, proficiency of learners, etc.

Complexity

Operationalisation or measurement of complexity was done through: a. amount of subordination, number of different verb forms used, type-token ration and number of different word types. Results are more mixed than those of fluency studies, although there is plenty of evidence that strategic planning helps up complexity of production. 13 of the 19 studies reported here showed positive effects, while the rest (6) did not find any effect. Overall, strategic planning appears to have a greater effect on grammatical complexity. 

Both second and foreign language learners recorded benefit from planning on complexity. 

Proficiency: Advanced learners may not benefit from planning in terms of complexity (Kawauchi, 2005).

Working memory was found to be significantly related to complexity in planning group, not in no-planning group (Guara´-Tavares, 2008).

Complexity doesn't depend on laboratory/class setting. In testing, there is very less or no effect on complexity.

Task factors interact with planning time to have an effect on complexity, but 'how' is not clear. Most dialogic and monologic tasks showed a positive effect of planning on complexity. 

Task complexity: Foster and Skehan (1996) found that more complex decision making tasks did not have an effect on complexity while personal information and narrative tasks have higher grammatical richness with planning condition. But other studies give different results. More studies are needed to reach conclusions. 

Planning time variable: Mehnert's (1998) is the only study in planning time variable reported no effect on complexity in all of her different conditions. Different types of planning also did not have any effect on complexity. The degree of guidance might have an effect on the outcomes, however. 

Accuracy 

Accuracy also showed mixed results in various results. Thirteen of the studies found that planning enhanced accuracy but six reported no effect. 

Learner’s proficiency has an effect on accuracy. Kawauchi (2005) found less effect on advanced level learners than low proficiency learners. Learner proficiency needs to be controlled while investigating effects of planning on accuracy. 

Learners’ attitudes towards planning also have an effect on accuracy. 

Working memory wasn't found to be related to accuracy.

Task type is a potential influence factor. Interaction between task type and planning where accuracy is concerned. Foster and Skehan (1996) reported that planning led to greater accuracy in personal and decision making tasks (low and high difficulty tasks). No conclusion is really possible about how task type influences the effect that planning has on accuracy. How learners orientate to the task is also an important task. 

Type of planning failed to find effects in the studies listed here. Mochizuki, N. and L. Ortega. (2008) showed that when the guided planning is focused it can have an effect on accuracy. 

General Comments

Overall, strategic planning has clear effects on task performance. In the light of research, it is possible to say that conclusions about learner and task variables' influence on performance can be made. But we need more data and more studies that carefully control and manipulate these variables. 

Ellis questions: The three dimensions of performance we chose are not perfect descriptors of performance. Are ratings based on these these dimensions sensitive to the influence of strategic planning? This doubt is natural since no study reported an effect of planning on the ratings even when there were differences in discourse analysis results! It may thus seem that ratings are sensitive to influence of strategic planning on fluency in learning context, not in testing. 

Thus testing context seems insulated to the influence of strategic planning in terms of fluency. Only Wigglesworth (1997) showed that there is a significant effect (in discourse analysis). Why such insulation? Ellis speculates that it is because the test-taker knows that he/she is being tested, which leads to focus on accuracy at the expense of fluency and complexity. That is, testing context neutralizes the influence of planning otherwise available in learning contexts. Or may be the length of planning must be longer (althouth Mehnert (1998) proves this wrong). 

But in learning studies, strategic planning is found to have consistent influence on fluency throughout almost all studies. Prior planning of what to say saves on attentional resources enabling learner to speak fluently. Skehan proposes this trade-off between aspects of performance when learners focus on one aspect of performance. Depending on what is prioritized, the other aspects suffer. This explains the variable results of many studies. More studies are needed to say definite things about the influence of learner characteristics and task characteristics. 

Proficiency is one such variable that must be studied. Except in testing context, it has shown variations. 

Task variables like degree of structure need to be studied in detail. Existing studies have shown effects. But what is less clear is how task structure interacts with planning. 

Other considerations are: 

  • Whether planning is guided or unguided.
  • Whether focus is invited to form or meaning.

3. Studies in Within-task planning time

In fact there is only few studies that report this. Yuan and Ellis (2003) compared the performance of learners in pressurised (no planning) and unpressurised (within-task planning) planning time conditions. Learners with the within-task planning condition spoke longer. Fluency was not significantly different for both groups. Within-task planning resulted in greater syntactical complexity (without much differences in syntactical and lexical variety). It also generated more accurate speech in terms of error-free clauses and verbs. Surprisingly, there was no effect on fluency, probably due to the pressure to speak without prior planning. Thus the study concludes that when given ample time to plan, we can expect more accurate and complex speech in learning contexts. 

One problem with this kind of planning is that we do not know what the learners were doing during the planning stage. Other studies have shown that the effect of unpressurised planning is most beneficial in the initial stages of the task performance. 

Summary

Strategic planning is the most researched planning type so far. This preference is not theory-based. The other two types of planning must be examined with equal vigour. 

Rehearsal results in greater fluency and complexity. But these effects do not transfer to the performance of a new task unless there is some kind of additional intervention. That is, simple repetition of a task may not have a measurable impact on acquisition.
Strategic planning clearly benefits fluency. Results are mixed with complexity and accuracy. The possible trade-off between these two aspects must be the reason (i.e. learners will tend to prioritize either complexity or accuracy). Other variables which have an impact on the effect of strategic planning are the learners’proficiency (the effects are less evident in very advanced learners),the degree of structure of the information in the task and working memory. 
Within-task planning may benefit complexity and accuracy without having a detrimental effect on fluency.

Theoretical Perspectives

Theories of variability
Levelt's model of speaking
Skehan's theory
Robinson's theory

Levelt's Model 

There are three overlapping processes: conceptualisation, forumlation and articulation. The utterances can be monitored prior to and after the production. 
Two characteristics of speech production: a) controlled and automatic processing, b) incremental production. Conceptualiser and monitor operate under controlled processing. Formulator and articulator operate under automated conditions. However, the cases of native speakers and learners might be different in terms of formulation and articulation. 

Putting Levelt's model along with the limited attentional processes of learners explains many findings of the studies discussed above. The relationships between planning and aspects of performances can be understood with this theoretical background. 
  • Rehearsal and strategic planning are likely to assist conceptualization and thus facilitate fluency
  • They may have effects on formulation and articulation as well since the linguistic resources necessary are already accessed during conceptualisation. 
  • Fluency oriented learners may not benefit in terms of complexity and accuracy. In other words, since only limited resources are available, focus on one aspects leads to problems in other aspects of performance. 
  • Advanced learners who have lesser problems in formulation and articulation will have an effect on fluency, but much less on complexity and accuracy. 
  • In unpressurised within-task planning, complexity and accuracy may increase due to benefits to formulation. 
These accounts are helpful in understanding how things work. But we also need to know how individual differences interact with different aspects of performance and variables we study. All learners do not engage in all the processes to the same extent. Orientation, working memory, language aptitude, willingness to communicate, and anxiety are a few such individual differences. 

A framework suggested by Ellis has four sets of variables. 



TTT










The model hypothesises that task and individual variables influence how learners plan and mediate the effect planning has on production. We need to answer the question 'how does planning assist development of fluency and acquisition of linguistic knowledge'. This has yet to be studied. Fluency development and acquisition must be two different phenomena. Like Skehan says, may be fluency is an outcome of development of exemplar-based system. So fluency can develop independent of acquisition. Rehearsal and strategic planning helps learners to develop exemplar-based system. Therefore, such planning helps build fluency. Within-task planning might influence automatization of grammatical knowledge. 

There are three senses of acquisition: a) acquisition of new linguistic features, b) restructuring of existing linguistic resources and c) development of greater control or accuracy over existing linguistic features. This understanding is necessary to theorize the relationship between planning and acquisition. The studies have showed that planning has very less influence on first kind of acquisition. Second and third kind of acquisition experiences effects of planning. Planning affects restructuring by its effect on complexity (Skehan, 1998). These assumptions are based on the condition that more complex production leads to acquisition. 

Limitations and Directions for Future Research

There is lack of information on what learners do during planning. 
Within-task planning and the combined effects of within-task and pre-task planning haven't been studied yet. 
Longitudinal study of the effects of planning has not been done yet. The maximum duration so far is ten weeks!
Studies listed did not collect baseline data of native speakers performing the same tasks. 
Stuidies did not give proficiency level data of the learners. 
No study has investigated the extended performance of learners on a task (not just the early stage of the task, but the entirety of the task performance). 
How individual learner factors affect performance is not studied yet.




References
Bygate, M. (1996). ‘Effects of task repetition: Appraising the developing language of learners’ in J. Willis and D. Willis (eds): Challenge and Change in Language Teaching. Heinemann.

Bygate, M. (2001). ‘Effects of task repetition on the structure and control of oral language’ in M. Bygate, P. Skehan, and M. Swain (eds): Researching Pedagogic Tasks, Second Language Learning, Teaching and Testing. Longman.

Elder, C. and N. Iwashita. (2005). ‘Planning for test performance: Does it make a difference?’ in R. Ellis (ed.): Planning and Task Performance in a Second Language. John Benjamins.

Ellis, R. (2005). ‘Planning and task-based research: theory and research’ in R. Ellis (ed.): Planning and Task-Performance in a Second Language. John Benjamins.

Foster, P. and P. Skehan. (1996). ‘The influence of planning on performance in task-based learning,’ Studies in Second Language Acquisition 18/3: 299–324.

Gass, S., A. Mackey, M. Fernandez and M. Alvarez-Torres. (1999). ‘The effects of task repetition on linguistic output,’ Language Learning 49: 549–80.

Kawauchi, C. (2005). ‘The effects of strategic planning on the oral narratives of learners with low and high intermediate proficiency’ in R. Ellis (ed.): Planning and Task-Performance in a Second Language. John Benjamins.

Mehnert, U. (1998). ‘The effects of different lengths of time for planning on second language performance,’ Studies in Second Language Acquisition 20: 52–83.

Mochizuki, N. and L. Ortega. (2008). ‘Balancing communication and grammar in beginninglevel foreign language classrooms: A study of guided planning and relativization,’ Language Teaching Research 12: 11–37.

Sheppard, C. 2006. The Effects of Instruction Directed at the Gaps Second Language Learners Noticed in their Oral Production. Unpublished PhD Thesis, University of Auckland.

Skehan, P. (1998). A Cognitive Approach to Language Learning. Oxford: Oxford University Press.

Skehan, P. and P. Foster. (1999). ‘The influence of task structure and processing conditions on narrative retellings,’ Language Learning 49/1:93–120.  

Tavakoli, P. and S. Skehan. (2005). ‘Strategic planning, task structure, and performance testing’ in R. Ellis (ed.): Planning and TaskPerformance in a Second Language. John Benjamins.

Wigglesworth, G. (1997). ‘An investigation of planning time and proficiency level on oral test discourse,’ Language Testing 14/1: 21–44.

Yuan, F. and R. Ellis. (2003). ‘The effects of pre-task and on-line planning on fluency,complexity and accuracy in L2 monologic oral production,’Applied Linguistics 24/1: 1–27.

Monday, 8 May 2017

How does Planning time affect oral test performance- a multifaceted approach

"A multifaceted approach to investigating pre-task planning effects on paired oral test performance" is a paper written by Ryo Nitta and Fumiyo Nakatsuhara in the year 2014. The paper explores the effect of planning time on oral test task performance using a multifaceted approach. Most of the earlier studies looked at this issue by considering the performance of a group of test-takers as a whole. The problem with such an approach is that the fine differences between test-takers' performances, and the differences within particular test-taker's individual and collaborative performance will be lost. Ryo and Nakatsuhara's approach makes sure that intra-test-taker differences in performances are noticed too.

This study looked at 32 foreign language learners' performance on decision making tasks under planned and unplanned conditions. The study used rating scores, discourse analysis and conversation analysis to understand co-constructed performance apart from a questionnaire to understand test-taker attitudes to planning time. Conversation analysis provided valuable insights and implications for teaching and testing.

In teaching research, planning time is seen beneficial because it cognitively helps the limited attentional capacity/resources of the learner. Planning time activates rule based system. Therefore while performing the task, the learner can use the preactivated rule system, and concentrate on the use of memory based system. It also encourages learners to access explicit analytic knowledge since automatised implicit knowledge is lesser available during live performance. Nature of planning, task type, proficiency level of learners, etc. affects performance. Generally, longer planning time is found to be beneficial to fluency, but lesser useful for accuracy and complexity in teaching research.

Testing Research
In standardised testing, planning time is provided mainly for the sake of fairness. That is, to control the level of cognitive demand imposed by potentially unfamiliar topics and enabling test takers to produce their best performance.

The use of unguided planning for shorter periods in tests has found mixed results in earlier research. The limited effects observed might be due to the high stakes nature of testing context, which focuses test-takers' attention on accuracy. This results in careful online planning. Thus the possible effects of planning might be overridden, says Ellis in his edited book "Planning And Task Performance In A Second Language (2005). Measurement methods used in such research might have influenced results.

Dialogic Tasks
The nature of task used is very important. There are monologic and dialogic tasks. They differ vastly. Monologic tasks do not have an interlocutor. Talking to a microphone in response to a voice prompt or written question is very different from interacting with a live interlocutor in person. They both use very different performance processes. In a monologue, test-takers use their own resources. They solve problems on their own. They construct whole/entire performance. In a dialogic task, the discourse is co-constructed. Language, ideas, vocabulary, constructions, etc. are exchanged. The process is constantly open and is dependent on both (or all) parties involved in the dialogue. 

Galaczi in his paper titled Peer–peer interaction in a speaking test: The case of the First Certificate in
English examination" published in Language Assessment Quarterly, 2008 says that there are three patterns of interaction in pair discussions. They are, collaborative, parallel and asymmetric interaction. Collaborative pattern has exchange of roles of listener and speaker. They support each other's topic and develop each other ideas. Parallel pattern involves each one developing their own argument. There is less agreement on each others' ideas. Asymmetric pattern involves unbalanced contributions. One person takes secondary role while the other speaks the most. It is observed that in tests, it was collaborative pattern that received the highest scores, and parallel pattern, the least. 

Paired oral tasks are designed to measure interactional competence. But usually in research, we only look at cognitive complexity of different tasks, and linguistic demands of task design without attention to the 'co-construction' aspect of interaction. It is important to know how pre-task planning affects interactive patterns of dialogue because the kind of interaction pattern in a task has important implications for validity of the test. 

Multifaceted approach
Performance of particular test-taker could vary at different times within a task. But previous studies have looked at the collective performance of both/all parties involved in the task, assuming uniformity of performance throughout the task. This is not fair to the actual way performance takes place. Interactions involve non-linear processes as this study clearly shows us. 

This study is process-oriented (not summative). It studies differences and similarities in performance processes of interactions under different planning conditions. It used two decision making tasks. Tasks were something like this: which item in the picture is important for a happy life? Choose the most important two of them. Rating scales from previous researches were modified with extra bottom points to include the current subjects' performance. Fluency, accuracy, complexity and interaction were considered as targets to be measured. 

Results
Score analysis showed that there was a slight upgradation of fluency and complexity with planning time condition. 
Discourse analysis showed that there was an improvement in breakdown fluency and longer turn length with planning time condition. Also, planning time reduced the speed of fluency (number of words per minute). 

The increase in complexity under planned conditions might be due to presentation of planned language during planning stage. Such increase in complexity was observed only in the beginning of the task. During the advanced stage of the task performance, complexity levels were low. Also, the pattern of interaction was parallel. That is, each one gave their points without building up on each others's points. The utterances were longer, but more like a monologue instead of dialogue. There was no co-construction. When what is planned during planning time is exhausted, the discourse fell into a stagnant period.

But negative planning time condition led to gradual increase in turn lengths, incorporating each other's points collaboratively. 

Planned interaction ended clumsily which contrasted with unplanned interactions. Planned interactants only expressed their individual ideas during their turns.

Implications for Teaching and Testing
Use planning time for clear purposes, as planned by the task-designer. For example, It is better not to use planning time for tasks that are aimed at developing interactional competence.
Pre-task planning before a pair task is not advisable as it might change the interactional pattern of a task.
We ought to reconsider the duration of test tasks especially in pair oral tasks. As we have seen in this study, when planning time is provided, interaction became collaborative only after a while into the interaction. Therefore, we must reflect whether the given task performance time is sufficient to reach that threshold point where performance can become collaborative. 
There is scope to wonder if provision of planning time in oral pair tasks is wise at all. But further research is necessary before we reach any such conclusion. 

The paper is available for download HERE from Sage Publications' website.

Sunday, 7 May 2017

The effects careful online planning and task repetition on accuracy, complexity, and fluency in oral production: Paper by Ahmadian and Tavakoli

"The effects of simultaneous use of careful online planning and task repetition on accuracy, complexity, and fluency in EFL learners' oral production" is a paper published by Ahmadian and Tavakoli in the year 2010 in the journal Language Teaching Research (Vol. 15, Issue 1). The following is a summary of the outcome, and possible implications for language testing research. Original paper is available at this link: Access original paper at Sage Publications.

Oral Proficiency Interview


The goal of the paper is to study the effects of careful online planning (COLP) and task repetition (TR) on oral production of EFL learners' oral production. Planning time was operationalised as COLP and was studied against pressurized online planning (POLP). TR was operationalised as with and without repetition (+TR and -TR) conditions. COLP normally requires more time to complete a task, as there is no restriction to the amount of time that can be used to complete a task.

The results of the experiment showed that COLP has been successfully operationalised. That is, learners who used COLP used more time in completing the task than those learners who used POLP.

COLP produced more error-free clauses and correct verb forms than POLP group. Therefore, the conclusion is that COLP enhances accuracy of EFL learners' oral production.

COLP learners have outperformed POLP learners n terms of descriptive and inferential measures of complexity. Therefore, the conclusion is that COLP enhances complexity of EFL learners' oral production.

POLP group (-TR) produced more syllables per minute than COLP (-TR) group. Therefore, COLP causes disfluency in EFL learners' oral production. Also, Task repetition has positive effect on fluency.

Task repetition assists complexiy of EFL learners oral production. Both measures of complexity were higher in +TR condition than -TR condition.

Simultaneous use of COLP and +TR in learning tasks showed the following effects on accuracy, complexity and fluency. COLP +TR group produced more error-free clauses and correct verb forms than all other groups. Also, complex language was produced by learners performing narrative tasks in both measures of complexity.

Two interesting findings:
1. COLP +TR has outperformed all other groups in terms of complexity measures.
2. Engaging in COLP resulted in a degree of disfluency. But COLP +TR condition learners exceeded POLP _TR condition learners' fluency in both measures of fluency.

Implications for Testing
In learning research, it is easy to operationalise careful online planning. One could give enough time to complete the task. But in testing, we cannot provide unlimited time for planning because of constraints offered by testing context. Therefore, though COLP supports eliciting best performance, we cannot practice it in testing as it is. But we can increase or decrease the amount of time provided in some testing contexts. For example, in computer delivered tests, timed provision of planning time can be adjusted or varied according to specifications previously set by the test designer. Today most standardised tests provide planning time for fairness' sake. Test designers ought to go beyond this and provide research-based, sufficient amounts of planning time, and make sure that test-takers make use of it for planning performance. This is important.

Likewise, task repetition cannot be applied in testing, since it affects validity and readily encourages 'preparation' negatively. But test designers can select tasks that have specific characteristics that can encourage particular test strategies and processes we want test-takers to use in a test. This is very much possible and advisable. Especially in classroom-based tests of oral proficiency, teacher can do this with a bit of careful planning and thought.

Thursday, 4 May 2017

Co-constructed discourse in small group interaction

Merril Swain has given an elegant and thought provoking lecture at the 22nd Annual Language Testing Research Colloquium held in Vancouver, British Columbia in March, 2000. The following is a sort of summary of that lecture. The title of the lecture is: 'Examining dialogue: another approach to content specification and to validating inferences drawn from test scores.'

Image from HERE
Interaction in small groups is very important for pedagogic and testing research since it is very practical to use such interactions in classroom and in testing. Swain sees it as an interface between testing and learning research in second language. Small groups do not have the asymmetry of power as in an interview where the interlocutor is usually a senior, more powerful person. 

The way Swain looks at discourse and text analysis is different from the conventional. She looks at the content and underlying strategic and cognitive processes that generate and are generated by the discourse in the task. This approach studies real dialogues between people, not monologues or recorded 'think-alouds'. 

Small group tasks are included even in high stakes tests. The reasons why small group interactions are included in many tests these days are:
    1. dissatisfaction with oral interviews as the sole means of assessing oral proficiency
    2. search for other tasks that can elicit aspects of oral proficiency
    3. an attempt to mirror teaching practices in testing
    4. economic reasons- group tasks are less expensive than one-to-one oral interviews

Since small group tasks are used in high stakes tests and tests in general to make important decisions, validation is necessary (but not much validation work is done in this field). What is the basis of interpreting group performance as evidence for individual language ability? In this light McNamara asks, "whose performance is it anyway?". Interviewer behaviour is found to influence candidate performance either positively or negatively. In other words, performance is never solo. It is always jointly constructed in an oral interaction. Fulcher states that candidates found group interaction less anxiety generating than one-to-one interaction. This means, affective responses lead to differences in performance. Barry found complex relationship between test-taker's and other participants' characteristics (in a study of extroversion and performance). 

The above survey shows us that individual performance in a group task cannot be safely interpreted as evidence for underlying language abilities of individuals. A situated performance interpretation is needed. If not, interpretation of tests might lead to unfair biases and induced biases.

Sociocultural Theory of Mind
Swain's research sees how output serves L2 learning. Stretching present stage of language through earner's need to communicate successfully creates linguistic form and meaning, leading to noticing gaps in linguistic system, and thus learning process. Learning may happen through use of a dictionary, grammar book, or asking a peer or a teacher or generating and testing hypotheses, etc. Thus through attempts to communicate successfully, learning happens. 

She sees output as dialogue, as a response to criticisms against seeing language as input and output, as mechanical system. Such dialogues serve both communicative and cognitive functions. This view was developed from Vygotskian theory and its interpretation.

Orignins of cognitive functioning is primarily social. Two ideas she discusses are:
1. Higher cognitive processes are mediated activities and their source is interaction.
Dialogues generate strategies. Strategies become strategic patterns of reasoning at cognitive level. In dialogues during problem solving tasks, these strategic processes become visible. Thus in mediated learning, strategies employed by learners are visible and can be studied. 
2. Knowledge is constructed through dialogue. Dialogue can be with self or others. Dialogue mediates construction of knowledge. Co-construction of linguistic knowledge happens in dialogue. In dialogue, successful communication is important. Collaborative dialogues build knowledge in the course of problem solving using language.

Implications for Testing
1. Dialogue provides validation evidence. In dialogue, cognitive and strategic processes are visible. So by studying dialogues, we understand how participants approach task's demands. This understanding of strategies and processes used can inform understanding of constructs being measured. Therefore, dialogues help validation. 
2. The process and outcome (both) of interaction are a joint achievement. Therefore, caution must be exercised in interpretation of group performance as evidence for individual linguistic abilities. 

Swain takes neo-Vygotskian Socio-cognitive perspective for her research on dialogue. These are the questions she asks: How do we know that there is learning in Co-constructed dialogue? Would one type of task be more useful to focus students' attention on form than another?

Swain finds more than input and output in the tasks she studied. There is collaborative co-construction of learning. There is negotiation of meaning, hypothesis formation and hypothesis testing. Collaboratively, participants reach a solution. Therefore it is right to ask whose performance it is.

She used dictogloss and jigsaw tasks in her research. Information gap was embedded in jigsaw task. Jigsaw task did not give a language model, but comprehending the story from given pictures was easy. Dictogloss task invoved listening to a story read aloud at normal speed. Dictogloss task gave a language model. But without comprehending the story delivered orally, candidates could not proceed with the task. So, processing demands of both tasks are different. Therefore, she expected different strategies and performance levels on both tasks. But the prediction went wrong. Learners performed equally well on both the tasks. Also, proficiency interacted with performance. High proficiency students wrote well on dictogloss and jigsaw tasks. Low proficiency students came up because they had vocabulary help in the dictogloss task. Generally, jigsaw students performed better than dictogloss students. Another point she emphasizes is that we cannot predict what the test-takers would focus in a task- however controlled the task is. Therefore, she suggests to make use of task discourse for materials for measurement. 

Summary 
Small group is of interest to both language learning and testing research. Testers usually measure performance in small groups. In a group, performance is jointly constructed, and distributed among participants. Dialogues in groups foreground cognitive and strategic processes. Testing therefore must look for fair means for scoring this shared or co-constructed dialogue as individual ability. Interlocutor in a pair is very important. 

Dialogues lead to learning. Measuring implies measuring how much learning has taken place too. Therefore examining content of dialogues may be useful. We can understand cognitive and strategic processes involved in performance. Qualitative validation methods like expert judgement, introspective and retrospective accounts of test takers and raters, interviews and test and discourse analysis of performance can inform us of the important characteristics of tasks, discourse and interlocutors. 

Saturday, 29 April 2017

Measurement of Performance in Task-Based Assessment

Introduction
Task-based assessment (TBA) uses tasks to gather information about the linguistic abilities of test-takers. It employs either performance-referenced or construct-based approach to do this. Performance-referenced TBA looks at target language use (TLU) tasks to identify what aspects of performance to be tested, so that the test is useful for its intended purposes. This is called work-sample approach. Criterion-based approach selects a theory of language learning and use to define the underlying language abilities to be measured in the test. Whatever is the approach, after the test is conducted, we need to interpret or assess the performance to get statements regarding the test-taker's abilities to predict future performance, or to take decisions. For this, we have three different ways.

1. Direct assessment of task outcomes
Direct assessment is useful for closed tasks. Closed tasks are tasks that result in a solution that is either right or wrong. Scoring is dichotomous- either right or wrong. Such criterion-referenced tests directly let us judge whether the test-taker has passed or failed in the test.
Advantages: There is very little subjectivity involved. The test result very clearly states the outcome. It is easy. It is quick.
Disadvantages: Direct tests need direct observation of task performance. That is, a rater must observe each test-taker individually. This is time consuming and almost impossible. In case of direct written tests, this problem can be easily solved by evaluating the test scripts later. Another disadvantage is the lack of clarity regarding such test's ability to measure language ability as against non-linguistic or general knowledge. This problem more real for direct performance-referenced tests where work sampling approach is used. Construct-based tests of the direct nature do not face this problem.

2. Discourse Analytic Methods
As the name indicated, this method analyses the discourse produced in the test. It counts specific linguistic features appearing in a task discourse. There are many ways one can do this. First is to look for the test-taker's linguistic competence by looking for complexity, fluency and accuracy in the discourse. Second is to look for sociolinguistic competence- like the use of different strategies for eliciting information from interlocutors. Third is to look for discourse competence- like appropriate use of connectors, topic changers, etc. Fourth is to look for strategic competence- like use of different strategies to negotiate meaning, build discourse, or to overcome breakdowns in communication.

Discourse analytic methods are generally objective. It can be considered direct measurement. Some do not consider this as direct since in the real world, we do not analyse discourse to see if someone has succeeded in communicating something.

This method is used widely in task-based research. It is a time consuming process since the discourse requires to be transcribed for analysis. Therefore, this is not used widely in task-based language teaching. But in research, this is a very useful tool. Most of the times, one uses discourse analysis along with external rating in order to compare the assessment afforded by both.

3. External Ratings
External ratings make use of an external rater/observer to assess performance. Judgment comes from this rater. But this is not direct assessment in the nature of judgment made. Here, judgment is made by the rater subjectively, while in direct assessment, judgment is objective since it comes as part of the performance itself.

The most common external rating type uses rating scales. Both performance-referenced and system-referenced tests make use of rating scales. What do rating scales do? They specify the ability/competency being measured, and provide different levels of performance as bands of performance. Usually this scale is on a spectrum ranging from 0 to 5. The highest level usually is native-like proficiency.

Another kind of scale used makes use of checklists. This checklist lists different abilities/aspects of performance. The rater observes the performance and checks off those that are present in performance.

In TBA, there are two ways of specifying competence.
1. In Behavioural terms. 2. In Linguistic terms. Choice of behavioural or linguistic specification depends on whether you want to know about the general language proficiency or specific language proficiency of the test-taker.

Behavioural specification of competence uses rating scales that provide us with levels of test-taker performance. For example: instructions are correctly given with clarity. Such scales are essentially task dependent. Every task needs its own rating scale.

Linguistic specification of competence is problematic. How do we specify competence? What aspects of linguistic competence must be measured? There are two choices a) holistic measure which looks at general language proficiency to fulfill functions required, without attention to specific linguistic features. b) analytic measure which looks for specific linguistic abilities in terms of the four language skills. Many tests uses both holistic and analytic measures. Analytic scale requires us to specify what aspect of performance are we looking at. These aspects need to be connected to an underlying theory of language. Skehan's framework suggest 'accuracy, fluency and complexity' as three dimensions. One can also make use of functional definitions of ability instead of this. The basic rule is to use a theory of communicative competence or performance to determine the competencies to be measured.

How does one determine the checklists or performance levels used in rating scales? McNamara proposes two approaches- theoretically driven and empirically driven. Theoretically driven approach uses context-driven descriptors. Criterion behaviours are specified. Performance is observed to see if criterion behaviour is satisfied. Here, descriptors are developed in terms of linguistic skill displayed. In empirically driven approach,  Rasch scaling procedure is used to statistically study the relationship between different items in a test and general abilities to be deduced to form a scale. Such scales are not very common since the process involved is complex. An alternative way is to develop scales using the insights received from discourse analysis.

Criterion level in a scale has to be defined. This is a problematic process. Most often, this decision is arbitrary or relative. For example, some universities choose IELTS score 6 as criterion for admission
criterion, while some others choose 7 or 7.5. Such decisions are not theoretically or empirically explained.

A very important aspect of rating scale is its interaction with the rater. This is another field of study altogether.

4. Self Assessment
Self assessment has many problems and advantages. Advantages are that they are easy, less time consuming, less expensive, helps learners to be self-regulated learners, develop reflective learning, etc. Disadvantages are about the reliability and validity of self assessment. How able are learners to judge language proficiency. Bachman has found that self assessment is more meaningful when questions are about actual performance or needs, rather than about linguistic abilities. According to some studies, it lacks predictive and concurrent validity (criterion-related validity). Reliability is high when it is about internal consistency. But it is not the same case with test-retest reliability.

Friday, 30 December 2016

David Nunan’s Framework for Task Based Language Teaching

Nunan’s framework for Task-Based Language Teaching (TBLT) is built upon the concept of pedagogical tasks. Pedagogical tasks are real world tasks used for learning. They range from rehearsal task to activation tasks. Rehearsal tasks give learners some practice in doing a real world task, while an activation task activates cognitive skills and strategies required for such activities. It its strongest form, TBLT is very much like Communicative Language Teaching where language acquisition is a subconscious process, and conscious grammar teaching is unnecessary. Thus strong form of TBLT advocates the replication of natural processes of language acquisition in the classroom.

Nunan advocates a sort of form focus in the early stages of second language acquisition. So he includes what he calls as enabling skills in his framework. Enabling skills facilitate the processes of authentic communication. They are of two kinds- language exercises and communicative activities. Language exercises are judged with the linguistic outcomes (In contrast, pedagogic tasks are evaluated with their completion in terms of goal.). Communicative activities take manipulative practice from language exercises, and meaningful communication from pedagogical tasks. These activities are authentic and free because of the meaningful communication involved.

This framework is chosen from many others available because of the simplicity of its organisation and its effectiveness in creating task-based lessons and syllabii. Nunan himself suggests ways of organizing syllabii and developing units. 


Tuesday, 27 December 2016

Four marking properties of a task

What are the four marking properties of a Language-learning Task?

  1. Tasks are meaning focused communicative activities
  2. Tasks have some kind of a gap (information/opinion/reasoning)
  3. Tasks require learners to use their own linguistic and non-linguistic resources to complete the task.
  4. Tasks have a communicative outcome. (Even in an input task, there is some kind of an outcome.)


Friday, 14 October 2016

Focus on form and focus on meaning: two approaches to language teaching


There are two approaches to language teaching- form-focused and meaning-focused approaches.

Form-focused approach
This approach believes in carefully training language learners in the structures of the language by providing grammatical forms/structures and their usage. In this approach, the teacher identifies the forms that should be taught in class. Usually, by the end of a particular lesson, learners are expected to ‘accurately’ produce these target language structures in language. Therefore, accuracy is the aim of form-focused approach.

Teachers exercises control over such classrooms. Presentation of the form and its practice is highly controlled. During the production stage when learners produce target language using the focus forms, teacher slowly relaxes control so that learners can freely produce language on their own. There is explicit correction involved in this process, because accuracy is very important.

The salient features of this approach are:
  1. Focus on forms selected by the teacher
  2.  Introduction of the form before communicative activity
  3. Teacher controlled classroom which is gradually relaxed
  4. Success of form-focused instruction is learners’ ability to accurately produce target forms in language

Meaning-focused approach
In this approach, the focus is on meaningful communication in target language. Stress is on fluency, not on accuracy. Therefore, target language use is highly encouraged in a meaning-focused classroom, even though it involves errors.

But this approach is not entirely without focus on form/language. During meaning-focused communicative activities, learners themselves would naturally think about certain language items to be used in order to communicate effectively. This is called a focus on language. This focus on language can also involve the teacher. For example, learners may ask the teacher for some clarification. In such occasions, teacher acts as a participant, not as a controlling agent.

Later, the teacher can draw attention to particular forms which are used in communication or came up in a text used. The learner can be lead to look at and practice a form in this context. Even here, focus is on meaning. Focus on form comes after the communicative activity. It can be incidental correction of a form to build awareness. Corrections are done by the teacher who stands as an objective outsider to communicative activity. For this we need a teacher who is confident who can comment on the language of the learner.
Teacher doesn’t control learners’ language use. If learners complete the activity, it is a successful procedure. During the procedure, there emerges a focus on language. Teacher participates and helps by clarifying doubts. Finally it ends with a focus on form. The idea is that these classes should help the learner to improve fluency and accuracy. The focus on form is for accuracy.
This approach can have these elements chronologically:
  1. Focus on meaning: this is the communicative activity. Communication is aimed at fluent use of language.
  2. Focus on language: here learners themselves reflect on their appropriate language use. This might also involve the teacher as a participant.
  3. Focus on form: specific forms are focused upon. Teacher corrects or comment upon learner’s language use.


What is important is that a focus on form comes after the communicative activity (focus on meaning) and focus on language. 

Thursday, 6 October 2016

Validity of a Test

A test is said to be valid if it measures what it is intended to measure.
Construct validity refers to a general notion of validity.
When we talk about validity, we need to have evidence to prove it. Therefore, we have different types of validity. They are described below.

1.       Content validity
A test has content validity if its content constitutes a fair representative sample of the language skills, structures, etc. with which it is meant to be concerned. It is like a proper sample of relevant structures dealt with in the teaching programme. The idea is that the test should not be based on any particular section of the syllabus. If it is so, then there would be negative backwash effect.

A document specifying skills and structures to be tested is necessary to create a good test. What is expected in a test for content validity is a wise representation of this test specification. Can compare test specifications and test content to judge on the content validity of the test. Usually this validation is done by someone not related to the construction of the test itself.

What is the significance of content validity? First, greater the content validity, the more accurate will the measurement be. Second, if there is no content validity, there would be negative backwash. Therefore, writing full-test-specification is a necessary step to ensure a test’s content validity.

2.       Criterion-related validity
Criterion-related validity is defined as the degree to which the results of the test agree with another set of results provided by some independent, highly dependable assessment of the candidate’s language ability. This independent assessment is the criterion against which the test is validated here.
There are two kinds of criterion-related validity.
a.       Concurrent validity.
Concurrent validity is established when a test and its criterion are administered at about the same time. This is done in many ways. It may not be practical to test all the items in the test specification. In such cases, short tests are conducted on large scale to save time, money and effort. In that case, in order to ensure validity, conduct a full-test on a selected sample set of students using four scorers to ensure reliable scoring. This is the criterion against which the shorter test would be validated. Then, we need to compare the two scores using a correlation coefficient. If there is a great deal of agreement, then the shorter test is said to be valid. Here, the purpose of the test determines what level of agreement is to be expected. A high stakes test should look for a greater agreement, while a low stakes test can look for a lower agreement with the criterion. In informal situations, the teachers’ judgment can also be the criterion against which concurrent validation is done.
b.       Predictive validity
Predictive validity a kind of criterion-referenced validity, which is the degree to which a test can predict a candidate’s future performance. For example, a test like GRE predicts whether a candidate is able to undertake graduate studies in a university.
Here, criterion can be an assessment of the candidates’ language ability done by his/her supervisor in a university or the result/outcome of a course. Depending on the criterion chosen, the correlation coefficient is adjusted. Then the test score is validated against the chosen criterion.



Friday, 16 September 2016

Lev Semyonovich Vygotsky’s Sociocultural Theory of Human Development in ELT

Vygotsky was a Soviet and Russian developmental and educational psychologist. He is the originator of ‘sociocultural theory of human development’. This theory influenced instruction in various domains as it helped scholars understand how human beings learn. Related to the field of Second Language Acquisition, the sociocultural theory afforded conceptualizations about how human beings learn a new language and what factors promote this learning. Vygotsky conceptualizes a social plain of interaction related to language. According to him, unique human mental functions or the higher order psychological processes happen in this plain. This is the same plain where actual human interactions take place. The cultural development of a child happens first in this intermental plain where people interact. Only in the second stage does it appear in the intramental plain of the individuals ‘within’. Language development also follows this route as cultural development for internalization.

Lev Semyonovich Vygotsky (1896-1934)
Lev Semyonovich VygotskyWhat is the mechanism by which the transition from intermental to intramental plain happens? Explanation given by Vygotsky and like-minded scholars is of ‘mediation’. By mediation, they mean transformation of impulsive, natural and non-mediated behavior into higher order mental processes by using symbolic, technological or human tools like language/computers/human beings. This happens primarily through socially meaningful activities. Mediation presupposes human participation. This is because ‘meaning’ can only be communicated when there is participation and communicability. These qualities are afforded only by human beings. In short, this can be summarized as ‘human behavior is mediated by language.’

Point to Ponder: This then leads us to question whether technology-enabled communication which sometimes claims that machines can teach language is effective at all.




Such language mediation is the reason behind the existence of polysemy. In interaction, meaning exchanges guide mediation, and lead to internalization. How does this transformation from inter to intramental space happen? It is not a linear one-step process. It is a complicated process that involves the construction of the inner plain itself. In communicational exchanges, people share their inner plains. Such shared spaces later afford the internalization of language. According to Vygotsky, when experts guide non-experts towards internalization, something called a Zone of Proximal Development (ZPD) is created. Instruction is supposed to be easier in this zone. 

Notes prepared from: Barohny, E. (2016). Bringing Vygotsky and Bakhtin into the second Language classroom: A focus on the unfinalized nature of communication. The Journal of Language Teaching and Learning, 6(1), 114-125. 

Amazon.in