Diberdayakan oleh Blogger.
RSS
Tampilkan postingan dengan label CALL RELATED TO LINGUISTICS. Tampilkan semua postingan
Tampilkan postingan dengan label CALL RELATED TO LINGUISTICS. Tampilkan semua postingan

Spelling in Computer-Assisted Language Learning

Rimrott & Heift 2005
96
Spelling in Computer-Assisted Language Learning - Background
Spell Checking in CALL
• The use of word processors has become an integral part of the language learning
classroom
• spell checkers have turned into a highly desirable tool for non-native writers
• however, the success of generic spell checkers in correcting misspellings by non-native
writers has not been studied extensively
Generic spell checkers
• are designed for native speakers
• yet in CALL, we are dealing with non-native writers
• assume that most misspellings are performance-based (typos)
• yet non-native writers also make competence-related errors because they do not
know the foreign language that well (e.g. for )
• correction rate for native speakers’ misspellings is above 90%
• no comparable studies for non-native speakers’ misspellings
The algorithms of generic spell checkers are based on empirical findings
• most (80 – 95%) misspellings contain only a single error of omission, addition,
substitution or transposition (e.g. , , , )
• the first letter of a misspelled word is usually correct

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS

WEB-BASED INTELLIGENT COMPUTER-ASSISTED LANGUAGE LEARNING SYSTEM FOR YORÙBÁ(YiCALL)

953
WEB-BASED INTELLIGENT COMPUTER-ASSISTED
LANGUAGE LEARNING SYSTEM FOR YORÙBÁ(YiCALL)
Odetunji A. Odejobi1 & Tony J. Beaumont2
Computer Science Department
School of Engineering and Applied Science
Aston University
Birmingham, B4 7ET
United Kingdom
ABSTRACT
In this presentation, we describe a web-based Intelligent Computer Assisted Language Learning (iCALL) system for the
learning of Yoruba Language (YiCALL). YiCALL development is based on the integration of ideas from computer aided
education; computer mediated communication as well as techniques in artificial intelligence. The system is designed for
access over the internet. Various design and implementation issues with respect to components of the system are here
discussed and the direction of ongoing work highlighted.
KEYWORDS
e-learning, speech synthesis and recognition, CALL
1. INTRODUCTION
Computer-Assisted Language Learning (CALL) provides the basic technology for assisting language learners
to acquire important communication skills in a given language. Recent advances in Computer Mediated
Communication (CMC), CALL, and World Wide Web (WWW) facilitates the integration of these
technologies in the development of powerful language education systems . We present a general framework
underlying a pioneering research work -to the best of our knowledge this is first such project on Yorùbá
language- focused on the development of a web-based intelligent CALL (iCALL) system for Yorùbá
language. It is important that a CALL system possess the ability to adapt its behaviour to the goals, tasks,
interests, and specific needs of individual users or groups of users (Brusilovsky, 2002). The central goal of
modern approached to language learning and teaching includes, communicative language teaching, goaloriented
learning and process approach to writing. Basically, language learning strategies seek to enhance
student’s autonomy and control over the learning process (Warshauer, et al, 1996). Since speech and writing
are the basic media of human communication, a CALL system that exploits them would provide a better
language learning environment. This paper provides an overview of ongoing research into the development
of an intelligent web-based iCALL for Yorùbá.
Yorùbá is one of the four major languages spoken in Africa. Other languages in this category include
Arabic, Hausa, and Swahili. In Nigeria, Yorùbá is one of the three major native languages (Hausa and Igbo)
spoken alongside English, which is the official language. In Nigeria, the homeland of Yorùbá lies between
longitudes 20 30’and 6030’ East of the Meridian and Latitudes 60 and 90 North of the Equator (CIA, 2001).
Yorùbá is the native language of people in Lagos, Oyo, Ogun, Ondo, Ekiti, and Osun states of Nigeria. It is
also spoken in some part of Edo, and Kogi states of Nigeria as well as in Central Togo, East Central part of
Republic of Benin and in Sierra Leone (where it is called Aku). There are 25 letters in the Yorùbá language
alphabet. This is made up of 18 consonants (b, d, f, g, gb, h, j, k,l , m, n, p, r, s, s, t, w, y) and seven vowels (a,
e,?, i, o, o, u). There are five nasalized vowels in the language (an, en, in, on, un). Yorùbá is a tone language
with 3 contrastive tone and 2 allotones. There are about 30 million speakers of Yorùbá language in the South
Western part of Nigeria. Students cite many reasons for studying Yoruba, including personal interest in West
IADIS International Conference e-Society 2003
954
African cultures, research interests, and fulfilment of foreign language requirements (CIA, 2001). African-
American students often study Yorùbá out of interest in their own heritage, since many of the slaves brought
to North America during the 18th and 19th centuries came from Yorùbá -speaking areas (Ajolore, 1974).
2. OVERVIEW OF YiCALL ARCHITECTURE
The basic configuration of YiCALL is as shown in Figure 1. There are three basic modules in the system
architecture namely; user interface, language resource, and intelligent learning control modules. The user
interface module comprises; (1) Automatic Speech Recognition (ASR), (2) Text -to-Speech (TTS) synthesis
and (3) Natural Language Processing (NLP) sub-modules. The language resource module comprise of the
orthography (or written) and voice knowledge base and language curriculum. The intelligent control module
control and coordinates the learning process based on some evaluation criteria that takes account of student’s
ability. Each of the technologies applied in this work have been used to develop commercial applications, but
they still have some limitations (Kohler, 2001).
Figure 1. The overview of architecture Yorùbá iCALL prototype
2.1 YiCALL User interface
The TTS sub-module implements the text -to-speech conversion task in the user interface by converting
Yoruba text , typed by learners, into synthetic speech. Spoken equivalent of system response could also be
generated while displaying corresponding texts. To provide a flexible learning environment, the speech
synthesis process is based on the concatenation of tone-syllable units from a pre-recorded and annotated
speech corpus (Lee and Vox, 2002). Information extracted in respect of the syntax and semantics of input
sentences are used to generate the intonation and rhythm of the synthesised utterance. Since semantics and
pragmatic analysis of unrestricted text is difficult, heuristic methods are being applied in determining the
accent and phrase structure which are important in determining prosodic parameter of the synthetic speech
(Portele and Barbara, 1997; Black et al, 1996).
The ASR sub-module serves as the voice interface to capture learner’s pronunciations. It also provides
voice feedback during learning. ASR technology provides the means for the tutoring system to capture the
voice of the learner. Features are extracted from captured voice signal and used to determine what was
spoken. The substantial progress achieved in automatic speech recognition in the past two decades has led to
User Interface
Yorùbá Speech
Synthesis Module
Speaker Independent
Speech Recognition
Module
Data and
Knowledgebase
Interface
Speech
Knowledge
Database
Language Learning
Curriculum Database
Intelligent learning Control
Learner
model
Natural Language
Processing Module
Learning Evaluation
WEB-BASED INTELLIGENT COMPUTER-ASSISTED LANGUAGE LEARNING SYSTEM FOR
YORÙBÁ(YiCALL)
955
a variety of successful demos and some commercial products using speech technology. In a CALL
environment, where potential users may be non-native speakers of the language, ASR systems have to deal
with variety of speaker accents. Result of research in multilingual recognition and spoken dialog systems
(Adda-Decker, 2001; Kohler, 2001) is being exploited for solving this problem.
The Natural Language Processing (NLP) module provides the formal framework for modelling aspects of
the syntactic and semantic structure of Yorùbá language. A trigram language model of Yorùbá based on the
Hidden Markov Model (HMM) is being developed. NLP techniques, such as parsing and semantic analysis,
play important role within language tutoring systems (Kupiec, 1992). Holland and Kaplan (1995) have
discussed the significant trends in the exploitation of these techniques, design issues and tradeoffs, as well as
current and potential contributions of NLP technology with respect to instructional theory and educational
practice. We intend to annex NLP tools and techniques in providing an effective language model for
YiCALL.
2.2 Language resource and intelligent control
Text and speech corpuses emanating from two local newspapers and their spoken equivalents, recorded by an
adult male native-speaker, form the basic language resource. The content area selected for the learning is the
Yorùbá greeting environment. The basic structure for greeting is ; {Situation/event}. That is , the word
before an event or situation signifies a greeting. The context and situation of various Yorùbá greeting
were compiled into the curriculum of six lessons. Each lesson has four levels and the level selected for
learning is dependent on learner’s profile. The Knowledge and Database Interface (KDI) selects learning
module and exercises, interpret learner’s input, and compiles appropriate response to guide the learner. The
KDI is design around object oriented model. It contains a structured curriculum for Yorùbá language as well
as those for learning the alphabet, phonology, morphology and phonetics of the language. The language
resource is designed in line with standard speech application language resource requirement (Holland and
Kaplan, 1995).
The activities of the KDI and the user interface are under the control of the Intelligent Learning Control
(ILC) sub-module. The ICL controls the learning process based on the learner’s model and level of
proficiency. It determines what module to present to learner and control the activity of the speech generation
and recognition process.
3. LEARNING AND LEARNER’S MODEL
The learner model stores the characteristics of the learner relevant to the system’s tutoring strategies. The
learner model defined in the system specifications comprised of data objects which describe the following
parameters; (i) personal details of the learner, (ii) the system estimates of learner’s grammatical and oral
proficiency in Yorùbá and (iii) a function describing the relative stable characteristics of the learner. The
learner’s model is updated via the parsing and analysis of contextual information which includes error
classes, potential causes of errors, the response strategies selected by the tutoring module and the level of
help that was sought by the learner.
To implement the tutoring process, the learning prototype is based on two strategies, namely;
Reinforcement learning, for drilling and proficient learning stage and Learning by analogy, for introductory
and intermediate learning stage. A five-tuple finite state automaton was used to model these learning process
as follows; Learner:= < SL, Sp, Su, SF, Ss >. Where: Sp is the present state; Su is the set of possible learning
units, SL is the set of possible learning states; SF is the final or desirable learning state, Ss is a step in the
strategy. In this context then, Sp, SFÎSL; Ss: Sp × Su¾¾® Sp and Sp=Ss(Sp, Su). Thus, using the present
learning unit and applying the next unit step in the leaning strategy to the present state will make the learning
to move to another state in the learning process, say Spi. If Spi = Sf then the learning process is complete and
the learner is expected to have achieved a predefined communication proficiency in the language.
IADIS International Conference e-Society 2003
956
4. IMPLEMENTATION
Implementation of the learning, language and user models are specified using finite-state compilers and
algorithms, and the results are stored as finite-state transducers. Creating, validating and verifying the
proposed implementation specification is an ongoing work. The information flow in the web-based
implementation of YiCALL is as shown in Figure 2. The aim is to make learners have access to YiCALL
using any WWW browser. At present we are experiment with SALT (Intel, 2002:http://www.saltforum.org),
Speech Application Language Tag, which is a mark up language for implementing speech interface. Other
optimization and customization programmes to make YiCALL easily accessible via a WWW browser would
be developed around Java Jdeveloper toolkits.
Figure 2. Information flow in web-based- implementation of YiCALL.
5. SUMMARY, CONCLUSION, AND ONGOING WORK
To facilitate a flexible and yet user friendly CALL, the system should, as much as possible, exploit available
medium of communication in natural learning process. There are two media of communication in natural
learning environment, namely; speech and writing. In a flexible and goal oriented learning environment, it
should be possible for the learner to interact with the computer using speech and writing. The focus of
current work is the computational analysis, design and implementation of YiCALL based on the integration
of ideas from speech synthesis , speech recognition and artificial intelligence. At the same time we seek to
make the system widely available via the internet. The limitations of speech recognition, speech synthesis
and natural language processing as well as the inherent problem of integrating the system with AI techniques
is generating unique challenge in the design and implementation of the proposed system.
ACKNOWLEDGEMENT
The contributions of the Commonwealth Scholarship Commission in United Kingdom and The British
Council to this research are hereby acknowledged.
REFERENCES
Adda_Decker, M.(2001) Towards Multilingual Interoperability in Automatic Speech Recognition, Speech Communication, Vol. 35.
pp.5-20.
Ajolore, O.(1974) Learning to Use Yorùbá focus sentence in a multilingual setting, Ph.D Thesis, University of Ilinouis, USA.
Black, A.W. and Taylor, P., and Caley, R.(1996). The FESTIVAL Speech Synthesis System. The URL:
http//www.cstr.ed.ac.uk/projects/festival.html.
Brusilovsky, P.(2002) From Adaptive Hypermedia to the Adaptive Web, Keynote address, Proceedings of IADIS, International
Conference WWW/Internet 2002, Lisbon, Portugal, Isaias, P., (ed.).
CIA(2001) CIA World Factbook 2001, the URL:http://www.cia.gov/cia/publucations/factbook.
Holland, V.M. and Kaplan. J.D.(1995) Natural Language Processing Techniques in Computer-Assisted Language Learning: Status and
Instructional Issues, Instructional Science, Vol. 23, Iss. 5-6, pp 351-380.
Kohler, J.(2001) Multilingual Phone Models for Vocabulary-Independent Speech Recognition Tasks, Speech Communication, Vol. 35.
pp.21-30.
Kupiec, J.(1992) Robust Part -of-Speech Tagging Using a Hidden Markov Model, Computer Speech and Language, No 6. pp.225-242.
Laniran, Y.(1992) Intonation in tone languages: the phonetic implementation of tone in Yoruba, PhD thesis, Cornell University, USA.
Lee, K. and Vox, R.V.(2002) A Segmental Speech Coder Based on Concatenative TTS, Speech Communication, Vol. 38. pp. 89-100.
Portele, T. and Barbara, H. (1997). Towards a Prominence-base Synthesis System, Speech Communication, Vol. 21. pp.61-72.
Warshauer, M., Turbee, L., Roberts, B.(1996) Computer Learning Network and Student empowerment, System s, Vol. 24 No. 1, pp 1-14.
WEB
Browser
(Java)
YiCALL
Yoruba
Learner

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS

Morphological Processing and Computer-Assisted Language Learning

Morphological Processing and
Computer-Assisted Language Learning
John Nerbonne, Duco Dokter and Petra Smit
Alfa-informatica, BCN, P.O. Box 716,
University of Groningen
NL 9700 AS Groningen
The Netherlands
Tel +31 50 363 59 74
email: nerbonne@let.rug.nl
Abstract
Contrary to most current practice and contrary to the explicit comments of some practitioners,
natural language processing (NLP) can now play a valuable role in computer-assisted language
learning (CALL). This paper reports on GLOSSER, and discusses the position of NLP within
CALL using GLOSSER as an example. GLOSSER is an intelligent assistant for Dutch students
learning to reading in French. It has been fully implemented and tested, and it offers information on
approximately 30,000 different words (or rather: lexemes), which may be taken from any text (no
special preparation is required). The assistance takes the form of:
· information on the grammatical meaning of morphology;
· entries in a bilingual dictionary;
· examples of word use taken from over one million words of text (including some bilingual text).
The application has received a warm welcome in user-studies, and has been found a useful tool by
students. It relies essentially on lemmatization, part-of-speech (POS) disambiguation, lexeme
indexing, and bilingual text alignment-–all elements of NLP technology.
Keywords: Vocabulary, Reading, On-line Dictionary Access.
The CALL Perspective
While computer-assisted language learning (CALL) is often cited as an application area for natural
language processing (NLP) (Zaenen and Nunberg, 1996), Zock (1996) and others have noted the
following discrepancy: most CALL programs in fact make little use of the language technology
developed by NLP specialists.
This point may sound paradoxical: Aren't CALL programs by definition language technology? It is
crucial to the general aims of this paper that we clarify what is understood under the term language
technology.
Language technology is the panoply of techniques that are used to carry out tasks that are specific
to language, in contrast to techniques that can be used for more general purposes. Examples of
language technology are: speech recognition, lemmatization (finding the stem or lemma of an
inflected word), parsing, text generation, speech synthesis (converting text to speech), or part-ofspeech
(POS) disambiguation (finding the syntactic category of a word)--even when this is
ambiguous as in the word left, which can be a noun (on the left), verb (She left), adjective (the left
side), or adverb (Turn left!).
Although CALL employs the computer to assist in language teaching and in language self-study,
most CALL programs make little essential use of language technology, exploiting instead
hypertext, digital audio and video, (simple) database technology and network communication.
There are several reasons for this. Many current CALL programs focus on drills and exercises,
answer keys and grammar explanations, thus, putting self-study courses into electronic form (Last
1992). This implies that existing resources (books, exercises) be reworked for computer
deployment, and the language data used for this approach can be hand-coded and “hard-wired” into
an application, obviating the need for language technology with its language processing charter.
Drills and exercises require that any processing of user input make allowance for
learners’ errors and ideally that the processing be able to recognize and diagnose these errors, which
is beyond NLP today (and it likely to remain beyond it in the foreseeable future).
The use of non-language techniques is appropriate and relatively successful, which just poses the
question more insistently: shouldn't language technology be applied to CALL? This paper proposes
a positive answer to this question, illustrating the advantages of NLP in CALL in a modest
application, which, however, relies essentially on NLP.
What can NLP do for CALL?
A further reason for the minor impact of NLP in CALL may be that it is misunderstood by many
language learning specialists. Salaberry (1996, p.12) assesses the suitability of language technology
for CALL quite negatively in one of the leading journals:
“Linguistics has not been able to encode the complexity of natural language [...] That
problem has been acknowledged by the most adamant proponents of Intelligent CALL
[ICALL (NDS)]. Holland (1995) lists the reasons that have prevented ICALL from
becoming an alternative to CALL. The most important reason for this failure is that NLP
(Natural Language Processing) programs--which underlie the development of ICALL--
cannot account for the full complexity of natural human languages.”1
Although indeed linguistics has not yet been able to encode the entire complexity of natural
language, this does not imply that NLP cannot be useful to CALL.
Firstly, we see Salaberry as guilty of a fallacy of division--assuming that what is true of the whole
must be true of the parts. So while it is true that faithful models of human linguistic behavior are
likely to remain beyond the reach of language technology for many years, perhaps decades, the same
is not true of many subdisciplines. Phonological and morphological descriptions of many languages
1 Salaberry's reference to Holland (1995) is not accompanied by bibliographic information.
are quite complete--and much more reliable than the analyses of most language teachers, so that
their accuracy cannot be the stumbling block to effective CALL. Also, although CALL in the area of
feedback for drills and exercises with free input needs to be fully consistent and error-proof, this is
not the case for all subtasks in CALL.
Secondly, even if language technology might lack the power to deal effectively with some of
traditional CALL, it seems reasonable to see where it can be useful. Here we agree with Salaberry
and others about how the issue should be decided: it is not the technology per se, but the
contribution it can make to teaching and learning that determines its usefulness for CALL.
It is difficult to argue from first pedagogical principles for the need for NLP in CALL. An important
obstacle to such an argument is that language pedagogy experts, while agreeing on the importance
of holding the attention of learners, allowing repetition, and aiming for a range of practical exercise,
still differ on many further points (Larsen-Freeman 1991). Under these circumstances it is wise for
CALL developers not to embrace any pedagogical theory too exclusively, but to provide tools or
modules which might prove useful from various perspectives (cf. Lantolf 1996).
GLOSSER and CALL
This paper supports Last's claim that relatively straightforward programs, inter alia GLOSSER, can
achieve a great deal of success within the general framework of CALL (Last, 1992). Trust is a key:
students must trust their CALL systems to be right about the information they provide. That trust is
threatened when systems make errors, and students lose confidence in the learning system. This
means that CALL systems should be reliable linguistically, stable technically, and predictable in the
specific educational support they provide. It also implies that complex systems are more likely to
cause problems in the student-teacher (program) interaction. The further one moves from the level
of individual words and phrases to the semantic level, the greater is the danger that students’ trust
will be disappointed, given the current state of the art in linguistic technology.
The focus of GLOSSER, therefore, is on words. Learners are given the task of understanding texts
and may call on GLOSSER for information on the words in the text. The educational value of this
focus is widely supported by research. The reading of texts significantly improves the learners'
vocabulary by providing lexical context, even without the use of additional sources like dictionaries
(Krantz 1990). The context provided by full text not only contributes to a better understanding of
the possible uses of a specific word, but it also creates a framework in which words are more easily
remembered (Mondria 1996).
GLOSSER’s stance is pedagogically sound--even assessed from the wide variety of pedagogies now
current in language learning. GLOSSER, the vocabulary learning assistant described in more detail
below, allows students to learn language in a communication task (namely, that of reading). This
approach enables the use of support tools, not merely exercises and drills, and thus shares some of
the motivation for Communicative CALL within the CALL paradigm (Warschauer 1996). The
choice of reading material is entirely up to the student and/or teacher, but it may include authentic
materials, which Widdowson (1990) and others have argued improves the quality of learning by
involving the learner more directly in the community in which the target language is spoken. Krantz
(1990) emphasizes the importance of learning vocabulary words in context, and GLOSSER supports
exactly and only that. Our more general point is a simple consequence of GLOSSER’s success:
NLP can improve CALL now.
The Use of GLOSSER
GLOSSER is a support tool, not a language course, and there are several ways in which support
tools may be deployed. GLOSSER may be used in individualized instruction, following the wish for
more student-centered learning, and accommodating learner differences.
So, on the one hand GLOSSER is a real CALL application, in that it facilitates language learning by
providing on-line information on individual words of French texts thus helping students improve
their comprehension of French texts and improve their vocabulary. On the other hand, however, it is
also a tool for text comprehension, in that it assists people who know some French but cannot read
it quickly or reliably due to the presence of a number of unknown words in the text (Nerbonne and
Smit 1996). Therefore, it can be used not only in educational tasks, but also as an on-line tool for
reading assistance, creating a multitude of potential applications in educational and professional use.
Finally, its easy-to-use nature suits GLOSSER for unsupervised use (with or without accompanying
instruction).
GLOSSER--Technical Realization
GLOSSER is implemented in UNIX, and facilitates the reading of French texts by Dutch students.
Four sources of information are available on words: morphological analysis, POS-disambiguation, a
dictionary and examples of word use in especially collected corpora. All sources rely heavily on
morphological analysis and indexing techniques, which are implemented in the programming
language C. Other modules, including the interface and communication are implemented in the
Tcl/Tk scripting language (Ousterhout 1994), ensuring easy rewriting, rapid prototyping and
portability. Script language code is slow, but overall speed is still good: a single lookup of all
sources of information takes approximately 2 seconds (see the section on performance for details),
mostly in morphological analysis.
Another version of GLOSSER was created for the World-Wide Web, as proof of concept and as a
to demonstrate the flexibility of the support for CALL (Dokter 1997b). This version has limited
functionality however, due to restrictions on its data.
Fig. 1 Front-end and Morphological Analysis The user normally views GLOSSER in this form (we
have, however, translated some labels from Dutch into English for this presentation). The large
window on the left contains the text being read, in this case Jules Vernes De la terre à la lune. The
user has clicked on the word égalerent, asking for information. The smaller windows on the right
show, from top to bottom, the dictionary entry for the word in a French-Dutch dictionary;
the morphological analysis, including the grammatical meaning of the inflection, namely that the
word is a third-person plural passé simple form of égaler, and finally, in the bottom window, a
further example of the word as used in another text. Note that the other example is a different
inflectional form.
The front-end of GLOSSER, displayed in Figure 1, consists mainly of four separate windows. The
main window (left) provides the general control, a browser (read-only editor) and three on/offswitches
for controlling the other three sources of information provided (these switches open the
other windows). When in use, these other windows display (for any one word) a dictionary entry,
morphological analysis with POS-disambiguation, and (possibly bilingual) examples. The window
providing examples actually consists of two windows, one for display of the example, the other for
the related translation. Finally, there is a separate help window, that provides some information on
the use and interpretation of the different parts of GLOSSER.
GLOSSER's user interface tries to be helpful. First, words in the text that are currently under the
cursor (available for look-up) are automatically highlighted. The only thing the user needs to do for
a look-up, is to click the highlighted word. Second, users can add notes to the original text (insert
translations), to avoid the need for more than one lookup.
Morphological analysis/POS-disambiguation is directly informative to the user but also crucial to
other processes. It is used to find the underlying lexemes (dictionary forms) of words, since in
general dictionaries do not provide entries for inflected forms such as crois, croyons, crurent, cru
(all forms found under croire). The part-of-speech, also provided by this analysis (verb, noun, etc.),
allows the program to choose the right dictionary entry in the case of syntactic ambiguity, which is
very common. Finally, morphological analysis is also used in providing extra examples from other
texts--this allows a lexeme-based index (instead of string-based) and increases the efficiency of the
corpus. The efficiency improves because more examples of lexemes are found when all inflected
forms are found (see below on examples).
GLOSSER was fortunate in having state-of-the-art software for morphological analysis and POSdisambiguation
from Rank Xerox Research Centre: Locolex (Bauer and Zaenen 1995). A sample
analysis is shown in Figure 1, middle right window. Locolex incorporates a stochastic POS tagger
for disambiguation. In case Locolex disambiguates incorrectly (quite infrequently), the alternatives
are listed so that the user may specify another morphological analysis, which is then used for look-up
in the dictionary and examples index.
Dictionary and Examples: GLOSSER was likewise fortunate in obtaining the Van Dale dictionary
Hedendaags Frans (Van Dale 1993). Figure 1, upper right window illustrates the front-end of the
dictionary within GLOSSER. For dictionary lookup lexemes and POS are used (as generated by the
morphological analysis). The availability of the POS of words greatly improves the accuracy of
dictionary lookup, which otherwise may suffer from grammatical ambiguity, leading to many
candidate dictionary entries, obscuring the dictionary’s value.
A very rudimentary sense of word-sense disambiguation has been built in: if one of the examples in
the dictionary matches the word context in the original text, this translation is highlighted. This is
often the case in fixed expressions such as guerre mondiale `world war’, for the lexeme mondial.
The user can select any translation from the dictionary, and insert it into the original text. Selection
is done in the same user-friendly way as word selection in the text. To provide a rich selection of
examples, a large and varied corpus was needed, including colloquial, literary, technical, political,
and other prose. Bilingual texts were particularly attractive. GLOSSER relied partly on specialized
corpus projects, such as the ECI and MULTEXT (see references for URL’s) for bilingual corpora.
A partner project developed a tool for aligning bilingual corpora (Paskaleva and Mihov 1998).
Monolingual corpora were mainly found on the World-Wide Web, for example the Gutenberg
project (see references of this paper for URL).
The current corpus size for GLOSSER is 5 MB in monolingual, 3 MB in bilingual text (that is, the
size of the French text), including 16,701 different lexemes. The texts are indexed by determining
the lemmata and POS of the individual words using the same morphological software described
above. An index (Dokter 1997a) links lemmata to full, possibly inflected forms in the original
corpora. This way, a specific instance of a lexeme can be retrieved, in its original lexical context. In
the case of GLOSSER, this is in general two or three sentences, depending on the length. Lexemebased
indexing relates inflectional variants to a single lexeme (dictionary form). A search for
examples of croire turns up the nearly 100 possible inflected variants. This improves the chance of
finding examples of a given lexeme immensely.
Examples are displayed, with a reference to the source (if available), in the Examples window, as
shown in Figure 2.
Figure 2 The window showing examples. If the example has been found in a bilingual text (shown
here), the user can pop-up the translation in the parallel text directly.
Performance
The next chapter will describe the functional performance of GLOSSER, which allows some space
for processing performance to be mentioned here. Processing times for the different modules are
based on an average calculated by processing about 100 words. The words were take from a single
text, but, because the program does not cache previously looked up words, this should not introduce
any bias. Times in the table are given in seconds.
Time Used
Morphology/POS Disambiguation 1.579
Dictionary 0.134
Examples 0.139
Total 1.852
The process of morphological analysis and part of speech disambiguation clearly consumes most of
the time needed to process one look-up, which is not surprising, since it is the most complex and
important part of the application. The other processes, although implemented by means of scripting,
take advantage of indexing techniques. A full look-up takes less 2 seconds, which users found
satisfactory.
Functionality
The intended functionality of GLOSSER was to provide robust text-independent support for Dutch
students of French. Once GLOSSER was sufficiently stable to support reading of essentially all
non-specialized texts, the demonstrator was subjected to performance analysis and user studies.2
We review this to demonstrate the maturity of available NLP technology.
2 We thank Dr. Maria Stambolieva and Dr. Aneta Dineva of the Bulgarian Academy of Science, who
collected data and began this analysis at the University of Groningen in April 1997.
Performance Analysis
In order to analyze performance, we selected 500 words in 100-word samples, taken at randomly
chosen points in five different texts. These were checked word by word for accuracy in analysis. The
texts varied in genre: official European Commission prose, (soft) pornography, poetry, and political
opinion.
There were four types of mistakes, distributed unevenly in the text.
1. Mistakes in input due to incorrect selection by testers or input errors in the text itself (misspelled
words and incorrect coding schemes);
2. Missing words in the morphology or the dictionary;
3. Incorrect linguistic analysis;
4. Irrelevant corpus examples.
None of these resulted in unexpected program responses, unrecoverable errors, or failures to
respond. We illustrate and discuss each of the error types in turn.
(1) Input. Testers reasonably tried to view cliticized elements such as the d' in d'argent as words
which might be looked up, but the program does not treat these as independent words. This decision
was motivated by convenience, but also by the consideration that the intermediate level of user we
aimed to help would have no need of assistance for these words.
We also include misspelled words in this category. Recalling that GLOSSER is intended for
students, it might reasonably be expected that spelling had been checked, but error-correcting
capabilities are still lacking in the current realization. Several errors resulted in applying the
application to ASCII text, because the program expects Latin8 encoding. Some of the errors were
invisible to the eye, e.g., one in which accented capitals had been encoded as unaccented (which is a
common typeface), e.g. Église.
(2) Missing Words. Only 12 correctly analyzed words did not appear in the dictionary or
morphological analysis. Seven of the missing words were brand names and the like, e.g., Collier's,
Vargas and Life. It is difficult to know what to make of this circumstance. The morphology will
never be so comprehensive that all such words are processed The problem clearly cannot be
regarded as a shortcoming of the dictionary, which was chosen for its limited coverage. A more
comprehensive dictionary would be less useful to intermediate-level students. Two missing words
were fréquemment `frequently’ and généreusement `generously’, which in fact are in the dictionary,
but listed under the adjectives they are derived from, fréquent and généreux, respectively.
Morphological analysis does not resolve these cases. This suggests that a second level of dictionary
indexing would be useful. The remaining three missing words were not in the dictionary.
(3) Incorrect Analyses. Of 500 words, a total of 17 incorrect analyses were encountered, rather
more than we had encountered in trials with users (where none were reported). There were no
errors of morphological analysis in the sense that the preferred analyses were in every case
morphologically possible analyses of the words. All the errors were faulty assignment of POS
categories (which could result in a preference for a possible, but incorrect analysis). These result in
incorrect dictionary look-ups 30% of the time (5 cases). Virtually none of these placed the user in
the incorrect dictionary entry--they involved using an adjective as a noun, etc., so that the dictionary
look up was the same. The rather higher number of errors in POS assignment here has to do with the
fact that very frequent words tend to be ambiguous and difficult to categorize, while users are
relatively untroubled by them: they don't look up very frequent words. This is a point where the
intended application very naturally tolerates shortcomings in the underlying technology.
(4) Faulty Corpus Treatment. There were 47 errors, or nearly 10%. These ranged from finding no
examples (most frequent) to finding irrelevant examples, most frequently in connection with
derivational morphology (which was allowed). In fact, this is a point where the performance analysis
seems too forgiving. Since the random sample of words tested included substantially more frequent
words than would a random sample of words users would select for look up, the problem is actually
greater. This naturally suggests that the GLOSSER's corpora were too small, and indeed they were.
The 4.2 MB of text contained only 16,701 different word stems. The difficulty is that, to provide
coverage of, say three occurrences of the most frequent 30,000 words, we should need a much
larger corpus (at least ten times as larger). This went beyond the project’s charter as a prototype.
To sum up, four major mistake types appear rather more frequently than one would wish. None of
these mistakes surfaced in extensive user experimentation however.
A User Study
To determine the value of GLOSSER in actual education, a user study was conducted on 22 adult
students. Dokter et al. (1998) provides a more complete report on this study, which we summarize
here. The students were all in their second or third course (each course takes three months), and
were comparable in proficiency to second- or third-year high school students. The goal of this study
was to evaluate the program in comparison to the traditional method of text reading and
comprehension using a hand-held dictionary. The group was divided randomly in two. The specific
factors that were considered relevant and could be accounted for in this study included the overall
judgment of the program, a simple measurement of the effect which the use of GLOSSER had on
text comprehension and the functionality of the program. Apart from the above factors, the subjects
using GLOSSER were asked to comment on the system, so as to get a clear picture of users’
demands of the application and suggestions for improvement.
Setup
Each session started with an introduction that explained the major purpose of the experiment, and a
short demonstration of the program. At first all the subjects were given some time to get acquainted
with GLOSSER. This was done to make the subjects more comfortable with the experimental
environment. Then the students were randomly assigned to two groups. Both groups were presented
the same text; their task was to read this text within a limited time (20 minutes) and answer
questions about this text afterwards. The text was extracted from Jules Verne's De la terre à la lune
1865, and contained approximately 250 words. The first group had this text displayed in the
GLOSSER browser, the other group used a version on paper and were provided with the same
dictionary in a hand-held format.
After the time for reading was up, the text was taken away and the subjects were given a
questionnaire. This questionnaire consisted of two parts. The first, eleven questions on the text, was
identical for both groups. The second part consisted of several questions concerning the evaluation
of the program for the group using GLOSSER, and the hand-held dictionary for the other subjects.
All questions were answered on a Likkert scale of 1-5 to facilitate statistical analysis.
Results
The results from this study can be divided into three classes, according to the specific issue
addressed:
· comprehension
· functionality of GLOSSER vs. dictionary
· subject evaluation of GLOSSER vs. dictionary
Although the group was too small for very sensitive statistical analysis, we used results for further
development of the prototype, and in a more intuitive way. Comprehension was evaluated by
posing questions on the text. Other means of measuring comprehension examined might include the
time necessary for reading the text and also vocabulary comprehension after reading. Although
results showed that GLOSSER users were faster and understood better, the differences were not
significant. The real time differences seemed to be obscured by the fact that users were given more
time than necessary, and nearly all of them used their excess time to check and recheck their work.
We expect that speed could be shown to be significantly faster with a more tightly controlled task.
GLOSSER users scored significantly higher on a question assessing confidence. They felt more
certain that they’d completed the task well, perhaps because the software made the task easier.
Functionality concerned the quality and usefulness of the sources used for GLOSSER, being a
dictionary, morphological analysis and examples, and of the application as a whole. Among other
things, we compared the number of lookups and the number of words not found in the dictionary.
The average number of lookups with GLOSSER in relation to the average number of lookups with
the dictionary was 45/14, which is a significant difference (p < 0.001). The number of lookup events
with GLOSSER was still higher, constituting a ratio of 53/14, but some words were looked up
more than once, which was uncontrollable for the dictionary lookups. This figure clearly shows that
users of GLOSSER managed to look up many more words (and read the given information) within
the same amount of time.
Forms Looked Up Words Looked Up
GLOSSER 53 45
Hand-held Dictionary 14 13
Although the hand-held and on-line dictionaries were identical in content, the number of lookups
displayed by GLOSSER could be influenced by POS-disambiguation. For completeness, we note
that the lay-out of the displayed information was nearly, but not completely identical.
A further issue concerning the functionality of the program is the specific use which the subjects
made of the informational sources:
Searches= 629
Source Used % of total
Dictionary 623 99.0
Morphological analysis 276 43.9
Examples 261 41.5
Clearly, the dictionary was taken to be the most important source for support in reading texts. An
interesting point here, which is not obvious in these numbers, is that users often consulted other
sources immediately after looking up the word in the dictionary. This indicates that when the
information a dictionary provides is regarded as insufficient for direct comprehension of a word, or a
part of the text, other sources are consulted.
GLOSSER was judged superior to the hand-held dictionary in ease of use (although this result was
not significant). All users were keen on using future versions of GLOSSER (or using this version
further). The overall judgment of the program was very positive, 4.2 on a scale of 1-5.
Additional comments of users mainly concerned the interface. One comment which resulted in
reimplementation suggested desirability of making annotations in the text. This eliminated the need
for double look-ups.
Conclusions of the User Study
The user study points to one dangerous feature of any well-working tool: overuse. It seems that
students tend to overuse a tool such as GLOSSER. Even though the dictionary interrupts the
reading process, and even though the students did not need to look up nearly 20% of the words,
they did. This suggests that students ought to be warned against this.
On the other hand, if overuse is a failing, it’s the failing of an attractive system: GLOSSER
improves the ease with which language students can approach a foreign language text. The most
important difference is simply the number of words that can be looked up and the subsequent
decrease in time needed for reading the text. Both of these may be expected to improve vocabulary
acquisition (Krantz 1990). Although the difference in text comprehension shown by the two groups
was not significant, we expect that a more tightly controlled task would show a modest difference.
Future study would also be profitable on short- vs. long-time retention effects. The overall
reception of GLOSSER was positive. All subjects judged the information GLOSSER provides to be
sufficient, and the program in general to be user-friendly.
Previous Work
The idea of applying morphological analysis to aid learners or translators, although not new, has not
been the subject of extensive experimentation. Antworth (1992) applied morphological analysis
software to create glossed text. But the focus was on technical realization, and the application was
the formatting of inter-linearly glossed texts for scholarly purposes. The example was Bloomfield's
Tagalog texts.
The work of the COMPASS project (Breidt and Feldweg 1997) had a similar focus to our own---
that of providing ``COMPprehsion ASSistance'' to less than fully competent foreign language
readers. Their motivation seems to have stemmed less from the situation in which language
learning is essential and more from situations in which one must cope with foreign language. In
addition, they focused especially on the problems of multi-word lexemes, examples such as English
call up which has a specific meaning `to telephone' but whose parts need not occur adjacently in
text, see call someone or other up.
Conclusions
Morphological processing is sufficiently mature to support nearly error-free lemmatization. This
functionality can be used to automate dictionary access, to explain the grammatical meaning of
morphology, and to provide further examples of the word in use (perhaps in different forms).
Second language learners look up words faster and more accurately using systems built on
morphological processing. It is to be expect that this will improve their acquisition of vocabulary.
This is but one instance where NLP has matured sufficiently to be of service in CALL. Nerbonne,
Jager and van Essen (1998) discuss experiments with other technologies, in particular speech
recognition (see Rothenberg (1998) and Witt and Young (1998)) and parsing (articles by van
Heuven (1998) and Murphy, Krüger and Grieszl (1998)).
Prospects
There are many prospects for technical improvement in the GLOSSER system. Some of the ideas of
the COMPASS (Breidt and Feldweg 1997) system on looking up and indexing multi-world lexemes
would be a valuable addition, as would certainly be a dictionary which contained actual
pronunciations. The present set of examples demonstrates the concept sufficiently, but there are too
few texts and many are inappropriate. We view all of these as less pressing than finding an
experimental deployment for the system in actual language teaching. This would probably call for a
number of pedagogical improvements in the manner in which information is presented, in the
opportunity for teachers to add to material, and to monitor use, and in the addition of facilities to
provide for review. All of these would be potentially interesting technically and pedagogically.
The prospects vis-à-vis the more general point of this paper, the use of NLP in CALL will
undoubtedly include disappointments. The identification of important barriers is a favorite pastime
among CALL aficionados, the obstacles including exaggerated claims (and subsequent
disappointments), insufficient infrastructure, competition with a staff who feels threatened by CALL,
need for staff training, incompatibility of poor fit with other materials, and complex decision and
purchasing structures. We refrain from developing these points beyond this list of simple reminders
(but Salaberry (1996), discussed above, develops many). The obstacles are genuine, and may not be
ignored, but they have received substantial comment elsewhere.
We remain confident that the long-term advantages are substantial enough for all concerned that
CALL will continue growing. More extensive exploitation of language technology should make
CALL more useful sooner.
Acknowledgments
The Copernicus program of the European commission supported the GLOSSER project in grant
343 (1994). The authors were the members of the project in Groningen. Lauri Karttunen, Elena
Paskaleva, Gábor Prószéky and Tiit Roosmaa joined in a common design for UNIX and Windows95
versions of the program. Valuable criticism has come from Poul Andersen, Susan Armstrong and
Serge Yablonsky; Edwin Kuipers joined us to make a web version of the program; and Lili
Schurcks-Grozeva assisted in the user study.
References
Antworth, E. (1992), “Glossing Text with the PC-KIMMO Morphological Parser”. In: Computers
and the Humanities, 26(5-6), pp.389-98.
Bauer, F.S.D. and A. Zaenen (1995), “Locolex: Translation rolls off your tongue”. Proceedings of
the Conference of the ACH-ALLC ’95, Santa Barbara, USA.
Breidt, E. and H. Feldweg (1997), Accessing Foreign Languages with COMPASS. Machine
Translation, special issue on New Tools for Human Translators, pp.153-174.
Van Dale (1993), Handwoordenboek Frans-Nederlands, 2nd ed. Van Dale Lexicografie, Utrecht.
Dokter, D.A. (1997a), “Indexing Corpora for GLOSSER”. Techreport, Alfa-informatica,
Groningen University, Groningen.
Dokter, D.A. (1997b), From GLOSSER to Glosser-WeB. Techreport, Alfa-informatica, Groningen
University, Groningen.
Dokter, D.A., J.Nerbonne, L. Schurcks-Grozeva and P. Smit, (1998), GLOSSER: a User Study. In
S. Jager, J. Nerbonne, and A. van Essen (eds.),. pp.167-176.
ECI, European Corpus Initiative Multilingual Corpus I,
http://www.elsnet.org/resources/eciCorpus.html
S. Jager, J. Nerbonne, and A. van Essen (eds.), (1998), Language Teaching and Language
Technology. Swets and Zeitlinger, Lisse.
van Heuven, V. Computer-Assisted Learning to Parse in Dutch. In S. Jager, J. Nerbonne, and A.
van Essen (eds.), pp.74-81.
Krantz, G. (1990), Learning Vocabulary in a Foreign Language; A Study in Reading Strategies.
Ph.D.-Thesis, University of Göteborg, Göteborg, Sweden.
Lantolf, J.P. (1996), SLA theory building: “Letting all the flowers bloom!”. Language Learning
46(4), 713-749.
Larsen-Freeman, D. and M.H. Long, (1991), An Introduction to Second Language Acquisition
Research. Longman, London.
Last, R. (1992), Computers and Language Learning: Past, Present - and Future? In C. Butler (ed.),
Computers and Written Texts, Blackwell, Oxford. pp.227-245.
Mondria, J.-A. (1996), Vocabulaireverwerving in het vreemde-talenonderwijs: De effecten van
context en raden op de retentie, Ph.D.-Thesis, University of Groningen, Groningen, The
Netherlands.
MULTEXT, Multilingual Text Tools and Corpora, http://www.lpl.univ-aix.fr/projects/multext/
Murphy, M, A.Krüger, and A.Grieszl, RECALL--Providing an Individualized CALL Environment.
In S. Jager, J. Nerbonne, and A. van Essen (eds.),. pp.62-73.
Nerbonne, J. and P. Smit, (1996), GLOSSER: in Support of Reading. COLING ‘96, Copenhagen.
pp.830-35.
Nerbonne, J., S. Jager, and A. van Essen, Introduction. In S. Jager, J. Nerbonne, and A. van Essen,
(eds.),. pp.1-11.
Ousterhout, J.K. (1994), Tcl and the Tk Toolkit, Addison-Wesley Publishers Ltd.
Paskaleva, E. and S. Mihov (1998), Second Language Acquisition from Aligned Corpora. In S.
Jager, J. Nerbonne, and A. van Essen (eds.),. pp.43-52.
Project Gutenberg, http://www.etext.org/Gutenberg/
Rothenberg, M. (1998), The New Face of Distance Learning. In S. Jager, J. Nerbonne, and A. van
Essen (eds.), pp. 146-48.
Salaberry, M.R. (1996), A theoretical foundation for the development of pedagogical tasks in
computer-mediated communication, CALICO Journal 14(1), pp. 5-34.
Warschauer, M. (1996), Computer-Assisted Language Learning: An Introduction. In Fotos, S.
(ed.), Multimedia Language Leaching, Logos International, Tokyo. p3-20
Widdowson, H.G. (1990), Aspects of Language Teaching, Oxford University Press, Oxford.
Witt, S. and S. Young, Computer-Assisted Pronunciation Teaching Based on Automatic Speech
Recognition. In S. Jager, J. Nerbonne, and A. van Essen (eds.), pp.25-35.
Zaenen, A. & G.Nunberg (1996) Communication Technology, Linguistic Technology and the
Multilingual Individual. In T.Andernach, M.Moll & A.Nijholt (eds.) Proc. of Computational
Linguistics in The Netherlands V, Parlevink: Twente. pp. 1-12.
Zock, M. (1996), Computational Linguistics and its Use in the Real World: the Case of Computer-
Assisted Language Learning. In COLING 1996, Copenhagen. pp.1,002-1,004.

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS

Computational linguistics

Computational linguistics
From Wikipedia, the free encyclopedia
Jump to: navigation, search
This article is about the scientific field. For the journal, see Computational Linguistics (journal).
Question book-new.svg
This article needs additional citations for verification.
Please help improve this article by adding reliable references. Unsourced material may be challenged and removed. (February 2010)

Linguistics
Languages of the world
Theoretical linguistics
Cognitive linguistics
Generative linguistics
Quantitative linguistics
Phonology · Morphology
Syntax · Lexis
Semantics · Pragmatics
Descriptive linguistics
Anthropological linguistics
Comparative linguistics
Historical linguistics
Etymology · Phonetics
Sociolinguistics
Applied linguistics
Computational linguistics
Forensic linguistics
Internet linguistics
Language acquisition
Language assessment
Language development
Language education
Linguistic prescription
Linguistic anthropology
Neurolinguistics
Psycholinguistics
Second language acquisition
Related articles
History of linguistics
List of linguists
List of unsolved problems
in linguistics
Portal
v • d • e

Computational linguistics is an interdisciplinary field dealing with the statistical and/or rule-based modeling of natural language from a computational perspective. This modeling is not limited to any particular field of linguistics. Traditionally, computational linguistics was usually performed by computer scientists who had specialized in the application of computers to the processing of a natural language. Computational linguists often work as members of interdisciplinary teams, including linguists (specifically trained in linguistics), language experts (persons with some level of ability in the languages relevant to a given project), and computer scientists. In general, computational linguistics draws upon the involvement of linguists, computer scientists, experts in artificial intelligence, mathematicians, logicians, philosophers, cognitive scientists, cognitive psychologists, psycholinguists, anthropologists and neuroscientists, among others.
Contents
[hide]

* 1 Origins
* 2 Subfields
* 3 See also
* 4 References
* 5 External links

[edit] Origins

Computational linguistics as a field predates artificial intelligence, a field under which it is often grouped. Computational linguistics originated with efforts in the United States in the 1950s to use computers to automatically translate texts from foreign languages, particularly Russian scientific journals, into English.[1] Since computers can make arithmetic calculations much faster and more accurately than humans, it was thought to be only a short matter of time before the technical details could be taken care of that would allow them the same remarkable capacity to process language.[2]

When machine translation (also known as mechanical translation) failed to yield accurate translations right away, automated processing of human languages was recognized as far more complex than had originally been assumed. Computational linguistics was born as the name of the new field of study devoted to developing algorithms and software for intelligently processing language data. When artificial intelligence came into existence in the 1960s, the field of computational linguistics became that sub-division of artificial intelligence dealing with human-level comprehension and production of natural languages.[citation needed]

In order to translate one language into another, it was observed that one had to understand the grammar of both languages, including both morphology (the grammar of word forms) and syntax (the grammar of sentence structure). In order to understand syntax, one had to also understand the semantics and the lexicon (or 'vocabulary'), and even to understand something of the pragmatics of language use. Thus, what started as an effort to translate between languages evolved into an entire discipline devoted to understanding how to represent and process natural languages using computers.[citation needed]
[edit] Subfields

Computational linguistics can be divided into major areas depending upon the medium of the language being processed, whether spoken or textual; and upon the task being performed, whether analyzing language (recognition) or synthesizing language (generation).

Speech recognition and speech synthesis deal with how spoken language can be understood or created using computers. Parsing and generation are sub-divisions of computational linguistics dealing respectively with taking language apart and putting it together. Machine translation remains the sub-division of computational linguistics dealing with having computers translate between languages.

Some of the areas of research that are studied by computational linguistics include:

* Computational complexity of natural language, largely modeled on automata theory, with the application of context-sensitive grammar and linearly-bounded Turing machines.
* Computational semantics comprises defining suitable logics for linguistic meaning representation, automatically constructing them and reasoning with them
* Computer-aided corpus linguistics
* Design of parsers or chunkers for natural languages
* Design of taggers like POS-taggers (part-of-speech taggers)
* Machine translation as one of the earliest and least successful applications of computational linguistics draws on many subfields.

The Association for Computational Linguistics defines computational linguistics as:

...the scientific study of language from a computational perspective. Computational linguists are interested in providing computational models of various kinds of linguistic phenomena.[3]

[edit] See also

* Association for Computational Linguistics
* Collostructional analysis
* Computational lexicology
* Computational Linguistics (journal)
* Computational science
* Computational semiotics
* Computer-assisted reviewing
* Dialog systems
* Grammar induction



* Human speechome project
* Internet linguistics
* National Centre for Text Mining
* Natural language processing
* North American Computational Linguistics Olympiad
* Quantitative linguistics
* Semantic relatedness
* Translation memory
* Ubiquitous Knowledge Processing Lab



* Universal Networking Language

[edit] References

1. ^ John Hutchins: Retrospect and prospect in computer-based translation.

Proceedings of MT Summit VII, 1999, pp. 30–44.
2. ^ Arnold B. Barach: Translating Machine

1975: And the Changes To Come.
3. ^ The Association for Computational Linguistics What is Computational Linguistics?

Published online, Feb, 2005.

[edit] External links
At Wikiversity you can learn more and teach others about Computational linguistics at:
The Department of Computational linguistics

* Association for Computational Linguistics (ACL)

o ACL Anthology of research papers

o ACL Wiki for Computational Linguistics

* CICLing annual conferences on Computational Linguistics

* Computational Linguistics – Applications workshop

* Free online introductory book on Computational Linguistics

(Internet Archive copy)
* Language Technology World

* Resources for Text, Speech and Language Processing


Retrieved from "http://en.wikipedia.org/wiki/Computational_linguistics"
Categories: Computational linguistics | Applied linguistics | Linguistics | Formal sciences

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS

TESL/Applied Linguistics: Ph. D. Program Programs

Iowa State University
Department of English
Contact Us | Search

TESL/Applied Linguistics: Ph. D. Program
Programs

* Ph.D. in ALT
* M.A. in TESL/AL
* TESL Certificate
* Linguistics Program
* K-12 ESL Endorsement
* English Placement Test
* ESL courses
* Intensive English

About

* Faculty
* Students
* Resources
* Research
* Projects
* CALL Research Lab
* International TAs
* Financial Aid
* What's New
* English Department
* TESL/AL Links
* Homepage



Autumn
Overview

Our doctoral program in Applied Linguistics & Technology was launched in the fall semester of 2005. The program aims to meet the growing need for professionals in areas of applied linguistics that intersect with computer technology. Learn more about this new program by reading more about:

* Why Study AL&T at ISU?
* What is Applied Linguistics?
* Applied Linguistics and Technology
* AL&T Program Goals and Methods of Evaluation
* AL&T Curriculum
* Information for Applicants

Why Study AL&T at ISU?

* Join an expanding area of research and practice at the interface of language and technology, including corpus linguistics, computer-assisted language learning, and computer-assisted language assessment.
* Enter a growing profession with job opportunities in academics and business. Study with internationally known faculty who will mentor you as you develop professional expertise in your areas of interest.
* Develop your expertise in technology for English language analysis, learning, and assessment.
* Acquire practical teaching and research experience in a technology-rich environment where you can learn and experiment.
* Take advantage of technology-related assistantship opportunities, including teaching ESL, teaching first-year composition, and assisting professors in research.
* Be part of a diverse, stimulating English Department which includes a doctoral program in Rhetoric and Professional Communication and other areas of English, including an MA program in TESL/Applied Linguistics.
* Take your place in the history of AL&T at ISU by helping to shape its future.

Back to Top
What is Applied Linguistics?

The name "applied linguistics" is known world wide to denote analytic and empirical linguistic approaches for investigating topics related to second language acquisition and language use. This field employs a variety of distinctive analytical and empirical methods to find solutions to such language problems as how best to teach English as a second language, how to evaluate language ability fairly, how to program a computer to recognize linguistic input, or how to analyze the linguistic structure of professional prose (e.g., the experimental scientific article) so that the structure can be taught effectively. As these examples indicate, applied linguistics can denote the linguistic study of a range of language-in-use phenomena. Any particular doctoral program in applied linguistics typically focuses on a narrower set of the broad issues falling within the scope of applied linguistics.

The name applied linguistics is used to denote professional organizations such as the American Association of Applied Linguistics and the International Association of Applied Linguistics, a national research center and information clearinghouse called the Center for Applied Linguistics, and journals focusing on issues of second language acquisition and language use such as Applied Linguistics, Issues in Applied Linguistics, and International Review of Applied Linguistics.

You can learn more about the broad discipline of applied linguistics through texts that introduce the field such as the following:

Hinkel, E. (2005). Handbook of Research in Second Language Teaching and Learning. Mawah, NJ: Erlaum Associates.
Schmitt, N. (Ed.). (2001). Handbook of applied linguistics. London: Edward Arnold.
Kaplan, R. (Ed.). (2001). Handbook of applied linguistics. Oxford: Oxford University Press.

Back to Top
Applied Linguistics and Technology

Over the past years changes have occurred in the computer hardware and software technologies implicated in second language teaching, second language assessment, language analysis and many aspects of language use. As a consequence, at the heart of current research in applied linguistics are questions concerning many technology-related issues:

How can technology intersect with language teaching practices in beneficial ways? What links should be made between second language acquisition research and technology-based language learning? How can learning accomplished through technology be evaluated? How does technology change basic issues of construct definition, validation, and fairness in language assessment? How does it affect the pragmatics of interpersonal communication? How does it change linguists' perspectives on grammatical and lexical patterns in language? How can technology expand and sharpen research across all areas of applied linguistics? These questions illustrate that the issues at the intersection of applied linguistics and technology are as complex as they are important.

Despite the significance of technology-related issues in applied linguistics, it seems that in many places of the English-speaking world, technology is taken for granted as it becomes integrated into everyday practices. As a consequence, the dramatic changes it offers for second language learners, teachers, and the profession have not been sufficiently investigated. Applied linguists need to engage more consciously and proactively with today's complex language-technology reality, which creates new opportunities and challenges for language learners as well as applied linguists engaged in language teaching, assessment and research. The Applied Linguistics & Technology program at Iowa State University was developed in response to the need for focused inquiry through a combination of knowledge about applied linguistics and technology.

The Applied Linguistics & Technology program at Iowa State University focuses on applied linguistics and technology. With its ever-expanding role in international communication in general and its prominent role as the language of technology, English is arguably the language tied up with technology in the most multi-faceted ways. Moreover, technology plays an important role in teaching and assessment of virtually all languages today.

To learn more about applied linguistics and technology, read books such as

Chapelle, C. A. (2003). English language learning and technology: Lectures on applied linguistics in the age of information and communication technology. Amsterdam: John Benjamins Publishing.

Chapelle, C. (2001). Computer applications in second language acquisition: Foundations for teaching, testing, and research. Cambridge: Cambridge University Press.

Crystal, D. (2001). Language and the Internet. Cambridge: Cambridge University Press.

Posteguillo, S. (2003). Netlinguistics: An analytic framework to study language, discourse and ideology in Internet. Castello de la Plana, Spain: Universitat Jaume.

Back toTop
Program Goals and Methods of Evaluation
Program Goals

Graduates of the doctoral program in Applied Linguistics & Technology, should be able to

* synthesize fundamental issues and concepts in applied linguistics
* use computer technology for constructing and implementing materials for teaching and assessing English
* conduct empirical research and engage in critical analysis to evaluate computer applications for English language teaching and assessment
* engage in innovative teaching and assessments through the use of technology
* evaluate multiple perspectives on the spread of technology and its roles throughout the world, particularly as they relate to English language teaching

Evaluation Methods

Measures for evaluating students' success in meeting program goals include evidence of their

* reformulation of knowledge from coursework in applied linguistics to course projects, portfolio materials, and dissertation
* development of software for course assignments and for technology in teaching
* participation in evaluation projects in technology courses and in dissertation research
* use of technology in teaching and in technology practicum courses
* reasoned selection of critical perspectives toward technology in course papers and the dissertation

Back to Top

AL&T Currlculum

The curriculum for the Applied Linguistics & Technology program consists of coursework in the following areas: Foundation Courses; Core Courses in Applied Linguistics; Technology in Applied Linguistics; Research Methods; Electives; and Dissertation Research. Students will also be required to fulfill a foreign language requirement and pass a portfolio assessment, a preliminary examination, and a final oral examination. The program consists of 72 credits.

Foundation Courses (12 credits):

* Computer Methods in Applied Linguistics (English 510)
* Introduction to Linguistic Analysis (English 511)
* Grammatical Analysis (English 516/537)
* Sociolinguistics (English 514)

Core Courses in Applied Linguistics (15 credits):

* Second Language Acquisition (English 517)
* Second Language Assessment (English 519)
* Literacy: Issues and Methods for Nonnative Speakers of English (English 524)
* Methods in Teaching Listening and Speaking Skills to Nonnative Speakers of English (English 525)
* English for Specific Purposes (English 528)

Technology in Applied Linguistics (9 credits):

* Computer-Assisted Language Learning (English 526)
* Computational Analysis of English (English 520)
* Practicum in Technology and Applied Linguistics (English 688)

Research Methods (12 credits):

* Discourse Analysis (English 527)
* Research Methods in Applied Linguistics (English 623)
* Qualitative research methods (e.g., Soc 513)
* Quantitative research methods (e.g., Stat 401)

Electives (12 credits)

Four courses, two of which must be Seminars in Applied Linguistics and one of which must be in technology. The Seminar in Applied Linguistics (English 630) is a repeatable course because topics will vary. Students must take this course twice to fulfill this requirement. The course in technology may be either a seminar in applied linguistics or a course in another discipline.

Other electives may be taken in areas such as English Literature, Curriculum & Instruction, Anthropology, Foreign Language Literature and Linguistics, Rhetoric and Professional Communication, Computer Science

Dissertation (12 credits)

Back to Top
Information for Applicants

Ph.D. applicants must have completed a Master's degree prior to their first semester in the program.
Minimum TOEFL scores for Ph.D. applicants: 111 (IBT)/ 273 (CBT)/ 640 (PBT)

Application Deadline: January 5 (fall entry only). Find more information about applying to the graduate programs in the English Department


Back to Top



Becoming the Best

TESL/Applied Linguistics, Department of English• 203 Ross Hall • Ames, IA 50011 • Phone: 515-294-2180 • Fax: 515-294-6814

Copyright © 2004, Iowa State University of Science and Technology. All rights reserved.

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS

An Intelligent Computer Assisted Language Learning System for Arabic Learners

An Intelligent Computer Assisted
Language Learning System for Arabic
Learners
Khaled F. Shaalan*1
The British University in Dubai (BUiD), UAE
This paper describes the development of an intelligent computer-assisted language learning
(ICALL) system for learning Arabic. This system could be used for learning Arabic by students at
primary schools or by learners of Arabic as a second or foreign language. It explores the use of
Natural Language Processing (NLP) techniques for learning Arabic. The learners are encouraged
to produce sentences freely in various situations and contexts and guided to recognise by themselves
the erroneous or inappropriate functions of their misused expressions. In this system, we use NLP
tools (including morphological analyser and syntax analyser) and error analyser to issue feedback to
the learner. Furthermore, we propose a mechanism of correction by the learner which allows the
learner to correct the typed sentence independently, and allows the learner to realise that what the
error is.
Introduction
Computer-assisted language learning (CALL)2 addresses the use of computers for
language teaching and learning. CALL emerged in the early days of computers. Since
the early 1960s, CALL software was designed and implemented. The effectiveness of
CALL systems has been made obvious by many researchers (Lam & Pennington,
1995; McEnery, Baker, & Wilson, 1995). Until quite recently, computer-assisted
language learning was a topic of relevance mostly to those with a special interest in
that area. Recently, though, computers have become so widespread in schools and
homes and their uses have expanded so dramatically that the majority of language
teachers must now begin to think about the implications of computers for language
*Corresponding author. Institute of Informatics, The British University in Dubai (BUiD), PO Box
502216, Dubai, UAE. Email: khaled.shaalan@buid.ac.ae
Computer Assisted Language Learning
Vol. 18, Nos 1 & 2, February 2005, pp. 81 – 108
ISSN 0958-8221 (print)/ISSN 1744-3210 (online)/05/010081–28
# 2005 Taylor & Francis Group Ltd
DOI: 10.1080/09588220500132399
learning. Using computers provides a number of advantages for language learning
(Warschauer, 1996):
. Repeated exposure to the same material is beneficial or even essential to learning.
. A computer is ideal for carrying out repeated drills, since the machine does not get
bored with presenting the same material and since it can provide immediate nonjudgmental
feedback.
. A computer can present such material on an individualised basis, allowing
students to proceed at their own pace and freeing up class time for other activities.
. The process of finding the right answer involves a fair amount of student choice,
control, and interaction.
. The computer can create a realistic learning environment, since listening can be
combined with seeing, just as in the real world.
. Multimedia and hypermedia technologies allow a variety of media (text, graphics,
sound, animation, and video) to be accessed on a single machine. Hence, skills are
easily integrated, since the variety of media makes it natural to combine reading,
writing, speaking and listening in a single activity.
. Internet technology facilitates communications among the teacher and the
language learners. It allows a teacher or student to share a message with a small
group, the whole class, a partner class, or an international discussion list of
hundreds or thousands of people.
. Incorporating NLP techniques provide learners with more flexible—indeed, more
‘intelligent’—feedback and guidance in their language learning process.
More than a decade ago, Intelligent Computer-Assisted Language Learning (ICALL)
started as a separate research field, when Artificial Intelligence (AI) technologies were
mature enough to be included in language learning systems. The beginning of the
new research field was characterised by Intelligent Tutoring Systems (ITS), which
embedded some NLP features to extend the functionality of traditional language
learning systems. The continuous advances in ICALL systems have been
documented in several publications (Cameron, 1999; Gamper & Knapp, 2002;
Holland, Kaplan, & Sama, 1995; Swartz & Yazdani, 1992).
By far the majority language learning programs have been developed for English,
followed by Japanese, French, and German (Gamper & Knapp, 2002). Current
Arabic (I)CALL systems have the weakness that learners cannot key in an Arabic
sentence freely. Similarly, they cannot guide the learner to correct the most likely
ill-formed input sentences. The learner just accepts the information which follows
the programmed instruction that is pre-installed in the computer. For these
systems to be useful, more research to combine NLP techniques with language
learning systems is needed. Parsing, the core component in ICALL systems,
allows the system both to analyse the learner’s input and to generate responses to
that input (Holland, Maisano, & Alderks, 1993). Allowing learners to phrase their
own sentences freely without following any pre-fixed rules can improve the
effectiveness of ICALL systems, especially when the expected learner answers are
82 K. F. Shaalan
relatively short and well-focused (Boytcheva, Vitanova, Strupchanska, Yankova, &
Angelova, 2004). Both the well- and ill-formed structure of the input sentence can
be recognised. The learner should be allowed to correct the typed sentence
independently.
This paper describes an ICALL system for Arabic using NLP techniques, called
Arabic ICALL, which can solve the weaknesses of current Arabic (I)CALL
systems. In Arabic ICALL, there are two main types of test items for interaction
with the learner—selection-type that tends to elicit answers easily classified as right
or wrong and supply-type requiring the learners to write a few words. The objective
test method is used to assess the learner’s knowledge or skills where each question
has one (and only one) correct answer—and there is no ambiguity about what that
correct answer should be. The present system guides learners to recognise by
themselves the erroneous or inappropriate functions of their misused expressions.
In other words, it helps learners to make use of their errors. It doesn’t give them
the correct answer directly but it enables them to try over and over again. In this
system, we use NLP tools (including morphological analyser and syntax analyser)
and error analyser to issue feedback to the learner. Furthermore, we propose a
mechanism of correction by learners which allows the learner to correct the typed
sentence independently, and allows learners to realise what the error is. Arabic
ICALL follows the curriculum of Arabic grammar at the Egyptian primary
schools.
The rest of this paper is structured as follows. First, related work on Arabic
language learning programs is given. We also discuss limitations of current Arabic
language learning systems as well as limitations of the Arabic ICALL system, and
briefly describe our proposed Arabic ICALL system. The sections following on from
this present the main components of the Arabic ICALL system. Finally, we conclude
the paper and give directions for future work.
Related Work
The linguistic computation of an Arabic sentence is a difficult task (Othman, Shaalan,
& Rafea, 2003). The difficulty comes from several sources: (1) the length of sentences
and the complexity of Arabic syntax; (2) the omission of diacritics (vowels) in written
Arabic ‘‘altashkiil’’; (3) the free word order nature of Arabic sentence; and (4) the
presence of an elliptic personal pronoun ‘‘alDamiir almustatir’’. For these reasons,
there is very little research involving Arabic (I)CALL (Ditters, Oostdijk, & Cameron,
1993).
Research into Arabic (I)CALL can be classified by two approaches: the
Computer as a tool and the Computer as a tutor. In the Computer as a tool approach,
some computer programs can be used as a tool that does not necessarily provide
any language material at all, but rather empowers the learner (usually a native
speaker) to use or understand language. In the Computer as a tutor approach, the
process of finding the right answer involves a fair amount of student choice,
control, and interaction.
Intelligent Computer Assisted Language learning 83
Computer as Tool
Hegazi, Ali, Abed and Hamada (1989) presented a way of representing Arabic syntax
in Prolog as production rules. The system can detect some errors concerning Arabic
syntax, and so can be used for an educational environment.
Abou Ela (1994) developed an expert system, the Arabic Syntax Analyzer
(ESASA), which can be used as a tool to assist Arabic linguists in building Arabic
grammar rules. The grammar is expressed using a declarative language called
Grammar Writing Language (GWL). This tool is aimed at building Arabic natural
language applications including CALL.
Using the Internet for publishing web-based CALL materials that contain non-
Latin alphabets requires the solution of various technical problems. There are so
many unknown factors associated with the operating system of a distant user that
affect the browsing characteristics of these materials. Cushion and He´mard (2002)
described how recent technological developments have provided the possibility of
overcoming these technical problems in conjunction with the Java programming
language and the Unicode character numbering system.
Shaalan (2003) developed an Arabic grammar checker, called Arabic GramCheck.
Arabic GramCheck looks for common Arabic grammatical problems, describes the
problem, and offers suggestions for improvement. This program is useful in pointing
to problems believed typical of native speaker writing. Thus, the learner can avoid
such problems in future.
Computer as Tutor
Gheith, Dawa, and Afifty (1996) developed Instructional Software for Teaching the
Arabic Language (ISTAL) for grade one prep school. The system presents the
curriculum as a simple concept associated with a set of generated sentences. Then,
the system generates an exercise for the student and the student’s answer is
automatically evaluated by comparing it to the system’s solution.
Recently Nielsen (2001) and Nielsen and Carlsen (2003) developed a system for
learning Arabic, ArabVISL, at the University of Southern Denmark. The system is an
interactive Web-based application. It allows students of Arabic as a foreign language
to analyse Arabic sentences by using Arabic script and specific Arabic grammatical
terminology.
The Interactive Language Learning Project at London Guildhall University has
produced course materials for the University’s Arabic classes (Cushion & He´mard,
2003). The system is designed for learning Arabic at the beginner level. This study
focused on problems associated with learning a language with an unfamiliar alphabet.
It discussed the possible use of CALL authoring as part of the learning process.
Mote, Johnson, Sethy, Silva, and Narayanan (2004) developed a speech-enabled
computer learning environment designed to teach Arabic spoken communication to
American English speakers, called Tactical Language Training System (TLTS). This
system can detect errors in learner speech. The TLTS incorporates two speech-
84 K. F. Shaalan
enabled learning environments: an interactive game called the Mission Practice
Environment (MPE) that simulates conversations with native speakers, and an
intelligent tutoring system called the Mission Skill Builder (MSB) for acquiring and
practising communicative skills.
Limitations of Current Arabic Language Learning Systems
In Egypt, where there is a growing demand in using computers for teaching and
learning, some publishers of off-the-shelf school textbooks provide students with
either CD’s or web sites that contain vocabulary and grammar practice. However,
most of these systems have some common limitations, which are:
1. They often resemble the traditional workbook exercises from which they were
adapted.
2. From a pedagogical perspective, the definition of acceptable answers to exercises
is highly constrained. For instance, in the linguistic analysis ( إعراب ) questions,
the learner can type his answer as follows: ‘‘ مبتدأ مرفوع وعلامة رفعة الضمة ’’ (inchoative
is in nominal case and the diacritic sign is dam-mah). Nevertheless, the system
would consider this response as a wrong answer since it stores the answer of this
question as: ‘‘ مبتدأ مرفوع بالضمة ’’ (inchoative is in nominal case and the diacritic sign
is dam-mah).
3. Error feedback commonly does not address the source of an error. For instance,
the system displays the correct answer without any explanation of the student’s
mistake. This makes the system’s feedback a generic catchall response.
4. For vocabulary exercises, the student is referred to the corresponding page in the
textbook, which displays the word in question in a word list. In addition to the
pedagogical limitations, the student has to consult the textbook, which is an
unnecessary inconvenience given the potential of the Web.
Limitations of Arabic ICALL
Arabic ICALL system has been successfully implemented using SICStus Prolog on
an IBM PC. The system has some limitations:
. The system as described is targeted at a particularly well-formed subset of Arabic,
which would not extend well to more colloquial dialects. Even standard newswire
is likely to frequently include pre-verbal subjects and adverbials which are not
considered in this work. This restriction to a well-formed subset might be
appropriate for people trying to learn Arabic in a formal style.
. As vowels are usually omitted in written Arabic, our system does not handle the
vowled Arabic text where letters are written with diacritic signs.
. Although ordering words to form a sentence is a type of question normally
classified under the objective test method, it is not included into our system due to
the free word order nature of Arabic that is usually dependent on semantics.
Intelligent Computer Assisted Language learning 85
. Since the task of automatic processing of free natural language in ICALL is hard,
the objective test method is used such that the expected learner’s answer is
relatively short and well-focused.
. Thepresent system does not diagnose spelling errors. It accepts only answers that are
free of typographical errors. We have designed, but only partially implemented due
to lack of time and fund, an Arabic Spell Checker (Shaalan, Allam,&Gomah, 2003).
As the Arabic Spell Checker is not fully implemented, it was not integrated with
ArabicICALL.Tobe part of the final ArabicICALL’s error analyser, this integration
should distinguish typographical (misspelling) errors from wrong answers.
The Proposed System Architecture
Figure 1 shows the overall architecture of the proposed Arabic ICALL system. This
system consists of the following components: user interface, course material, sentence
analysis, and feedback. The user interface provides the means of communications
between the learner and the Arabic ICALL system. The course material includes
educational units, an item (question) bank, a test generator, and an acquisition tool.
The sentence analysis includes a morphological analyser, syntax analyser (parser),
grammar rules, and lexicon. The feedback component includes an error analyser that
is used to parse ill-formed learner input and to issue feedback to the learner.
Figure 1. Overall architecture of the proposed Arabic ICALL system
86 K. F. Shaalan
User Interface
The user interface provides the means of communications between the learner and
the Arabic ICALL system. It is used to present multimedia lessons (text, graphics,
sound, and animation), to present the test items, to allow the learner to take the test,
and to present feedback to the learner.
In (I)CALL, there are three possible approaches for reacting to the learner’s
response in order to give appropriate feedback to the learner: pattern matching-based
approach, statistical-based approach, and rule-based approach.
The pattern matching-based approach requires that exercise authors enter many
different correct and wrong answers with their associated feedback. This is a very
tedious and time consuming task yet only provides appropriate feedback when the
learner types in one of the expected answers. Thus, this approach requires a great deal
of up front teacher knowledge, experience, and effort.
Both statistical-based approach and rule-based approach give some freedom to the
language learners in the way they phrase their answers, while enabling the exercise
author to enter only one possible correct answer, thus saving much time compared to
the previous pattern matching answer coding approach.
The statistical-based approach uses statistical methods to acquire knowledge.
Parameters are automatically learned (estimated) from a corpus that is labeled with
the properties needed. Both a statistical model and language parameters should be
specified by humans. The characteristics of the statistical-based approach are: (1) no
strict sense of well-formedness in mind; and (2) have a large parameter space (e.g.,
100,000 words using Tri-gram model requires 105*105*105 parameters). The
advantages of this approach are that it does not need computational models to be
established and computation is easy.
The disadvantages of this approach are: (1) cannot restrict computation using
heuristics of the linguistic theory; (2) requires a large amount of data to train the
statistical model. Parsers and error diagnosis tools cannot be trained on raw data. The
data must be tagged by humans, which is costly, time consuming, and sometimes
ambiguous. Meaning, hand-tagged corpora is very expensive; (3) even ungrammatical
permutation (i.e., sequence) of words of the ill-formed learner’s answer are still
probable (i.e., have some probability to occur). It is hard to decide whether a string of
words is grammatical or not. Actually, strings that are not words at all still will have
some probability in the mass probability of the model; (4) the large parameter space
of statistical models is a serious problem when decoding (i.e., searching for the most
probable structure to assign to a string of words). Statistical models do not distinguish
between different forms of words, for example, play, plays, played, playing are not
treated as related words originating from the verb play.
The rule-based approach provides detailed analysis of the learner’s answer using
linguistic (morphological and syntactic) knowledge. The characteristics of the rulebased
approach are: (1) has a strict sense of well-formedness in mind; (2) imposes
linguistic constraints to satisfy well-formedness; (3) allows the use of heuristics (such
as a verb cannot be preceded by a preposition); and (4) Relies on hand-constructed
Intelligent Computer Assisted Language learning 87
rules rather than automatic training from data. The advantages of this approach are
that it is easy to incorporate the linguistic knowledge, and it is easy to augment the
grammar rules with heuristic rules, which are capable of detecting ill-formed input
and providing appropriate feedback. The disadvantage of this approach is that it is not
easy to obtain high coverage (completeness) of the linguistic knowledge. However, it
could be useful for limited domain where errors in the input can be expected.
It is well-established that feedback is an essential prerequisite for effective learning.
Both the pattern matching-based approach and statistical-based approach lack a
systematic and automatic way in diagnosing the learner’s ill-formed input and
providing appropriate feedback. Data collection is costly and time consuming. On the
contrary, the rule-based approach has the advantage of providing appropriate
feedback because it performs detailed analysis for both well-formed and ill-formed
answers. It is easy to acquire linguistic knowledge, and to specify linguistic constraints
and heuristics. For these reasons, we decided to follow the rule-based approach in
developing Arabic ICALL.
Course Material
The course material includes educational units, item bank, test generator, and
acquisition tool (see Figure 2). Each educational unit is a collection of Arabic
grammar lessons that addresses a common topic. The item bank (question bank) is a
Figure 2. The proposed course material architecture
88 K. F. Shaalan
database of test items. The test generator withdraws test items as needed to develop a
test. The acquisition tool allows the instructor to author and maintain lessons, and to
create and maintain test items.
Primary Level Lessons of Arabic Language
The educational units include Arabic grammar lessons for the primary level.
Specifically, they cover the following:3
الأسماء (nouns) – النعت (adjective) – المثنى والجمع (dual and plural) – الأفعال (verbs) – الضمائر
(pronouns) – الجملة الأسمية (nominal sentence) – المبتدأ والخبر (inchoative and enunciative) – أنواع
الخبر (types of the enunciative) – الجملة الفعلية (verbal sentence) – العطف (conjunctions) – الظرف
(adverb) – الحال (circumstantial accusative) – النداء (interjection) – المفعول لأجله (causative object) –
المفعول المطلق (unrestricted object) – إن وأخواتها (Inna and her sisters) – أنواع خبر إن وأخواتها (types of
predicate of Inna and her sisters) – كان وأخواتها (Kana and her sisters) – أنواع خبر كان وأخواتها (types
of predicate of Kana and her sisters) – الحروف (articles)
Figure 3 shows an example of a lesson explaining the unrestricted object ‘‘ المفعول
المطلق ’’. It consists of an explanation of this grammar rule, an example, sound
functionality, lesson test, and some navigation aids.
The lessons are stored in a database. The system includes some instructional
templates to allow for quick generation of instructional material. The structure of
lessons consists of two database relations, namely, lesson relation and example
relation.
Figure 3. A lesson
Intelligent Computer Assisted Language learning 89
Item Bank
The item bank is a database of test items. This component is used to generate
different types of test items each time the learner is allowed to take a test. The test
generator selects test items in random order. The instructor determines the selection
criterion, and all the test items that match this criterion are collected. Then, we apply
a random function to present the selected test items to the learner.
Figure 4 shows an example of a test item. It consists of an explanation of a question
header (identify the inchoative and enunciative, and the type of the enunciative in the
following), a given sentence (the brave soldiers fought to victory), a learner input area
(a pull down menu and two text boxes), an ‘‘Answer button’’ to generate the model
answer, a ‘‘Check button’’ to check the learner’s answer, a hyperlink to the relevant
grammar lesson, and some navigation aids. It is worth noting that the learner’s
answer, whether correct or wrong, is compared against the system-generated answer.
In Arabic ICALL, there are two main types of test items for interaction with the
learner: supply-type (short-answer/fill-in-the-blank) or selection-type (matching, true/
false, identify, or multiple-choice) interactive questions. The objective test method is
used to assess the learner’s knowledge or skills. From the linguistic point of view, the
type of questions used in our Arabic ICALL system can be classified as follows:
1. Identify words according to certain morphological features or identifying
constituents according to certain syntactic features
. Examples:
. Identify the verb, subject, and object in the following sentence
’’عين الفعل والفاعل والمفعول به في الجملة الآتية‘‘
Figure 4. A test item
90 K. F. Shaalan
. Extract the adjective and the described noun in the following sentence
’’أستخرج النعت و المنعوت في الجملة الآتية:‘‘
2. Verb conjugation
. Examples:
. Give the correct present and imperative tense of the following verbs
’’أكتب الفعل المضارع والأمر للأفعال الآتية ‘‘
. Present tense—fill in the blank with the correctly conjugated form of the weak verb
in parentheses
’’الفعل المضارع – أكمل باستخدام الفعل المعتل المناسب مما بين القوسين‘‘
3. Noun morphology
. Examples:
. What pronoun would you use to talk to the following people?
’’ما هي الضمائر التي يمكن استخدامها مع الأشخاص التاليين‘‘
. Fill in the blanks with the correct form of the demonstrative noun
’’أكمل باستخدام أسم الإشارة المناسب‘‘
4. Identify the grammatical relation
. Examples:
. Identify the negation or prohibition in the following sentence
’’عين النفي أو النهى في الجملة الآتية:‘‘
. Put the connective particle in the correct place in the following sentences
’’ ضع حرف العطف في مكانه الصحيح من الجمل الآتية:‘‘
5. Linguistic analysis of words between brackets or a sentence
. Examples:
. What’s the difference in linguistic analysis between the following? Give the
reason ‘‘ ’’هل تري فرقا إعرابيا في التالي؟ اذكر السبب
. Give the reason for the accusative end case of the words between brackets
’’ اذكر سبب نصب الكلمات التي بين القوسين:‘‘
6. Transform the sentence category
. Examples:
. Change the following nominal sentence into verbal sentence
’’الجملة الآتية أسمية أجعلها فعلية‘‘
. Advise your colleague of the following using imperative verb
’’انصح زميلك بما يلي مستخدمًا فعل أمر:‘‘
7. Agreement
. Examples:
. Rewrite the following sentences, changing the demonstrative noun from the
singular form into the plural form, and change what is necessary to make it
grammatically correct sentence
’’أجعل الإشارة للجمع وغير مايلزم‘‘
. Is the agreement between the adjective and the described noun in
the following sentence correct or incorrect? Give the reason
’’هل التطابق بين النعت والمنعوت في الجمل الآتية صحيح؟ اذكر السبب ‘‘
8. Review test
Intelligent Computer Assisted Language learning 91
. Examples:
. Complete the following passage by selecting the correct verb/correct adjective
for the context
’’أكمل القطعة التالية باختيار الفعل أو النعت المناسب للسياق‘‘
. Are the following sentences grammatically correct?
’’هل الجمل التالية صحيحة إعرابيا‘‘
The structure of the item bank consists of three database relations, namely, question
title relation, question content relation, and answer relation.
Sentence Analysis
Logic programming plays an essential role in NLP because it attempts to use logic
to express grammar rules and to formalise the process of parsing (Gazdar &
Mellish, 1990). A grammar specified this way is known as logic grammar since it
represents rules as Horn clauses (Dougherty, 1994). Logic grammars can be
conveniently implemented in Prolog. Prolog-based grammars can be quite efficient
in practice (Allen, 1995). The Prolog interpretation algorithm uses exactly the same
search strategy as the depth-first top-down parsing algorithm, so all that is needed is
a way to reformulate grammar rules as clauses in Prolog. Definite clause grammars
(DCGs) notation was developed as a result of research in natural language parsing
and understanding (Pereira, Sheiber, & David, 1986). DCGs allow one to write
grammar rules directly in Prolog, producing a simple recursive descent parser.
During the construction of the Arabic parser, feature-structures are translated into
Prolog terms. Because of this translation step, parsing can make use of Prolog’s
built-in term-unification, instead of the more expensive feature-unification. Prologs
that conform to the Edinburgh standard have DCGs as part of their implementations.
In the current system, grammar rules of Arabic are written in the DCG
formalism, which is automatically translated into executable code in SICStus
Prolog.
The sentence analysis in Arabic ICALL includes a morphological analyser, syntax
analyser (parser), grammar rules and lexicon (see Figure 5). The sentence analysis
works as follows. Learner written input is first fed into the interactive preprocessor,
where it is grouped into words. The words in the input are then decomposed into
roots and affixes by the morphological analyser, which obtains information about
the subparts from the lexicon (for example, part of speech, number, case). The
subparts so identified are then reunified into whole words and passed along to the
syntactic parser. The Arabic parser, which is based on DCG formalism, tries to
build a structure (usually, ‘‘parse tree’’) based on the information from the lexicon
concerning the grammatical relations between the words. The parser then applies a
set of descriptive rules representing the grammar of Arabic until it finds the
structure represented by the input sentence. This structure is passed to the feedback
component that is equipped with an error analyser that identifies and records any
errors made in the structural description that is generated.
92 K. F. Shaalan
Morphological Analysis
The Arabic language is based on the Semitic root-and-pattern scheme of forming word
roots, as well as the concatenation of root and affixes. We need a sophisticated
morphological analyser that is capable of transforming the inflected Arabic word into its
origin. To achieve this function we developed a morphological analyser for inflected
Arabic words (see Rafea & Shaalan, 1993). The morphological analyser analyses the
inflected Arabic word to extract the root and its features. An Augmented Transition
Network (ATN) (Woods, 1970) technique was successfully used to represent the
context-sensitive knowledge about the relation between a root and inflectional
additions. The ATN consists of arcs. Each of which is a link from a departure node to
a destination node, called states (see Figure 6).ATNsadditionally employ registers which
hold linguistic information. ATNs also allow actions to be associated with each arc, for
instance, the setting of the register with an omitted affix, the conversion/addition of a
weak letter. An exhaustive-search to traverse the ATN generates all the possible
interpretations of an inflected Arabic word. Themorphological analyser is implemented
in Prolog and integrated with the Arabic DCG parser.
Figure 7 shows an example of analysing the verb ‘‘ شاهدتك ’’ (I saw you) using the
ATN shown in Figure 6. This verb is broken down into the verb ‘‘ شاهد ’’ (saw), the first
person pronoun ‘‘ ت’’, and the second person pronoun ‘‘ .’’ ك
Figure 5. The proposed sentence analysis architecture
Intelligent Computer Assisted Language learning 93
An Arabic monolingual lexicon was also needed to successfully implement the
morphological analyser. The lexicon is designed to reflect the word categories in
Arabic—each with a different set of features.
The morphological analysis in Arabic ICALL system analyses the learner’s answer
in response to a generation question, such as fill-in-the-blank. This answer should
meet certain morphological rules. These rules are used to guide the analysis of the
learner’s answer. This has the following advantages: minimising the ambiguity,
facilitating the generation of the feedback in case of ill-formed input, and speeding up
the analysis phase.
The Grammar Formalism
The grammar for Arabic contains the grammatical knowledge required to analyse a
grammatically correct sentence. The grammar is being developed especially for
learning Arabic. Currently, it concerns Arabic grammar at the primary level. We
adopted general solutions as much as possible, as this increases the chances that the
grammar can be used in other domains as well. Thus, in designing the grammar we
seek a balance between short-term goals (a grammar which covers sentences typical
for learning Arabic and is reasonably robust and efficient) and long-term goals (a
grammar which covers the major constructions of Arabic in a general way).
Figure 7. Analysis of the Arabic verb ‘‘ شاهدتك ’’ (I saw you) using ATN technique
Figure 6. ATN representing the relation between the additions and root of an inflected Arabic word
94 K. F. Shaalan
Arabic grammar in Arabic ICALL is written in DCG formalism. The
central formal operation in DCG is the unification of feature-structures. Table 1
describes the features used in the current grammar along with their possible
values.
Grammar rules for grammatically correct sentence. A DCG rule has the following
form:
Nonterminal-symbol ? body
Where ‘‘body’’ is a sequence of one or more items separated by commas. Each item is
either a non-terminal symbol or a sequence of terminal symbols written within square
Table 1. Features and their values for the Arabic grammar
Feature Possible values Comments
Gender Masculine/feminine
Number Singular/dual/plural
Definiteness Definite/indefinite Applied only to nouns
Special noun Yes/no Determine whether or
not the noun is Inna and
its sisters ( (إن وأخواتها
Pattern Form of pattern (wazen)
End case Accusative/nominative/genitive Iarab الأعراب
Transitivity Transitive/ intransitive Applied only to verbs
Special verb Yes/no Determine whether or
not the verb is Kan
and its sisters ( (كان و أخواتها
Affix Affixes of the inflected word
Current category Category of the grammatical symbol
being parsed
Next category Category the grammatical symbol
that follows the current symbol
Noun as adjective Yes/no Determine whether or
not the noun can be
used as an adjective.
Noun as annexation Yes/no Determine whether or
not the noun can be
annexed ( (إضافة
Verb tense Past/present/imperative
Single word Single form of the broken plural
Infinitive Verb Infinitive form
Person First/second/third Applied only to pronouns
Noun refers to time
or place
Time/place Determine whether the
accusative ( ظرف ) is related to
time, related to place, or both
Intelligent Computer Assisted Language learning 95
brackets ([and]). The meaning of the rules is that ‘‘body’’ is a possible form for a
phrase of type non-terminal symbol. In the right side of a rule, in addition to nonterminals
and lists of terminals, there may also be sequences of procedure calls,
written within curly brackets ({and}). These are used to express extra conditions that
must be satisfied for the rule to be valid.
In the following, we show an extraction of DCG rules used for parsing a
grammatically correct Arabic verbal sentence.
verbal_sentence ? simple_verbal_sentence (1)
verbal_sentence ? prefixed_verbal_sentence (2)
verbal_sentence ? special_verbal_sentence (3)
simple_verbal_sentence ? verb, subject, object, unrestricted_object (4)
For simplicity, these rules do not include linguistic features such as gender, number
and definition, which are assigned to each non-terminal.
Rule (4) illustrates a grammar rule for parsing a simple verbal sentence that consists
of four constituents: a verb, a subject, an object and an unrestricted object ‘‘ .’’مفعول مطلق
An unrestricted object is a noun that originates from the infinitive verb. This kind of
repetition is considered a mark of good style. In Arabic, repeating the verbal noun
after the verb makes the sentence more emphatic. This is explained by the following
example:
He helped me a great deal of help
ساعدنى مساعدة عظيمة
Grammar rules for linguistic analysis.We have also developed another grammar that is
used to parse the learner’s answer in response to a question about the linguistic
analysis of a given Arabic sentence. The parser takes the learner’s answer and
converts it into a quadruple abstract representation of the canonical form:
الموقع الإعرابي + الإعراب + علامة الإعراب + السبب
Reason + Analytic sign + End case + Analytic location
For example, the linguistic analysis of the word between brackets in the sentence
استمتعت بجو الريف (استمتاعاً)
I enjoyed the rural weather (very much)
is:
مفعول مطلق منصوب بالفتحة لأنه مفرد.
Unrestricted object is in accusative case and the diacritic sign is fat-hah because it is in
singular form.
96 K. F. Shaalan
The following lists all of the possible learner’s answers:
مفعول مطلق منصوب وعلامة نصبه الفتحة لأنه مفرد .
مفعول مطلق منصوب وعلامة نصبه الفتحة وهو مفرد .
.مفعول مطلق منصوب بالفتحة لأنه مفرد .
This will be parsed into the abstract representation:
مفعول مطلق + منصوب + فتحة + مفرد
singular + fat-hah + accusative + unrestricted object
This abstract representation is unique and unambiguous such that it facilitates
comparing the learner’s answer with the correct answer generated by the system.
To show how the linguistic analysis is generated, consider the following question as
an example.
أعرب ما بين القوسين في الجملة الآتية:
& ألف العقاد (كتبا كثيرة).
Give the linguistic analysis of the words between brackets in the following sentence:
Al-aakad authored (many books).
Consider also the following DCG rule that describes the linguistic analysis of the
words between brackets. These words are an object that is followed by an adjective
such that the adjective agrees with the object (the noun it modifies) in number,
gender, definition, and end case.
object(Words_bet_brackets, Rest, Analysis) ? [Object], [Adjective],
{get_analysis(Object, Gender, Num, Def, Words_bet_brackets, Rest1, End_case,
Analysis ,(’مفعول به‘ , 1
get_analysis(Adjective, Gender, Num, Def, Rest1, Rest, End_case, Analysis ,('نعت‘ , 2
append(Analysis1, Analysis2, Analysis)
}.
This rule takes the list of words to be analysed as input (the words between brackets)
and produces as output both the rest of this list, if any, and the linguistic analysis of
these words. Parsing of these words proceeds as follows. The first word is bound to
the variable Object. The get_analysis/9 procedure yields both the features and the
linguistic analysis of this word. Similarly, the second word is bound to the variable
Adjective. The get_analysis/9 takes the recognised features of the first word and checks
their agreement in number, gender, and definition features with features of the
second word, and yields the linguistic analysis of this word. The end case of the
adjective ( كثيرة -many) is determined by the noun it modifies which is the object
كتبا) -books) in this example. Finally, the linguistic analyses of both words are
Intelligent Computer Assisted Language learning 97
concatenated into a list that constitutes the answer to the question. The generated
linguistic analysis of words ‘‘ كتبا كثيرة '' (many books) is the following list:
[[’مفرد‘ ,’فتحة‘ ,’منصوب‘ ,’نعت‘] ,[’جمع تكثير‘ ,’فتحة‘ ,’منصوب‘ ,’مفعول به‘]]
[[object, accusative, fat-hah, broken-plural], [adjective, accusative, fat-hah, singular]]
The definition of get_analysis/9 procedure is as follows:
get_analysis(Word_being_parsed, Gender, Num, Def, [Word_bet_bracketj Rest],
Rest_words, End_case ,Analysis, Location):-
morph(Word_being_parsed, lex(_,noun,Gender,Num,Def,Adj,_,_)),
(Word_bet_bracket = = Word_being_parsed -4
get_Irab(Location,Num, End_case, Word_analysis),
Rest_words =Rest, Analysis = [Word_analysis])
; Rest_words = [WordjRest], Analysis = []).
This procedure returns the linguistic analysis of the word being analysed if it is one of
the words that occur between brackets in the given question. Otherwise,
an empty list is returned. The procedure get_analysis/9 calls the procedure morph/2
to get features of the input word. Then, it calls get_Irab/3 to generate the quadruple
abstract representation form. The features that result from the morphological analysis
of the words between brackets in the above example are as follows:
كتبا . (books): noun, female, broken-plural, indefinite,. . .
كثيرة . (many): noun, female, singular, indefinite,adj,. . .
get_Irab(Location, Num, End_case, Word_analysis):-
get_end_case(Location, End_case),
get_analytic_sign(Num, End_case, Analytic_sign),
Word_analysis = [Location, End_case, Analytic_sign, Reason].
The procedure get_Irab/3 takes the location and number of the input word and
generates its linguistic analysis. It uses two facts get_end_case/2 and get_analytic_sign/3.
The fact get_end_case/2 determines the end case of the word from its location within
the sentence. The fact get_analytic_sign/3 determines the analytic sign of the word
from its number and end case.
The linguistic analyses that result from processing get_Irab/3 to the words between
brackets in the above question are as follows:
كتبا . (books): [‘ [’جمع تكثير‘ ,’فتحة‘ ,’منصوب‘ ,’مفعول به
[object, accusative, fat-hah, broken-plural]
كثيرة . (many): [‘ نعت ’, End_case, ‘ [’مفرد‘ ,’فتحة
[adjective, End_case, fat-hah, singular]
Where End_case will be determined by the agreement between the adjective and the
noun it modifies. This agreement takes place by the unification of the variable
End_case in the body of the above DCG rule.
98 K. F. Shaalan
get_end_case (‘ .(’منصوب‘ ,’مفعول به
get_end_case (‘ .(’مرفوع‘ ,’مبتدأ
get_end_case (‘ .(_ ,’نعت
. . .
get_analytic_sign (‘ .(’ألف‘ ,’مرفوع‘ ,’مثنى
get_analytic sign (‘ .(’فتحة‘ ,’منصوب‘ ,’جمع تكثير
get_analytic sign(‘ .(’فتحة‘ ,’منصوب‘ ,’مفرد
. . .
The Feedback System
Feedback is the computer’s response to answers made by learners. Feedback gives
students a feel of how well they are progressing through a lesson, thereby increasing
their confidence levels. It also reinforces the subject matter. In Arabic ICALL, the
feedback component includes an error analyser that is used to parse ill-formed learner
input and to issue feedback to the learner (see Figure 8). We have augmented the
Arabic grammar with heuristic rules (buggy rules) which are capable of parsing illformed
input and which apply if the grammatical rules fail. The feedback component
is implemented using SICStus Prolog. The feedback system compares the analysis of
the learner’s answer with the correct answer that is generated by the system. If there is
a match, a positive message will be sent to the learner. Otherwise, a feedback message
Figure 8. The proposed feedback architecture
Intelligent Computer Assisted Language learning 99
will be sent to the learner. The learner can either read the feedback message and
correct the typed sentence instantly, or restudy the related grammar items and then
correct the sentence without further assistnace. In the following subsections, we show
how the system catches the learner’s errors and how it handles the ill-formed natural
language input.
Rules for Error Analysis
In our implementation, we have augmented the Arabic grammar with heuristic rules
which are capable of parsing ill-formed input (buggy rules) and which apply if the
grammatical rules fail. As an example, consider the following question to complete a
sentence with a suitable unrestricted object ‘‘ :’’المفعول المطلق
أكمل بمفعول مطلق مناسب:
. •أبر أبى ________.
Complete the following with the correct unrestricted object
. ________.I am kind to my father
The following is an analysis of the possible learner’s answer along with the
corresponding feedback:
. A word that is not a noun. Issue a message describing that the unrestricted object
should be a noun.
. A word that is a noun but does not originate from the infinitive verb. Issue a
message describing that the unrestricted object should be the infinitive of the
verb.
. A word that is both a noun and originates from the infinitive verb but is defined.
Issue a message describing that the unrestricted object should be undefined.
. A word that is a noun, originates from the infinitive verb but needs the end case
‘‘Alef Tanween’’, and is undefined. Issue a message describing that a missing end
case of the unrestricted object.
. A Correct answer. Issue a positive message.
From the above analysis of the possible learner’s answer, we augmented the grammar
rule of the unrestricted object by the heuristic rules (buggy rules), defined by
check_uo_correctness/4, that handle every possible ill-formed construction as follows.
unrestricted_object(UO,Infinitive) – 4[Word],
{morph(Word, lex(Stem,Category, _, _, Def, _, _, _)),
check_uo_correctness(Infinitive,Feedback,Word,[Stem,Category,Def]),
(Feedback = = إجابة صحيحة ’- 4 UO= unrestricted_object(Word)
; UO= incorrect_unrestricted_object(Word,Feedack) )
}.
check_uo_correctness(_, Feedback, Word, [_,Category,_]):- not Category= noun,!,
error_flagging(not(noun),‘‘ المفعول المطلق ’’, Feedback,[Word]).
100 K. F. Shaalan
check_uo_correctness(Infinitive, Feedback, Word,[Stem, noun,_]):- not Stem=Infinitive,
!,
error_flagging(not(infinitive), ‘‘ المفعول المطلق ’’, Feedback,[Word]).
check_uo_correctness(Infinitive, Feedback, Word,[Infinitive, noun, defined]):-
error_flagging(not(undefined_noun), ‘‘ المفعول المطلق ’’, Feedback, [Word]),!.
check_uo_correctness(Infinitive, Feedback, Word,[Infinitive, noun, undefined]):-
name(Word, Str),
name(Infinitive,Str1),
(need_alaf_tanween(Infinitive) -4
(append(Str ا‘‘, 1 ’’,Str) -4Feedback = ‘‘ ’’إجابة صحيحة
; error_flagging(need_alaf_tanween, ‘‘ المفعول المطلق ’’, Feedback, [Word])
)
; Feedback = ‘‘ ’’إجابة صحيحة
).
Error Handling Mechanism
Learner’s responses which have special handling mechanisms in case of ill-formed
learner input are: linguistic analysis, classification into categories, sentence
transformation, and completing a sentence. They are discussed in the following
subsections.
Handling of linguistic analysis. Linguistic analysis questions can apply either for an
entire sentence or a part of it. The latter is usually a sequence of words between
brackets. The following description outlines the steps for handling of linguistic
analysis:
. Parse the given sentence (or the sequence of words between brackets) and
generate its linguistic analysis in a quadruple abstract representation form.
. Convert learner answer into the abstract representation form.
. Compare the learner’s answer with the generated answer to issue the appropriate
feedback message.
Example:
هل تري فرقا إعرابيا فيما بين القوسين؟
* استمتعت بجو الريف (استمتاعاً).
* أذهب إلى الريف (استمتاعاً) بجوه.
What’s the difference in linguistic analysis of the words between brackets?
& I enjoyed the country weather (very much)
& I go to the countryside (to enjoy) its weather
The parser is used to analyse each of the input sentences. The generated correct
linguistic analyses of the words between brackets are the following:
& First word: [‘ [’مفرد‘ ,’فتحة‘ ,’منصوب‘ ,’مفعول مطلق
[unrestricted object, accusative, fat-hah, singular]
Intelligent Computer Assisted Language learning 101
& Second word: [‘ [’مفرد‘ ,’فتحة‘ ,’منصوب‘ ,’مفعول لأجله
[causative object, accusative, fat-hah, singular]
The difference, in this case, is in the analytic location (i.e., the first argument in the
quadruple abstract representation). The learner’s answer is also converted to the
quadruple abstract representation. The comparison between these representations
will issue the appropriate feedback that describes the source of the error. The possible
source of the errors could be: incorrect analytic location, incorrect end case, incorrect
reason, or a partially correct answer.
Handling of classification into categories. Classification into categories questions can
apply either for identifying morphological categories or for identifying syntactic
constituents (possibly, a complete sentence). The following description outlines the
steps for handling classification into morphological categories:
. Morphologically analyse the words in the given sentence and determine the words
features.
. Generate an N lists, a classification of the words according to the questions words.
. Assign the learner answer to N Lists.
. Compare the learner’s answer with the generated answer to issue the appropriate
feedback message.
Example:
عين الاسم والفعل والحرف في الجملة الآتية:
* وقف التلاميذ احتراماً للمعلم
Identify the category of each of the words in following sentence
& The students stood up respecting the teacher
The morphological analyser is used to analyse each inflected Arabic word to recognise
its category. Then, according to the word category, the words are classified into three
lists. The following is the generated correct answer:
& Verb: [[‘ وقف ’, verb, male, singular, past, . . .]] (stood)
& Noun: [[‘ المعلم ’, noun, male, singular, def, . . .], [‘ احتراما ’, noun, male, singular, undef, . . .],
التلاميذ‘] ’, noun, male, plural, def, . . .]] (the teacher, honoring, the students)
& Particle: [[‘ ال ,’ particle, def_article, . . .]] (the)
The learner’s answer is also assigned to three lists containing verbs, nouns, and
particles, respectively. The comparison between the corresponding lists will issue the
appropriate feedback that describes the source of the error. The possible source of the
errors could be: missing words from the respective morphological category, or
assigning a word to an incorrect morphological category.
The following description outlines the steps for handling of classification into
syntactic constituents:
102 K. F. Shaalan
. Parse the given sentence and determine the parse tree.
. Generate an N lists, a classification of the sentence’s constituents according to the
questions words.
. Assign the learner answer to N Lists.
. Compare the learner’s answer with the generated answer to issue the appropriate
feedback message.
Example:
عين المبتدأ والخبر مبينا نوعة في الجملة الآتية:
* الجنود الشجعان ينتصرون
Identify the inchoative and enunciative, and the type of the enunciative in the following
sentence.
& The brave soldiers fought to victory
The parser is used to analyse the input sentence into a parse tree as follows:
nominal_sentence(inchoative(noun(' الجنود ' noun, male, plural, def, . . .),
adj(' الشجعان ', noun, male, plural, def, adj, . . .)),
enunciative(verbal_sentence(verb(‘ ينتصرون ’, verb, male, plural, present, . . .))))
Then, according to the parse tree, the words are classified into three lists. The
following is the generated correct answer:
& inchoative: [noun(' الجنود ', noun, male, plural, def, . . .), adj(' الشجعان ', noun,
male, plural, def, adj, . . .)] (the brave soldiers)
& enunciative: [verbal_sentence(verb(' ينتصرون ', verb, male, plural, present, . . .))]
(make victory)
& enunciative type: [verbal_sentence]
The learner’s answer is also assigned to three lists containing inchoative,
enunciative, and enunciative type, respectively. The comparison between the
corresponding lists will issue the appropriate feedback that describes the source of
the error. The possible source of the errors could be: incorrect constituent
type (analytic location), or assigning a syntactic constituent to an incorrect
category.
Handling of transformation of a sentence. Transformation questions require the learner
to change/rewrite the form of a sentence. The following description outlines the steps
for handling of transformation of a sentence:
. Parse the given sentence and determine the parse tree; apply a tree-to-tree
transformation to generate the transformed parse tree.
. Parse the learner’s answer to determine the parse tree.
. Compare the parse tree of the learner’s answer with the parse tree of the generated
answer to issue the appropriate feedback message.
Intelligent Computer Assisted Language learning 103
Example:
الجملة الآتية أسمية أجعلها فعلية:
* المعلم يشرح الدرس
Change the following nominal sentence into verbal sentence
* The teacher teaches the lesson
The parser is used to analyse the input sentence into a parse tree as follows:
nominal_sentence(inchoative(noun(' المعلم ', noun, male, singular, def, . . .)),
enunciative(verbal_sentence(verb(' يشرح ', verb, male, singular, present, . . .),
object(noun(' الدرس ', noun, male, singular, def, . . .)))))
Then, the parse tree of the nominal sentence is transformed to the following verbal
sentence.
verbal_sentence(verb(' يشرح ', verb, male, singular, present, . . .)),
subject(noun(' المعلم ', noun, male, singular, def, . . .)),
object(noun(' الدرس ', noun, male, singular, def, . . .)))
In addition, words of the transformed parse tree is grouped in a list as follows:
& List of words: [verb(' يشرح ', verb, male, singular, present, . . .),
noun(' المعلم ', noun, male, singular, def, . . .), noun (' الدرس ', noun, male, singular, def, . . .)]
The learner’s answer is also analysed into a parse tree and words are grouped into a
list. The comparison between the corresponding representations will issue the
appropriate feedback that describes the source of the error. The possible source of
errors could be: extra words, missing words, grammatically incorrect sentence, or
incorrect transformation of a word (incorrect verb: tense, number, gender, etc;
incorrect noun: number, gender, definition, etc).
Handling of Fill-in-the-blanks. Fill-in-the-blank questions can apply for the generation
of isolated words with a particular form, or to the generation of words to
complete a sentence that achieves feature agreement among its constituents. The
following description outlines the steps for handling of rewriting of isolated words
with particular morphological form:
. Morphologically generate the given words and determine their features.
. Morphologically analyse the learner’s answer and determine the words features.
. Compare the parse tree of the learner’s answer with the parse tree of the generated
answer to issue the appropriate feedback message.
Example:
ثن واجمع سالما الكلمات الآتية:
* المهندس
* الثمرة
* صحراء
104 K. F. Shaalan
What is the correct dual and regular plural of the following words?
& Engineer
& Fruit
& Desert
The morphological generator is used to synthesise the given words into the dual and
plural forms taking into consideration the possible end case. The following is the
generated correct answer:
& Dual list: [[' المهندسان ,' noun, male, dual, def, nominal, . . .], [[' المهندسين ', noun, male,
dual, def, accusative, . . .], [' الثمرتان ', noun, female, dual, def, nominal, . . .], [' ,'الثمرتين
noun, female, dual, def, accusative, . . .], [' صحراوان ', noun, female, dual, undef,
nominal, . . .], [' صحراوين ', noun, female, dual, undef, accusative, . . .]]
& Plural list: [[' المهندسون ', noun, male, plural, def, nominal, . . .], [[' المهندسين ', noun, male,
plural, def, accusative, . . .], [' الثمرات ', noun, female, plural, def, . . .], [' صحراوات ', noun,
female, plural, undef, . . .]]
The morphological analyser is used to analyse each inflected Arabic word of the
learner’s answer. These words are classified into two lists containing dual and plural
forms, respectively. The comparison between the corresponding lists will issue the
appropriate feedback that describe the source of the error. The possible source of
errors could be: different word, incorrect word category (switch dual forms with
plural forms), incorrect generation of a word (incorrect verb: tense, number, gender,
etc; incorrect noun: number, gender, definition, etc).
The following description outlines the steps for handling of fill-in-the-blank to
achieve agreement among sentence constituents:
. Analyse the sentence and determine its constituents.
. Morphologically generate the missed words.
. Parse the learner’s answer to determine the parse tree.
. Compare the parse tree of the learner’s answer with the parse tree of the generated
answer to issue the appropriate feedback message.
Example:
أكمل بمفعول مطلق مؤكد للفعل:
* أبر أبي ________.
Complete the sentence with an unrestricted noun that makes the sentence more
emphatic?
& I am kind with my father ________.
The parser is used to analyse the partial input sentence and determine its constituents
as follows:
verbal_sentence(verb(' أبر ', verb, male, singular, present, intrans, ' ,(. . . 'بر
subject(noun(' أبي ', noun, male, singular, undef, . . .)))
Intelligent Computer Assisted Language learning 105
The morphological generator uses the infinitive verb ‘ بر ' (kindness) of the main verb to
synthesise the unrestricted object ` برا ' (extremely kind). The learner's answer is also
analysed into a parse tree. The comparison between the corresponding representations will
issue the appropriate feedback that describe the source of the error. The possible source of
errors could be: different word (sense or category), incorrect morphological generation of a
word (incorrect verb: tense, number, gender, etc; incorrect noun: number, gender,
definition, etc), incorrect syntactic generation of a word (`s' not in emphatic form, does not
originates from the infinitive verb, etc)
Conclusions
In this paper, we described the development of an ICALL system for learning Arabic
by students at primary schools or by learners of Arabic as a second or foreign
language. NLP tools can be useful for ICALL and hence usefully used in Arabic
ICALL for reacting to the learner’s response in order to give appropriate feedback to
the learner. Learner-system communication in free natural language is computationally
the most challenging and pedagogically the most valuable scenario in Arabic
ICALL. The deep syntactic analysis of the learner’s answer, whether correct or
wrong, is compared against a system-generated answer. This enables feedback
elaboration that helps learners to understand better their knowledge gab.
Arabic ICALL has two main types of test items for interaction with the learner:
selection-type and supply-type requiring the learners to write a few words. The
objective test method is used such that there is no ambiguity about what that correct
answer should be.
The rule-based approach is used to give some freedom to the language learners in
the way they phrase their answers, while enabling the exercise author to enter only
one possible correct answer, thus saving much time compared to the previous pattern
matching answer coding approach. While the statistical-based approach is costly and
has some difficulties in distinguishing between well-formed and ill-formed input, the
rule-based approach has the advantage of providing detailed analysis of the learner’s
answer using linguistic (morphological and syntactic) knowledge. NLP tools in
Arabic ICALL include a morphological analyser and a syntax analyser. These tools
are used to analyse inflected Arabic words and Arabic sentences, and provide a
linguistic analysis of an Arabic sentence. The error analyser is responsible for
diagnosing and handling of ill-formed input.
Arabic ICALL was implemented using SICStus Prolog, Visual Basic, Flash and
Microsoft Access. The system is transportable and capable of running on an IBM
PC which allows the learner to use it to learn Arabic language anywhere and
anytime.
We plan to enrich the present system—for example, add a student or course
management facility that allows us to have full performance and record-keeping
features, add more multimedia instructional material to allow learners fully
comprehend what they learn in natural settings, make the system available on the
Internet to serve remote learners worldwide (especially learners of Arabic as a second
106 K. F. Shaalan
language), and extend the grammar coverage to include more advanced grammar
levels.
Notes
1. The author is on leave of absence from the Faculty of Computers and Information, Cairo
University.
2. CALL is also known as computer-assisted instruction (CAI), computer-aided instruction (CAI),
or computer-aided language learning.
3. We refer to Cachia (1973) for the translation of Arabic terminology.
References
Abou Ela, M. (1994). Knowledge-based techniques in an Arabic syntax analysis environment.
Unpublished doctoral thesis, Institute of Statistical Studies and Research, Cairo University,
Egypt.
Allen, J. (1995). Natural language understanding (2nd ed.). Redwood City, CA: The Benjamin/
Cummings Publishing Company.
Boytcheva, S., Vitanova, I., Strupchanska, A., Yankova, M., & Angelova, G. (2004). Towards the
assessment of free learner’s utterances in CALL. Proceedings of InSTIL/ICALL2004—NLP and
Speech Technologies in Advanced Language Learning Systems, 17–19 June 2004, Universita Ca
Fascari, Venice, Italy, 187–190.
Cachia, P. (1973). The monitor: A dictionary of Arabic grammatical terms. Beirut: Librairie du Liban.
Cameron, K. (Ed.). (1999). CALL-media, design and applications. Lisse, The Netherlands: Swets &
Zeitlinger.
Cushion, S., & He´mard, D. (2002). Applying new technological developments to CALL for Arabic.
Computer Assisted Language Learning (CALL): An International Journal, 15(5), 501 – 508.
Cushion, S., & He´mard, D. (2003). Designing a CALL package for Arabic while learning the
language ab initio. Computer Assisted Language Learning (CALL): An International Journal, 16
(2/3), 259 – 266.
Ditters, E., Oostdijk, N., & Cameron, K. (1993). Processing Arabic: Computer assisted language
learning. Oxford, England: Intellect Books.
Dougherty, R. (1994). Natural language computing: An English generative grammar in prolog.
Mahwah, NJ: Lawrence Erlbaum Associates.
Gamper, J., & Knapp, J. (2002). A review of intelligent CALL systems. Computer Assisted Language
Learning (CALL): An International Journal, 15(4), 329 – 342.
Gazdar, G., & Mellish, C. (1990). Natural language processing in prolog: An introduction to
computational linguistics. Massachusetts, MA: Addison-Wesley.
Gheith, M., Dawa, I., & Afifty, M. (1996). ISTAL: Instructional software for teaching the Arabic
language. Proceedings of the First International Conference on Computer and Advanced Technology
in Education, Egypt, 41 – 66.
Hegazi, N., Ali, G., Abed, E. M., & Hamada, S. (1989). Arabic expert system for syntax education.
Proceedings of the Second Conference on Arabic Computational Linguistics, Kuwait, 596 – 614.
Holland, M. V., Maisano, R., & Alderks, C. (1993). Parsers in tutors: What are they good for?
CALICO Journal, 11(1), 28 – 46.
Holland, M. V., Kaplan, J. D., & Sama, M. R. (Eds.). (1995). Intelligent language tutors: Theory
shaping technology. Mahwah, NJ: Lawrence Erlbaum Associates.
Lam, F. S., & Pennington, M. C. (1995). The computer vs the pen: A comparative study of word
processing in a Hong Kong secondary classroom. Computer Assisted Language Learning
(CALL): An International Journal, 8(1), 75 – 92.
Intelligent Computer Assisted Language learning 107
McEnery, T., Baker, J. P., & Wilson, A. (1995). A statistical analysis of corpus based computer vs.
traditional human teaching methods of part of speech analysis. Computer Assisted Language
Learning (CALL): An International Journal, 8(2), 259 – 274.
Mote, N., Johnson, L., Sethy, A., Silva, J., & Narayanan, S. (2004). Tactical language detection
and modeling of learner speech errors: The case of Arabic tactical language training for
American English speakers. Proceedings of InSTIL/ICALL2004—NLP and Speech Technologies
in Advanced Language Learning Systems, Italy. Retrieved May 2005 from http://sisley.cgm.
unive.it/ICALL2004/link13.htm.
Nielsen, H. (2001). ArabVISL: Arabic grammar on the Internet. Proceedings of the Fifth Nordic
Conference on Middle Eastern Studies, Lund. Retrieved May 2005 from http://www.hf-fak.vib.
no/institutter/smi/pal/paltoc.html
Nielsen, H., & Carlsen, M. (2003). Interactive Arabic grammar on the Internet: Problems and
solutions. Computer Assisted Language Learning (CALL): An International Journal, 16(1), 95 –
112.
Othman, E., Shaalan, K., & Rafea, A. (2003). A chart parser for analyzing modern standard Arabic
sentence. Proceedings of the MT Summit IX Workshop on Machine Translation for Semitic
Languages: Issues and Approaches, USA. Retrieved May 2005 from http://www-2.cs.cmu.
edu~alavie/semitic-MT-wshp.html
Pereira, F., Shieber, C., & David, H. (1986). Definite clause grammars for language analysis—a
survey of the formalism and a comparison with augmented transition networks. In B. Grosz,
K. Jones & B. Webber (Eds.), Readings in natural language processing (pp. 24 – 101). San
Francisco, CA: Morgan Kuffmann.
Rafea, A., & Shaalan, K. (1993). Lexical analysis of inflected Arabic words using exhaustive search
of an augmented transition network. Software Practice and Experience, 23(6), 567 – 588.
Shaalan, K. (2003). Arabic GramCheck: a grammar checker for Arabic. Egyptian Informatics
Journal, Faculty of Computers & Information, 4(1), 94 – 111.
Shaalan, K., Allam, A., & Gomah, A. (2003). Towards automatic spell checking for Arabic.
Proceedings of the Fourth Conference on Language Engineering, Egyptian Society of Language
Engineering (ELSE), Egypt, 240 – 247.
Swartz, M., & Yazdani, M. (Eds.). (1992). Intelligent tutoring systems for foreign language
learning, chapter introduction. New York: Springer-Verlag.
Warschauer, M. (1996). Computer-assisted language learning: An introduction. In S. Fotos (Ed.),
Multimedia language teaching (pp. 3 – 20). Tokyo, Japan: Logos International.
Woods, W. (1970). Transition network grammar for natural language analysis. comm. ACM, 10,
591 – 66.
108 K. F. Shaalan

  • Digg
  • Del.icio.us
  • StumbleUpon
  • Reddit
  • RSS