Youssef Zaghloul is Egyptian and has a background in computer engineering. After studying at the Egypt-Japan University of Science and Technology in Alexandria and spending an exchange year at Kyushu University in Japan, he discovered computational linguistics and natural language processing. Lately, his academic path has taken him to Spain and Italy, within the framework of a European Master in Linguistic Data Science (EMLDS), promoted by the Catholic University of Milan, the University of Zaragoza, and the Nova University of Lisbon.
While the Arab world is generally represented as dominated by the nostalgia of its past, some countries are in fact heavily investing in technological innovation and artificial intelligence
Last update: 2026-09-29 12:46:21
Martino Diez interviews Youssef Zaghloul, a young Egyptian researcher in computational linguistics
Youssef, you started as a computer engineer, and you ended up studying computational linguistics. How did it happen?
During my undergraduate studies I approached problems mainly as an engineer, that is computationally: you have a dataset, you have a model, and you try to adapt the model to the task. As the saying goes, if all you have is a hammer, everything looks like a nail. In particular, my graduation thesis focused on machine translation for texts that switch between Egyptian Arabic and English, which is very common in everyday communication in Egypt. The difficulty was twofold: code-switching itself, and the fact that Egyptian Arabic was much less represented in available data than English or Modern Standard Arabic, at least some years ago. At some point, I realized that improving the model was not enough. I also needed to understand how the language works and why the model was making specific errors. Trying to analyze these errors pushed me beyond model fitting and made me more attentive to the linguistic and humanities side of the problem. This coincided with my experience in Japan.
Japan? This is not a usual destination for Egyptian students.
Indeed. After the secondary school, I enrolled at the Egypt-Japan University of Science and Technology. This is a very recent reality, launched as a partnership between the Egyptian and the Japanese government in 2009, and it is now one of the leading technological universities in my country. When I enrolled, though, it was still little known. Luckily, some friends mentioned it to me, and I decided to attend it, even if it was far away from my town and a little more expensive than the average. As this university is a joint initiative with Japan, we had to take two courses in Japanese, and I had one exchange year in Kyushu. It was my first experience abroad and it was very enriching but at the same time challenging, even though I stayed mostly within the international community. This experience stimulated my interest in humanities and linguistics, although at that time I didn’t know that there was a field called computational linguistics. In the master program that I am currently attending we are continually exposed to these subjects, and it is amazing to me, because I didn’t expect myself to be enjoying this part of the program. Such a mixing of science and humanities, as I find in the courses of Professor Passarotti and Mambrini for instance, is not common in Egypt.
How did your family react to your choice?
I come from Beni Suef in Upper Egypt. My older siblings all went into medical fields (medicine and pharmacy), so there was a natural expectation that I might follow a similar path. But I didn’t. I scored 98/100 in the final high school exams, which is considered a high score in Egypt, and I chose engineering instead.
Engineering is a respectable subject, though!
Well, the worst part was that, after choosing computer engineering, I gradually moved toward becoming a linguistic engineer… This was really difficult to explain. In fact, it is mainly a social thing. For many families, especially outside Cairo, having a doctor or an engineer in your family means economic security, while humanities are often seen as less useful. Now, however, since the economy is not performing at its best, even jobs like pharmacist or lawyer cannot guarantee a future for you. Perhaps only medicine still can. At any rate, in my country there is a strong separation between sciences and humanities. This is one reason why interdisciplinary fields such as computational linguistics are still not widely known. It may look like a luxury for a developing country.
Other parts of the Arab world are investing heavily in this field.
Yes, for example, in the Gulf, there are several important initiatives. At New York University Abu Dhabi, Professor Nizar Habash established the CAMeL Lab, the Laboratory for Computational Approaches to Modeling Language, which is specifically devoted to Arabic and with which I have also had the opportunity to work. The Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) is another important example, working on language models, datasets, and dedicated infrastructure. Qatar too is investing significantly in this field. A few years ago, people often described Arabic as a low-resource language. I think that is no longer true, at least for Modern Standard Arabic. Data, tools, and models are growing quickly, and the situation is also improving for many dialects, although some remain less well resourced than others.
Why does Arabic require particular attention? Couldn’t you simply apply the models developed for English?
Arabic is morphologically very rich. A single word can carry information that would be distributed across several words in English, for example tense, person, negation, voice. If you apply methods developed for English too mechanically, you can lose part of that structure. I have understood this problem much better thanks to a shared task (AbjadNLP 2026, co-located with EACL 2026) in which I participated last January. It was about stylometry.
What is this?
It is about author recognition. It is very useful to attribute texts but also to teach language models how to mimic the style of a particular author. I first encountered the subject in a class with Professor Mambrini: we were working on digital humanities, and he mentioned that writers can leave unconscious stylistic signals in their pages. This didn’t make sense to me at first and the engineer that was in me rebelled. I used to think that what matters is the message that you want to transmit, while the form is irrelevant. Studies on stylometry were mainly on English. OK, I said to myself, maybe it can work for English, but for Arabic?
I participated in a challenge on authorship attribution. We were asked to identify the author of an Arabic text using stylometric features within a closed set of 21 candidates. In English, quite simple techniques like the frequency of stop words worked extremely well. In Arabic, without morphological processing, performance was much lower because Arabic root-and-pattern morphology generates over 10,000 verb forms per root. When I introduced morphological analysis, the results improved dramatically. My system reached about 96 percent accuracy, and I won the first place. What interested me was not only the score. Because my model was small, I could understand why it was working.
I also wanted to see whether an LLM could imitate the style of particular writers. In my experiments, the model came much closer to the target style in English. In Arabic, so far, the generated texts remain much further away, and the styles tend to overlap.
There is something almost paradoxical about linguistic engineering leading you back to Arabic morphology.
And syntax, and semantics, and prosody... At school most of my education was in English. In Arabic I was only taught religion and some literature. Because of that, I was not always fully aware of the linguistic and rhetorical richness of my own mother tongue. I did not have a particularly deep knowledge of Arabic grammar or rhetoric. Many Egyptian students consider these subjects difficult and boring. Through computational linguistics I began to study morphology, syntax, semantics, and related aspects of Arabic much more seriously. There is a paradox here: we could end up in a situation in which a computer knows Arabic morphology and the average user does not. It is weird, but this is what is happening right now, I would say.
How important were your experiences abroad?
As I said, Japan was my first international experience, and it changed me a great deal. I would say it has shaped a lot of who I am today, and I am really grateful for this experience. I had studied Japanese, but everyday communication and building relationships were still difficult. At the same time, meeting students from many countries made me more curious about similarities and differences between languages and cultures. I remember a Tunisian friend with whom I often spoke English: she could understand Egyptian Arabic quite well, while Tunisian Arabic was much harder for me.
And between Italy and Spain, where did you feel closer to home?
In Egypt, we often feel that Italians, Spaniards, and Greeks are somehow closer to us than people from other European countries, perhaps because of a shared Mediterranean culture. So far I have felt slightly closer to Spain, although I think the city matters a lot. Milan is large and international, while Zaragoza is smaller and it is easier to come into contact with people.
What is the main challenge for AI and language technology in the Arab world?
The main challenge for me is to bring technical expertise and linguistic knowledge into closer dialogue. The Arab world now has growing investment, research centers, and many qualified people. In Egypt in particular, there is a lot of talent and potential, but stronger mentorship, research opportunities, and interdisciplinary training could help develop it further. I believe that we should also be more aware of the importance of the humanities in the computational world. This could, in turn, increase the motivation to study them.
The opinions expressed in this article are those of the authors and do not necessarily reflect the position of the Oasis International Foundation
© ALL RIGHTS RESERVED