Skip to content
View emansafwatm's full-sized avatar

Block or report emansafwatm

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
emansafwatm/README.md

Hi, I’m Eman Khater

I am an AI researcher specializing in Large Language Models, Natural Language Processing, multilingual machine translation, multimodal artificial intelligence, knowledge representation, semantic AI, and intelligent multimodal systems.

My research focuses on transformer model development, Arabic language technologies, distributed GPU training on high-performance computing platforms, model optimization, and efficient AI deployment.

My long-term research objective is to develop efficient, trustworthy, and multimodal AI systems that integrate language understanding, knowledge representation, and human-centered interaction.

Research Interests

  • Large Language Models
  • Natural Language Processing
  • Machine Translation
  • Multilingual and low-resource language technologies
  • Arabic language technologies
  • Multimodal artificial intelligence
  • Knowledge-enhanced AI
  • Knowledge representation and semantic technologies
  • Embodied artificial intelligence
  • Human–computer interaction
  • Extended Reality
  • Transformer optimization and efficient deployment
  • High-performance and distributed computing

Selected Research Projects

Vocalized Arabic–English Parallel Dataset

A reproducible pipeline for filtering, processing, and automatically vocalizing Arabic–English parallel data for machine translation, Arabic diacritization, and Arabic NLP research.

View repository

Resolving Lexical Ambiguity in Arabic Neural Machine Translation

A public research overview of an ongoing study examining the effects of Arabic diacritization and character-level versus word-level tokenization on bidirectional Arabic–English neural machine translation. The manuscript is currently under peer review, while the complete experimental materials remain private during the review process.

View repository

Arabic–English Machine Translation Assessment

A reproducible research project evaluating Arabic–English machine translation systems using lexical and semantic metrics, including BLEU, chrF++, BERTScore, SBERT, and COMET.

View repository

Transformer Training Strategies for Morphologically Rich Machine Translation

A work-in-progress revision of a study comparing models trained from scratch with multilingual pretrained Transformer models for Arabic–English machine translation. The experimental design, results, and conclusions are currently being reassessed and validated.

View repository

Ontology-Based Adaptive Learning Research

A historical research portfolio documenting my MSc research on ontology-based adaptive assessment, semantic reasoning, Arabic financial-accounting knowledge representation, and comparison with Item Response Theory. The work resulted in three peer-reviewed publications.

View repository

AI/ML Teaching Portfolio

A collection of implementation-oriented notebooks covering artificial intelligence, machine learning, data science, natural language processing, deep learning, and computer vision.

View repository

Additional Research Directions

Embodied, Low-Latency Language Agents for Extended Reality

A proposed hybrid edge–cloud framework for context-aware language agents in Extended Reality environments. The framework combines on-device multimodal perception with cloud-based semantic reasoning using Large Language Models.

The research investigates model quantization, knowledge distillation, anticipatory dialogue prediction, multimodal interaction, response latency, user trust, presence, and cognitive load.

Knowledge-Based Adaptive Learning Systems

Research on ontology-driven adaptive examination systems using OWL, SPARQL, semantic reasoning, and domain-specific knowledge representation for personalized computer-based assessment.

Publications and Research Output

  • Author of peer-reviewed research in Arabic NLP, machine translation, knowledge representation, semantic technologies, and adaptive learning systems.
  • Published and accepted research on vocalized Arabic–English parallel data for machine translation and diacritization.
  • Two first-author journal manuscripts are currently under peer review, and an additional machine translation study is undergoing substantial revision before resubmission.
  • Conducted large-scale distributed GPU experiments using the Bibliotheca Alexandrina High-Performance Computing cluster.

Selected Publications

  • Vocalized Arabic–English Parallel Corpus: A High-Quality Resource for Machine Translation and Diacritization
    E. Khater, M. E. Shehab, M. Aborezqa, and K. Mahar.
    International Conference on Computer Technology and Applications (ICCTA 2025).

    View publication

  • Arabic Ontology Model for Financial Accounting
    E. Khater, A. Hegazy, and M. Sakre.
    International Conference on Soft Computing and Software Engineering (SCSE 2015).
    View publication

  • Ontology-Based Adaptive Examination System in E-Learning Management Systems
    E. Khater, A. Hegazy, and M. E. Shehab.
    IEEE Seventh International Conference on Intelligent Computing and Information Systems (ICICIS 2015).
    View publication

  • Comparing Ontology-Based and Item Response Theory in Computer Adaptive Testing
    E. Khater, A. Hegazy, and M. E. Shehab.
    IEEE Seventh International Conference on Intelligent Computing and Information Systems (ICICIS 2015).

    View publication

Technical Skills

Programming Languages: Python, C++, C#, SQL, JavaScript

AI and Machine Learning: PyTorch, TensorFlow, Hugging Face Transformers, scikit-learn

Natural Language Processing: MarianMT, NLLB-200, multilingual Transformers, Sentence Transformers

Evaluation: BLEU, chrF, BERTScore, SBERT, COMET

Model Optimization: Quantization, knowledge distillation, BitsAndBytes, SparseGPT, AWQ

High-Performance Computing: Linux, CUDA, SLURM, distributed GPU training

Deployment and Development: Docker, Git, GitHub, ONNX Runtime, TensorRT, Jupyter Notebook, Google Colab

Knowledge and Semantic Technologies: RDF, OWL, SPARQL, Protégé, Neo4j

Data Analytics: pandas, NumPy, Matplotlib, Power BI, R, statistical modeling

Additional Technologies: Angular, Node.js, ArcGIS, Google Earth Engine

Teaching Experience

I have taught and supported undergraduate and professional courses in:

  • Natural Language Processing
  • Machine Learning
  • Artificial Intelligence
  • Deep Learning for NLP
  • Python programming
  • Data exploration
  • Data structures
  • Object-oriented programming
  • Computer architecture

I have also supervised undergraduate capstone and research projects in NLP, Large Language Models, and intelligent systems.

Current Research Direction

My current work investigates multilingual Large Language Models, transformer optimization, Arabic–English machine translation, knowledge-enhanced AI, and multimodal language systems.

I am particularly interested in developing efficient and trustworthy AI models that combine semantic understanding, multimodal perception, knowledge representation, and human-centered interaction.

Research Opportunities

I am interested in funded PhD and research opportunities related to:

  • Large Language Models
  • Natural Language Processing
  • Machine Translation
  • Multilingual and Arabic language technologies
  • Multimodal artificial intelligence
  • Knowledge-enhanced and semantic AI
  • Embodied AI and Extended Reality
  • Human–computer interaction
  • Efficient model training and deployment

Contact

Pinned Loading

  1. Vocalized-Arabic-Parallel-Dataset Vocalized-Arabic-Parallel-Dataset Public

    Reproducible pipeline for filtering and vocalizing Arabic-English parallel data, with full and partial diacritization variants for NLP and machine translation research.

    Python

  2. From-Syntax-to-Semantics-A-Comprehensive-Assessment-of-Arabic-English-Machine-Translation From-Syntax-to-Semantics-A-Comprehensive-Assessment-of-Arabic-English-Machine-Translation Public

    Reproducible evaluation of Arabic–English neural machine translation using lexical, character-based, and semantic quality metrics.

    Python

  3. Resolving-Lexical-Ambiguity-in-Arabic-Neural-Machine-Translation-Overview Resolving-Lexical-Ambiguity-in-Arabic-Neural-Machine-Translation-Overview Public

    Public research overview of a study on Arabic diacritization, lexical ambiguity, and tokenization strategies in Arabic-English neural machine translation.

  4. Ontology-Based-Adaptive-Learning-Research Ontology-Based-Adaptive-Learning-Research Public

    Historical research portfolio on ontology-based adaptive assessment, semantic reasoning, and Arabic financial-accounting knowledge representation.

    C#

  5. Evaluating-Transformer-Training-Strategies-for-Morphologically-Rich-Machine-Translation Evaluating-Transformer-Training-Strategies-for-Morphologically-Rich-Machine-Translation Public

    Work-in-progress research repository for a revised study comparing Transformer training strategies in Arabic–English machine translation.

  6. AI-ML-Teaching-Portfolio AI-ML-Teaching-Portfolio Public

    Curated teaching portfolio covering artificial intelligence, machine learning, data science, NLP, deep learning, and computer vision through implementation-oriented Jupyter notebooks.

    Jupyter Notebook 1