Quote of the day

A person who never made a mistake never tried anything new.

- Albert Einstein

Contact

Detection of AI-Generated Social Engineering in Code-Switched Darija-French Communications

Open Cybersecurity

Detection of AI-Generated Social Engineering in Code-Switched Darija-French Communications

Detection of AI-Generated Social Engineering in Code-Switched Darija-French Communications

Description

Investigate the detectability of AI-generated phishing and social-engineering messages in communications that mix Moroccan Darija, French, and Modern Standard Arabic, and develop a robust detection model.

Details

Context and Problem Statement

Phishing detectors often rely on linguistic cues learned from English corpora. Generative AI removes many of these cues, while Moroccan communication frequently mixes Darija, French, and Modern Standard Arabic, a pattern rarely represented in training data.

Research Question

Which linguistic and structural signals remain useful for distinguishing AI-generated social-engineering messages from legitimate communication under code-switching, and do they remain robust under adversarial paraphrasing?

Proposed Approach

Construct a balanced corpus of anonymized legitimate messages, real phishing samples, and generated variants annotated for manipulation strategies. Compare multilingual encoders, stylometry, and structural features such as headers and link reputation, followed by adversarial paraphrase testing.

Expected Contribution

A North-African multilingual social-engineering benchmark and a robustness analysis under AI-assisted paraphrasing.

Expected Prototype

A message-scoring service suitable for integration into email or messaging gateways, with explanation and adjustable risk thresholds.

Datasets

Nazario, Enron, and a consent-based collected multilingual corpus.

Challenges

Privacy protection, imperfect labels, potential bias against legitimate language varieties, and the evolving generation-detection arms race.

Research Question

Does AI-generated social-engineering detection remain reliable in Darija-French code-switched communication, and how much performance survives adversarial paraphrasing?

Innovation

The project addresses a concrete North-African security blind spot by modeling code-switching explicitly rather than assuming monolingual communication.

Expected Deliverable

An annotated multilingual corpus, an adversarially evaluated detector, and a prototype filtering gateway.

Technologies

Multilingual NLP and Code-Switching Transformers (AraBERT, XLM-R, Atlas-Chat) Stylometry and AI-Generated Text Detection Adversarial NLP Messaging Security

Required Skills

  • Multilingual Natural Language Processing
  • Email and Messaging Security
  • Data Collection and Anonymization Methodology
  • Adversarial Model Evaluation

Datasets

  • Nazario Phishing Corpus
  • Enron Email Dataset
  • Collected and anonymized multilingual corpus

Morocco & Africa Relevance

Phishing campaigns targeting Morocco often exploit local language mixing and messaging-platform use, making context-specific defenses necessary.